6 + 7 Advanced Visualisation

Author

Klára Friesová

You already know how to make basic plots using ggplot functions. We’ll now dive deeper into data visualisation using ggplot2 package and its extensions.

Every plot in ggplot consists of 7 parts that together serve as instructions on how to draw a plot:

To produce any plot, ggplot needs data, mapping and at least one layer. The other parts have some defaults, but often need adjustments for a better appearance.

We will again use the penguin dataset and play with the individual plot parts. In addition we will use some other packages, so please make sure you have them installed and loaded.

library(palmerpenguins)
library(gapminder)
library(readxl)
library(ggrepel)
library(tidyverse)
library(patchwork)

glimpse(penguins)
Rows: 344
Columns: 8
$ species           <fct> Adelie, Adelie, Adelie, Adelie, Adelie, Adelie, Adel…
$ island            <fct> Torgersen, Torgersen, Torgersen, Torgersen, Torgerse…
$ bill_length_mm    <dbl> 39.1, 39.5, 40.3, NA, 36.7, 39.3, 38.9, 39.2, 34.1, …
$ bill_depth_mm     <dbl> 18.7, 17.4, 18.0, NA, 19.3, 20.6, 17.8, 19.6, 18.1, …
$ flipper_length_mm <int> 181, 186, 195, NA, 193, 190, 181, 195, 193, 190, 186…
$ body_mass_g       <int> 3750, 3800, 3250, NA, 3450, 3650, 3625, 4675, 3475, …
$ sex               <fct> male, female, female, NA, female, male, female, male…
$ year              <int> 2007, 2007, 2007, 2007, 2007, 2007, 2007, 2007, 2007…

6.1 Faceting

From Chapter 3, you already know how to visualise the relationship between the penguin bill length and bill depth.

penguins |> 
  ggplot(aes(x = bill_length_mm, y = bill_depth_mm, fill = species, shape = species)) +
  geom_point() + 
  geom_smooth(aes(colour = species), method = 'lm', show.legend = F) + 
  scale_shape_manual(values = c(21, 22, 24)) + 
  labs(x = 'Bill length (mm)', y = 'Bill depth (mm)', fill = 'Species', shape = 'Species') + 
  theme_bw()

But what if we now want to look at the differences between males and females in addition? We could change the definition of the fill and colour aesthetics, so that they display sex and keep just shape for species identity:

penguins |> 
  ggplot(aes(x = bill_length_mm, y = bill_depth_mm, fill = sex, shape = species)) +
  geom_point() + 
  geom_smooth(aes(colour = sex), method = 'lm', show.legend = F) + 
  scale_shape_manual(values = c(21, 22, 24)) + 
  labs(x = 'Bill length (mm)', y = 'Bill depth (mm)', fill = 'Sex', shape = 'Species') +
  theme_bw()

But as you see, it gets a bit hard to read the plot. There are some NA values in the sex column, which would be better removed so that they do not add a mess to the plot. And what about splitting the plot into three smaller plots, one for each species? This is where facetting comes into play.

penguins |> 
  filter(!is.na(sex)) |> 
  ggplot(aes(x = bill_length_mm, y = bill_depth_mm, fill = sex, shape = species)) +
  geom_point() + 
  geom_smooth(aes(colour = sex), method = 'lm', show.legend = F) + 
  scale_shape_manual(values = c(21, 22, 24)) + 
  labs(x = 'Bill length (mm)', y = 'Bill depth (mm)', fill = 'Sex', shape = 'Species') + 
  facet_wrap(~species)+
  theme_bw()

Much better, but do we really need the different shapes for different species if we have them now in different facets? It would be better to add shape differentiation to the sex variable, as not everyone is able to distinguish colours. It might also help in the legend, where we now have just black dots for both levels.

penguins |> 
  filter(!is.na(sex)) |> 
  ggplot(aes(x = bill_length_mm, y = bill_depth_mm, fill = sex, shape = sex)) +
  geom_point() + 
  geom_smooth(aes(colour = sex), method = 'lm', show.legend = F) + 
  scale_shape_manual(values = c(21, 22)) + 
  labs(x = 'Bill length (mm)', y = 'Bill depth (mm)', fill = 'Sex', shape = 'Sex') + 
  facet_wrap(~species)+
  theme_bw()

We can now clearly see that for Adélie penguins, there is a positive relationship between bill length and bill depth for females, but not for males. On the other hand, for Chinstrap penguins, the relationship is stronger for males than for females.

6.2 Data argument

Till now, we have worked just with one data frame for the whole plot. You might remember that it is possible to place aesthetic mappings (aes()) either in the ggplot() call, and then it works for the whole plot, or in a geom function, which then overwrites the global mappings for that layer only. The same is possible for the data argument. We can define a different dataset for a certain layer to add, e.g., points from another dataset or highlight a subset of the data.

As an example, we will draw a scatterplot of penguin body mass vs flipper length, where we highlight penguins from the Dream Island with bigger red points:

penguins |> 
  ggplot(aes(body_mass_g, flipper_length_mm))+
  geom_point()+
  geom_point(data = penguins |> filter(island == 'Dream'), colour = 'red', size = 3)+
  theme_bw()+
  labs(x = 'Body mass (g)', y = 'Flipper length (mm)')
Warning: Removed 2 rows containing missing values or values outside the scale range
(`geom_point()`).

Note that we added specifications of the point appearance of the highlighted point in the geom_point() function outside the aes(). This means the characteristics are fixed for that layer and do not change according to the data. All the points of that layer will always be red and of size 3.

6.3 Scales

Scale functions always follow the pattern scale_{aesthetic}_{type}. We already used scale_shape_manual() to manually rewrite the default shapes used in the plot. We can use scales to modify values, limits, breaks or labels of any aesthetics in our plot, typically colour, fill, and shape. To define desired values on your own, you can use manual scale definition, but there are many other types, e.g. continuous (useful, e.g., for defining labels for continuous variables), discrete (e.g. labels for discrete variables), gradient (e.g. for defining colour gradient for continuous variables). scale_x_{type} and scale_y_{type} are useful to define axis limits, ticks and labels.

We will use the example of penguin body mass vs. flipper length and modify different scales of this plot. First of all, we will show with different colours on which island the penguins were measured. To change the default colours, we can define our own colours using scale_fill_manual() and the values argument inside, similarly to what we did for the shape before. In the values argument, we can define any color using its HEX reference. Some colors also have a name, which might be used instead, see this website for their list.

penguins |> 
  ggplot(aes(body_mass_g, flipper_length_mm, fill = island))+
  geom_point(pch = 21)+
  scale_fill_manual(values = c(Dream = '#7FFFD4', Biscoe = 'orchid3', Torgersen = 'deepskyblue2'))+
  theme_bw()+
  labs(x = 'Body mass (g)', y = 'Flipper length (mm)', fill = 'Island')
Warning: Removed 2 rows containing missing values or values outside the scale range
(`geom_point()`).

But we can also use a predefined colour scale, e.g. from colour brewer using scale_fill_brewer(). We can also modify the labels in the legend by using the argument labels. Let’s say we want to add the word ‘Island’ to each label:

penguins |> 
  ggplot(aes(body_mass_g, flipper_length_mm, fill = island))+
  geom_point(pch = 21)+
  scale_fill_brewer(type = 'qual', palette = 7, labels = c('Dream' = 'Dream Island', 'Biscoe' = 'Biscoe Island', 'Torgersen' = 'Torgersen Island'))+
  theme_bw()+
  labs(x = 'Body mass (g)', y = 'Flipper length (mm)', fill = 'Island')
Warning: Removed 2 rows containing missing values or values outside the scale range
(`geom_point()`).

We used type = 'qual' inside scale_color_brewer() to use colours appropriate for discrete values. You can experiment with many more such colour scales that are already available, e.g., here.

Instead of colouring the points according to the island, we can also use colours to show a numerical variable, e.g. year when the measurement was taken:

penguins |> 
  ggplot(aes(body_mass_g, flipper_length_mm, fill = year))+
  geom_point(pch = 21)+
  theme_bw()+
  labs(x = 'Body mass (g)', y = 'Flipper length (mm)', fill = 'Year')
Warning: Removed 2 rows containing missing values or values outside the scale range
(`geom_point()`).

Note that the plot automatically chooses a continuous colour scale instead of the discrete one before, and the legend now shows a range from the lowest to the highest value. But for the year, it doesn’t really make sense to show decimals. We can modify this using the argument breaks inside scale_fill_continuous():

penguins |> 
  ggplot(aes(body_mass_g, flipper_length_mm, fill = year))+
  geom_point(pch = 21)+
  scale_fill_continuous(breaks = c(2007, 2008, 2009))+
  theme_bw()+
  labs(x = 'Body mass (g)', y = 'Flipper length (mm)', fill = 'Year')
Warning: Removed 2 rows containing missing values or values outside the scale range
(`geom_point()`).

Scales are also useful to set limits, breaks or labels on the plot axes. For example, we might want to zoom only to the penguins with body mass between 3000 and 5000 g:

penguins |> 
  ggplot(aes(body_mass_g, flipper_length_mm, fill = year))+
  geom_point(pch = 21)+
  scale_fill_continuous(breaks = c(2007, 2008, 2009))+
  scale_x_continuous(limits = c(3000, 5000))+
  theme_bw()+
  labs(x = 'Body mass (g)', y = 'Flipper length (mm)', fill = 'Year')
Warning: Removed 72 rows containing missing values or values outside the scale range
(`geom_point()`).

Note that when we do this, we get a warning message, indicating how many points were not displayed because they are outside the defined range. That’s just to prevent us from accidentally missing them.

It is even possible to transform the axis using the transform argument or shortcuts for commonly used transformations, e.g. scale_x_log10().

penguins |> 
  ggplot(aes(body_mass_g, flipper_length_mm, fill = year))+
  geom_point(pch = 21)+
  scale_fill_continuous(breaks = c(2007, 2008, 2009))+
  scale_x_continuous(transform = 'log')+
  theme_bw()+
  labs(x = 'Body mass (g)', y = 'Flipper length (mm)', fill = 'Year')
Warning: Removed 2 rows containing missing values or values outside the scale range
(`geom_point()`).

The scale_x_ and scale_y functions might also be used to change which ticks are displayed on the axes:

penguins |> 
  ggplot(aes(body_mass_g, flipper_length_mm, fill = year))+
  geom_point(pch = 21)+
  scale_fill_continuous(breaks = c(2007, 2008, 2009))+
  scale_y_continuous(breaks = seq(170, 230, 5))+
  theme_bw()+
  labs(x = 'Body mass (g)', y = 'Flipper length (mm)', fill = 'Year')
Warning: Removed 2 rows containing missing values or values outside the scale range
(`geom_point()`).

In addition to color, fill and shape of the points, we can also modify their size, or opacity. Here, we plot the penguing bill length vs bill depth and adjust the size of the points according to their body mass. In addition, we make the points a bit transparent using the alpha argument so that we can see the overlapping points.

penguins |> 
  ggplot(aes(bill_length_mm, bill_depth_mm, fill = island, size = body_mass_g))+
  geom_point(pch = 21, alpha = 0.6)+
  theme_bw()+
  labs(x = 'Body mass (g)', y = 'Flipper length (mm)', fill = 'Island', size = 'Body mass (g)')
Warning: Removed 2 rows containing missing values or values outside the scale range
(`geom_point()`).

We can again play with the size aesthetics, e.g., modify the size of the smallest and largest points, using the scale_size_continuous() function.

6.4 Legend modifications

In all the plots we made so far, the legend was by default placed on the right side next to the plot. Sometimes it is helpful to move it somewhere else, for example, due to space constraints. Legend position might be modified inside the theme() function. We can move it to the bottom of the plot:

penguins |> 
  ggplot(aes(body_mass_g, flipper_length_mm, fill = island))+
  geom_point(pch = 21)+
  theme_bw()+
  theme(legend.position = 'bottom')+
  labs(x = 'Body mass (g)', y = 'Flipper length (mm)', fill = 'Island')
Warning: Removed 2 rows containing missing values or values outside the scale range
(`geom_point()`).

Sometimes, there is a space inside the plot, so it is possible to place the legend there and reduce the image size. In this case, we have to set legend.position = 'inside' and specify two arguments, legend.position.inside() and legend.justification.inside(). The values should indicate coordinates on the two axes and have values of 0-1, where 0 means the beginning of the axis and 1 the end. For example, 0, 0 means the bottom left corner of the plot, 1, 1 the top right corner.

penguins |> 
  ggplot(aes(body_mass_g, flipper_length_mm, fill = island))+
  geom_point(pch = 21)+
  theme_bw()+
  theme(legend.position = 'inside', legend.position.inside = c(1, 0), legend.justification.inside = c(1, 0))+
  labs(x = 'Body mass (g)', y = 'Flipper length (mm)', fill = 'Island')
Warning: Removed 2 rows containing missing values or values outside the scale range
(`geom_point()`).

To remove the legend completely, set legend.position = 'none'.

The theme() function allows us to modify many more plot elements. We can remove any plot component by setting the argument for it to element_blank(), modify the text size, label angle, ticks length, background grid colour and much more. See the documentation for more details.

penguins |> 
  ggplot(aes(body_mass_g, flipper_length_mm, fill = island))+
  geom_point(pch = 21)+
  theme_bw()+
  theme(panel.grid = element_blank(), axis.title = element_text(size = 14, face  = 'bold'), axis.text.x = element_text(angle = 45))+
  labs(x = 'Body mass (g)', y = 'Flipper length (mm)', fill = 'Island')
Warning: Removed 2 rows containing missing values or values outside the scale range
(`geom_point()`).

6.5 Factors and reordering

When using a categorical variable in a plot, the default order of the categories in alphabetical. However, that is not always what we want and it might be helpful to reorder the categories for the final presentation of our plots. For this, turning categorical variables to factors becomes especially useful. We can then use functions from the forcats package, whoch is also part of tidyverse to reorder the categories as desired.

For illustration, we will now use data on the characters from Star Wars coming with the dplyr package.

glimpse(starwars)
Rows: 87
Columns: 14
$ name       <chr> "Luke Skywalker", "C-3PO", "R2-D2", "Darth Vader", "Leia Or…
$ height     <int> 172, 167, 96, 202, 150, 178, 165, 97, 183, 182, 188, 180, 2…
$ mass       <dbl> 77.0, 75.0, 32.0, 136.0, 49.0, 120.0, 75.0, 32.0, 84.0, 77.…
$ hair_color <chr> "blond", NA, NA, "none", "brown", "brown, grey", "brown", N…
$ skin_color <chr> "fair", "gold", "white, blue", "white", "light", "light", "…
$ eye_color  <chr> "blue", "yellow", "red", "yellow", "brown", "blue", "blue",…
$ birth_year <dbl> 19.0, 112.0, 33.0, 41.9, 19.0, 52.0, 47.0, NA, 24.0, 57.0, …
$ sex        <chr> "male", "none", "none", "male", "female", "male", "female",…
$ gender     <chr> "masculine", "masculine", "masculine", "masculine", "femini…
$ homeworld  <chr> "Tatooine", "Tatooine", "Naboo", "Tatooine", "Alderaan", "T…
$ species    <chr> "Human", "Droid", "Droid", "Human", "Human", "Human", "Huma…
$ films      <list> <"A New Hope", "The Empire Strikes Back", "Return of the J…
$ vehicles   <list> <"Snowspeeder", "Imperial Speeder Bike">, <>, <>, <>, "Imp…
$ starships  <list> <"X-wing", "Imperial shuttle">, <>, <>, "TIE Advanced x1",…

First, we can look at the frequency of different eye colors among the characters.

ggplot(starwars, aes(y = eye_color)) + 
  geom_bar()

We might want to reorder the categories according to their frequency. We can do that using the fct_infreq function.

ggplot(starwars, aes(y = fct_infreq(eye_color))) + 
  geom_bar()

We can see here, that there are some very rare eye colors appearing only once. Using the fct_lump_min() function, we can group them together in a category ‘other’.

starwars |> 
  mutate(eye_color = fct_lump_min(eye_color, min = 2) |> 
           fct_infreq()) |> 
  ggplot(aes(y = eye_color)) + 
  geom_bar()

Now, we can look at the relationship between the eye color and total body mass (kg). We will first calculate the mean body mass for each category.

starwars_mean_mass <- starwars |> 
  mutate(eye_color = fct_lump_min(eye_color, min = 2)) |> 
  group_by(eye_color) |>
  summarise(mean_mass = mean(mass, na.rm = TRUE))
starwars_mean_mass
# A tibble: 9 × 2
  eye_color mean_mass
  <fct>         <dbl>
1 black          76.3
2 blue           86.5
3 brown          66.1
4 hazel          66  
5 orange        282. 
6 red            81.4
7 unknown        31.5
8 yellow         81.1
9 Other          94.7

And then visualize the relationship. We will use geom_col(), where we specify the height of the bars in the plot by a numerical variable in our data. Note the difference between geom_bar() used previously, where we only specified one of x or y aesthetics and the other was calculated as a count of indivials in the given category. To order the columns from the smallest to the highest, we will use the fct_reorder() function. Inside the fct_reorder() function, we specify two arguments: the first specifies which variable should be reordered, and the second specifies which values should be taken to make the order. You can imagine this as if you first arrange the table according to the mean body mass, and then fix the order of the categorical variable (eye color) from this ordered table.

starwars_mean_mass |> 
  mutate(eye_color = fct_reorder(eye_color, mean_mass)) |> 
  ggplot(aes(x = eye_color, y = mean_mass)) + 
  geom_col()

We can also order the categories in descending order using desc() inside fct_reorder().

starwars_mean_mass |> 
  mutate(eye_color = fct_reorder(eye_color, desc(mean_mass))) |> 
  ggplot(aes(x = eye_color, y = mean_mass)) + 
  geom_col()

Sometimes, it might be useful to reorder the categories manually. For this, we can use the function fct_relevel().

starwars_mean_mass |> 
  mutate(eye_color = fct_relevel(eye_color, c('yellow', 'orange', 'red', 
                                              'blue', 'hazel', 'brown',
                                              'black', 'unknown', 'Other'))) |> 
  ggplot(aes(x = eye_color, y = mean_mass)) + 
  geom_col()

6.6 Composing plots together

Sometimes it is useful to combine multiple plots in one figure with multiple panels. This might be easily done using the patchwork package.

library(patchwork)

(Please place this line at the top of your script to the other library() calls.)

Let’s say we want to make one figure showing the relationship between the penguin bill length and bill depth for the three penguin species, and a boxplot showing the distribution of body mass for individual species next to each other.

We will first create these two plots and save them to an object:

p1 <- penguins |> 
  ggplot(aes(x = bill_length_mm, y = bill_depth_mm, fill = species, shape = species)) +
  geom_point() + 
  geom_smooth(aes(colour = species), method = 'lm', show.legend = F) + 
  scale_shape_manual(values = c(21, 22, 24)) + 
  labs(x = 'Bill length (mm)', y = 'Bill depth (mm)', fill = 'Species', shape = 'Species') + 
  theme_bw()

p2 <- ggplot(penguins, aes(x = species, y = body_mass_g, fill = species)) +
  geom_boxplot() + 
  labs(y = 'Body mass (g)', fill = 'Species') + 
  theme_bw() + 
  theme(axis.title.x = element_blank())

p1
`geom_smooth()` using formula = 'y ~ x'
Warning: Removed 2 rows containing non-finite outside the scale range
(`stat_smooth()`).
Warning: Removed 2 rows containing missing values or values outside the scale range
(`geom_point()`).

p2
Warning: Removed 2 rows containing non-finite outside the scale range
(`stat_boxplot()`).

And then combine the two plots together:

p1 + p2
`geom_smooth()` using formula = 'y ~ x'
Warning: Removed 2 rows containing non-finite outside the scale range
(`stat_smooth()`).
Warning: Removed 2 rows containing missing values or values outside the scale range
(`geom_point()`).
Warning: Removed 2 rows containing non-finite outside the scale range
(`stat_boxplot()`).

So easy it is. And we can play more with the plot layout. For example, we can adjust the width of each panel in the combined plot.

p1 + p2 + plot_layout(widths = c(2, 1))
`geom_smooth()` using formula = 'y ~ x'
Warning: Removed 2 rows containing non-finite outside the scale range
(`stat_smooth()`).
Warning: Removed 2 rows containing missing values or values outside the scale range
(`geom_point()`).
Warning: Removed 2 rows containing non-finite outside the scale range
(`stat_boxplot()`).

Some useful functionalities of the patchwork package are that you can add annotations to the plots to identify individual panels:

p1 + p2 + plot_annotation(tag_levels = 'a')
`geom_smooth()` using formula = 'y ~ x'
Warning: Removed 2 rows containing non-finite outside the scale range
(`stat_smooth()`).
Warning: Removed 2 rows containing missing values or values outside the scale range
(`geom_point()`).
Warning: Removed 2 rows containing non-finite outside the scale range
(`stat_boxplot()`).

Or collect all legends of the subplots and place them in one place:

p1 + p2 + plot_layout(guides = 'collect')
`geom_smooth()` using formula = 'y ~ x'
Warning: Removed 2 rows containing non-finite outside the scale range
(`stat_smooth()`).
Warning: Removed 2 rows containing missing values or values outside the scale range
(`geom_point()`).
Warning: Removed 2 rows containing non-finite outside the scale range
(`stat_boxplot()`).

And there are many more options to customise your plot layout, annotations etc.

6.7 Other examples of geom functions

Some other useful geom functions include, geom_pointrage() for showing summary characteristics, or model results.

penguins |> 
  group_by(species) |> 
  summarise(min_body_mass = min(body_mass_g, na.rm = T), median_body_mass = median(body_mass_g, na.rm = T), max_body_mass = max(body_mass_g, na.rm = T)) |> ggplot()+
  geom_pointrange(aes(x = species, y = median_body_mass, ymin = min_body_mass, ymax = max_body_mass))

geom_text() or geom_label() display labels instead of points when we want to compare e.g. specific individuals.

starwars |> filter(species == 'Human') |> 
  ggplot(aes(x = mass, y = height, label = name)) + 
  geom_text(alpha = 0.6) 
Warning: Removed 15 rows containing missing values or values outside the scale range
(`geom_text()`).

And geom_text_repel() from the ggrepel package is helpful to distinguish overlapping points.

starwars |> filter(species == 'Human') |> 
  ggplot(aes(x = mass, y = height, label = name)) + 
  geom_text_repel(alpha = 0.6, max.overlaps = Inf) 
Warning: Removed 15 rows containing missing values or values outside the scale range
(`geom_text_repel()`).

6.8 Useful extensions

It is possible to do almost everything you would imagine using ggplot2 and its extensions. We do not have time to cover all of that during this semester, but you are very welcome to experiment and explore this world on your own.

Just a few tips for packages that you might useful for your work:

  • ggpubr to make publication-ready plots - some easy-to-use functions for creating and customising ggplot plots

  • ggeffects to calculate model predictions and plot them (will be replaced by modelbased, but the functionality shouldn’t be lost)

  • gordi a ggplot-based package for making ordination plots, developed in our group

  • ggtree for visualisation of phylogenetic trees

  • ggnewscale to use multiple colour or fill scales in the same plot, especially useful for maps

  • shiny for creating web apps from R

6.9 Exercises

  1. facet_wrap() facets the plot according to one variable, look in the documentation, what the related function facet_grid() does and use it to make a plot of penguin bill length and bill depth faceted by species and island. Use custom colours and point shapes to distinguish between males and females.

  2. Use the gapminder dataset, from the gapminder package (the dataset will be loaded in the environment once you load the package), which provides values for life expectancy, GDP per capita, and population size for each country of the world. Visualise the relationship between GDP per capita (on a log scale) and life expectancy in 2007. Make the size of the points proportional to the country population size. Highlight the Czech Republic and Slovakia with different colours.

  3. Use the same data as in the previous exercise and create a plot showing the development of life expectancy in the Czech Republic and Slovakia over time.

  4. Use the Pokémon dataset (read_csv('https://raw.githubusercontent.com/rfordatascience/tidytuesday/main/data/2025/2025-04-01/pokemon_df.csv')) and visualise the relationship between Pokémon weight and height. Use logarithmic transformation for both axes. Make all points partly transparent. Highlight 20 Pokémons with the highest attack power by different color and larger point size.

  5. Facet the previous plot according to Pokémon type *and color them using colors from the column color_1.

  6. Use the dataset Axmanova-Forest-understory-diversity-analyses.xlsx from the folder forest_understory (Link to the Github folder) and visualise the relationship between radiation and mean Ellenberg-type indicator values for light. Use custom colours to distinguish forest types and change the labels in the legend to ‘Alluvial’, ‘Oak’, ‘Oak hornbeam’ and ‘Ravine’.

  7. Modify axis labels, legend, colours, etc. of all plots you created so far, so that you like their appearance and save them to the folder plots.

Chapter 7

  1. Combine the two plots from exercises 2 and 3 using p1 / p2. What is the difference compared to p1 + p2? Add annotations ‘A’ and ‘B’ to the panels. Modify the legend so that there is only one legend with coloured points at the bottom.

  2. Calculate mean, minimum and maximum of penguin bill length for each species and visualize it using geom_pointrange(). Order the species according to their mean bill length.

  3. Take the plot from exercise 6, order the categories in legend according to the column forest_type. Move legend to the bottom of the plot.

  4. Use the dataset Axmanova-Forest-understory-diversity-analyses.xlsx from the folder forest_understory (Link to the Github folder) and recreate the following plot:

  5. Use the Pokémon dataset (read_csv('https://raw.githubusercontent.com/rfordatascience/tidytuesday/main/data/2025/2025-04-01/pokemon_df.csv')), calculate the mean defense power for each Pokémon type and draw a barplot showing the mean defense power for each type. Order the bars from the highest to the shortest. *Colour the bars according to the colour defined for each Pokémon type in the column color_1. Add errorbars showing the minimum and maximum for each type.

  6. Select electric Pokémons and visualise the relationship between their weight and height. Use their names as text labels instead of points. Do the same for ice Pokémons and then combine both plots together. Add annotations ‘A’ and ‘B’ to the panels.

  7. Modify axis labels, legend, colours, etc. of all plots you created so far, so that you like their appearance and save them to the folder plots.

6.10 Further reading

ggplot2 vignette: https://ggplot2.tidyverse.org

R 4 Data Science: https://r4ds.hadley.nz/visualize.html

Cheatsheet: https://rstudio.github.io/cheatsheets/html/data-visualization.html

Aesthetics specifications vignette: https://ggplot2.tidyverse.org/articles/ggplot2-specs.html

ggplot2: Elegant Graphics for Data Analysis (3e): https://ggplot2-book.org

R Graphics Cookbook: https://r-graphics.org

Fundamentals of Data Visualization: https://clauswilke.com/dataviz/

patchwork vignette: https://patchwork.data-imaginist.com/index.html

ggplot Extensions: https://exts.ggplot2.tidyverse.org/gallery

Mastering Shiny: https://mastering-shiny.org

Tutorial for effective visual science communication: https://ascpt.onlinelibrary.wiley.com/doi/full/10.1002/psp4.12455

Graphics principles cheatsheet: https://github.com/GraphicsPrinciples/CheatSheet/blob/master/NVSCheatSheet.pdf