Workshop 2: Data visualisation

BI3010 Statistics for Biologists

Author

Roos & Pinard

Published

8 September 2026

Learning Objectives

TipLearning objectives
  1. Learn how to use the R package ggplot2 to do data visualisation.
  2. Consider what constitutes an effective visualisation versus an ineffective one.
  3. Learn some techniques to make your figures more visually appealing to an audience.

These skills will be used throughout the course.

Workshop structure

NoteThe Analytical Workflow

These workshops follow an analytical workflow. Today we focus on steps 2 and 3, expanding the toolkit you built in Workshop 1 using data visualisation.

  1. Formulate a research question.
  2. Perform exploratory data analysis [Focus of today].
  3. Identify any hidden assumptions [Focus of today].
  4. Fit an appropriate model.
  5. Diagnose the model and check assumptions.
  6. Summarise the results.
  7. Interpret and provide inferences.

Introduction to ggplot2

ggplot2 is an implementation of the “Grammar of Graphics”, a framework for data visualisation. It allows you to build plots layer by layer, starting with a base plot and then adding additional elements such as scales, themes, and geometries. Using this framework, you can create pretty sophisticated figures with surprisingly little effort. It’s such a popular tool that familiarity with ggplot2 is sometimes asked for by employers; we’ve even seen job adverts from companies such as Spotify listing this an essential criterion.

It is important to note, however, that ggplot2 is not the only way to create visualisations. It’s just an increasingly preferred one. Many of the figures you will make on this course can be near-perfectly recreated using the visualisation framework offered by base R. We simply prefer ggplot2 as we feel it is a more intuitive way to create figures; indeed, nearly every single figure you have and will see on this course were produced in ggplot2.

Load the data

For the first part of this workshop, we’ll be working with data on the Amazon River dolphin, or the boto (Inia geoffrensis, I wonder if the person who named the species had a friend called Geoff…). For this first interaction with the data, we’ll broadly say we’re interested in seeing if there is sexual dimorphism in asymptotic length (i.e. once an animal has stopped growing) between females and males, as well as how the species respond to droughts. As always, before we do any analysis, we need to ensure there is nothing weird in the data. In Workshop 1 you did this using simple summary statistics (e.g. mean()). While that is useful, complementing it with data visualisations is really powerful.

R sessions sometimes reopen with objects from earlier work, including hidden ones (often because you clicked “save workspace” when you quit RStudio). Leftover data or functions can cause headaches and confusion, so if you think this might be affecting you, you can begin by clearing your global environment by running:

rm(list = ls())

This removes all objects in the global environment. It does not delete any files, so don’t worry about losing any of your work. Alternatively, in RStudio you can do the same from the Environment pane by clicking the broom icon (Clear All).

Download the data file below and save it to your BI3010/data/ folder. Open your BI3010.Rproj file to start RStudio with the working directory already set, then open a new script and load the data:

Download boto.txt

boto <- read.table("data/boto.txt", header = TRUE)

Given how easy it is to do, let’s run a quick summary() and str() to see if anything stands out:

summary(boto)
str(boto)

If nothing immediately seems off, we can continue.

Building a first plot

We’ll build this figure up slowly, one layer at a time, to make the underlying philosophy of ggplot2 clear. But before we can do that, we first need to install the package (if you haven’t already) and then load it ready to use:

install.packages("ggplot2")  # Install the software (run once)
library(ggplot2)             # Load the software

Once we have ggplot2 installed and loaded, we can make a simple “ggplot”:

ggplot()

Cool… a completely empty canvas. Well, it shouldn’t really come as a surprise that there’s nothing in this figure. For starters, we haven’t even told it which dataset to use. Let’s correct that now:

ggplot(boto)

Still nothing. What we’re missing here is where we tell ggplot2 a couple of things. First, we need to say in what form we want the data shown. Do we want points? Or lines? Or errorbars? Secondly, we need to say which specific data, or “aesthetics” (shortened to aes() in ggplot2), we want shown and where. Do we want length included? On the x or y axis?

Let’s have length on the y-axis and drought on the x-axis, with the form (or geometry) that we want the data shown as being points:

ggplot(boto) +
  geom_point(aes(y = length, x = drought))

Now we’re getting somewhere. From an early data analysis perspective, what have we already learnt, just from the quick summary stats and from this figure?

First impressions

Question 1. What variables are available in this dataset?

Three variables: sex, drought, and length.

Question 2. How many observations does the dataset contain?

71 boto dolphins.

Question 3. Based on the scatterplot, is the relationship between boto length and number of drought years positive, neutral, or negative?

Very likely negative: longer drought histories appear to correspond to shorter dolphins.

Question 4. What range of lengths do boto dolphins reach at maturity?

From summary(): roughly 170 to 238 cm. The scatterplot gives the same impression.

Additional geom_ layers

The above figure is great. It’s simple and clear (though a bit ugly, we’ll come back to that later), but how would we include a categorical variable like sex? Can we just copy our code from above and replace drought with sex? Let’s see:

ggplot(boto) +
  geom_point(aes(y = length, x = sex))

Kind of? It’s a bit hard to see the spread of data for each sex. A better way to show this is to use some tool that helps the reader understand the range of length for each sex, but also how the data is spread. There are two common options people use for this: boxplots and violin plots. It’s trivially easy to swap between them, because all we need to do is change the geometry function from geom_point() to either geom_boxplot() or geom_violin():

# Boxplot
ggplot(boto) +
  geom_boxplot(aes(y = length, x = sex))

# Violin plot
ggplot(boto) +
  geom_violin(aes(y = length, x = sex))

Both of these present the same information in different ways. The boxplot uses boxes and lines to indicate the range (the first and third quantiles) and the median. The violin plot also shows the spread of data, but by creating density plots which are mirrored. Boxplots have a long history and use in science, but people are increasingly using violin plots, simply because they’re more transparent and clearer than boxplots.

With either boxplots or violin plots, it’s fairly common that people add data points to the figure as well, to really hammer home the spread of data. But we’ve seen before that using geom_point() basically gives us a line of points that doesn’t help us understand the spread at all.

An alternative is to use something called “jittering”. When we jitter points, we add a small random value to either or both the x and y position to spread the data out a little. With our boto data we wouldn’t want to add anything to y (we’re interested in the length of the dolphins, so we don’t want to mess that up), but we can happily add something to x (or sex). To do so, we use geom_jitter() and specify that “noise” should only be added to x and not y, which we do by setting height = 0 (i.e. zero noise added to “height”, or the y-axis) and width = 0.1 (i.e. points can move 0.1 units left and right on the x-axis):

ggplot(boto) +
  geom_violin(aes(y = length, x = sex)) +
  geom_jitter(aes(y = length, x = sex), height = 0, width = 0.1)

There are two things to note with the above code. First, geom_jitter() comes after geom_violin(). If we put the violin after the jitter, we wouldn’t see any of the jittered points; they’d be under the geom_violin() layer. Secondly, in geom_jitter() we don’t add the height = and width = arguments to aes(), because aes() is used to reference information within the boto dataset. Whenever we want to make changes to a geom_ that are not related to the data, we do so outside of aes().

Debugging ggplot2 code

The following code snippets contain errors that will either produce an error message or silently fail to do what was intended. Try to spot each problem before running the code.

Puzzle 1. What is wrong with this code?

ggplot(boto) +
  geom_jitter(aes(y = length, x = sex, height = 0, width = 0.1))

The code runs but produces a warning: “Ignoring unknown aesthetics: height and width”. The arguments height and width are fixed values, not variables from the dataset, so they belong outside aes().

Puzzle 2. Find the error:

ggplot(boto) +
  geom_point(aes(x = length, x = drought))

Error: “formal argument ‘x’ matched by multiple actual arguments”. Two variables are both mapped to x. One should be y = drought.

Puzzle 3. Find the error:

ggplot(boto) +
  geom_point(aes(x = length y = drought))

Error: “unexpected symbol in: … geom_point(aes(x = length y”“. There is a missing comma between x = length and y = drought.

Question 4. How would you produce a histogram of boto lengths? (Hint: type geom_ in your script and wait for RStudio’s autocomplete, browse the ggplot2 reference, or search online.)

ggplot(boto) +
  geom_histogram(aes(x = length))

Question 5. From the boxplot, what features tell you how symmetrical the distribution is for each sex?

The box spans the interquartile range (Q1 to Q3) and the line inside shows the median. If the median sits centrally in the box and the whiskers are roughly equal in length, the distribution is roughly symmetric. Here, females have a slightly longer upper whisker and males a slightly longer lower whisker, suggesting modest asymmetry in both groups, though the differences are small.

Question 6. What is the difference between the interquartile range shown in the boxplot and the range?

Interquartile range (IQR): the difference between the 75th and 25th percentiles, i.e. the width of the box. It describes the spread of the central 50% of observations.

Range: the difference between the maximum and minimum values in the data.

Question 7. What is the difference between the interquartile range and variance?

Both measure spread, but differently. The IQR captures how wide the central half of the data is, ignoring the extremes. Variance is calculated from all observations and reflects how far each value deviates from the mean on average. A dataset with extreme outliers can have a large variance but a modest IQR.

Adding more information

With these basic figures in place, we can begin expanding on them. There’s no shortage of options available to us. We can change the colours of most geom_s, we can have the colour of a geom_ change based on another variable, the shape, the size, the transparency, and so on. More than we have time to cover in this workshop.

With that in mind, let’s use a few of these options, starting with colour. We’ll keep x = drought and y = length, and colour the points according to sex. Because sex is a variable within our boto dataset, we include this within aes(). Remember, anything we want to change in the figure based on a variable in our dataset is included via aes().

ggplot(boto) +
  geom_point(aes(y = length, x = drought, colour = sex))

With a really minor addition to the code, we now have points coloured red and blue for female and male. In doing so, we can see some evidence for sexual dimorphism, but also that drought has a negative relationship with length for both sexes.

Let’s assume we think the points are a bit too small for our audience to read. In that case, we can specify a value for size =. Should size = go inside aes() or outside? We’re not using a variable in our dataset to determine the size of the point, so we include size = outside of aes():

ggplot(boto) +
  geom_point(aes(y = length, x = drought, colour = sex), size = 2)

Easy enough, but the figure still looks kind of ugly. Let’s make it more visually appealing. A really easy way to do this is to select one of the many themes built into ggplot2. Personally, I tend to use theme_bw() or theme_minimal(), but have a look online for other options if you don’t like these. How do we include a theme? As an additional layer. Because themes change the overall look of a figure, we can place them anywhere in our layers (though traditionally they go at the end). At the same time, let’s change the labels to make it look less like “code” and more polished, using the labs() (short for labels) function:

ggplot(boto) +
  geom_point(aes(y = length, x = drought, colour = sex), size = 2) +
  labs(y = "Boto length (cm)",
       x = "Number of droughts in past 20 years",
       colour = "Boto sex") +
  theme_bw()

The last graphing option we’ll cover in this workshop is facetting. Facetting allows us to split a figure into multiple smaller figures (technically called “multiples”) based on a variable in our dataset. There are two options: facet_wrap() and facet_grid(). The difference is that facet_grid() gives you fairly precise control, especially when you want to facet by two variables, whereas facet_wrap() automates some of the decisions for you. Generally, facet_wrap() works well and is a good option to start with. In either, we specify how we want to split the figure using ~ (called “tilde”). For instance, facet_wrap(~sex) is like saying “split the figure into multiple facets based on sex”:

ggplot(boto) +
  geom_point(aes(y = length, x = drought, colour = sex), size = 2) +
  labs(y = "Boto length (cm)",
       x = "Number of droughts in past 20 years",
       colour = "Boto sex") +
  facet_wrap(~sex) +
  theme_bw()

Interpreting the faceted figure

Question 1. Imagine drawing a line through the points in each panel. Would those lines have the same slope?

Probably not. Male boto length appears to have a stronger negative relationship with drought than females, suggesting the slopes differ between sexes.

Question 2. Does your answer to Question 1 suggest any assumption might be violated?

Yes, this raises a potential violation of additivity. If drought affects males more strongly than females, the two predictors aren’t acting independently; their combined effect matters. If the slopes were identical for both sexes, the additivity assumption would hold.

Question 3. Having explored the data with plots and summary statistics, do you notice anything suspicious?

Nothing obviously wrong, but the precision is worth questioning. Lengths are recorded to the nearest centimetre. How do you measure a dolphin in the Amazon to 1 cm accuracy? Something worth verifying with whoever collected the data.

Choosing geoms

Have a look at the ggplot2 reference, specifically at the layers section. Using the information there, see if you can figure out which geom_ to use to recreate the figure described below. Two things to note. First, this figure is created with a single geom_ with only x = and y = specified, and no other options. Secondly, if you figure out which geom_ to use, you’ll be prompted to install a new package when you run the code (it’s fine to install it, just type 1 and Enter in the console when prompted), and if you’re asked “Do you want to install from sources the package which needs compilation?”, just say no.

The figure divides the plotting area into hexagonal cells and colours each one by the number of data points that fall within it.

Challenge. Which geom_ produces a hexagonal bin plot? What does it show, and when might you prefer it over a regular scatterplot?

geom_hex(). It divides the plot area into hexagonal bins and colours each bin by how many points fall inside it. It is useful when there are so many overlapping points that a regular scatterplot becomes a smear, sometimes called “overplotting”.

ggplot(boto) +
  geom_hex(aes(x = drought, y = length)) +
  theme_bw()

Interpreting common plots

Scatterplots, histograms, density plots, boxplots and violin plots are all commonly used for EDA. Each type of plot has its own particular strengths for exploring certain aspects of a dataset, so it is important to understand how to interpret them. We highlight some features here, but be sure you’re confident about how to interpret all of the plots you produce in the workshops. We’ll assume you’re comfortable interpreting scatterplots and boxplots at this point (but feel free to ask if you need a reminder); it’s worth taking a second to talk about histograms and density plots.

Histograms

Frequency histograms are useful for visualising the distribution of data. The histogram represents the number of observations (count or frequency) per interval (or “bin”) of the variable. Earlier in the workshop you produced a violin plot comparing boto dolphin body lengths for males and females. You may have noticed that the shapes of the violins (i.e. density plots) for the two sexes were different: the female one was like a vase, wider at the base than in the middle or top, whereas the male group was wider at the top than the base and had less curvy sides. The shape of the density plot in the violin reflects a smoothed version of a histogram, and in a frequency histogram we can see these patterns too.

Create a histogram using the boto length data (you did this in Question Set Two, Question 4). We can improve the appearance of the graph by adding outlines to the bars and labels for the axes:

ggplot(boto) +
  geom_histogram(aes(x = length), colour = "white") +
  labs(x = "Boto dolphin length",
       y = "Frequency") +
  theme_bw()
`stat_bin()` using `bins = 30`. Pick better value `binwidth`.

From the plot it is relatively easy to see how many dolphins are in each bin, but it is not so easy to see what lengths of dolphin each bin represents. In ggplot2, the default behaviour is to create 30 bins. That means ggplot divides the range of the x variable (in this case length) into 30 bins of equal width, so each bin covers roughly (max - min) / 30, or about 2.3 cm here. There is no ideal bin width, but usually you’d set one to give a reasonably smooth presentation of the distribution, and not so thin that each bin has a single observation. You can set the bin width manually by adding the binwidth argument to geom_histogram() (or specify how many bins you want with bins):

ggplot(boto) +
  geom_histogram(aes(x = length), colour = "white", binwidth = 5) +
  labs(x = "Boto dolphin length",
       y = "Frequency") +
  theme_bw()

So, for a histogram where the range of values goes from 170 to 238 cm, with a binwidth of 5, each bin represents an interval of 5 cm. To interpret the middle bin (the one that sits on top of 200 cm), you can read off the plot that there are 10 observations (i.e. 10 dolphins) with lengths between 197.5 and 202.5 cm. For the bin at the far right, there is 1 observation between 237.5 and 242.5 cm (i.e. only one dolphin around 240 cm long).

Density plots

Density plots are a similar idea to histograms, in that they’re very useful for understanding the spread, or distribution, of your data. In truth they’re exceptionally similar to histograms, but differ in two important ways: (1) the y-axis doesn’t show counts but proportion (i.e. what proportion of observations are at this value), and (2) they don’t use bins and bars, but instead use a smoothed line. Otherwise you use them in the same way: “what’s the spread of my data?”

Here’s what the equivalent density plot looks like for the boto dolphins:

ggplot(boto) +
  geom_density(aes(x = length)) +
  labs(x = "Boto dolphin length",
       y = "Proportion") +
  theme_bw()

If you compare it to the histogram, you’d probably reach the same interpretation with either figure: most dolphins are below about 210 cm long, a decent number are between 210 and 230 cm, and very few are more than 230 cm.

It’s possible to combine a histogram and a density plot, though it’s a bit fiddly because you need to convert proportion into counts (by multiplying proportion by the total number of dolphins and by the binwidth). I wouldn’t normally do this in my own work, but it might help you see how the two relate to each other. Here’s what the code could look like:

ggplot(boto) +
  geom_histogram(aes(x = length), colour = "white", binwidth = 5) +
  geom_density(aes(x = length, y = after_stat(density * n * 5))) +
  labs(x = "Boto dolphin length",
       y = "Frequency") +
  theme_bw()

Before we move on, remember those violin plots from earlier? They’re basically just density plots rotated onto their side and split by whichever factor you add to the x-axis. They don’t show the proportion, but you read them the same way: “where is my data most concentrated?”

The Good and The Bad

The second dataset for this workshop is abdn_flats.txt. Download it below and save it to your BI3010/data/ folder. It contains more than 500 records of rental properties advertised across Aberdeenshire. The variables are:

  1. Price: monthly rent (£)
  2. Bedrooms: number of bedrooms
  3. PublicRooms: number of living rooms
  4. Rooms: total bedrooms plus living rooms
  5. Bathrooms: number of bathrooms
  6. FloorArea: total floor area (m²)
  7. EPC: Energy Performance Certificate (A is most efficient, G is worst)
  8. Tax: council tax band (A is lowest, H is highest; based on property value in April 1991)
  9. DayAdded: when the record was added to the dataset (not when the property was listed)
  10. Latitude: latitude coordinate (North to South; think of a “ladder”)
  11. Longitude: longitude coordinate (East to West; “the long way around”)
  12. UniCommuteTime: estimated walk time to the Zoology Building (minutes)
  13. HouseType: detached, semi-detached, terrace, or flat
  14. Furnished: fully furnished, part furnished, or unfurnished

Download abdn_flats.txt

Do a quick EDA to get familiar with the data and check for anything unusual. If you spot what looks like an error, make a note of it. We’ll return to this dataset later in the course.

The Good

Using whichever variables you like, make the most informative and visually appealing figure you can. Imagine you have been hired by a real estate company to create a figure that helps potential renters understand what may be determining rent prices, using R, ggplot2 and the abdn_flats data. Create your most informative and visually appealing figure now, noting that your “audience” here are not scientists, they’re members of the public, so we want something that effectively communicates the information but is also eye catching. Keep the lecture on effective communication in your mind and avoid some of the pitfalls we discussed.

TipCombining multiple plots with patchwork

Sometimes the clearest way to show several aspects of a dataset is to arrange multiple plots side by side or stacked. The patchwork package makes this straightforward:

install.packages("patchwork")  # run once
library(patchwork)

p1 <- ggplot(abdn) + geom_histogram(aes(x = Price))
p2 <- ggplot(abdn) + geom_point(aes(x = UniCommuteTime, y = Price))

p1 + p2        # side by side
p1 / p2        # stacked
(p1 + p2) / p3 # two on top, one below

Each object (p1, p2, etc.) is just a ggplot saved to a variable. You combine them using + (side by side) and / (stacked). That’s all there is to it.

If you need help figuring out how to implement one of your ideas, see the Intro2R chapter on graphics for a slightly more detailed walk through on ggplot2, or else use your LLM of choice to help. If using an LLM, just remember to mention that you are using R, ggplot2, and do not want to use any additional packages. There’s nothing wrong with using additional packages, but they can add lots of extra moving parts to the mix (and cause conflict issues). We’d suggest you try to limit the number of additional packages you have loaded while you’re still getting the hang of things, but if you feel comfortable using R, then go ahead by all means.

Your figure. Produce your best figure. Describe what you think makes it effective.

There’s no single right answer. Here’s one approach combining several views of the data:

abdn <- read.table("data/abdn_flats.txt", header = TRUE)

library(patchwork)  # lets you stitch ggplots together

p1 <- ggplot(abdn) +
  geom_histogram(aes(x = Price), fill = "#EC1365", colour = "white", binwidth = 25) +
  theme_classic() +
  labs(x = "Rental price (£)", y = "Frequency")

p2 <- ggplot(abdn) +
  geom_density(aes(x = Price, fill = Tax), alpha = 0.7) +
  scale_fill_brewer(palette = "Set2") +
  theme_classic() +
  theme(legend.position = "top") +
  labs(x = "Rental price (£)", fill = "Tax band", y = "Density")

p3 <- ggplot(abdn) +
  geom_point(aes(x = UniCommuteTime, y = Price), colour = "#EC1365") +
  theme_classic() +
  labs(x = "Walk time to Uni (mins)", y = "Rental price (£)")

p4 <- ggplot(abdn) +
  geom_violin(aes(x = HouseType, y = Price), fill = "#EC1365",
              draw_quantiles = c(0.25, 0.75)) +
  theme_classic() +
  labs(x = "Type of property", y = "Rental price (£)")

(p1 + p2) / (p3 + p4)

The Bad

Now for the part I’m excited about. Do the same as above, but this time make it a miserably ugly figure. It still needs to contain the information (i.e. no blank figures) and be somewhat readable. If you need inspiration, search for “graph crimes” or “chart crimes” in your search engine, there are lots of examples you can use for “inspiration”. Once you’re happy with your masterpieces, add your figures onto the MyAberdeen discussion board, and we’ll have an “Art Exhibit” to finish the workshop :D

Your disaster figure. Describe what you did to make it so terrible.

One approach: abuse the theme() function to override every colour and size setting.

ggplot(abdn) +
  geom_point(aes(x = PublicRooms, y = FloorArea, colour = Latitude),
             shape = 24, size = 10) +
  theme(
    panel.background = element_rect(fill = "yellow"),
    plot.background  = element_rect(fill = "pink"),
    panel.grid.major = element_line(colour = "red", linewidth = 2),
    panel.grid.minor = element_line(colour = "blue", linewidth = 1),
    axis.title.x     = element_text(size = 20, angle = 90, colour = "purple"),
    axis.title.y     = element_text(size = 4,  angle = 0,  colour = "orange"),
    axis.text        = element_text(size = 15, colour = "green")
  ) +
  scale_colour_gradient(low = "pink", high = "green") +
  labs(x = "Y-axis", y = "X-axis", colour = "[?]") +
  facet_wrap(~EPC, scales = "free_y")

Figures from the Real World

Below are a series of figures seen “out in the wild”. Have a look through them and consider what works, what doesn’t work, and why they made the decisions they did when producing the figure. Also give some thought to how you might create these figures in ggplot2.

The y-axis!

This figure comes from the Financial Times and shows bus usage inside and outside London. Pay close attention to the y-axis. (The Financial Times produces their figures in ggplot2 and even has their own theme package, ftplottools.)

Discussion. Is there anything inappropriate about this figure?

This figure made a lot of people angry online, mostly because the y-axis runs from -50% to +100% rather than from 0, and because the values have been log-transformed, which some argued exaggerated the differences.

In truth, it’s a pretty clever axis. A 100% increase means doubling (5 buses to 10). A -50% decrease means halving (10 buses to 5). On a log scale, doubling and halving take up the same vertical distance, which makes the axis symmetric and honest. Without the transformation, a -50% decline would look much smaller than an equivalent +100% increase, which would actually be misleading.

People get super annoyed by y-axes, especially when the bottom is not set to zero. Honestly? I don’t care. There are times when it’s entirely appropriate to do that, and other times when it’s not. But I dislike it when people treat it as though there’s some rule that there’s only a single correct way to set up a figure’s axes. There’s not. Use your brain. Think about what the best way is to visualise your data, and do that.

Are you even happy?

This figure comes from the World Happiness Report. Most figures in the full report are reasonably well done; this one is an exception.

Discussion. What is wrong with this figure?

The bars show rankings, not magnitudes. Finland is ranked 1st, New Zealand is 10th. The figure is just a glorified sorted list; the bars add no information beyond the order.

More subtly, I’m always sceptical of “happiness” ratings as data. How do you assign a number to how happy you feel? Would you and I give the same score if we felt the same way? What if I’m just having an off day? How many people do you ask? The data underlying this figure is most likely pretty imprecise. At the very least, it doesn’t pass a validity check. Apply that check before treating numbers like these as meaningful.

Birds and light

This figure comes from Pease et al. (2025), published in Science. The study used more than 60 million bird detections linked to satellite measures of light pollution. The figure shows three stages of data processing: unfiltered detections, filtered detections, and the response variable used in the model.

Discussion. What works and what doesn’t in this figure?

Several problems, published in one of the most prestigious journals in science:

  1. Stacked bars make comparisons hard. Can you tell whether “Unfiltered detections” in August 2023 differs from September 2023? I genuinely can’t. And the “Response variable” slice at the bottom is so thin it’s almost invisible. That’s the variable that was actually used in the model. What’s the point of including it if it can’t be read?
  2. The colour palette is probably not colour-blind friendly. It costs nothing to check, and there’s no reason to exclude part of the audience.
  3. The x-axis is overloaded. Printing every month label makes it cluttered and hard to read. Three labels (start, middle, end) would be much cleaner. This probably happened because the author stored date as a factor in R, which prints every level.

A better approach: points for observation counts connected by a line, split into three separate panels with independent y-axes, so the trend in the “Response variable” is actually visible.

Marmosets and ketamine (?!)

The next figure comes from Wood et al. (2025), also published in Science. Honestly, I don’t really get this paper and I don’t know how to summarise it. The authors did something to monkey brains. There was ketamine. There were vehicles. I don’t know man. Anyway, here’s the figure…

Discussion. What problems do you see?

A pretty miserable figure all things considered. Here are my issues with it:

  1. The colour palette. The hell is with the weird colour palette and pattern that looks like it was stolen from a 1990s bus seat? Who was like “Yeaaaaaah bruh, this looks siiiiiiick dude.” Why is one bar darker than the other?
  2. Abbreviations. Do you know what “Sal/Veh” means? Great figures don’t need a legend or a paper for you to make sense of them; you can see amazing figures with zero context and they still make total sense. I don’t know what this figure is telling me even with context.
  3. Inconsistent data point styling. All the randomly shaped points. What do they mean? Why are some filled and others empty? Are those all the observations they had, because if so, what’s that? Like 10 observations? Get outta here with that.
  4. Bars are the wrong choice. The bars are completely unnecessary here and probably misleading. What the authors are trying to show are specific values (called “point estimates”). For example, the category “Sal/DCZ” is estimated at roughly -50 on the y-axis. That’s the value that’s important. By using a bar you’re implying there was also a -49, and -48, and -47, all the way to 0. That’s not the case. It’s -50. Just -50. Using bars is, for some reason, super common in some fields. Medics looooooove them. They’re overused and rarely the best choice. A lot of the time (but not always), bars should be replaced by points.

How I’d make this figure: I’d use points for the estimated values instead of bars. I probably wouldn’t add points to show the data, but if I did, I’d make those points much smaller (and be consistent). I’d write the full category name and rotate them 90 degrees so it’d fit. (I’d also get rid of the P-value “bars” and asterisks at the top of the figure.)

Glossary / 词汇表
ggplot2 framework
ggplot2 ggplot2
An R package for data visualisation based on the Grammar of Graphics. Plots are built by combining layers (data, aesthetics, geometries, themes).
基于图形语法的 R 数据可视化包。通过组合数据、美学映射、几何对象和主题等图层来构建图表。
Grammar of Graphics 图形语法
A framework that describes plots in terms of data, aesthetics, and geometric objects. ggplot2 implements this framework in R.
一种用数据、美学属性和几何对象来描述图形的框架,ggplot2 在 R 中实现了这一框架。
Aesthetic mapping (aes()) 美学映射
The link between a variable in the dataset and a visual property of the plot, such as position (x, y), colour, shape, or size. Defined inside aes().
数据集中变量与图形视觉属性(如位置、颜色、形状、大小)之间的对应关系,在 aes() 中定义。
Geometry (geom_) 几何对象
The visual form used to represent data: points (geom_point), bars (geom_bar), lines (geom_line), etc. Each geom_ function adds a layer to the plot.
数据的视觉呈现形式,如点(geom_point)、柱(geom_bar)、线(geom_line)等。每个 geom_ 函数向图形添加一个图层。
Layer 图层
One component added to a ggplot, such as a geometry, a set of labels, or a theme. Layers are stacked in the order they are written; later layers appear on top.
添加到 ggplot 图形中的一个组件,如几何对象、标签或主题。图层按照书写顺序叠加,后写的图层显示在上方。
Theme 主题
Controls the non-data elements of a plot: background colour, grid lines, axis text, and so on. Common options include theme_bw() and theme_minimal().
控制图形中非数据元素的样式,如背景颜色、网格线、坐标轴文字等。常用选项包括 theme_bw() 和 theme_minimal()。
Facet 分面
A way to split one plot into multiple panels based on a categorical variable. facet_wrap(~variable) creates one panel per level of the variable.
根据分类变量将一张图拆分为多个子图的方法。facet_wrap(~变量) 为变量的每个水平创建一个子图。
Plot types
Histogram 频率直方图
A plot that counts how many observations fall within each equal-width bin of a continuous variable. Useful for understanding the shape and spread of a distribution.
统计连续变量在每个等宽区间(组距箱)内观测值数量的图表,用于了解数据分布的形状和范围。
Binwidth 组距
The width of each bar in a histogram. Wider bins produce a smoother picture; narrower bins show more detail. Set with the binwidth argument in geom_histogram().
直方图中每个柱的宽度。组距越宽,图形越平滑;组距越窄,细节越多。通过 geom_histogram() 的 binwidth 参数设置。
Density plot 密度图
A smoothed version of a histogram where the y-axis shows proportion rather than count. Read it the same way: taller regions contain more observations.
直方图的平滑版本,纵轴表示比例而非计数。解读方式相同:较高的区域包含更多观测值。
Boxplot 箱形图
A plot that summarises a distribution using the median (central line), interquartile range (box), and range (whiskers). Outliers may appear as individual points beyond the whiskers.
用中位数(中线)、四分位距(箱体)和范围(须线)来概括数据分布的图表。离群值可能以须线外的单独点呈现。
Interquartile range (IQR) 四分位距
The difference between the 75th and 25th percentiles. It captures the spread of the central 50% of observations and equals the box height in a boxplot.
第75百分位数与第25百分位数之差,反映中间50%数据的分散程度,等于箱形图中箱体的高度。
Violin plot 小提琴图
A plot that shows the distribution of a variable as a mirrored density curve. Wider sections contain more observations. Often combined with jittered points.
以对称密度曲线展示变量分布的图表,较宽处包含更多观测值,常与抖动点结合使用。
Scatterplot 散点图
A plot that shows the relationship between two continuous variables using points. Each point represents one observation.
用点展示两个连续变量之间关系的图表,每个点代表一个观测值。
Plotting techniques
Jitter 抖动
A small random offset added to data points to prevent them overlapping. In ggplot2, use geom_jitter() and control the amount with the height and width arguments.
为防止数据点重叠而添加的微小随机偏移量。在 ggplot2 中使用 geom_jitter(),通过 height 和 width 参数控制抖动幅度。
Overplotting 过度叠加
When many data points overlap in a scatterplot, making the density of points hard to read. Solutions include jittering, transparency (alpha), or hexagonal bins (geom_hex).
散点图中大量数据点重叠导致难以判断密度的问题。解决方法包括抖动、透明度(alpha)或六边形分箱(geom_hex)。
Dataset terms
Sexual dimorphism 性别二型性
Systematic differences in body size, shape, or colour between males and females of the same species. The boto data explores whether males and females differ in asymptotic length.
同一物种雌雄个体在体型、形态或颜色上的系统性差异。本次数据探讨雌雄亚马逊河豚的渐近体长是否存在差异。
Colour-blind safe palette 色盲友好调色板
A set of colours that remain distinguishable for people with common forms of colour blindness. In ggplot2, scale_fill_brewer() and scale_colour_brewer() offer several such palettes.
对常见色盲类型人群仍可区分的颜色组合。在 ggplot2 中,scale_fill_brewer() 和 scale_colour_brewer() 提供多种此类调色板。
Energy Performance Certificate (EPC) 能源效能证书
A rating from A (most efficient) to G (least efficient) that indicates how energy efficient a building is. Used in the Aberdeen rental dataset.
评估建筑能源效率的评级,从 A(最高效)到 G(最低效),在阿伯丁租房数据集中使用。
Asymptotic length 渐近体长
The maximum body length an animal approaches as it ages, after growth has effectively stopped. Used in growth models and in this workshop to describe adult boto dolphins.
动物随年龄增长趋于停止生长时所接近的最大体长,用于生长模型,在本次工作坊中用于描述成年亚马逊河豚的体长。