boto <- read.table("data/boto.txt", header = TRUE)Workshop 2: Data visualisation
BI3010 Statistics for Biologists
Learning Objectives
- Learn how to use the
Rpackageggplot2to do data visualisation. - Consider what constitutes an effective visualisation versus an ineffective one.
- Learn some techniques to make your figures more visually appealing to an audience.
These skills will be used throughout the course.
Workshop structure
These workshops follow an analytical workflow. Today we focus on steps 2 and 3, expanding the toolkit you built in Workshop 1 using data visualisation.
- Formulate a research question.
- Perform exploratory data analysis [Focus of today].
- Identify any hidden assumptions [Focus of today].
- Fit an appropriate model.
- Diagnose the model and check assumptions.
- Summarise the results.
- Interpret and provide inferences.
Introduction to ggplot2
ggplot2 is an implementation of the “Grammar of Graphics”, a framework for data visualisation. It allows you to build plots layer by layer, starting with a base plot and then adding additional elements such as scales, themes, and geometries. Using this framework, you can create pretty sophisticated figures with surprisingly little effort. It’s such a popular tool that familiarity with ggplot2 is sometimes asked for by employers; we’ve even seen job adverts from companies such as Spotify listing this an essential criterion.
It is important to note, however, that ggplot2 is not the only way to create visualisations. It’s just an increasingly preferred one. Many of the figures you will make on this course can be near-perfectly recreated using the visualisation framework offered by base R. We simply prefer ggplot2 as we feel it is a more intuitive way to create figures; indeed, nearly every single figure you have and will see on this course were produced in ggplot2.
Load the data
For the first part of this workshop, we’ll be working with data on the Amazon River dolphin, or the boto (Inia geoffrensis, I wonder if the person who named the species had a friend called Geoff…). For this first interaction with the data, we’ll broadly say we’re interested in seeing if there is sexual dimorphism in asymptotic length (i.e. once an animal has stopped growing) between females and males, as well as how the species respond to droughts. As always, before we do any analysis, we need to ensure there is nothing weird in the data. In Workshop 1 you did this using simple summary statistics (e.g. mean()). While that is useful, complementing it with data visualisations is really powerful.
R sessions sometimes reopen with objects from earlier work, including hidden ones (often because you clicked “save workspace” when you quit RStudio). Leftover data or functions can cause headaches and confusion, so if you think this might be affecting you, you can begin by clearing your global environment by running:
rm(list = ls())This removes all objects in the global environment. It does not delete any files, so don’t worry about losing any of your work. Alternatively, in RStudio you can do the same from the Environment pane by clicking the broom icon (Clear All).
Download the data file below and save it to your BI3010/data/ folder. Open your BI3010.Rproj file to start RStudio with the working directory already set, then open a new script and load the data:
Given how easy it is to do, let’s run a quick summary() and str() to see if anything stands out:
summary(boto)
str(boto)If nothing immediately seems off, we can continue.
Building a first plot
We’ll build this figure up slowly, one layer at a time, to make the underlying philosophy of ggplot2 clear. But before we can do that, we first need to install the package (if you haven’t already) and then load it ready to use:
install.packages("ggplot2") # Install the software (run once)
library(ggplot2) # Load the softwareOnce we have ggplot2 installed and loaded, we can make a simple “ggplot”:
ggplot()Cool… a completely empty canvas. Well, it shouldn’t really come as a surprise that there’s nothing in this figure. For starters, we haven’t even told it which dataset to use. Let’s correct that now:
ggplot(boto)Still nothing. What we’re missing here is where we tell ggplot2 a couple of things. First, we need to say in what form we want the data shown. Do we want points? Or lines? Or errorbars? Secondly, we need to say which specific data, or “aesthetics” (shortened to aes() in ggplot2), we want shown and where. Do we want length included? On the x or y axis?
Let’s have length on the y-axis and drought on the x-axis, with the form (or geometry) that we want the data shown as being points:
ggplot(boto) +
geom_point(aes(y = length, x = drought))Now we’re getting somewhere. From an early data analysis perspective, what have we already learnt, just from the quick summary stats and from this figure?
First impressions
Question 1. What variables are available in this dataset?
Three variables: sex, drought, and length.
Question 2. How many observations does the dataset contain?
71 boto dolphins.
Question 3. Based on the scatterplot, is the relationship between boto length and number of drought years positive, neutral, or negative?
Very likely negative: longer drought histories appear to correspond to shorter dolphins.
Question 4. What range of lengths do boto dolphins reach at maturity?
From summary(): roughly 170 to 238 cm. The scatterplot gives the same impression.
Additional geom_ layers
The above figure is great. It’s simple and clear (though a bit ugly, we’ll come back to that later), but how would we include a categorical variable like sex? Can we just copy our code from above and replace drought with sex? Let’s see:
ggplot(boto) +
geom_point(aes(y = length, x = sex))Kind of? It’s a bit hard to see the spread of data for each sex. A better way to show this is to use some tool that helps the reader understand the range of length for each sex, but also how the data is spread. There are two common options people use for this: boxplots and violin plots. It’s trivially easy to swap between them, because all we need to do is change the geometry function from geom_point() to either geom_boxplot() or geom_violin():
# Boxplot
ggplot(boto) +
geom_boxplot(aes(y = length, x = sex))# Violin plot
ggplot(boto) +
geom_violin(aes(y = length, x = sex))Both of these present the same information in different ways. The boxplot uses boxes and lines to indicate the range (the first and third quantiles) and the median. The violin plot also shows the spread of data, but by creating density plots which are mirrored. Boxplots have a long history and use in science, but people are increasingly using violin plots, simply because they’re more transparent and clearer than boxplots.
With either boxplots or violin plots, it’s fairly common that people add data points to the figure as well, to really hammer home the spread of data. But we’ve seen before that using geom_point() basically gives us a line of points that doesn’t help us understand the spread at all.
An alternative is to use something called “jittering”. When we jitter points, we add a small random value to either or both the x and y position to spread the data out a little. With our boto data we wouldn’t want to add anything to y (we’re interested in the length of the dolphins, so we don’t want to mess that up), but we can happily add something to x (or sex). To do so, we use geom_jitter() and specify that “noise” should only be added to x and not y, which we do by setting height = 0 (i.e. zero noise added to “height”, or the y-axis) and width = 0.1 (i.e. points can move 0.1 units left and right on the x-axis):
ggplot(boto) +
geom_violin(aes(y = length, x = sex)) +
geom_jitter(aes(y = length, x = sex), height = 0, width = 0.1)There are two things to note with the above code. First, geom_jitter() comes after geom_violin(). If we put the violin after the jitter, we wouldn’t see any of the jittered points; they’d be under the geom_violin() layer. Secondly, in geom_jitter() we don’t add the height = and width = arguments to aes(), because aes() is used to reference information within the boto dataset. Whenever we want to make changes to a geom_ that are not related to the data, we do so outside of aes().
Debugging ggplot2 code
The following code snippets contain errors that will either produce an error message or silently fail to do what was intended. Try to spot each problem before running the code.
Puzzle 1. What is wrong with this code?
ggplot(boto) +
geom_jitter(aes(y = length, x = sex, height = 0, width = 0.1))
The code runs but produces a warning: “Ignoring unknown aesthetics: height and width”. The arguments height and width are fixed values, not variables from the dataset, so they belong outside aes().
Puzzle 2. Find the error:
ggplot(boto) +
geom_point(aes(x = length, x = drought))
Error: “formal argument ‘x’ matched by multiple actual arguments”. Two variables are both mapped to x. One should be y = drought.
Puzzle 3. Find the error:
ggplot(boto) +
geom_point(aes(x = length y = drought))
Error: “unexpected symbol in: … geom_point(aes(x = length y”“. There is a missing comma between x = length and y = drought.
Question 4. How would you produce a histogram of boto lengths? (Hint: type geom_ in your script and wait for RStudio’s autocomplete, browse the ggplot2 reference, or search online.)
ggplot(boto) +
geom_histogram(aes(x = length))
Question 5. From the boxplot, what features tell you how symmetrical the distribution is for each sex?
The box spans the interquartile range (Q1 to Q3) and the line inside shows the median. If the median sits centrally in the box and the whiskers are roughly equal in length, the distribution is roughly symmetric. Here, females have a slightly longer upper whisker and males a slightly longer lower whisker, suggesting modest asymmetry in both groups, though the differences are small.
Question 6. What is the difference between the interquartile range shown in the boxplot and the range?
Interquartile range (IQR): the difference between the 75th and 25th percentiles, i.e. the width of the box. It describes the spread of the central 50% of observations.
Range: the difference between the maximum and minimum values in the data.
Question 7. What is the difference between the interquartile range and variance?
Both measure spread, but differently. The IQR captures how wide the central half of the data is, ignoring the extremes. Variance is calculated from all observations and reflects how far each value deviates from the mean on average. A dataset with extreme outliers can have a large variance but a modest IQR.
Adding more information
With these basic figures in place, we can begin expanding on them. There’s no shortage of options available to us. We can change the colours of most geom_s, we can have the colour of a geom_ change based on another variable, the shape, the size, the transparency, and so on. More than we have time to cover in this workshop.
With that in mind, let’s use a few of these options, starting with colour. We’ll keep x = drought and y = length, and colour the points according to sex. Because sex is a variable within our boto dataset, we include this within aes(). Remember, anything we want to change in the figure based on a variable in our dataset is included via aes().
ggplot(boto) +
geom_point(aes(y = length, x = drought, colour = sex))With a really minor addition to the code, we now have points coloured red and blue for female and male. In doing so, we can see some evidence for sexual dimorphism, but also that drought has a negative relationship with length for both sexes.
Let’s assume we think the points are a bit too small for our audience to read. In that case, we can specify a value for size =. Should size = go inside aes() or outside? We’re not using a variable in our dataset to determine the size of the point, so we include size = outside of aes():
ggplot(boto) +
geom_point(aes(y = length, x = drought, colour = sex), size = 2)Easy enough, but the figure still looks kind of ugly. Let’s make it more visually appealing. A really easy way to do this is to select one of the many themes built into ggplot2. Personally, I tend to use theme_bw() or theme_minimal(), but have a look online for other options if you don’t like these. How do we include a theme? As an additional layer. Because themes change the overall look of a figure, we can place them anywhere in our layers (though traditionally they go at the end). At the same time, let’s change the labels to make it look less like “code” and more polished, using the labs() (short for labels) function:
ggplot(boto) +
geom_point(aes(y = length, x = drought, colour = sex), size = 2) +
labs(y = "Boto length (cm)",
x = "Number of droughts in past 20 years",
colour = "Boto sex") +
theme_bw()The last graphing option we’ll cover in this workshop is facetting. Facetting allows us to split a figure into multiple smaller figures (technically called “multiples”) based on a variable in our dataset. There are two options: facet_wrap() and facet_grid(). The difference is that facet_grid() gives you fairly precise control, especially when you want to facet by two variables, whereas facet_wrap() automates some of the decisions for you. Generally, facet_wrap() works well and is a good option to start with. In either, we specify how we want to split the figure using ~ (called “tilde”). For instance, facet_wrap(~sex) is like saying “split the figure into multiple facets based on sex”:
ggplot(boto) +
geom_point(aes(y = length, x = drought, colour = sex), size = 2) +
labs(y = "Boto length (cm)",
x = "Number of droughts in past 20 years",
colour = "Boto sex") +
facet_wrap(~sex) +
theme_bw()Interpreting the faceted figure
Question 1. Imagine drawing a line through the points in each panel. Would those lines have the same slope?
Probably not. Male boto length appears to have a stronger negative relationship with drought than females, suggesting the slopes differ between sexes.
Question 2. Does your answer to Question 1 suggest any assumption might be violated?
Yes, this raises a potential violation of additivity. If drought affects males more strongly than females, the two predictors aren’t acting independently; their combined effect matters. If the slopes were identical for both sexes, the additivity assumption would hold.
Question 3. Having explored the data with plots and summary statistics, do you notice anything suspicious?
Nothing obviously wrong, but the precision is worth questioning. Lengths are recorded to the nearest centimetre. How do you measure a dolphin in the Amazon to 1 cm accuracy? Something worth verifying with whoever collected the data.
Choosing geoms
Have a look at the ggplot2 reference, specifically at the layers section. Using the information there, see if you can figure out which geom_ to use to recreate the figure described below. Two things to note. First, this figure is created with a single geom_ with only x = and y = specified, and no other options. Secondly, if you figure out which geom_ to use, you’ll be prompted to install a new package when you run the code (it’s fine to install it, just type 1 and Enter in the console when prompted), and if you’re asked “Do you want to install from sources the package which needs compilation?”, just say no.
The figure divides the plotting area into hexagonal cells and colours each one by the number of data points that fall within it.
Challenge. Which geom_ produces a hexagonal bin plot? What does it show, and when might you prefer it over a regular scatterplot?
geom_hex(). It divides the plot area into hexagonal bins and colours each bin by how many points fall inside it. It is useful when there are so many overlapping points that a regular scatterplot becomes a smear, sometimes called “overplotting”.
ggplot(boto) +
geom_hex(aes(x = drought, y = length)) +
theme_bw()
Interpreting common plots
Scatterplots, histograms, density plots, boxplots and violin plots are all commonly used for EDA. Each type of plot has its own particular strengths for exploring certain aspects of a dataset, so it is important to understand how to interpret them. We highlight some features here, but be sure you’re confident about how to interpret all of the plots you produce in the workshops. We’ll assume you’re comfortable interpreting scatterplots and boxplots at this point (but feel free to ask if you need a reminder); it’s worth taking a second to talk about histograms and density plots.
Histograms
Frequency histograms are useful for visualising the distribution of data. The histogram represents the number of observations (count or frequency) per interval (or “bin”) of the variable. Earlier in the workshop you produced a violin plot comparing boto dolphin body lengths for males and females. You may have noticed that the shapes of the violins (i.e. density plots) for the two sexes were different: the female one was like a vase, wider at the base than in the middle or top, whereas the male group was wider at the top than the base and had less curvy sides. The shape of the density plot in the violin reflects a smoothed version of a histogram, and in a frequency histogram we can see these patterns too.
Create a histogram using the boto length data (you did this in Question Set Two, Question 4). We can improve the appearance of the graph by adding outlines to the bars and labels for the axes:
ggplot(boto) +
geom_histogram(aes(x = length), colour = "white") +
labs(x = "Boto dolphin length",
y = "Frequency") +
theme_bw()`stat_bin()` using `bins = 30`. Pick better value `binwidth`.
From the plot it is relatively easy to see how many dolphins are in each bin, but it is not so easy to see what lengths of dolphin each bin represents. In ggplot2, the default behaviour is to create 30 bins. That means ggplot divides the range of the x variable (in this case length) into 30 bins of equal width, so each bin covers roughly (max - min) / 30, or about 2.3 cm here. There is no ideal bin width, but usually you’d set one to give a reasonably smooth presentation of the distribution, and not so thin that each bin has a single observation. You can set the bin width manually by adding the binwidth argument to geom_histogram() (or specify how many bins you want with bins):
ggplot(boto) +
geom_histogram(aes(x = length), colour = "white", binwidth = 5) +
labs(x = "Boto dolphin length",
y = "Frequency") +
theme_bw()So, for a histogram where the range of values goes from 170 to 238 cm, with a binwidth of 5, each bin represents an interval of 5 cm. To interpret the middle bin (the one that sits on top of 200 cm), you can read off the plot that there are 10 observations (i.e. 10 dolphins) with lengths between 197.5 and 202.5 cm. For the bin at the far right, there is 1 observation between 237.5 and 242.5 cm (i.e. only one dolphin around 240 cm long).
Density plots
Density plots are a similar idea to histograms, in that they’re very useful for understanding the spread, or distribution, of your data. In truth they’re exceptionally similar to histograms, but differ in two important ways: (1) the y-axis doesn’t show counts but proportion (i.e. what proportion of observations are at this value), and (2) they don’t use bins and bars, but instead use a smoothed line. Otherwise you use them in the same way: “what’s the spread of my data?”
Here’s what the equivalent density plot looks like for the boto dolphins:
ggplot(boto) +
geom_density(aes(x = length)) +
labs(x = "Boto dolphin length",
y = "Proportion") +
theme_bw()If you compare it to the histogram, you’d probably reach the same interpretation with either figure: most dolphins are below about 210 cm long, a decent number are between 210 and 230 cm, and very few are more than 230 cm.
It’s possible to combine a histogram and a density plot, though it’s a bit fiddly because you need to convert proportion into counts (by multiplying proportion by the total number of dolphins and by the binwidth). I wouldn’t normally do this in my own work, but it might help you see how the two relate to each other. Here’s what the code could look like:
ggplot(boto) +
geom_histogram(aes(x = length), colour = "white", binwidth = 5) +
geom_density(aes(x = length, y = after_stat(density * n * 5))) +
labs(x = "Boto dolphin length",
y = "Frequency") +
theme_bw()Before we move on, remember those violin plots from earlier? They’re basically just density plots rotated onto their side and split by whichever factor you add to the x-axis. They don’t show the proportion, but you read them the same way: “where is my data most concentrated?”
The Good and The Bad
The second dataset for this workshop is abdn_flats.txt. Download it below and save it to your BI3010/data/ folder. It contains more than 500 records of rental properties advertised across Aberdeenshire. The variables are:
Price: monthly rent (£)Bedrooms: number of bedroomsPublicRooms: number of living roomsRooms: total bedrooms plus living roomsBathrooms: number of bathroomsFloorArea: total floor area (m²)EPC: Energy Performance Certificate (A is most efficient, G is worst)Tax: council tax band (A is lowest, H is highest; based on property value in April 1991)DayAdded: when the record was added to the dataset (not when the property was listed)Latitude: latitude coordinate (North to South; think of a “ladder”)Longitude: longitude coordinate (East to West; “the long way around”)UniCommuteTime: estimated walk time to the Zoology Building (minutes)HouseType: detached, semi-detached, terrace, or flatFurnished: fully furnished, part furnished, or unfurnished
Do a quick EDA to get familiar with the data and check for anything unusual. If you spot what looks like an error, make a note of it. We’ll return to this dataset later in the course.
The Good
Using whichever variables you like, make the most informative and visually appealing figure you can. Imagine you have been hired by a real estate company to create a figure that helps potential renters understand what may be determining rent prices, using R, ggplot2 and the abdn_flats data. Create your most informative and visually appealing figure now, noting that your “audience” here are not scientists, they’re members of the public, so we want something that effectively communicates the information but is also eye catching. Keep the lecture on effective communication in your mind and avoid some of the pitfalls we discussed.
patchwork
Sometimes the clearest way to show several aspects of a dataset is to arrange multiple plots side by side or stacked. The patchwork package makes this straightforward:
install.packages("patchwork") # run once
library(patchwork)
p1 <- ggplot(abdn) + geom_histogram(aes(x = Price))
p2 <- ggplot(abdn) + geom_point(aes(x = UniCommuteTime, y = Price))
p1 + p2 # side by side
p1 / p2 # stacked
(p1 + p2) / p3 # two on top, one belowEach object (p1, p2, etc.) is just a ggplot saved to a variable. You combine them using + (side by side) and / (stacked). That’s all there is to it.
If you need help figuring out how to implement one of your ideas, see the Intro2R chapter on graphics for a slightly more detailed walk through on ggplot2, or else use your LLM of choice to help. If using an LLM, just remember to mention that you are using R, ggplot2, and do not want to use any additional packages. There’s nothing wrong with using additional packages, but they can add lots of extra moving parts to the mix (and cause conflict issues). We’d suggest you try to limit the number of additional packages you have loaded while you’re still getting the hang of things, but if you feel comfortable using R, then go ahead by all means.
Your figure. Produce your best figure. Describe what you think makes it effective.
There’s no single right answer. Here’s one approach combining several views of the data:
abdn <- read.table("data/abdn_flats.txt", header = TRUE)
library(patchwork) # lets you stitch ggplots together
p1 <- ggplot(abdn) +
geom_histogram(aes(x = Price), fill = "#EC1365", colour = "white", binwidth = 25) +
theme_classic() +
labs(x = "Rental price (£)", y = "Frequency")
p2 <- ggplot(abdn) +
geom_density(aes(x = Price, fill = Tax), alpha = 0.7) +
scale_fill_brewer(palette = "Set2") +
theme_classic() +
theme(legend.position = "top") +
labs(x = "Rental price (£)", fill = "Tax band", y = "Density")
p3 <- ggplot(abdn) +
geom_point(aes(x = UniCommuteTime, y = Price), colour = "#EC1365") +
theme_classic() +
labs(x = "Walk time to Uni (mins)", y = "Rental price (£)")
p4 <- ggplot(abdn) +
geom_violin(aes(x = HouseType, y = Price), fill = "#EC1365",
draw_quantiles = c(0.25, 0.75)) +
theme_classic() +
labs(x = "Type of property", y = "Rental price (£)")
(p1 + p2) / (p3 + p4)
The Bad
Now for the part I’m excited about. Do the same as above, but this time make it a miserably ugly figure. It still needs to contain the information (i.e. no blank figures) and be somewhat readable. If you need inspiration, search for “graph crimes” or “chart crimes” in your search engine, there are lots of examples you can use for “inspiration”. Once you’re happy with your masterpieces, add your figures onto the MyAberdeen discussion board, and we’ll have an “Art Exhibit” to finish the workshop :D
Your disaster figure. Describe what you did to make it so terrible.
One approach: abuse the theme() function to override every colour and size setting.
ggplot(abdn) +
geom_point(aes(x = PublicRooms, y = FloorArea, colour = Latitude),
shape = 24, size = 10) +
theme(
panel.background = element_rect(fill = "yellow"),
plot.background = element_rect(fill = "pink"),
panel.grid.major = element_line(colour = "red", linewidth = 2),
panel.grid.minor = element_line(colour = "blue", linewidth = 1),
axis.title.x = element_text(size = 20, angle = 90, colour = "purple"),
axis.title.y = element_text(size = 4, angle = 0, colour = "orange"),
axis.text = element_text(size = 15, colour = "green")
) +
scale_colour_gradient(low = "pink", high = "green") +
labs(x = "Y-axis", y = "X-axis", colour = "[?]") +
facet_wrap(~EPC, scales = "free_y")
Figures from the Real World
Below are a series of figures seen “out in the wild”. Have a look through them and consider what works, what doesn’t work, and why they made the decisions they did when producing the figure. Also give some thought to how you might create these figures in ggplot2.
The y-axis!
This figure comes from the Financial Times and shows bus usage inside and outside London. Pay close attention to the y-axis. (The Financial Times produces their figures in ggplot2 and even has their own theme package, ftplottools.)
Discussion. Is there anything inappropriate about this figure?
This figure made a lot of people angry online, mostly because the y-axis runs from -50% to +100% rather than from 0, and because the values have been log-transformed, which some argued exaggerated the differences.
In truth, it’s a pretty clever axis. A 100% increase means doubling (5 buses to 10). A -50% decrease means halving (10 buses to 5). On a log scale, doubling and halving take up the same vertical distance, which makes the axis symmetric and honest. Without the transformation, a -50% decline would look much smaller than an equivalent +100% increase, which would actually be misleading.
People get super annoyed by y-axes, especially when the bottom is not set to zero. Honestly? I don’t care. There are times when it’s entirely appropriate to do that, and other times when it’s not. But I dislike it when people treat it as though there’s some rule that there’s only a single correct way to set up a figure’s axes. There’s not. Use your brain. Think about what the best way is to visualise your data, and do that.
Are you even happy?
This figure comes from the World Happiness Report. Most figures in the full report are reasonably well done; this one is an exception.
Discussion. What is wrong with this figure?
The bars show rankings, not magnitudes. Finland is ranked 1st, New Zealand is 10th. The figure is just a glorified sorted list; the bars add no information beyond the order.
More subtly, I’m always sceptical of “happiness” ratings as data. How do you assign a number to how happy you feel? Would you and I give the same score if we felt the same way? What if I’m just having an off day? How many people do you ask? The data underlying this figure is most likely pretty imprecise. At the very least, it doesn’t pass a validity check. Apply that check before treating numbers like these as meaningful.
Birds and light
This figure comes from Pease et al. (2025), published in Science. The study used more than 60 million bird detections linked to satellite measures of light pollution. The figure shows three stages of data processing: unfiltered detections, filtered detections, and the response variable used in the model.
Discussion. What works and what doesn’t in this figure?
Several problems, published in one of the most prestigious journals in science:
- Stacked bars make comparisons hard. Can you tell whether “Unfiltered detections” in August 2023 differs from September 2023? I genuinely can’t. And the “Response variable” slice at the bottom is so thin it’s almost invisible. That’s the variable that was actually used in the model. What’s the point of including it if it can’t be read?
- The colour palette is probably not colour-blind friendly. It costs nothing to check, and there’s no reason to exclude part of the audience.
- The x-axis is overloaded. Printing every month label makes it cluttered and hard to read. Three labels (start, middle, end) would be much cleaner. This probably happened because the author stored date as a factor in R, which prints every level.
A better approach: points for observation counts connected by a line, split into three separate panels with independent y-axes, so the trend in the “Response variable” is actually visible.
Marmosets and ketamine (?!)
The next figure comes from Wood et al. (2025), also published in Science. Honestly, I don’t really get this paper and I don’t know how to summarise it. The authors did something to monkey brains. There was ketamine. There were vehicles. I don’t know man. Anyway, here’s the figure…
Discussion. What problems do you see?
A pretty miserable figure all things considered. Here are my issues with it:
- The colour palette. The hell is with the weird colour palette and pattern that looks like it was stolen from a 1990s bus seat? Who was like “Yeaaaaaah bruh, this looks siiiiiiick dude.” Why is one bar darker than the other?
- Abbreviations. Do you know what “Sal/Veh” means? Great figures don’t need a legend or a paper for you to make sense of them; you can see amazing figures with zero context and they still make total sense. I don’t know what this figure is telling me even with context.
- Inconsistent data point styling. All the randomly shaped points. What do they mean? Why are some filled and others empty? Are those all the observations they had, because if so, what’s that? Like 10 observations? Get outta here with that.
- Bars are the wrong choice. The bars are completely unnecessary here and probably misleading. What the authors are trying to show are specific values (called “point estimates”). For example, the category “Sal/DCZ” is estimated at roughly -50 on the y-axis. That’s the value that’s important. By using a bar you’re implying there was also a -49, and -48, and -47, all the way to 0. That’s not the case. It’s -50. Just -50. Using bars is, for some reason, super common in some fields. Medics looooooove them. They’re overused and rarely the best choice. A lot of the time (but not always), bars should be replaced by points.
How I’d make this figure: I’d use points for the estimated values instead of bars. I probably wouldn’t add points to show the data, but if I did, I’d make those points much smaller (and be consistent). I’d write the full category name and rotate them 90 degrees so it’d fit. (I’d also get rid of the P-value “bars” and asterisks at the top of the figure.)