# Get R set up
library(ggplot2) # For plots
setwd("my/file/path")
my_data <- read.csv("data/sales_data.csv")
# Plot the data
ggplot(my_data, aes(x = country, y = total_sales)) +
geom_bar()Workshop 1: Getting Back into R
BI3010 Statistics for Biologists
Welcome to the first BI3010 workshop. Over the next ten workshop sessions, you’ll go from working with raw data to fitting complex statistical models, building up a workflow you can apply to almost any dataset you encounter, including every single one of your honour’s project next year. Today’s workshop focuses on getting comfortable with R and learning to spot problems in data before analysis begins, a step that sounds unglamorous but is where scientists spend 80% of their time.
Learning Objectives
- (Re)familiarise yourself with the basics of
R; - Use
Rto perform simple descriptive summaries of data and data processing; - Use the results of the EDA (Exploratory Data Analysis) to identify any issues in the data;
- Begin thinking about the assumptions we make with a specific dataset even before we begin using statistics.
All of these objectives link to useful skills for initial, quick explorations and screening of datasets. Because it’s so important to ensure that the data we subsequently use in analysis is suitable for answering our question, the importance of such EDA cannot be overstated.
These skills will be used throughout the course.
Workshop structure
The workshops on this course will have you learn an analytical workflow, focussing mainly on analysing data using linear models (and eventually generalised linear models). That is, you’re going to learn how to go from getting a dataset to extracting evidence to answer your research question. Importantly though, this workflow is not unique to linear models (or GLMs). You would (or at least should) use this same general workflow, regardless the dataset, the scenario, or the analytical method you may find yourself using.
The general workflow we’ll be using can be broken down into:
- Formulate a research question.
- Perform exploratory data analysis [Focus of today].
- Identify any hidden assumptions.
- Fit an appropriate model.
- Diagnose the model and check assumptions.
- Summarise the results.
- Interpret and provide inferences.
For the BI3010 workshops we are generally going to skip Step 1 in the workflow for convenience sake, but in your own work Step 1 is key as every subsequent step in the workflow is dependent on your question. In your own work (e.g. your Honour’s project) it’s important to take the time to think carefully about what the question you want to answer is, how best to design an experiment or data collection protocol to get data suitable for answering it, which statistical method would be suitable for analysing the data, and what caveats and limitations come with all of these. Basically, don’t take the omission of Step 1 in this course as evidence that it’s not important. It is.
Getting Started with R
Before we begin the workshop, let’s make sure that you’re set up to work with R and RStudio. If you’re working on your own machine, make sure you have R and RStudio installed. If you don’t, you can download R from CRAN and RStudio from RStudio’s website. On university computers, these should already be installed.
Once you have R and RStudio set up, open RStudio. You should see a window with four main panes:
- Source: This is where you write your scripts.
- Console: This is where you can type and execute R code directly.
- Environment/History: This pane shows you the objects that are currently stored in R’s memory.
- Files/Plots/Packages/Help/Viewer: This is a multi-tabbed pane that allows you to navigate files, view plots, manage packages, and access documentation.
Running Code
There are a few ways to run R code:
- Directly in the console: You can type
Rcommands directly into the Console pane and press Enter to run them. This is useful for very quick tests and calculations but avoid using this for project work (none of it’s saved). - From a script: Scripts are simple text files that contain
Rcode. You can write multiple lines of code in a script and run them together or one by one. This is particularly useful for more complex analyses where you want to keep a record of your code.
To run code from a script, you can:
- Click the “Run” button at the top of the Source pane to run the selected line or code chunk.
- Press
Ctrl + Enter(orCmd + Enteron a Mac) to run the current line or selected code.
Instruction Orientation
Commenting guide
It’s good practice to make comments to yourself (and others) while writing code. Comments generally act as notes that help explain your general intent for a section of code, which in three weeks’ time can be an absolute lifesaver. They’re also super useful for anyone else reading your code (and your commenting may be evaluated for some jobs).
There are many different opinions as to how to write good comments (see here for a particularly stern view), but a general principle I subscribe to is: aim to very concisely describe your general purpose for the next few lines of code. What are you trying to do, and not, what does each line of code do.
For example, I would write:
And not:
# Load the ggplot2 package into the R environment so we can create plots later
# What happened to ggplot1?
library(ggplot2)
# Use the read.csv function to read the sales_data.csv file from the data directory
# and store the resulting data frame in a data frame called 'my_data'
# Remember that it's called my_data and not data. Calling it data made things break :(
my_data <- read.csv("data/sales_data.csv")
# my_data <- read.table("data/sales_data.csv") # DONT USE THIS IT DOESNT WORK
# mi_data <- read_csv("data/sales_data.csv") # ALSO DIDN'T WORK - DUNNO WHY
# Use ggplot to create a bar plot of the total sales by country for 2024.
# The x-axis represents the categories, and the y-axis represents the total sales.
ggplot(my_data, aes(x = country, y = total_sales)) + # using ggplot, take my_data, then plot country on the x and sales on the y
geom_bar() # then have the data visualised as bar, using the structure described in the previous line of code
# When I see comments like this IRL, I die a little bit inside. Keep it clean and concise.The latter is so overly descriptive that it becomes as much of a chore to read the comments as the actual code itself. In fact, I think the code does a better job of quickly and efficiently explaining its purpose than these comments do! Another way to think of how to best write comments is: think of comments as post-it notes. You want them to be short and to act as a signal for sections that you want to dive deeper into (e.g. for debugging).
On using LLMs
LLMs (Large Language Models, or AI) are now commonly used within programming and data analysis. They can be good for boring or fiddly tasks, one-off problems, and tidying code (see here for a pragmatic take on this). The catch is that you need to know enough about programming to tell when the code they produce is correct and when it’s plausible-sounding nonsense. That line blurs fast as the code gets more complex, especially when the code is more complex than you are capable of producing yourself. For analysis specifically, without a sound theoretical understanding of statistics and data, they tend to get it wrong 60% of the time if the user can’t specify their data and question properly (see here). To make sure that doesn’t happen to you, you need to understand statistical theory and assumptions.
Within BI3010, you are allowed to use an LLM (or genAI) during workshops but not during tests. I would advise using them with caution though. Keep in mind you are here to learn. Offloading that task to a LLM means you aren’t learning. I’d suggest you try to solve any problems yourself first. Make a mistake and then try to learn from the mistake. If you can’t solve it, you are welcome to have an LLM try help to solve the problem. If it does solve the problem, try to understand how it solved it.
Using them effectively and responsibly is a skill in itself. Make your prompt specific: describe what you want, copy and paste the relevant code that gives an error, include any error messages, and say you’re working in R (otherwise you might get Python thrown back at you). Mention any packages you want to use, or specify base R.
Workshop Instructions
In this workshop, we’re going to load two datasets, starting with alligator bite strength, and carry out Exploratory Data Analysis (EDA) along with some simple data quality assurance. Broadly, our aim in such cases is to identify weird things in the data that might indicate mistakes, such as data entry errors. In real-world analysis, this step is incredibly important because mistakes are easily made, especially during data entry. Imagine entering 5,000 nearly identical observations into a spreadsheet. People get bored, start thinking about other things, get mixed up about which observation they should be adding, among many other distractions. It’s hard to stay super focused during mundane work for long periods. It would be unrealistic to expect or assume that mistakes would never happen, so we need to check for them. The problem is that if we don’t spot these errors, they can significantly influence the results when it comes time to analyse the data.
The good news is that we don’t need any complex statistics to check for these issues, and the R coding is relatively straightforward.
Setting Up Your Project
We’ll use an R Project to organise our work. An R Project is just a folder somewhere on your PC with a small .Rproj file inside it. When you open that file, RStudio automatically sets the working directory to the project folder (basically saying “go here to find the data and scripts”), with no manual path-setting required, and it will work the same on any computer or operating system.
Step 1: Create the folder structure
Create a folder called BI3010 somewhere sensible:
- University computers: the H drive works well, e.g.
H:\BI3010 - Personal computers: anywhere in your Documents folder is fine
Inside BI3010, create two subfolders:
data: all data files for the course go herescripts: your R scripts, one per workshop
Your folder should look like this:
BI3010/
├── data/
└── scripts/
Step 2: Create the R Project
Open RStudio and go to File > New Project > Existing Directory. Navigate to your BI3010 folder and click Create Project. RStudio will create a BI3010.Rproj file inside the folder and restart with the project open.
From now on, you can either open RStudio by double-clicking BI3010.Rproj or by opening RStudio directly and using the Project drop down button at the top right to select BI3010.Rproj. This ensures the working directory is always set correctly.
Step 3: Create a script for this workshop
Go to File > New File > R Script (or Ctrl + Shift + N). Save it immediately into your scripts folder as workshop1.R. To run code, press the Run button or use Ctrl + Enter.
Download the Data
The data file for this section is embedded in this page. Click the button below to download it, then save it to your BI3010/data/ folder before continuing.
Load File into RStudio
Now we can load the data. The .txt extension tells us it’s a plain text file, so we use read.table(). We need two arguments: the filename (including the data/ subfolder path), and header = TRUE to tell R that the first row contains column names rather than data.
Because the project sets the working directory to BI3010/, we can refer to files using “relative paths” (e.g. data/alligator_bite.txt) rather than “absolute paths” (e.g. H:/BI3010/data/alligator_bite.txt). Doing this saves yourself a bit of typing but also means you can share the entire folder with someone else and they don’t need to change the path to where ever they saved the project.
Tip: ?read.table (or ? followed by any function name) opens the help page for that function. For example, you can use this to figure out what header = TRUE is doing in the read.table() code below.
ali <- read.table("data/alligator_bite.txt", header = TRUE)The above fits on one line, but you can split it for readability:
ali <- read.table("data/alligator_bite.txt",
header = TRUE)Checking It Worked
If you hit an error when trying to run code, don’t panic. Nothing in R is permanent (unless you very explicitly make it so). If you hit an error, read the error message, trouble shoot solutions, run the code again, and move on (assuming it was fixed). The only thing that’s permanent is your script and original data, so save your script often and write comments as you go. If you think you might have some redundant code that you’re not sure whether removing it will break your entire script, then comment it out with # first rather than deleting it (remember that comments are not read as code).
If you get an error loading data, common causes are:
- The file isn’t in the
datafolder inside yourBI3010project, or you didn’t open RStudio viaBI3010.Rproj. - You may have misspelled one of the words (including using a capital letter instead of lower case).
- Your file may be missing a column heading.
- The file isn’t in the correct format (e.g., the file saved as a
.csvinstead of a.txt). - The quotation marks are in the wrong location.
- You may have missed the final bracket
)in theread.table()function.
To confirm that the file has been read properly, you should see it in the Environment panel (normally in the upper right of RStudio).
To do an initial quick inspect the data, you have a few options. The following shows the first six rows (use n to change how many, e.g. head(ali, n = 10)):
head(ali)Or open the full dataset in a new tab; View() gives a spreadsheet-like view with sorting and filtering:
View(ali)Or print everything to the console:
aliNot recommended: it floods the console and is hard to read.
Question Set One
Question 1. How many variables do we have in this dataset?
There are 2 variables. There are several ways to check this:
summary(ali)
ncol(ali) # Number of columns (i.e., variables) in the ali dataset
str(ali) # Structure of the ali dataset
head(ali) # Displays the first few rows
View(ali) # Opens the dataset in a spreadsheet-like view
Question 2. What are the variable names?
The variable names are bite and length. The functions above will show you this (except ncol(), which only counts columns).
Question 3. Are the variable names capitalised?
No, bite and length are both lower case. A general tip: keep variable names short and all lower case. You will type them a lot, so save yourself the effort. Avoid names like Alligator_Bite_Strength_PSI.
These questions may seem trivial, but honestly, they’re the first questions we ask ourselves when we load a dataset into R for the first time.
Getting to Know the Data
Here’s what we know about the data. The data contains 78 observations (sometimes referred to as samples) of alligator bite strength, measured in pounds per square inch. All 78 alligators were measured at Parker Island Gator Farm, Florida, USA. For each bite strength, the length of the alligator, from nose to tail, was measured and is provided in centimeters (to one decimal place).
Before we move ahead with the data, let’s take a second to become familiar with alligator bite strength. Using your search engine of preference, do a quick search to find out the bite strength of alligators. Do the same for the length of an average alligator.
When we do exploratory data analysis, we’re specifically looking for things that look weird. In order to know if something is weird, we need to know what “normal” is; hence why it’s a good idea to have some sort of notion of what a “normal” alligator’s bite strength and length may be.
Write down the average bite strength of an alligator, measured in pounds per square inch, and the average length of an alligator in centimeters (ideally as a comment in your script).
With that done, let’s move on to understanding our data. For starters, let’s check if the data is sensibly structured. In R, this means that numbers are numbers, words are characters (or sometimes factors), and so on. Basically, we want to make sure there aren’t any bite measurements like 2174.94BIGBITE, which would mess us up by causing R to think that all bite values are characters and not numbers.
To do this quickly, we can use the str() function (short for “structure”):
str(ali)We see that ali is a data.frame (great, as it should be!), with 78 observations (or samples) and two variables. The $ operator is used to extract specific variables from a dataframe, so ali$bite will give you just the bite variable from the ali dataset. From this, we can see that both the length and bite variables are being treated as numbers (indicated by num), and we get a handful of example values for both length and bite.
So far, we’re not seeing anything unusual, which is great, it means we haven’t identified any problems that we need to fix.
If we wanted to, we could extract the length variable from the dataset by doing:
ali$lengthHow would you extract the bite variable?
Let’s use these extracted variables to calculate some summary statistics, starting with the mean (i.e. average). To do so, we use the mean() function. To find the average length of our alligators, we do:
mean(ali$length)We can see that the average alligator in Parker Island Gator Farm is 374 cm long.
Compare this value with the mean length of alligators you found from a quick internet search. It might be a bit longer than expected, but this could be due to these alligators coming from a farm rather than the wild. Perhaps they’re given lots of food to help them grow? While our numbers may not match exactly with online data, significant discrepancies could be worth investigating.
Similarly, calculate the average bite strength of our alligators and see if there’s anything unusual there.
mean(ali$bite)Assuming you found that alligator bite strength is roughly 2,000 pounds per square inch, our alligators appear to be biting with an extra 1,500 pounds per square inch on average. That seems like a lot more.
This is something that should make us sit up and pay attention. We don’t know for sure if there’s a problem, but we have reason to believe there might be one. It’s now up to us to prove that there isn’t a problem.
There are many descriptive techniques to check for anomalies in our data, but a convenient “short cut” is to use the summary() function. This function provides a generic summary of our entire dataset when we supply ali as the sole argument. Try this now and see if you can identify any issues with the data:
summary(ali)Keep in mind, with EDA we’re looking for anything weird in the data.
Question Set Two
Question 1. How small is the smallest alligator in the farm for which we have data, and what’s the biggest?
The smallest alligator is 153.5 cm and the largest is 2293.0 cm. Multiple approaches work:
range(ali$length) # minimum and maximum
min(ali$length) # minimum
max(ali$length) # maximum
summary(ali) # summary for all variables
summary(ali$length) # summary for length only
Question 2. What are the minimum and maximum values for bite strength?
The minimum bite strength is 832.9 PSI and the maximum is 123,632.0 PSI.
range(ali$bite)
Question 3. Do any of these minimums or maximums seem strange?
Yes. A 2293 cm alligator is longer than a bus. And 123,632 PSI is about 60 times the average bite strength of 2,000 PSI. At least one of these values is almost certainly a data entry error.
Data Detective
Hopefully, you’ve identified at least one questionable value in our dataset. What should we do now? Throw out the entire dataset? (No. The answer is “No, you do not throw out the dataset”, at least not yet.) Let’s try to figure out what’s going on, starting with finding the problem observations.
How do we do that?
Let’s start by considering what we think a weird observation might look like. We can use the summary(ali) output to help us here. For instance, we know we have an alligator that was apparently 2293 cm long (23 m, that’s longer than a bus!). How do we find that value in our dataset to see if there was anything else unusual about it?
This is where square brackets [] come into play. In R, square brackets are used for “indexing” to extract specific rows or columns from a dataframe. Think of square brackets as a way to point to or select parts of your data.
To use them effectively, first consider how you would find the information manually (e.g. in Excel):
- We’d start by looking at the
alidataset. - We’d then skim through the
alidataset, looking for any rows where alligators are 2293 cm long. - Once we’ve found these rows, examine all the information related to those alligators.
How does that translate into R? Almost exactly; we just need to write those three steps using code.
First, we’d start by looking at the ali dataset.
aliNext, we’d skim through the ali dataset, looking for any rows where alligators are 2293 cm long. This is where we can use square brackets [] in conjunction with conditional statements. Conditional statements allow us to specify criteria like “when this condition is met, do something”.
In our case, the condition is finding alligators that are 2293 cm long. In R, we write this condition using double equal signs (==) to denote “equal to” (a single = assigns a value; == tests whether two things match). This is basically the same as your eyes going down the Excel column looking for 2293, noting any instances where it was true or not. For example:
ali$length == 2293.0The output is a list of 78 values, each being either TRUE or FALSE, corresponding to each row in the ali dataset. TRUE values indicate rows where the alligator is 2293 cm long, and FALSE values indicate rows where it is not (in R, “not equal to” is denoted as !=).
Finally, once we’ve found these rows, we examine all the information related to those alligators. When using square brackets, remember that the row number comes before the comma, and the column number comes after (e.g. ali[row, column]). If we omit either the row or column number, R will return all rows or all columns respectively. For instance, to view all the information (i.e. all columns) for the 30th row, you would use:
ali[30,]To make life a bit easier for ourselves, we can use these square brackets in combination with the conditional statement to extract the entire row for the alligator that’s 2293 cm long. The TRUE and FALSE values act as indicators for which rows we want to display: if a value is TRUE, we see that row; if it’s FALSE, we don’t. We can achieve this by writing:
ali[ali$length == 2293,]What do you make of this observation? We’re expecting length to be about 200 to 400 centimeters (2 to 4 m), but this alligator is apparently 10 times bigger than that. And bite is about 100 times bigger. Neither of the two values seems to have a decimal point (or at least R has given them zeros after the decimal point).
Why might that be the case? Let’s find a “normal” observation to compare this giant alligator to. Assuming the first observation is a normal one, use the square brackets to extract the first row of data from the ali dataset:
ali[1, ]Well, they’re clearly different. The first alligator in our dataset looks pretty normal. Maybe a bit on the big side (just over 4 m long), but not weirdly so. There are only two possible explanations for this: either there are giant alligators among us, or there has been a data entry mistake.
Question Set Three
Question 1. What may have occurred to cause this apparent data entry mistake?
The decimal points were accidentally omitted during data entry. The correct values are 229.3 cm and 1236.32 PSI. If you look at the raw data file, neither value has a decimal point.
A More Refined Detective
The approach above works, but typing 2293 every time is awkward and can cause issues. Referencing specific numbers like this in code is called “magic numbers” because to anyone looking at your code, they have absolutely no idea why this number is special and what why you are working with it. Better to store the maximum in a clearly labelled variable and reference that instead:
# Get the maximum alligator length
max_length <- max(ali$length)
# Now use this value to find the problematic rows
ali[ali$length == max_length, ]By storing the result of max(ali$length) in a variable (max_length), we can simplify our code and make it more readable. This way, if we need to refer to the maximum length multiple times, we only need to use the max_length variable. Within the ali dataset, this code extracts all columns (since we have not specified anything after the comma in the square brackets) for any rows where the length of an alligator is equal to the maximum length of any alligator.
We can also adapt the conditional statement to find rows based on different criteria. For example, if we wanted to extract all rows with alligators that were greater than the average length, we could do:
# Calculate the average alligator length
average_length <- mean(ali$length)
# Extract all rows where the alligator length is greater than the average
ali[ali$length > average_length, ]In this case, ali[ali$length > average_length, ] will give us all columns for rows where the length of the alligator is greater than the calculated average length. We now get data from more than one alligator, as more than one alligator meets our condition of being longer than average. We can also see the “data entry mistake alligator” using this method.
Let’s refine our approach further by using a different function: quantile(). First, rerun the summary(ali) code to observe the quantiles (1st Qu. and 3rd Qu.) that it reports. These quantiles represent the values below which 25% of the data falls and above which 75% of the data falls. We can retrieve these quantile values ourselves by doing:
quantile(ali$length, probs = c(0.25,0.75))The probs argument in the quantile() function allows us to specify which probabilities we want to calculate quantiles for. In the previous example, we requested the 25% (0.25) and 75% (0.75) quantiles. Since we wanted to get these two values, we used the c() function, which combines (or “concatenates”) multiple values into a single vector. If we only need a single quantile, such as the 75th percentile, we can directly write quantile(ali$length, probs = 0.75). In this case, there’s no need to use c() because we’re only asking for one value.
Question Set Four
Question 1. Write the code that would extract all rows in the ali dataset where bite values are above the 3rd quantile.
ali[ali$bite > quantile(ali$bite, probs = 0.75),]
Question 2. Repeat the above, but for bite values less than or equal to the 33% quantile. If you’re unsure how to specify “less than or equal to”, have a quick search online or use your preferred LLM.
ali[ali$bite <= quantile(ali$bite, probs = 0.33),]
Question 3. Using the tools from above, can you identify any other possible data entry mistakes caused by missing decimal points?
Row 20 appears to be the only observation with this problem. The other bite and length values all look plausible.
Correcting Data
So, we’ve identified a problem. What now? Do we delete the observation? Declare that Sarcosuchus are not extinct and are alive and well in Parker Island Gator Farm, Florida, USA (and have actually doubled in size)? Set up a Facebook group to share our concrete and irrefutable evidence of Sarcosuchus and start producing a film to share this wonderful news with the world? Obviously not. That’d be silly.
But we do need to do something. Row 20 of our dataset is clearly a mistake, but correcting it is a really sensitive topic. Part of the reason for that sensitivity is because a non-trivial number of scientists have been found to actually fake their data in order to further their own careers and prestige (e.g. see here, here, and most ironically here for examples).
We want to be better than that. So before touching anything, go back to the source: not the downloaded file, but the original recorded values. Here that means contacting Parker Island Gator Farm to confirm our suspicion that row 20 is simply missing decimal points. (Please don’t actually contact them. They’ll be confused as hell.)
In the interest of time, I’ll confirm that it’s missing decimal points so we can move on to fixing the problem. Specifically, bite should have two decimal points, and length should have one.
We have two options: edit alligator_bite.txt directly and reload, or fix it in R. Do it in R. If you edit the raw file, there’s no record of what changed or why, and that starts to look a lot like the data manipulation in the examples I showed above.
Doing the correction in your code is totally transparent; it’s right there for anyone who looks at your code to see.
So how do we fix row 20? Well, we use the skills from above: extracting, indexing, assigning, and simple math.
Let’s start with ali$bite. We only want to change row 20. How do we select only row 20? ali$bite[20]. Note that we don’t need to specify a column in the square brackets, because we’re working with a single column, the bite column.
Next, we need to replace ali$bite[20] with an altered version of itself, so we would write:
ali$bite[20] <- ali$bite[20]This will just replace the bite value from row 20 with itself. We’re not actually doing anything here, but we’re getting closer.
Next, we need to “add” back in the decimal point that was lost during data entry. How do we shift the value two places to the right of the decimal point? We divide by 100, leaving us with:
# Resolving data entry issue with missing decimal point
ali$bite[20] <- ali$bite[20] / 100Using a hardcoded row number like [20] works here, but it’s fragile: if the dataset were ever re-sorted or a row added above this one, row 20 would point to the wrong observation (i.e. 20 is a “magic number”). A safer approach uses the same conditional indexing we’ve already learned, targeting the value itself rather than its position:
ali$bite[ali$bite > 100000] <- ali$bite[ali$bite > 100000] / 100This reads as: “for any bite value greater than 100,000, divide it by 100.” It will find the right row regardless of where it sits in the dataset.
Question Set Five
Question 1. Write (copy & paste) the code you used to resolve the data entry mistake for ali$length. Be sure to add a clear comment indicating that you have made this change, so that if anyone reviews your code, they understand the reason for the modification and don’t suspect any foul play.
# Correcting missing decimal point, confirmed by Parker Island Gator Farm on 1/1/2020
ali$length[ali$length > 2000] <- ali$length[ali$length > 2000] / 10
Question 2. Having made any changes to your data, you need to check it to make sure you didn’t create a new mistake. Do so now using any techniques you feel are appropriate. Copy and paste the code you used to check if the correction worked.
summary(ali)
Do it for Real
Download the data file below and save it to your BI3010/data/ folder.
The goal is the same as before: find all the data entry mistakes. This dataset is larger and the errors are harder to spot, so you’ll need to be more systematic. A few hints:
unique()lists every distinct value in a column, useful for spotting typos in text columns.mean()returnsNAif any values are missing. Addna.rm = TRUEto ignore them:mean(variable, na.rm = TRUE).- There are five mistakes in total.
Start with an EDA as before. Can you find them all?
| Variable | Description |
|---|---|
species |
One of three species: Peregrine Falcon, Saiga Antelope, or Cheetah |
weight |
Body weight in grams |
speed |
Speed over a 50 m distance, in km/h |
- A total of 500 individuals were measured, but these weren’t equally divided between the three species.
- Multiple people participated in data entry, and it’s believed that not all were confident in how to do this.
- Physical copies of the data are available should any data need to be verified (contact one of the teaching staff if you need a data entry confirmed).
fast <- read.table("data/animal_speed.txt", header = TRUE)# Start your exploration here
summary(fast)
str(fast)
unique(fast$species)Question 1. Identify all unusual observations in the dataset.
- One individual has a recorded weight of 0 g. It should be 45514.
- One individual has a recorded weight of 3 g. It should be 47255.
- One of the species names has been entered incorrectly. It should be Peregrine Falcon.
- One Peregrine Falcon has a recorded speed of 35,711 km/h. It is missing two decimal places.
-
One observation has a missing value (
NA, or “not available”) in the speed column. The data provider does not know what this value should be.
Question 2. Copy and paste the code you used to resolve the above issues.
# Weight of 0 g corrected to the value confirmed by the data provider
fast$weight[fast$weight == 0] <- 45514
# Weight of 3 g corrected to the value confirmed by the data provider
fast$weight[fast$weight == 3] <- 47255
# Species typo corrected, then re-factored
fast$species[fast$species == "Peregrin Took"] <- "Peregrine Falcon"
fast$species <- factor(fast$species)
# Peregrine Falcon speed missing two decimal places, so divide by 100
fast$speed[fast$speed > 10000 & !is.na(fast$speed)] <- fast$speed[fast$speed > 10000 & !is.na(fast$speed)] / 100
# The NA in speed is left as-is: the data provider cannot confirm the true value
# Final check
summary(fast)
Question Set Six
Make sure you have corrected all five data entry errors before answering these questions - the answers depend on the corrected dataset.
Question 1. What is the median weight in this dataset?
The median weight is 46,453 g.
median(fast$weight)
Question 2. What is the median speed in this dataset?
The median speed is 102.1 km/h.
median(fast$speed, na.rm = TRUE)
Question 3. Which species is most common in this dataset?
The most common species is the Saiga Antelope (177 individuals, compared with 167 Cheetahs and 156 Peregrine Falcons). The table() function gives a count per species, but there are other ways to check:
table(fast$species)
Chupacabras, Aliens and Beer
The final dataset is chupa.txt. We’ll work through it together in the last 30 minutes, but start early if you’re ahead. It contains the following variables:
year: The year in which data was collectedcity: The city where the data was collectedalc: The average alcohol consumption in that city, in that yearufo: The number of unidentified flying objects (i.e. aliens) reported in that city, in that yearbelief: The most common response from a selection of members of the public when asked if they believe in the supernatural, in that city, in that yearchupa: The growth rate of chupacabra sightings from the previous year to the current year, in that city (where a value of zero means the number of sightings is constant, -0.5 means the number of sightings has declined by 50%, and 0.5 means they have increased by 50%).
There are data entry errors in this dataset. Find them all using EDA.
Download the data file below and save it to your BI3010/data/ folder.
As a group, we’ll look at snippets of code and identify the errors.
chupa <- read.table("data/chupa.txt", header = TRUE)head(chupa)
str(chupa)
unique(chupa$city)
summary(chupa)
unique(chupa$belief)
Find and fix all data entry errors in the chupa dataset. Describe each error and write the correction code.
Error 1 - Two spellings of Amarillo
One row has the city entered as “amarillo” (lower case a) while all other rows use “Amarillo”. Use unique(chupa$city) to spot this, then check the rows affected with indexing.
Error 2 - Placeholder value in ufo (9999)
The ufo column contains one value of 9999, which appears to be a placeholder for missing data rather than a real count. A histogram of ufo makes this obvious. This value should be replaced with NA.
Error 3 - Negative value in alc
Alcohol consumption can’t be negative. There’s one observation with a value of -5 in the alc column. A histogram will reveal it as a clear outlier on the left. The correct value is 5. Note: negative values in the chupa column are fine and expected, since it’s a growth rate.
Coding Puzzles
In the following snippets of code, can you identify the error? If you’re not sure, try running them and see what the error says.
Puzzle 1.
summary(alc)
# We have not specified that this is in the chupa dataset
summary(chupa$alc)
Puzzle 2.
chupa[chupa$ufo==9999]
# We're missing the comma to specify which columns to show
chupa[chupa$ufo == 9999, ]
Puzzle 3.
chupa[chupa$ufo = 9999, ]
# = assigns a value; use == for a conditional test
chupa[chupa$ufo == 9999, ]
Puzzle 4.
chupa[chupa$city = Houston,]
# Houston needs to be in quotation marks (it's a string, not an object)
# Also = should be ==
chupa[chupa$city == "Houston", ]
Puzzle 5.
ali[ali$Length > 200, ]
# Names are case sensitive. Should be `length`, not `Length`
ali[ali$length > 200, ]
Puzzle 6.
ali[ali$bite > mean(ali$bite), ]
# If there are NAs in bite this will fail - include na.rm = TRUE
ali[ali$bite > mean(ali$bite, na.rm = TRUE), ]
num or int in str() output.str() 输出中显示为 num 或 int。chr in str() output.str() 输出中显示为 chr。mean() return NA if any value is missing unless you add na.rm = TRUE.NA,mean() 等函数会返回 NA,除非添加 na.rm = TRUE。.R) containing a sequence of R commands. Run it top to bottom to reproduce an analysis exactly..R)。从上到下运行即可完整重现分析过程。mean(), summary()). You pass inputs inside the parentheses and get an output back.mean()、summary())。在括号内传入输入值,函数返回输出结果。na.rm = TRUE in mean(x, na.rm = TRUE). Arguments control how the function behaves.mean(x, na.rm = TRUE) 中的 na.rm = TRUE。参数控制函数的行为方式。df[rows, columns]. Leave a slot blank to mean "all".df[行, 列]。留空表示"全部"。TRUE or FALSE (e.g. x > 5). Used inside square brackets to filter rows that meet a condition.TRUE 或 FALSE 的逻辑判断(如 x > 5)。用于方括号内筛选满足条件的行。setwd() or via the RStudio menu.setwd() 或 RStudio 菜单设置。