HSCI 410 · Lesson 6

Exploratory Data Analysis and Visualization

Exploratory Data Analysis For Epidemiology

Learning objectives for this lesson:

  • Explain the purpose of exploratory data analysis and why summary statistics alone can mislead, using Anscombe's quartet as the illustration.
  • Choose an appropriate display for one variable, for two variables and for many variables, given the variable types involved.
  • Produce histograms, bar charts, boxplots and scatterplots in base R, and arrange several base R plots in one window.
  • Build the same plots in ggplot2 by mapping variables to aesthetics and adding geoms, scales, labels and themes.
  • Use facet_wrap() and facet_grid() to compare a display across groups.
  • Produce a correlation heatmap, a scatterplot with marginal histograms, and a combined multi-panel figure.
  • Read plots for data problems such as skew, outliers, heaping and impossible values, and write a figure caption that states what the figure shows.

This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University based on Dohoo, I. R., Martin, S. W., & Stryhn, H. (2012). Methods in Epidemiologic Research. VER Inc.

Lesson 6 · HSCI 410

Exploratory Data Analysis and Visualization

Looking at the data before modelling them, and turning what you see into a figure that others can read.

Exploratory Data Analysis For Epidemiology
The running case

A visual briefing for a loneliness working group

The request

The working group wants a one-page visual briefing on social connection before any modelling is commissioned.

The data

The briefing uses the 2021 wave of the Canadian Social Connection Survey (CSCS), a national online survey whose respondents were not sampled at random.

Each section adds plots to the briefing, and Section 4 assembles them into a single figure.

Lesson map

Four sections

SectionWhat it covers
1. Look Before You ModelExploratory analysis, Anscombe's quartet, choosing a display, and poor displays
2. Quick Looks in Base Rhist(), barplot(), boxplot(), plot(), pairs() and saving plots
3. The Grammar of Graphics with ggplot2Aesthetics, geoms, facets, colour, themes and ggsave()
4. Many Variables and Combined FiguresCorrelation heatmaps, marginal histograms, multi-panel figures and captions
Working in R

Every line gives you something to look at

  • Each activity gives the complete code, with a comment on every line.
  • The console output shown on the page is what R printed when the code was run.
  • The questions ask you to read and interpret the plots and the output.
  • A narrated walkthrough, Data Visualization in Base R and ggplot2, follows the same code.
Reference

Glossary: Key Terms, People & Concepts

📚 Reference page, available throughout the lesson

This glossary collects the key concepts, methods and people in this lesson. It can be used as a reference while working through the material or as a review before assessments. Typing in the search box filters the entries.

Key Concepts & Ideas
Exploratory Data Analysis Exploratory data analysis uses plots and simple summaries to learn the shape, quality and relationships of the data before any formal model is fitted. John Tukey promoted the approach from about 1970 and set it out in his book Exploratory Data Analysis (1977).
Confirmatory Analysis Confirmatory analysis tests a question that was set before the data were examined and reports estimates, confidence intervals and p-values. It usually follows exploration.
Anscombe's Quartet Anscombe's quartet is a set of four small datasets with nearly identical means, standard deviations, correlations and fitted lines but very different patterns when plotted. It is supplied with R as anscombe.
Shape of a Distribution The shape of a distribution describes whether values are symmetric or skewed, how many peaks they have, and whether they pile up at the lowest or highest possible value.
Outlier An outlier is a value far from most of the others. It may be a recording error or a genuine extreme, and it is checked before any decision is made about it.
Impossible or Implausible Value An impossible value lies outside the possible range of a variable, and an implausible value is possible but very unlikely, such as more than 112 hours of paid work a week.
Heaping (Digit Preference) Heaping occurs when answers pile up on round or standard numbers, such as 40 hours of work a week. It usually reflects rounding, and it reduces the information carried by small differences.
Ceiling and Floor A ceiling or floor is a pile-up of answers at the highest or lowest value a scale allows. On the three-item loneliness scale, 270 people in the briefing file scored the maximum of 9.
Top-Coding Top-coding records every value above a limit as the limit itself. In the CSCS data, household totals above 20 were recorded as 20, which produced a spike at that value.
Overplotting Overplotting occurs when many points are drawn on top of one another, so the plot cannot show where most of the data lie. Jitter and transparency reduce it.
Small Multiples Small multiples are a series of small panels that repeat the same display for different groups on shared axes, a term associated with Edward Tufte. In ggplot2 they are drawn with facets.
Chartjunk Chartjunk is Tufte's term for decoration in a graphic that carries no information, such as three-dimensional effects, heavy shading or ornaments.
Association and Causation An association means that two variables tend to vary together. It may arise because one causes the other, because of reverse causation, or because a common cause (a confounder) affects both.
Colour-Vision Deficiency Colour-vision deficiency is a reduced ability to tell some colours apart, most often red from green. It affects about 8% of men and 0.5% of women of northern European ancestry.
Figure Caption A figure caption states what a figure shows, the data source, the number of people included, and what each panel and mark means, so that the figure can be understood on its own.
Methods & Statistical Concepts
Histogram A histogram divides the range of a numeric variable into bins and draws a touching bar for the number of values in each. It is drawn with hist() in base R and geom_histogram() in ggplot2.
Bar Chart A bar chart draws one separate bar per category of a categorical variable. In base R it is drawn with barplot(table(x)), and in ggplot2 with geom_bar() or geom_col().
Boxplot A boxplot summarizes a numeric variable with its median, quartiles and whiskers, and draws values beyond the whiskers as separate points. Side-by-side boxplots compare groups.
Interquartile Range and Whiskers The interquartile range is the distance between the lower and upper quartiles, the length of the box. Whiskers reach the most extreme values within 1.5 times the interquartile range of the box.
Scatterplot A scatterplot places one point per observation for two numeric variables, showing the direction, form and strength of their relationship.
Scatterplot Matrix A scatterplot matrix draws a small scatterplot for every pair of several variables in a grid. In base R it is drawn with pairs().
Jitter Jitter adds a small random shift to each plotted value so that points sharing a value spread out. It changes only the picture and leaves the data unchanged.
Transparency (Alpha) Transparency draws points so that they can be seen through, so that areas where many points overlap look darker. It is set with alpha in ggplot2 or rgb(..., alpha) in base R.
Grammar of Graphics The grammar of graphics describes every statistical graphic as a combination of data, aesthetic mappings, geometric objects, scales, facets and a theme. ggplot2 implements a layered version of it.
Aesthetic Mapping An aesthetic mapping, written inside aes(), links a variable to a visual property such as horizontal position, colour or fill. A value given outside aes() is set for every mark.
Geom A geom is the type of mark that represents data in ggplot2, such as geom_histogram(), geom_boxplot(), geom_point() or geom_tile(). Each geom is a layer.
Facets Facets split a ggplot2 plot into panels by one or two categorical variables. facet_wrap() uses one variable, and facet_grid() places one variable in rows and another in columns.
Theme The theme controls the parts of a ggplot2 plot that do not depend on the data, such as fonts, text size, grid lines and background, for example theme_minimal().
Loess Smoother A loess smoother is a curve that follows the average of y across x by fitting many small weighted regressions, so it can bend. It is drawn with geom_smooth(method = "loess").
Pearson Correlation Coefficient The Pearson correlation coefficient, r, measures how closely two numeric variables follow a straight line, from −1 to +1. It is computed with cor().
Spearman's Rank Correlation Spearman's rank correlation is the Pearson correlation of the ranks of two variables. It suits ordinal or skewed variables and is less affected by extreme values.
Correlation Matrix A correlation matrix holds the correlation of every pair of several variables, with 1 on the diagonal. It is symmetric, so k variables give k(k − 1)/2 distinct correlations.
Correlation Heatmap A correlation heatmap draws a correlation matrix as coloured squares, usually with a diverging colour scale that is white at zero. It shows association and cannot identify causes.
Marginal Histogram A marginal histogram is drawn along the edge of a scatterplot to show the distribution of one of its variables. ggExtra::ggMarginal() adds them to a ggplot2 scatterplot.
patchwork patchwork is an R package that combines ggplot2 plots into one figure: + or | places plots side by side and / stacks them.
ggsave() and Resolution ggsave() writes a ggplot2 plot to a file of a chosen width and height. Resolution, in dots per inch (dpi), is usually 300 for print.
Okabe-Ito Palette The Okabe-Ito palette is a set of eight colours chosen so that people with common forms of colour-vision deficiency can tell them apart.
Bracket Recoding and Complete Cases Bracket recoding changes the rows of a variable that meet a condition, as in x[x == 20] <- NA. na.omit() keeps only rows with no missing values (a complete-case file).
Key People
John W. Tukey (1915–2000) John Tukey was an American statistician at Princeton University and Bell Laboratories. He introduced the boxplot, first called the schematic plot, in about 1970, and his book Exploratory Data Analysis (1977) set out the approach that gives the field its name.
Francis J. Anscombe (1918–2001) Francis Anscombe was a British-born statistician at Yale University. His 1973 paper in The American Statistician presented the quartet that bears his name.
Leland Wilkinson (1944–2021) Leland Wilkinson was a statistician and computer scientist who wrote The Grammar of Graphics (1999; second edition 2005) and developed the SYSTAT statistical software.
Hadley Wickham Hadley Wickham is a statistician and software developer who created ggplot2 and described the layered grammar of graphics (2010). He is chief scientist at Posit, formerly RStudio.
William S. Cleveland William Cleveland is a statistician who studied graphical perception with Robert McGill (1984) and developed the loess method of local regression smoothing.
Edward R. Tufte Edward Tufte is a statistician and political scientist and professor emeritus at Yale University. The Visual Display of Quantitative Information (1983) introduced the terms data-ink and chartjunk.
Masataka Okabe and Kei Ito Masataka Okabe and Kei Ito are Japanese scientists who proposed a colour palette for Color Universal Design, first published in 2002 and revised in 2008, that remains distinguishable for people with colour-vision deficiency.
No matching entries. Try a different search term.
Section 1 of 4

Look Before You Model

⏱ Estimated time: 45 minutes
Lesson 6 · Section 1

Look Before You Model

Plots show the shape of the data, and the shape decides which summaries and models make sense.

Exploratory data analysis

Two kinds of analysis

Exploratory

The analyst asks what the data look like.

Plots and simple summaries show shapes, unusual values and relationships, and they raise new questions.

Confirmatory

The analyst tests a question that was set in advance.

Models give estimates, confidence intervals and p-values for that question.

Tukey (1977) argued that exploration should come first.

The running case

The working group's three questions

  • The group asks how lonely people in the survey are.
  • The group asks whether loneliness differs by age and gender.
  • The group asks how loneliness goes together with social support and mental health.
4,045
respondents in the CSCS 2021 wave
3,083
respondents with complete data on the briefing variables
Anscombe's quartet

Four datasets with the same summary statistics

StatisticSet 1Set 2Set 3Set 4
Mean of x and of y9, 7.59, 7.59, 7.59, 7.5
SD of x and of y3.32, 2.033.32, 2.033.32, 2.033.32, 2.03
Correlation0.8160.8160.8160.817
Fitted liney = 3.00 + 0.500 x in every set
Anscombe's quartet

The plots tell four different stories

Four scatterplots of Anscombe's quartet, each with the same fitted line: set 1 a linear cloud, set 2 a curve, set 3 a line with one outlier, set 4 a vertical column of points with one far point.
The same fitted line (red) suits only set 1. Data: Anscombe (1973), supplied with R as anscombe.
What to look for

Features that plots reveal

One variable

The analyst checks the shape (symmetry, skew and peaks), unusual and impossible values, heaping on round or extreme values, and missing values.

Two or more variables

The analyst checks the direction of a relationship, its form (straight or curved), its strength, and whether it differs between groups.

Choosing a display

Match the display to the question and the variable types

Recall: Lesson 2, Section 3

A histogram shows one numeric variable, a bar chart the counts of a categorical variable, boxplots a numeric variable across groups, and a scatterplot two numeric variables.

New: many variables

A scatterplot matrix or a correlation heatmap gives an overview, and facets repeat a display for each group.

Check: loneliness across four age groups calls for boxplots side by side, with the points shown.

Poor displays

Four displays that conceal the data

Four poor displays: a bar chart with a truncated axis, a bar chart of means without spread, a pie chart with 14 slices, and an overplotted scatterplot.
A truncated axis, means without spread, too many slices and overplotting (CSCS 2021).
Carry forward

What to take into the next section

  • Identical summary statistics can describe very different data, so the analyst plots first.
  • The display is chosen from the question and the types of the variables involved.
  • Some displays conceal the data, and each has a better alternative.

Introduction and Overview

Lessons 4 and 5 fitted regression models to outcomes of many kinds. Each of those models rested on assumptions about the shape of the data: a straight-line relationship, residuals with an even spread, counts that are not more variable than the model allows, or observations that are independent. This lesson returns to the step that should come before any model is fitted, which is looking carefully at the data. Exploratory data analysis uses plots and simple summaries to find the shape of each variable, unusual or impossible values, and the relationships between variables. It tells the analyst whether the planned model suits the data, and it often suggests questions that the model should answer.

Learning Objectives

  • Explain the purpose of exploratory data analysis and how it differs from confirmatory analysis.
  • Use Anscombe's quartet to explain why summary statistics alone can mislead.
  • Choose an appropriate display for one variable, for two variables and for many variables, given the variable types involved.
  • Recognize four common poor displays and describe what each one conceals.
  • Compute summary statistics and draw a set of four plots in R for Anscombe's quartet.

Box 6.1 introduces the running case for the lesson, a one-page visual briefing on social connection that an analyst prepares for a regional working group on loneliness. Each section of the lesson returns to this briefing, so the box sets out the group's three questions and the survey data from which the plots are drawn.

📋 Box 6.1: The running case: a visual briefing on social connection

A regional health authority has formed a working group on loneliness. Before it commissions any statistical modelling, the group has asked an analyst for a one-page visual briefing on social connection, drawn from the 2021 wave of the Canadian Social Connection Survey (CSCS). The group has three questions: how lonely people in the survey are, whether loneliness differs by age and gender, and how loneliness goes together with social support and mental health. The 2021 wave has 4,045 respondents, one row per person, and 3,083 of them answered every question used in the briefing. Each section of this lesson adds plots to the briefing, and Section 4 assembles them into a single multi-panel figure with a caption. Section 2 describes how the survey recruited its respondents and what its sample can support.

The rest of this section explains what exploratory data analysis is, shows with Anscombe's quartet why summary statistics alone can mislead, matches displays to the questions they answer, and examines four common poor displays.

What Exploratory Data Analysis Is

The statistician John Tukey set out this kind of work in his book Exploratory Data Analysis (Tukey, 1977). He distinguished two activities. Confirmatory analysis tests a question that was set before the data were examined, and it reports estimates, confidence intervals and p-values. Exploratory analysis asks what the data look like, using plots and simple summaries, and it is open to finding things that nobody planned to look for. Tukey compared exploration to detective work: the analyst gathers clues about the data before deciding which formal test or model to apply.

Exploration serves three purposes in an epidemiological analysis. It checks data quality, by revealing values that are impossible, implausible or suspiciously common. It describes distributions, which decides whether a mean or a median is the better summary and whether a variable needs to be transformed or grouped. It shows relationships, including whether they are straight or curved and whether they differ between groups, which decides how variables should enter a model. Lesson 2 introduced numeric summaries, data cleaning and the basic plots for one and two variables (histograms, boxplots, bar charts and scatterplots). This lesson develops those plots into an exploratory workflow and adds displays for many variables.

Table 6.1 sets out five questions that the analyst asks during exploration, what a plot can show in answer to each, and how the answer changes the later analysis.

Table 6.1. Questions asked during exploration, what a plot can show, and what the answer changes in the analysis.

Question the analyst asksWhat the plot can showWhat it changes in the analysis
What shape does each variable have?Symmetry or skew, one peak or two, a floor or ceilingThe choice between a mean and a median; transformation or grouping
Are there unusual or impossible values?Outliers, values outside the possible range, spikes at a maximumChecking the codebook; recoding values to missing; a sensitivity analysis
Do answers pile up on particular values?Heaping on round numbers or on the end of a scaleCaution in interpreting small differences; grouping the variable
How do two variables go together?Direction, form (straight or curved) and strengthHow each variable enters a model; whether a curve is needed
Does a pattern differ between groups?Different shapes or slopes in different panelsWhether an interaction or a stratified analysis is needed

Each question in Table 6.1 is answered by looking at a plot. The next part shows, with four small datasets, why summary statistics cannot take the place of a plot.

Anscombe's Quartet: Why Summary Statistics Can Mislead

In 1973 the statistician Francis Anscombe published four small datasets to argue that graphs are an essential part of statistical analysis and that a computer should produce graphs as well as calculations (Anscombe, 1973). Each dataset has eleven pairs of values, labelled x and y, and the four datasets are supplied with R as a data frame called anscombe. Their summary statistics are almost identical, as Table 6.2 shows.

Table 6.2. Summary statistics for the four datasets in Anscombe's quartet (Anscombe, 1973).

StatisticSet 1Set 2Set 3Set 4
Mean of x9.09.09.09.0
Mean of y7.57.57.57.5
Standard deviation of x3.323.323.323.32
Standard deviation of y2.032.032.032.03
Correlation of x and y0.8160.8160.8160.817
Fitted liney = 3.00 + 0.500xy = 3.00 + 0.500xy = 3.00 + 0.500xy = 3.00 + 0.500x

A report that gave only these numbers would describe the four datasets as the same. The plots show that they are very different. Figure 6.1 draws each of the four sets as a scatterplot with its fitted line.

Four scatterplots labelled Set 1 to Set 4, each with the same red fitted line. Set 1 is a loose linear cloud; set 2 is a smooth arch; set 3 is a nearly perfect line with one high point; set 4 has ten points stacked at x equals 8 and one point at x equals 19.
Figure 6.1. Anscombe's quartet, with the least-squares line (red) added to each set. The line is the same in all four panels, and it describes only set 1 well. Data: Anscombe (1973), the anscombe data frame in R.

The accordion below takes the four panels of Figure 6.1 in turn and states, for each set, what the plot shows and whether the fitted line is a fair summary.

Set 1: the situation a straight line was built for

The points scatter evenly around a straight line. A correlation of 0.82 and a fitted slope of 0.5 are fair summaries of these data, and the assumptions of linear regression from Lesson 3, Section 2 (reviewed in Lesson 4, Section 1) look reasonable.

Set 2: a curve

The points follow a smooth arch with almost no scatter. The relationship is strong, but it is curved, so a straight line misdescribes it: the line overestimates y at both ends and underestimates it in the middle. A model with a squared term for x would fit these data almost exactly.

Set 3: one outlier

Ten points lie almost exactly on a line, and one point lies far above it. That single point pulls the fitted line upward and lowers the correlation from nearly 1 to 0.82. The analyst would check whether the point is a recording error before deciding how to handle it.

Set 4: one influential point

Ten of the eleven points share the same value of x (8), so within those points there is no relationship to estimate. The one point at x = 19 creates the whole correlation and decides the slope of the line on its own. In the terms used in Lesson 3, the point has high leverage and a large Cook's distance: it lies so far from the others along the x axis that the fitted line must pass through it.

The same lesson has been repeated with more recent examples. Starting from a dataset drawn in the outline of a dinosaur by Alberto Cairo, Matejka and Fitzmaurice (2017) generated the "Datasaurus Dozen", a set of further datasets that form stars, circles and lines, each with the same means, standard deviations and correlation as the dinosaur to two decimal places. The point is the same as Anscombe's: a summary statistic describes one feature of the data, and only a plot shows whether that feature is the one that matters.

Activity 6.1 reproduces Table 6.2 and Figure 6.1 in R from the anscombe data frame, and its first question asks what an analyst who saw only the summary statistics would conclude.

R Activity 6.1: Anscombe's quartet in R

This activity uses the anscombe data frame, which is supplied with R, so no file needs to be loaded. The first block prints the data and computes the mean and standard deviation of each column. The second computes the four correlations and two fitted lines. The third draws the four scatterplots in one window.

anscombe                          # four small datasets that come with R
round(colMeans(anscombe), 2)      # the mean of each column
round(sapply(anscombe, sd), 2)    # the standard deviation of each column
Console output
x1 x2 x3 x4 y1 y2 y3 y4 1 10 10 10 8 8.04 9.14 7.46 6.58 2 8 8 8 8 6.95 8.14 6.77 5.76 3 13 13 13 8 7.58 8.74 12.74 7.71 4 9 9 9 8 8.81 8.77 7.11 8.84 5 11 11 11 8 8.33 9.26 7.81 8.47 6 14 14 14 8 9.96 8.10 8.84 7.04 7 6 6 6 8 7.24 6.13 6.08 5.25 8 4 4 4 19 4.26 3.10 5.39 12.50 9 12 12 12 8 10.84 9.13 8.15 5.56 10 7 7 7 8 4.82 7.26 6.42 7.91 11 5 5 5 8 5.68 4.74 5.73 6.89 x1 x2 x3 x4 y1 y2 y3 y4 9.0 9.0 9.0 9.0 7.5 7.5 7.5 7.5 x1 x2 x3 x4 y1 y2 y3 y4 3.32 3.32 3.32 3.32 2.03 2.03 2.03 2.03

The data frame has eleven rows and eight columns: x1 and y1 form set 1, x2 and y2 form set 2, and so on. colMeans() gives the mean of every column, and sapply(anscombe, sd) applies the standard deviation function sd() to every column. Every x column has a mean of 9.0 and a standard deviation of 3.32, and every y column has a mean of 7.5 and a standard deviation of 2.03.

round(cor(anscombe$x1, anscombe$y1), 3)   # correlation in set 1
round(cor(anscombe$x2, anscombe$y2), 3)   # set 2
round(cor(anscombe$x3, anscombe$y3), 3)   # set 3
round(cor(anscombe$x4, anscombe$y4), 3)   # set 4
coef(lm(y1 ~ x1, data = anscombe))        # intercept and slope of the line, set 1
coef(lm(y4 ~ x4, data = anscombe))        # the same for set 4
Console output
[1] 0.816 [1] 0.816 [1] 0.816 [1] 0.817 (Intercept) x1 3.0000909 0.5000909 (Intercept) x4 3.0017273 0.4999091

The four correlations are 0.816, 0.816, 0.816 and 0.817. lm(y1 ~ x1, data = anscombe) fits a straight line, and coef() prints its intercept and slope: 3.00 and 0.500 for set 1, and 3.00 and 0.500 for set 4.

par(mfrow = c(2, 2))    # four plots in one window: two rows, two columns
plot(anscombe$x1, anscombe$y1, main = "Set 1", xlab = "x", ylab = "y")
plot(anscombe$x2, anscombe$y2, main = "Set 2", xlab = "x", ylab = "y")
plot(anscombe$x3, anscombe$y3, main = "Set 3", xlab = "x", ylab = "y")
plot(anscombe$x4, anscombe$y4, main = "Set 4", xlab = "x", ylab = "y")
par(mfrow = c(1, 1))    # back to one plot per window

These lines print nothing to the console. They draw four scatterplots in the Plots pane, arranged in two rows and two columns by par(mfrow = c(2, 2)). Your plots should match Figure 6.1, without the red lines. Section 2 explains par(mfrow = ...) in more detail.

R Reflect on what you just ran

Use the questions below to interpret the output you produced. Look at your console output and plots before answering.

1. Using the console output, list the statistics that the four sets share. Then state, in one sentence, what an analyst who saw only those statistics would conclude.

Model answerAll four sets have a mean of 9.0 for x and 7.5 for y, standard deviations of 3.32 for x and 2.03 for y, a correlation of about 0.82 (0.816 in sets 1 to 3 and 0.817 in set 4), and the same fitted line, with an intercept of 3.00 and a slope of 0.500. An analyst who saw only these numbers would conclude that the four sets show the same moderately strong, positive, straight-line relationship between x and y.

2. Describe each of your four plots in one sentence. For which set or sets is the fitted line y = 3.00 + 0.500x a fair summary?

Model answerSet 1 is a cloud of points scattered evenly around a rising straight line. Set 2 is a smooth arch that rises and then falls, so the relationship is curved. Set 3 is an almost perfect straight line with one point far above it. Set 4 has ten points stacked at x = 8 and one point far to the right at x = 19. The line is a fair summary only for set 1. In set 2 the shape is wrong, in set 3 one outlier distorts the slope, and in set 4 a single point creates the whole relationship.

3. The working group asks why the briefing should contain plots when a table of means and correlations would be shorter. Use this activity to write a two-sentence answer.

Model answerAnscombe's quartet shows that four datasets can share the same means, standard deviations, correlation and fitted line while having completely different patterns, including a curve, an outlier and a relationship created by a single point. Plots let the group see which pattern the survey data actually follow before anyone relies on a summary number or commissions a model built on it.
Saved.

Anscombe's quartet shows that almost identical summary statistics can describe very different data, so the analyst needs a plot that suits each question. The next part sets out how the number and the type of the variables point to a display.

Matching the Display to the Question

Choosing a display begins with two questions about the request: how many variables it involves, and what type each variable is. The variable types are those from Lesson 2. A numeric variable records an amount and is continuous (any value in a range, such as age) or discrete (whole numbers, such as the number of people in a household). A categorical variable places each person into a group and is binary, ordinal (ordered groups, such as self-rated mental health from poor to excellent) or nominal (unordered groups, such as province). Scale scores such as the loneliness score, which takes the whole-number values 3 to 9, are numeric for plotting purposes, although their small number of distinct values affects how they look, as later sections show.

Box 6.2 recalls the guide to the basic plots from Lesson 2 and poses a retrieval question about comparing loneliness across age groups.

Box 6.2: Recall: Lesson 2, Section 3 (Visualization Techniques)

Lesson 2, Section 3 matched the basic plots to the question and the variable type. A histogram shows the distribution of one numeric variable, a bar chart shows the frequencies of a categorical variable, boxplots compare a numeric variable across groups, and a scatterplot shows the relationship between two numeric variables. This section extends that guide to pairs of categorical variables and to many variables at once.

Retrieval question. Which display would you draw to see whether loneliness scores differ between four age groups, and why?

AnswerBoxplots side by side, one for each age group. Loneliness is numeric and age group is categorical, so boxplots compare the centre and the spread of the scores across the groups, and showing the individual points as well reveals how many people are in each group.

Figure 6.2 extends that guide into a decision tree, which starts from the number of variables in the question and then branches on their types.

How many variables does the question involve? One variable Two variables Many variables NumericHistogram or boxplot CategoricalBar chart of counts Numeric and numericScatterplot Categorical and numericBoxplots with points Two categoricalBar chart of percentages Several numericScatterplot matrix or heatmap Any display, by groupFacets (small multiples)
Figure 6.2. The number of variables and their types point to a display. Sections 2 and 3 draw each of these displays in R, and Section 4 covers the displays for many variables.

A histogram divides the range of a numeric variable into intervals (bins) and draws a bar for the number of people in each one, so the bars touch and their order follows the number line. A bar chart draws one bar for each category of a categorical variable, with gaps between the bars, and the categories can be placed in any sensible order, such as the order of an ordinal scale or from most to least common. Students often confuse the two because both use bars. The test is the variable on the horizontal axis: a histogram has a numeric axis, and a bar chart has categories.

For two variables, the display follows the pair of types. A scatterplot places one point per person for two numeric variables. Boxplots side by side compare a numeric variable across the groups of a categorical variable, and adding the individual points shows how many people are in each group and how their values are spread. For two categorical variables, a bar chart of percentages (for example, the percentage of each gender in each mental health category) compares the groups fairly when the groups differ in size. For many variables, a scatterplot matrix or a correlation heatmap gives an overview of every pair, and facets repeat one display for each group in a grid of small panels, an idea that Tufte (1983) called small multiples.

Interactive 6.1 applies the decision tree in Figure 6.2 to six requests from the working group. Each request is answered by choosing a display, and the feedback under the options explains why each choice suits the request or fails it.

📊 Interactive 6.1: Try it: choose a display for the briefing

Each request below comes from the working group. Choose the display you would draw first. Feedback appears under the options, and nothing is scored.

1. How are loneliness scores (a numeric score from 3 to 9) distributed across the 3,083 people in the briefing file?

2. How many people rated their mental health as poor, fair, good, very good or excellent?

3. Does loneliness (numeric) differ between the four age groups (categorical)?

4. Is higher social support (a numeric score from 1 to 7) associated with lower loneliness?

5. Do women, men and non-binary respondents differ in how they rate their mental health (both categorical)?

6. Which of six numeric briefing variables (loneliness, support, depression, anxiety, life satisfaction and age) go together?

Choosing the right type of display is the first step. A display of the right type can still mislead when it is drawn poorly, and the next part examines four common examples.

A Gallery of Poor Displays

Research on graphical perception has found that people judge positions along a common scale more accurately than lengths, angles or areas (Cleveland & McGill, 1984). A good display therefore encodes the comparison that matters as a position or a length on a shared axis, and it shows the data themselves wherever possible. The four displays below break these principles in ways that are common in reports and presentations. Each tab shows the poor version beside a better one, drawn from the briefing file.

The four tabs hold Figure 6.3 (a truncated axis), Figure 6.4 (means without spread), Figure 6.5 (a pie chart with many slices) and Figure 6.6 (overplotting), and the text under each figure states what the poor version conceals.

Left: bar chart of mean loneliness by gender with the axis starting at 5.3, making the bar for men look a fraction of the others. Right: the same bars with the axis starting at zero, where they look similar.
Figure 6.3. Mean loneliness score by gender (women 5.74, men 5.36, non-binary respondents 5.93; n = 3,083). Left: the axis starts at 5.3. Right: the axis starts at zero.

What it conceals. A bar encodes a value by its length, so a bar chart whose axis starts above zero misrepresents every comparison. On the left, the bar for men looks about one tenth of the height of the bar for non-binary respondents, although the means are 5.36 and 5.93, a ratio of about nine to ten. The right panel starts the axis at zero, so the lengths are proportional to the values. A related distortion is the dual axis, in which two series share a plot with two different vertical scales; the apparent crossing points and relative heights depend entirely on how the two scales were chosen, and two separate panels are usually clearer.

Note on axis ranges. The zero rule applies to bars. A dot or line display may use an axis that covers the possible range of the scale (3 to 9 for this loneliness score), provided the range is labelled.

Left: four bars showing mean loneliness by age group, all between 5.4 and 6.0. Right: boxplots for the same groups with every person drawn as a faint dot, showing scores from 3 to 9 in every group.
Figure 6.4. Loneliness score by age group (n = 3,083). Left: bars of the group means. Right: boxplots with each person shown as a jittered, transparent dot.

What it conceals. The bars show four means between 5.39 and 6.01 and nothing else. They hide the spread within each group (scores from 3 to 9 in every group), the shape of each distribution, and the number of people in each group, which ranges from 366 people aged 65 and older to 1,079 people aged 16 to 29. Weissgerber and colleagues (2015) reviewed 703 research articles in leading physiology journals, found that bar graphs were the most common way of presenting continuous data (85.6% of the articles included at least one), and recommended displays that show the distribution of the data, particularly dot plots of the individual values when samples are small. The right panel shows that the differences between age groups are small compared with the variation within each group.

Left: a pie chart of province of residence with fourteen slices, many of them thin. Right: the same counts as a horizontal bar chart sorted from British Columbia, the largest, to Yukon, the smallest.
Figure 6.5. Province or territory of residence in the 2021 wave (n = 4,045, including 50 people who did not answer). Left: a pie chart with 14 slices. Right: a bar chart sorted by count.

What it conceals. A pie chart asks the reader to compare angles and areas, which people judge less accurately than positions along a common scale. With fourteen slices, several of them thin, it is hard to tell whether Manitoba or New Brunswick has more respondents, or to read the smaller territories at all. The sorted bar chart places every category on a common scale, so the order and the size of the differences are clear. A pie chart can be acceptable for two or three categories that make up a whole; for more categories, a bar chart is clearer.

Left: a scatterplot of social support against loneliness in which the points form seven solid horizontal rows. Right: the same data with jittered, transparent points and a red fitted line sloping downward.
Figure 6.6. Loneliness score against social support (n = 3,083). Left: one solid point per person. Right: points jittered vertically by up to 0.2, drawn with transparency, and a least-squares line (red) with its 95% confidence band.

What it conceals. The loneliness score takes only seven values, so 3,083 points fall on seven horizontal lines and lie on top of one another. The left panel cannot show where most people are, and a row with five people looks the same as a row with five hundred. The right panel adds a small random vertical shift to each point (jitter) and draws the points with transparency, so dense areas look darker. The fitted line makes the downward trend easy to see: people with more social support tend to report less loneliness.

Several of these displays compare groups, such as genders or age groups. Box 6.3 sets out when groups are better shown by colour within one panel and when they are better shown in separate panels.

Box 6.3: Colour or facets?

When a display must show groups, the analyst can give each group its own colour within one panel, or draw one panel per group (facets). Colour works well for two or three groups whose values overlap little, and it lets the reader compare the groups directly. Facets work better when there are more groups, when the points overlap heavily, or when the groups differ greatly in size, because each group gets its own space on a shared scale. Section 3 draws both versions in ggplot2.

This section has explained why the analyst looks at the data before fitting a model, how the question and the variable types point to a display, and what four common poor displays conceal. The knowledge check and the reflection that follow review these ideas, and Section 2 draws the displays for the briefing file in base R.

Knowledge check: this section

1. What is the main purpose of exploratory data analysis?

Exploratory analysis asks what the data look like, including distributions, unusual values and relationships, so that the analyst can choose sensible summaries and models. Testing a prespecified hypothesis is confirmatory analysis, and outliers are investigated, which may or may not lead to removing them.

2. The four datasets in Anscombe's quartet have nearly identical means, standard deviations, correlations and fitted lines. What does the quartet demonstrate?

The four sets include a linear cloud, a curve, a line with one outlier and a relationship created by one point, yet they share the same summary statistics. Only plots reveal the differences.

3. A researcher wants to show the distribution of age, measured in years, among survey respondents. Which display is the natural first choice?

Age is a single numeric variable, so a histogram, which counts people in intervals along a number line, shows its shape. A pie chart and a bar chart are meant for categories.

4. Which feature distinguishes a histogram from a bar chart?

The horizontal axis of a histogram is numeric, divided into touching bins, while a bar chart shows separate categories with gaps between the bars. Both can show counts or percentages.

5. A bar chart of mean scores in four groups starts its vertical axis at 5.3. What is the main problem?

A bar encodes a value by its length. Starting the axis above zero cuts off part of every bar, so small differences look large. The axis of a bar chart should start at zero.

✎ Reflection

Exploratory data analysis looks at the data before modelling. This section described four features that plots reveal in a single variable (its shape, unusual or impossible values, heaping on particular values, and missing values) and showed, with Anscombe's quartet, that four datasets with identical means, standard deviations, correlations and fitted lines can follow completely different patterns. It also listed displays for different questions: a histogram for one numeric variable, a bar chart for one categorical variable, a scatterplot for two numeric variables, boxplots with points for a numeric variable across groups, and a bar chart of percentages for two categorical variables. Suppose a colleague has fitted a linear regression of weekly hours of physical activity (a numeric variable) on age group (four categories) in a community survey, and plans to report the model without any plots. Name two plots you would ask the colleague to draw first, explain what each plot could reveal, and describe how one possible finding would change the analysis.

Model answerI would first ask for a histogram of weekly hours of physical activity, because the outcome is a single numeric variable. It could show that the distribution is strongly skewed, with many people reporting zero hours and a few reporting very high values, and it could reveal impossible values such as 150 hours a week or heaping on round numbers such as 5 or 10 hours. I would then ask for boxplots of physical activity for each of the four age groups, with the individual points shown, because the question concerns a numeric variable across groups. These would show whether the groups differ in spread as well as in centre, how many people are in each group, and whether a few extreme values drive any difference. If the histogram showed a large spike at zero and a long right tail, a linear regression on the raw hours would be a poor fit, because its residuals would be skewed and their spread uneven. I would then suggest summarizing with medians, checking and recoding impossible values, and considering a model suited to the outcome, such as a count model for whole hours or a logistic model for meeting an activity guideline.
✓ Reflection saved!
● Complete the quiz and reflection to continue.
Section 2 of 4

Quick Looks in Base R

⏱ Estimated time: 50 minutes
Lesson 6 · Section 2

Quick Looks in Base R

Base R draws a useful diagnostic plot with one line of code.

Base R plotting

Four familiar functions and one new one

FunctionWhat it drawsVariables
hist(x)HistogramOne numeric
barplot(table(x))Bar chart of countsOne categorical
boxplot(y ~ group)Boxplots side by sideNumeric across groups
plot(x, y)ScatterplotTwo numeric
pairs(data)Scatterplot matrixSeveral numeric

Recall: the first four were used in the Lesson 2, Section 3 R activity, and pairs() is new.

Preparing the data

One briefing file for every plot

  • The code loads the CSCS file and keeps the 2021 wave, which has 4,045 people.
  • Short names replace the long survey names, for example loneliness and support.
  • Bracket recoding sets "Presented but no response" to missing and builds four age groups.
  • na.omit() keeps the 3,083 people with complete data on all nine briefing variables.
One variable

A histogram and a bar chart

Left: histogram of loneliness scores from 3 to 9, highest at 5 and 6, with a small rise at 9. Right: bar chart of self-rated mental health, highest for very good.
Loneliness score (left) and self-rated mental health (right), n = 3,083.
Finding a problem

A spike at the largest value

Panel A: histogram of household size falling from 1 person toward 19, with a tall bar at 20. Panel B: histogram of hours worked per week with tall bars at 0, 35 and 40 and a thin tail to 140.
Household size (A) has 116 people at its top value of 20. Hours worked per week (B) piles up on 40 and stretches to 140.
Responding to the problem

Check, decide and report

Check

The codebook and the source questions show that 20 means 20 or more.

Decide

A count analysis sets 20 to missing, and a grouped variable absorbs the problem.

Report

The methods section states the problem, the decision and the number of people affected.

Two variables

Boxplots and a scatterplot

Panel A: boxplots of loneliness by age group, medians 5, 5, 6 and 6. Panel B: scatterplot of support and loneliness forming seven solid stripes. Panel C: the same scatterplot with jittered, transparent points showing a downward trend.
Loneliness by age group (A) and against social support, before (B) and after (C) jitter and transparency.
Layouts and saving

Many plots at once, and plots saved to files

CodeWhat it does
pairs(brief[, c(...)])It draws a scatterplot for every pair of the chosen variables.
par(mfrow = c(1, 2))It places the next plots in one row of two panels.
png("file.png", width, height, res)It sends the next plots to an image file.
dev.off()It closes the file, which saves it.
Carry forward

What to take into the next section

  • Base R draws a quick diagnostic plot for each kind of question.
  • Quick plots reveal problems such as spikes at a top value and heaping on round numbers.
  • Jitter and transparency reveal patterns when many points share the same values.

Introduction and Overview

Section 1 explained why the analyst looks before modelling and how the question and the variable types point to a display. This section draws those displays for the survey data with base R, the plotting functions that are part of R itself. Base R plots need no extra packages and usually one line of code, which makes them the analyst's quick diagnostic tool: they are drawn, read and set aside, and only the useful ones are later rebuilt as polished figures (Section 3). The section also prepares the briefing file that every later plot uses, and it uses a quick plot to find a data problem that the cleaning lesson left behind.

Learning Objectives

  • Prepare an analysis file in base R with bracket recoding and na.omit().
  • Draw a histogram, a bar chart, side-by-side boxplots, a scatterplot and a scatterplot matrix in base R, with titles and axis labels.
  • Read a histogram for shape, unusual and impossible values, and heaping.
  • Use jitter and transparency to show data that share a small number of values.
  • Arrange several base R plots in one window with par(mfrow = ...) and save a plot to a file with png().

The section begins with the labelling arguments that every base R plot accepts. It then prepares the briefing file, reads plots of one variable and of two variables, and ends with layouts of several plots and plots saved to files.

Base R as the Analyst's Quick Diagnostic Tool

Base R draws a quick diagnostic plot with one line of code. Every plotting function accepts the same labelling arguments: main for the title, xlab and ylab for the axis labels. Every plot in the briefing should carry these labels from the start, because a plot without labels is easily misread, even by the person who drew it.

Box 6.4 recalls the four standard plotting functions from the Lesson 2 activity and names the briefing variables to which this section applies them.

Box 6.4: Recall: Lesson 2, Section 3 (R Activity: descriptives and standard plots)

The R activity in Lesson 2, Section 3 drew four standard plots with base R: hist(x) for a histogram of one numeric variable, barplot(table(x)) for a bar chart of one categorical variable, boxplot(y ~ group, data = ...) for boxplots of a numeric variable across groups, and plot(x, y) for a scatterplot of two numeric variables. This section applies the same four functions to the briefing file: the loneliness score, self-rated mental health, loneliness by age group, and loneliness against social support.

Retrieval question. Why is barplot() given table(x) and not the raw variable?

Answerbarplot() draws one bar for each number it is given, so it needs counts. table(x) counts the people in each category, and barplot() then draws one bar per count.

This section adds three tools that Lessons 1 and 2 did not use. The function pairs(data) draws a scatterplot matrix, one small scatterplot for every pair of several numeric variables, such as loneliness, support, depression and age. The setting par(mfrow = ...) places several plots in one window, and png() sends plots to a file; both are described at the end of this section.

Before any of these functions can be applied, the survey data must be prepared as a single analysis file, which the next part describes.

Preparing the Briefing File

All the plots in this lesson use one analysis file, called brief. The code that builds it follows the pattern of the course's data cleaning and regression walkthroughs. It loads the CSCS file from the course's GitHub repository, keeps the 2021 wave (which gives one row per person), and gives short names to six variables. It then uses bracket recoding, in which a condition inside square brackets selects the rows to change, to set the answer "Presented but no response" to missing (NA) for gender and self-rated mental health, and to build four age groups from age in years. The function factor() with a levels argument fixes the order in which categories appear in tables and plots. Finally, na.omit() keeps the people who have no missing value on any of the nine briefing variables.

Table 6.3 lists the nine briefing variables, the survey measure behind each one, its possible values and the way it is treated for plotting.

Table 6.3. The nine variables in the briefing file.

Short nameMeasure in the surveyValuesType for plotting
lonelinessThree-item UCLA Loneliness Scale (Hughes et al., 2004)3 to 9, higher is lonelierNumeric (seven whole values)
supportMultidimensional Scale of Perceived Social Support, mean of 12 items (Zimet et al., 1988)1 to 7, higher is more supportNumeric
depressionPatient Health Questionnaire-2 (Kroenke et al., 2003)0 to 6, higher is more symptomsNumeric (seven whole values)
anxietyGeneralized Anxiety Disorder-2 (Kroenke et al., 2007)0 to 6, higher is more symptomsNumeric (seven whole values)
life_satLife satisfaction, a single item1 to 10, higher is more satisfiedNumeric (ten whole values)
ageAge in years16 to 100Numeric
age_groupFour groups built from age16 to 29, 30 to 44, 45 to 64, 65 and olderCategorical (ordinal)
genderGenderWoman, Man, Non-binaryCategorical (nominal)
mental_healthSelf-rated mental healthPoor, Fair, Good, Very good, ExcellentCategorical (ordinal)

Two cautions apply to every result drawn from this file. Box 6.5 describes how the CSCS recruited its respondents and what its sample can support, and Box 6.6 describes the people whom the complete-case file leaves out.

Box 6.5: Background: What the CSCS sample can support

The Canadian Social Connection Survey (CSCS) is a national online survey of social connection, loneliness and health among people living in Canada. Respondents to the 2021 wave completed an anonymous online questionnaire between April and June 2021. They were recruited online and were not drawn at random from a sampling frame, so the chance that any person in Canada took part is unknown, and the public file contains no survey weights that could correct for this. Two consequences follow for the briefing. Percentages and means describe these respondents and should not be read as the prevalence of loneliness in the Canadian population, because people who answer an online survey about social connection may differ from those who do not. Associations between variables within the sample, such as the link between social support and loneliness, are the more useful product, although they also assume that taking part did not depend jointly on both variables. For optional reading, HSCI 207 Lesson 11, Section 2 makes the same point for CSCS respondents aged 65 and older.

Box 6.6: Who is left out?

The 2021 wave has 4,045 respondents and the briefing file has 3,083, so na.omit() removed 962 people (24%) who skipped at least one of the nine questions. Using a single complete-case file means that every panel of the briefing describes the same people, which makes the panels comparable. The cost is that the briefing describes only people who answered every question, and Lesson 2 described how such people can differ from those who skipped questions. The briefing should state the number of people included and the reason others were excluded.

Activity 6.2 carries out these steps in R. It loads the survey, builds the briefing file, draws a histogram of the loneliness score and a bar chart of self-rated mental health, and summarizes the household size variable that a later part of this section examines.

R Activity 6.2: preparing the briefing file and looking at one variable

This activity loads the survey, prepares the briefing file, draws a histogram and a bar chart, and checks the household size variable. The data are read directly from the course's GitHub repository, so an internet connection is needed the first time the code runs. The second block of code is the data preparation that every later activity in this lesson starts from.

github_url <- "https://raw.githubusercontent.com/jorgeandr3s/heal/main/cscs/public_data/CSCS2025_full_cleaned_deidentified_data_and_metadata.RData"
load(url(github_url))
data <- data[data$SURVEY_collection_year == 2021, ]   # the 2021 wave: one row per person
nrow(data)                                             # how many people?
Console output
[1] 4045

The file adds three objects to the environment, and the one called data holds the survey responses. The square brackets keep the rows from the 2021 wave and every column. The 2021 wave holds 4,045 people.

# Short names for the briefing variables
data$loneliness <- data$LONELY_ucla_loneliness_scale_score   # 3 to 9, higher = lonelier
data$support    <- data$PSYCH_zimet_multidimensional_social_support_scale_score   # 1 to 7
data$depression <- data$WELLNESS_phq_score                   # 0 to 6, higher = more symptoms
data$anxiety    <- data$WELLNESS_gad_score                   # 0 to 6, higher = more symptoms
data$life_sat   <- data$WELLNESS_life_satisfaction_num       # 1 to 10, higher = more satisfied
data$age        <- data$DEMO_age                             # years

# Gender: the answer "Presented but no response" becomes missing (NA)
data$gender <- as.character(data$DEMO_gender)
data$gender[data$gender == "Presented but no response"] <- NA
data$gender <- factor(data$gender, levels = c("Woman", "Man", "Non-binary"))

# Self-rated mental health, in order from Poor to Excellent
data$mental_health <- as.character(data$WELLNESS_self_rated_mental_health)
data$mental_health[data$mental_health == "Presented but no response"] <- NA
data$mental_health <- factor(data$mental_health,
  levels = c("Poor", "Fair", "Good", "Very good", "Excellent"))

# Four age groups
data$age_group[data$age < 30] <- "16 to 29"
data$age_group[data$age >= 30 & data$age < 45] <- "30 to 44"
data$age_group[data$age >= 45 & data$age < 65] <- "45 to 64"
data$age_group[data$age >= 65] <- "65 and older"
data$age_group <- factor(data$age_group,
  levels = c("16 to 29", "30 to 44", "45 to 64", "65 and older"))

# The briefing file: the briefing variables for people with no missing values
brief <- na.omit(data[, c("loneliness", "support", "depression", "anxiety",
                          "life_sat", "age", "age_group", "gender", "mental_health")])
nrow(brief)                                            # people in the briefing file
Console output
[1] 3083

Each bracket recode changes only the rows that meet the condition inside the brackets. The first age-group line, for example, writes "16 to 29" into age_group for every person whose age is below 30. na.omit() then keeps the 3,083 people with no missing value on the nine briefing variables.

summary(brief$loneliness)              # a numeric summary first
hist(brief$loneliness,
     breaks = seq(2.5, 9.5, by = 1),       # one bar for each score from 3 to 9
     main = "Loneliness (UCLA 3-item scale)",
     xlab = "Loneliness score (3 to 9)", ylab = "Number of people")
Console output
Min. 1st Qu. Median Mean 3rd Qu. Max. 3.00 4.00 5.00 5.57 7.00 9.00

The scores run from 3 to 9, the median is 5 and the mean is 5.57. The histogram appears in the Plots pane, with one bar for each whole score.

table(brief$mental_health)             # the counts that the bars will show
barplot(table(brief$mental_health),
        main = "Self-rated mental health",
        xlab = "Rating", ylab = "Number of people")
Console output
Poor Fair Good Very good Excellent 230 467 928 1024 434

The table gives the counts that the bars show, in the order set by factor(). Very good (1,024) is the most common rating and Poor (230) the least common.

summary(data$GEO_housing_household_size)   # people each respondent lives with
hist(data$GEO_housing_household_size,
     breaks = seq(-0.5, 20.5, by = 1),             # one bar for each value from 0 to 20
     main = "Household size",
     xlab = "Number of people the respondent lives with", ylab = "Number of people")
sum(data$GEO_housing_household_size == 20, na.rm = TRUE)   # how many at the top value?
Console output
Min. 1st Qu. Median Mean 3rd Qu. Max. NA's 0.000 1.000 3.000 4.463 6.000 20.000 618 [1] 116

The summary uses the full 2021 wave, in which 618 people did not answer the household questions. The maximum is 20, and sum(... == 20, na.rm = TRUE) counts the people at that value: 116. The argument na.rm = TRUE tells sum() to skip missing values, which would otherwise make the result NA.

R Reflect on what you just ran

Use the questions below to interpret the output you produced. Look at your console output and plots before answering.

1. Describe the shape of the loneliness histogram in two or three sentences, using the console output and the plot. Which score is most common, and what happens at the top of the scale?

Model answerThe most common score is 6 (755 people), closely followed by 5 (689), and the bars fall away on both sides, from 433 people at 3 and to 179 at 8. The bar at 9, the highest possible score, is taller than the bar at 8 (270 people), which is a pile-up at the ceiling of the scale. The mean (5.57) is slightly above the median (5), consistent with a slightly longer upper tail.

2. From the table of self-rated mental health, what percentage of the briefing file rated their mental health as poor or fair? Show the calculation.

Model answerPoor (230) plus Fair (467) gives 697 people. Dividing by the 3,083 people in the briefing file gives 697 / 3,083 = 0.226, or 22.6%, so a little more than one person in five rated their mental health as poor or fair.

3. The histogram of household size has a spike of 116 people at 20, its largest value. Explain why this spike is a problem, and write the bracket recode that would create a new variable household with the value 20 set to missing.

Model answerHousehold size is the sum of eight questions about who the respondent lives with, and totals above 20 were recorded as 20, so the value 20 means "20 or more" and, for most of these people, comes from implausible answers (105 of the 116 had totals above 20, up to 180). Treated as a count, it would distort a mean or a model. The recode is data$household <- data$GEO_housing_household_size followed by data$household[data$household == 20] <- NA. The decision and the number of people affected (116) should be reported in the methods.
Saved.

With the briefing file in place, the next part turns to plots of one variable.

Looking at One Variable

The first look at any variable is a plot of its distribution. Figure 6.7 shows the histogram of the loneliness score and the bar chart of self-rated mental health drawn in Activity 6.2, and the two paragraphs that follow read each plot in turn.

Left: base R histogram of loneliness scores, with bars for scores 3 to 9; the tallest bars are at 5 and 6, and the bar at 9 is taller than the bar at 8. Right: base R bar chart of self-rated mental health from Poor to Excellent; Very good is tallest and Poor is shortest.
Figure 6.7. The two plots drawn in Activity 6.2: a histogram of the loneliness score and a bar chart of self-rated mental health (n = 3,083).

Reading the histogram. The argument breaks = seq(2.5, 9.5, by = 1) sets the edges of the bins at 2.5, 3.5, 4.5 and so on, so each bar holds exactly one whole score. Without it, R chooses its own bins, which can merge two scores into one bar or leave empty gaps, and the picture of a variable with only seven values can change a good deal. The loneliness histogram rises from 433 people at the lowest score of 3 to a peak of 755 at a score of 6, and then falls to 179 at a score of 8. The bar at 9, with 270 people (8.8%), is taller than the bar at 8. A rise at the top of a scale is called a ceiling pile-up: a group of people chose the most extreme answer to all three items, and the scale cannot record anything lonelier. The mean (5.57) is a little above the median (5), which reflects the longer upper tail.

Reading the bar chart. The bars follow the order set by factor(), from Poor to Excellent. Very good is the most common rating (1,024 people, 33.2%), followed by Good (928). Together, Poor (230) and Fair (467) make up 697 people, or 22.6% of the briefing file. If the levels had not been set, R would have placed the categories in alphabetical order (Excellent, Fair, Good, Poor, Very good), which hides the shape of an ordinal scale.

Box 6.7 gathers the questions behind these two readings into a checklist that applies to any plot of one variable.

Box 6.7: A checklist for one-variable plots

For each variable, the analyst asks whether the distribution is symmetric or skewed, whether it has one peak or more, whether values pile up at the lowest or highest possible value (a floor or a ceiling), whether answers heap on round numbers, whether any value lies outside the possible range, and how many values are missing. Each answer is recorded in the analysis notes, because each one can change a later decision about summaries, recoding or models.

The checklist asks about values outside the possible range and about answers that pile up on particular values. The next part applies it to a variable whose numeric summary looked unremarkable.

Finding a Problem That Cleaning Left Behind

The household size variable, GEO_housing_household_size, records the number of people each respondent lives with, not counting the respondent, so 0 means living alone (503 people). The course's cleaning walkthrough used it to build a living-alone indicator, and its summary in Activity 6.2 looks unremarkable: a median of 3, a mean of 4.46 and a maximum of 20. The histogram tells a different story. Panel A of Figure 6.8 shows that histogram, and panel B shows a second variable from the 2021 wave, the hours worked per week.

Panel A: histogram of household size; bars fall from 723 at 1 person to about 16 at 19, then a bar of 116 at 20. Panel B: histogram of hours worked per week in 2-hour bins; tall bars near 0, 35 and 40, smaller peaks at 20 and 30, and a long thin tail to 140 hours.
Figure 6.8. Panel A: number of people the respondent lives with (n = 3,427; 618 missing). Panel B: hours worked per week during the COVID-19 pandemic (n = 2,978; 1,067 missing), in two-hour bins. Both variables are from the 2021 wave before any recoding.

In panel A, the bars fall smoothly from one other person toward larger households, and then a bar of 116 people appears at exactly 20, the largest value in the data. The real distribution of household sizes does not rise again at its maximum, so the spike signals that something about the value 20 differs from the other values. Tracing the variable back to the survey explains it. Household size is the sum of eight questions that ask how many partners, children, grandchildren, parents, in-laws, siblings, roommates and other people the respondent lives with, and totals above 20 were recorded as 20. For 105 of the 116 people at 20, the eight answers added up to more than 20, and one total reached 180. Some of these answers are implausible, such as five grandchildren, five in-laws and seven siblings in one household, and they may reflect careless or mistaken responses. The summary statistics did not reveal this problem, and the histogram made it obvious.

The analyst's response has three steps. The first is to check the codebook and the source questions, which is how the meaning of 20 was established. The second is to decide how the variable will be used. Used as a count, the value 20 cannot be trusted and can be set to missing with a bracket recode (Activity 6.2, question 3). Used as a grouped variable, such as living alone versus living with others, or 0, 1, 2 to 4 and 5 or more, the problem is absorbed by the top category and matters much less. The third step is to report the decision in the methods section, with the number of people affected, and, where the decision could change a result, to repeat the analysis both ways as a sensitivity analysis.

Panel B shows two further problems that the same plot can reveal. Answers to the question on hours worked per week heap on round numbers: 340 of the 2,978 people who answered reported exactly 40 hours, compared with one person at 39 and one at 41, and smaller peaks appear at 20, 30 and 35. Heaping of this kind, which Lesson 2 called digit preference, usually reflects rounding and standard working weeks, and it means that small differences in reported hours carry little information. The long right tail reaches 140 hours, and 13 people reported more than 112 hours a week, which would mean working more than 16 hours a day, every day. These values are implausible and would be checked and, in most analyses, set to missing.

Household size and working hours show that a histogram can reveal problems that summary statistics hide. Activity 6.3 draws boxplots and scatterplots of two variables and a scatterplot matrix, and it saves a two-panel figure to a file.

R Activity 6.3: two variables, layouts and saving a plot

This activity starts from the briefing file made in Activity 6.2. If you are starting a new R session, run the loading and preparation code from Activity 6.2 first (it is also at the top of the answer key). The code draws boxplots, two scatterplots and a scatterplot matrix, and then saves a two-panel figure to a file.

boxplot(loneliness ~ age_group, data = brief,
        main = "Loneliness by age group",
        xlab = "Age group", ylab = "Loneliness score (3 to 9)")
tapply(brief$loneliness, brief$age_group, median)   # the thick line in each box
Console output
16 to 29 30 to 44 45 to 64 65 and older 5 5 6 6

The formula loneliness ~ age_group draws one box for each age group. tapply() applies median() to the loneliness scores within each group and returns the values of the thick lines: 5, 5, 6 and 6.

plot(brief$support, brief$loneliness,              # every person as one point
     xlab = "Social support (1 to 7)", ylab = "Loneliness score (3 to 9)")
set.seed(2021)                                       # the same jitter every time
plot(jitter(brief$support), jitter(brief$loneliness),   # nudge each point a little
     pch = 16, col = rgb(0, 0, 0, 0.15),             # small grey see-through dots
     xlab = "Social support (1 to 7)", ylab = "Loneliness score (3 to 9)")

These lines draw two scatterplots and print nothing. The first shows seven solid stripes. The second uses jitter(), small filled dots (pch = 16) and a see-through grey (rgb(0, 0, 0, 0.15), where the last number is the opacity), so the dense areas look darker. set.seed(2021) makes the random jitter the same each time.

pairs(brief[, c("loneliness", "support", "depression", "age")],
      pch = 16, col = rgb(0, 0, 0, 0.1))   # every pair of variables in one grid

The scatterplot matrix has one panel for every pair of the four variables, with the variable names on the diagonal.

png("briefing_base_r.png", width = 1600, height = 800, res = 200)   # open a file
par(mfrow = c(1, 2))                       # one row, two plots
hist(brief$loneliness, breaks = seq(2.5, 9.5, by = 1),
     main = "A. Loneliness", xlab = "Loneliness score", ylab = "Number of people")
boxplot(loneliness ~ age_group, data = brief,
        main = "B. Loneliness by age group", xlab = "Age group", ylab = "Loneliness score")
par(mfrow = c(1, 1))
dev.off()                                  # close the file, which saves it
Console output
null device 1

Nothing appears in the Plots pane, because png() sent both plots to the file briefing_base_r.png in the working directory (getwd() shows where that is). The message after dev.off() names the graphics device R returned to. It may read null device 1 or RStudioGD 2, depending on whether other plots are open, and it can be ignored.

R Reflect on what you just ran

Use the questions below to interpret the output you produced. Look at your console output and plots before answering.

1. Using the boxplots and the medians printed by tapply(), compare loneliness across the four age groups. Mention the medians and the width of the boxes.

Model answerThe two older groups (45 to 64 and 65 and older) have a median loneliness score of 6, and the two younger groups (16 to 29 and 30 to 44) have a median of 5. The box for people aged 45 to 64 is the widest (4 to 8), so scores vary most in that group, while the box for people aged 16 to 29 is narrow (5 to 6) because most of them scored 5 or 6. Every group has people across the full range of the scale, so the differences between groups are small compared with the variation within them.

2. Compare the two scatterplots of social support and loneliness. What does the jittered, transparent version show that the first version hides?

Model answerThe first scatterplot shows only seven solid horizontal stripes, because loneliness takes seven values and thousands of points sit on top of each other, so it is impossible to tell where most people are. The jittered, transparent version spreads the points a little and darkens areas where many overlap. It shows that most people have support scores between about 4 and 6 and loneliness scores of 5 or 6, and that loneliness tends to be lower among people with more social support.

3. After running the last block of code, no new plot appears in the Plots pane. Explain why, and say how you would find the saved file.

Model answerThe function png() opened an image file, so the histogram and the boxplots were drawn into briefing_base_r.png and not on screen. dev.off() closed the file, which saved it. The file is in the working directory, which getwd() prints and which the Files pane in RStudio shows; opening it there displays the two-panel figure.
Saved.

The next part turns from single variables to pairs of variables.

Looking at Two Variables

A plot of two variables shows how the variables go together. Figure 6.9 shows the boxplots of loneliness by age group and the two scatterplots of loneliness against social support from Activity 6.3, and the paragraphs below read each panel.

Panel A: boxplots of loneliness for four age groups; medians 5, 5, 6 and 6; the 16 to 29 box is narrow, from 5 to 6, with points drawn at 3, 8 and 9. Panel B: scatterplot of social support against loneliness with seven solid horizontal stripes. Panel C: the same data jittered and transparent, darker between support 4 and 6 and loneliness 5 and 6, with a visible downward trend.
Figure 6.9. Loneliness by age group (A), and loneliness against social support drawn with solid points (B) and with jittered, transparent points (C), n = 3,083.

Boxplots. A boxplot summarizes a numeric variable with five numbers. The thick line is the median. The box runs from the lower to the upper quartile, so it holds the middle half of the people, and its length is the interquartile range (IQR). The whiskers extend to the most extreme values that lie within 1.5 times the IQR of the box, and any value beyond the whiskers is drawn as a separate point. Table 6.4 gives the numbers behind panel A.

Table 6.4. The numbers behind the boxplots in panel A of Figure 6.9: the number of people, the whiskers, the quartiles and the median in each age group.

Age groupPeopleLower whiskerLower quartileMedianUpper quartileUpper whisker
16 to 291,07945567
30 to 441,06034569
45 to 6457834689
65 and older36634679

The two older groups have a median of 6 and the two younger groups a median of 5, and the box for people aged 45 to 64 is the widest, so loneliness varies most in that group. The box for people aged 16 to 29 is narrow, because most of them scored 5 or 6. Its whiskers therefore stop at 4 and 7, and the scores 3, 8 and 9 in this group are drawn as separate points (212 people in all). These points are flagged by the 1.5 × IQR rule, and they are ordinary answers on the scale. With a variable that takes only a few whole values, the points beyond the whiskers say more about the narrowness of the box than about unusual people.

Scatterplots and overplotting. Panel B draws one solid point for each of the 3,083 people. Because the loneliness score takes only seven values, the points fall on seven horizontal lines and cover one another, a problem called overplotting, and the plot cannot show where most people are. Panel C applies two remedies. The function jitter() adds a small random amount to each value so that points sharing a value spread out a little, and col = rgb(0, 0, 0, 0.15) draws each point in black at 15% opacity, so that areas where many points overlap look darker. The function set.seed() fixes the random numbers, so the jittered plot is the same each time the code runs. Jitter changes only the picture: the data used in any calculation are unchanged. Panel C shows that most people have support scores between 4 and 6 and loneliness scores of 5 or 6, and that people with more support tend to report less loneliness.

The scatterplot matrix. The function pairs() draws a small scatterplot for every pair of the chosen variables, with the variable names on the diagonal. Each panel is a quick look at one pair, and the panel above the diagonal shows the same pair as the panel below it with the axes swapped. With 3,083 people and scores that take few values, the panels are crowded, so the matrix is best used to spot strong patterns or unexpected shapes, after which a single scatterplot or a correlation heatmap (Section 4) gives a clearer view.

Figure 6.10 shows the scatterplot matrix that Activity 6.3 draws for loneliness, support, depression and age.

A four by four scatterplot matrix for loneliness, support, depression and age. Panels involving loneliness and depression show grids of points because both take whole values; panels with age show vertical spread across ages 16 to 100.
Figure 6.10. The scatterplot matrix drawn by pairs() for four briefing variables (n = 3,083). Each off-diagonal panel is a scatterplot of the two variables named in its row and column.

Boxplots, scatterplots and the scatterplot matrix complete the base R displays that this section draws. The last part of this section shows how to arrange several plots in one window and how to keep them in a file.

Several Plots in One Window, and Saving Plots

The setting par(mfrow = c(rows, columns)) divides the plotting window into a grid, and each new plot fills the next cell, row by row. par(mfrow = c(1, 2)) places two plots side by side, and par(mfrow = c(2, 2)), used in the Anscombe activity, places four in a square. The setting stays in force until it is changed, so the code resets it with par(mfrow = c(1, 1)) afterwards.

Plots in the RStudio Plots pane disappear when the session ends. To keep a plot, the analyst sends it to a file. The function png("name.png", width = 1600, height = 800, res = 200) opens an image file 1,600 pixels wide and 800 pixels high, at a resolution of 200 pixels per inch, which prints at 8 by 4 inches. Every plot drawn afterwards goes into the file and does not appear on screen. The function dev.off() closes the file, which saves it in the working directory. The functions pdf() and jpeg() work in the same way for other formats, and the Export button in the Plots pane saves a plot by hand. Saving with code is better practice, because the code records exactly how the figure was made and remakes it in one step if the data change.

The card below links to a narrated R walkthrough of the base R plots in this section and the ggplot2 plots in Section 3.

Learn to do this in R

Narrated R walkthrough: Data Visualization in Base R and ggplot2

The walkthrough runs the code from this section and the next one line at a time, draws each plot as the line runs, and explains every argument. It is a good place to start if the activities feel fast.

Open the walkthrough

This section has prepared the briefing file and used base R to read one variable, to find a data problem that cleaning left behind, to show two variables together and to save a figure. The knowledge check and the reflection that follow review these steps, and Section 3 rebuilds the most useful plots as polished figures with ggplot2.

Knowledge check: this section

1. Which base R code draws a bar chart of the number of people in each category of self-rated mental health?

table() counts the people in each category, and barplot() draws one bar per count. hist() needs a numeric variable.

2. A histogram of household size falls smoothly from small to large households and then shows a tall bar at the largest value, 20. What is the most likely explanation?

A real distribution does not rise again at its maximum. A spike at the top value usually means that larger values were recorded as the maximum (top-coding) or that errors collect there, which the codebook can confirm. In the CSCS data, totals above 20 were recorded as 20.

3. In a base R boxplot, what does the thick line inside each box show?

The thick line is the median. The box runs from the lower to the upper quartile, and the whiskers reach the most extreme values within 1.5 times the interquartile range of the box.

4. A scatterplot of two survey scores shows only a few solid horizontal stripes because most points sit on top of each other. Which change would help most?

Jitter spreads points that share a value, and transparency makes dense areas darker, so the reader can see where most people are. Larger solid points would make the overplotting worse.

5. What does par(mfrow = c(1, 2)) do in base R?

mfrow divides the plotting window into a grid of rows and columns, here one row and two columns, which the next plots fill in turn. Saving to a file uses png() and dev.off().

✎ Reflection

This section used base R to take quick looks at the data and showed how a histogram can reveal problems that numeric summaries hide. Two examples came from the 2021 Canadian Social Connection Survey. Household size (the number of people a respondent lives with) had a median of 3, a mean of 4.46 and a maximum of 20, and its histogram showed a spike of 116 people at exactly 20, because totals above 20 had been recorded as 20 and many of those totals came from implausible answers. Hours worked per week showed heaping: 340 of 2,978 people reported exactly 40 hours, compared with one person each at 39 and 41, and 13 people reported more than 112 hours a week. Suppose an analysis plans to use hours worked per week as a numeric predictor of loneliness. Describe the plot you would draw first, the two problems it would reveal, and how you would handle each problem, including what you would report in the methods section.

Model answerI would first draw a histogram of hours worked per week with narrow bins, for example two-hour bins with hist(data$WORK_hours_per_week, breaks = seq(0, 140, by = 2)), because it is a single numeric variable and narrow bins show heaping that wide bins would hide. The first problem it would reveal is implausible values: 13 people reported more than 112 hours a week, which would mean more than 16 hours of work every day. I would check these against the codebook and, unless there were a clear explanation, set them to missing with a bracket recode such as data$hours[data$hours > 112] <- NA. The second problem is heaping, with 340 people at exactly 40 hours and smaller peaks at 20, 30 and 35. Heaping cannot be corrected, but it means that small differences in hours carry little information, so I would consider grouping the variable (for example, none, 1 to 34, 35 to 44 and 45 or more hours) or at least interpreting its coefficient per ten hours. In the methods I would state that 13 values above 112 hours were set to missing, give the number of people included in the analysis, and note that reported hours heap on round numbers, and I would repeat the main model with and without the recoded values as a sensitivity analysis.
✓ Reflection saved!
● Complete the quiz and reflection to continue.
Section 3 of 4

The Grammar of Graphics with ggplot2

⏱ Estimated time: 50 minutes
Lesson 6 · Section 3

The Grammar of Graphics with ggplot2

Every ggplot2 plot is built from the same parts, added in layers.

The layered grammar

Six parts of every plot

Data

A data frame with one row per person.

Aesthetics

Mappings from variables to x, y, colour or fill.

Geoms

The marks: bars, boxes, points and lines.

Scales

Rules that turn values into positions and colours.

Facets

Panels that repeat the plot for each group.

Theme

Fonts, grid lines and background.

The template

Data, a mapping, a geom and labels

Line of codePart of the grammar
ggplot(brief, aes(x = loneliness)) +Data and the aesthetic mapping
geom_histogram(binwidth = 1, fill = "#0B7B6B", colour = "white") +The geom, with fixed (set) colours
labs(title = "...", x = "...", y = "...")Labels
ggplot2 histogram of loneliness scores with teal bars.
Bars

geom_bar() counts; geom_col() draws values

Left: geom_bar chart of counts for self-rated mental health. Right: geom_col chart of mean loneliness for four age groups.
Left: R counts the rows in each category. Right: R draws the four means it was given.
Groups

A boxplot with every person shown

Boxplots of loneliness by age group with jittered transparent points for all 3,083 people.
Two layers: geom_boxplot() and geom_jitter() (n = 3,083).
A smoother

Loneliness across age, with a loess curve

Scatterplot of age against loneliness with a blue loess curve that dips near age 30, rises to a second high point near 55 and falls after 70, with a wide grey band at the oldest ages.
The loess curve dips near age 30 and rises to a second high point in the fifties; the band widens where data are sparse.
Facets

Small multiples with facet_grid()

Twelve panels of support against loneliness, rows for woman, man and non-binary, columns for four age groups, each with a downward line; the non-binary panels have few points and wide bands.
Rows: gender. Columns: age group. The non-binary panels hold between 6 and 26 people each.
Colour

A colour-blind-safe palette

Boxplots of loneliness by age group with three boxes per group filled orange for women, sky blue for men and green for non-binary respondents.
Gender mapped to fill, with three Okabe-Ito colours and theme_minimal().
Saving

ggsave() writes a figure at a chosen size

ArgumentWhat it controls
"loneliness_age_gender.png"The file name; the extension sets the file type.
pThe plot to save; without it, the last plot shown is saved.
width = 7, height = 4.5The size in inches.
dpi = 300The resolution in dots per inch, the usual standard for print.
Carry forward

What to take into the next section

  • A ggplot2 plot combines data, aesthetic mappings and geoms, with scales, facets and a theme.
  • Layers can be stacked, such as a boxplot with jittered points or a scatterplot with a smoother.
  • Groups are shown with colour or facets, using a palette that people with colour-vision deficiency can read.

Introduction and Overview

Base R plots are quick to draw, and Section 2 used them to check the data. A figure for the working group's briefing, or for a manuscript, needs more care: consistent labels, readable colours, panels that share a scale, and a file of the right size and resolution. The ggplot2 package (Wickham, 2016) is the most widely used tool in R for such figures. It belongs to the tidyverse, a family of packages that Lesson 1 loaded when setting up R. The walkthroughs and this lesson keep data preparation in base R, exactly as in Section 2, and ggplot2 is the tidyverse package that this lesson uses for figures.

Learning Objectives

  • Describe the layered grammar of graphics in terms of data, aesthetics, geoms, scales, facets and themes.
  • Build a histogram, a bar chart, a boxplot with jittered points, and a scatterplot with a smoother in ggplot2.
  • Distinguish geom_bar(), which counts rows, from geom_col(), which draws supplied values.
  • Compare a display across groups with facet_wrap() and facet_grid(), and choose between colour and facets.
  • Apply a colour-blind-safe palette and a theme, and save a figure with ggsave().

The section first describes the grammar on which ggplot2 is built. It then draws displays of one variable, of a numeric variable across groups and of two numeric variables, compares groups with facets and with colour, and ends with saving a figure to a file.

The Layered Grammar of Graphics

Leland Wilkinson described a grammar of graphics in which every statistical graphic is assembled from a small set of independent parts, in the way that a sentence is assembled from parts of speech (Wilkinson, 2005). Hadley Wickham adapted the idea into a layered grammar, in which a plot is built by adding layers one at a time, and implemented it as ggplot2 (Wickham, 2010). The practical benefit is that one template covers almost every display: once the parts are familiar, a histogram, a boxplot and a faceted scatterplot differ only in a line or two of code.

Figure 6.11 lists the six parts of a ggplot2 plot, each with an example taken from the briefing.

Data brief: one row per person (3,083 rows, 9 variables) Aesthetics aes(x = age_group, y = loneliness, fill = gender) Geoms geom_boxplot(), geom_jitter(), geom_smooth() Scales scale_fill_manual(values = c("#E69F00", "#56B4E9", "#009E73")) Facets facet_wrap(~ age_group) or facet_grid(gender ~ age_group) Theme labs(title = ..., x = ..., y = ...) + theme_minimal()
Figure 6.11. The six parts of a ggplot2 plot, each with an example from the briefing. The first three are needed for every plot; the last three have sensible defaults and are changed when needed.

The six flip cards below repeat the parts in Figure 6.11 and serve as a reference for the code in the activities of this section.

Click each card for a short description of the part and the code that controls it.

DataClick to explore
AestheticsClick to explore
GeomsClick to explore
ScalesClick to explore
FacetsClick to explore
Themes and labelsClick to explore

Equation 6.1 shows how the parts are joined into a single block of code.

Every ggplot2 plot follows the same template. The function ggplot() receives the data and the aesthetic mappings, and each further part is added with a plus sign at the end of the line before it:

The ggplot2 template
ggplot(data, aes(x = ..., y = ...)) +
  geom_...() +
  labs(title = ..., x = ..., y = ...) +
  theme_...()Eq 6.1

The accordion below describes three errors that beginners meet first when they use the template in Equation 6.1, with the remedy for each.

Three errors that beginners meet first

The package is not loaded. The message could not find function "ggplot" means that library(ggplot2) has not been run in this session. The package is installed once with install.packages("ggplot2"), and loaded with library(ggplot2) in every new session.

The plus sign starts a line. R reads each line as a complete command when it can. If a line ends without a plus sign, R draws the plot so far and then treats the next line, starting with +, as a separate command, which gives an error. The plus sign therefore goes at the end of a line, never at the start of the next.

A fixed colour is put inside aes(). Writing aes(colour = "blue") maps the word "blue" as if it were a variable, so the points are drawn in a default colour with a legend that says "blue". A fixed colour belongs outside aes(), as in geom_point(colour = "blue").

With the template in hand, the following parts apply it to the displays drawn in base R in Section 2, beginning with one variable.

One Variable in ggplot2

The histogram in Activity 6.4 maps loneliness to x and adds geom_histogram(binwidth = 1), which gives one bar per whole score, as breaks did in base R. The fill colour and the white bar outlines are set outside aes(), so they apply to every bar. Figure 6.12 carries the same information as the base R histogram in Section 2 (Figure 6.7), with labels that can be reused in the briefing.

ggplot2 histogram of loneliness scores from 3 to 9 with teal bars and white outlines on a grey background; the tallest bars are at 5 and 6.
Figure 6.12. Loneliness score drawn with geom_histogram(binwidth = 1) (n = 3,083).

Two geoms draw bars, and they are easily confused. geom_bar() counts the rows in each category, so it needs only an x aesthetic, such as mental_health. geom_col() draws bars whose heights are supplied in the data, so it needs both x and y. To draw mean loneliness by age group, the means are first computed with the base R function aggregate(), which returns a small data frame with one row per age group, and geom_col() then draws those values. Figure 6.13 draws one chart with each geom.

Left: geom_bar chart of the number of people in each self-rated mental health category, tallest for very good. Right: geom_col chart of mean loneliness score in four age groups, all between 5.4 and 6.0.
Figure 6.13. Left: geom_bar() counts the rows in each category of self-rated mental health. Right: geom_col() draws the four group means computed by aggregate(). The titles were added for this figure.

The right-hand panel is the "means without spread" display from Section 1, and it is shown here to illustrate geom_col(). For the briefing, loneliness by age group is better shown with boxplots and the individual points.

Activity 6.4 draws the plots in Figure 6.12 and Figure 6.13 and a boxplot with jittered points.

R Activity 6.4: first plots in ggplot2

This activity starts from the briefing file made in Activity 6.2 of Section 2. If you are starting a new R session, run the loading and preparation code from that activity first (it is also at the top of the answer key). The first line installs ggplot2 and is needed only once on each computer, so it is written as a comment; remove the # to run it.

# install.packages("ggplot2")   # run once, if not yet installed
library(ggplot2)
ggplot(brief, aes(x = loneliness)) +          # data and the aesthetic mapping
  geom_histogram(binwidth = 1,                # the geom: one bar per score
                 fill = "#0B7B6B", colour = "white") +
  labs(title = "Loneliness scores",           # labels
       x = "Loneliness score (3 to 9)", y = "Number of people")

This block prints nothing to the console and draws a histogram in the Plots pane. If R reports could not find function "ggplot", run library(ggplot2) again.

ggplot(brief, aes(x = mental_health)) +
  geom_bar() +                                # geom_bar() counts the rows in each category
  labs(x = "Self-rated mental health", y = "Number of people")

geom_bar() counts the people in each category of mental_health, in the order set by factor() in Section 2.

means <- aggregate(loneliness ~ age_group, data = brief, FUN = mean)
means                                         # one row per age group
ggplot(means, aes(x = age_group, y = loneliness)) +
  geom_col() +                                # geom_col() draws the values it is given
  labs(x = "Age group", y = "Mean loneliness score")
Console output
age_group loneliness 1 16 to 29 5.492122 2 30 to 44 5.388679 3 45 to 64 6.006920 4 65 and older 5.633880

aggregate() applies mean() to the loneliness scores within each age group and returns a data frame with one row per group. geom_col() then draws the four values in the loneliness column.

set.seed(2021)                                # the same jitter every time
ggplot(brief, aes(x = age_group, y = loneliness)) +
  geom_boxplot(outlier.shape = NA) +          # the box summarizes each group
  geom_jitter(width = 0.2, height = 0.2, alpha = 0.1) +   # every person as a faint dot
  labs(x = "Age group", y = "Loneliness score (3 to 9)")

The boxplot and the jittered points share the same axes. set.seed(2021) makes the jitter the same every time the code runs.

R Reflect on what you just ran

Use the questions below to interpret the output you produced. Look at your console output and plots before answering.

1. In the histogram code, identify the data, the aesthetic mapping, the geom and the labels. Which setting in geom_histogram() is set to a fixed value, and how would the plot change if it were mapped inside aes() instead?

Model answerThe data are brief; the aesthetic mapping is aes(x = loneliness); the geom is geom_histogram(binwidth = 1, ...); and the labels come from labs(). The fill ("#0B7B6B") and the outline colour ("white") are set to fixed values outside aes(), so every bar is teal with a white outline. If fill = "#0B7B6B" were written inside aes(), ggplot2 would treat the text as a variable with one value, draw the bars in its default colour (a pinkish red) and add a legend labelled with the text.

2. Using the table printed by aggregate(), report the mean loneliness score in each age group to two decimal places. Why does this table need geom_col() and not geom_bar()?

Model answerThe means are 5.49 (16 to 29), 5.39 (30 to 44), 6.01 (45 to 64) and 5.63 (65 and older). The table has one row per age group with the value to plot already computed, so the bar heights must come from the loneliness column. geom_col() draws supplied values, while geom_bar() counts rows, so with only x mapped it would give every age group a bar of height 1, because each group appears in one row of the table.

3. Compare the boxplot with jittered points to the bar chart of means. Name two things the boxplot display shows that the bars do not.

Model answerThe boxplot display shows the spread within each group: every age group has scores across the whole range from 3 to 9, and the box for 45 to 64 is wider than the others, whereas the bars show only four means between 5.39 and 6.01. It also shows the number of people in each group through the density of the points, with far fewer dots in the 65 and older group (366 people) than in the 16 to 29 group (1,079), which the bars hide completely. The medians (5, 5, 6 and 6) are also visible.
Saved.

The last plot in Activity 6.4 combines a boxplot with the individual points, a display that the next part examines more closely.

A Numeric Variable Across Groups

A boxplot with jittered points uses two geoms on the same axes. geom_boxplot(outlier.shape = NA) draws the median, the quartiles and the whiskers for each age group and suppresses the separate outlier points, because geom_jitter(width = 0.2, height = 0.2, alpha = 0.1) already draws every person. The width and height arguments set how far each point may be nudged sideways and up or down, and alpha sets the opacity, from 0 (invisible) to 1 (solid). Figure 6.14 shows the result for the four age groups.

Boxplots of loneliness score for four age groups, overlaid with faint jittered points forming horizontal bands at each whole score; the band of points is thinnest for the 65 and older group.
Figure 6.14. Loneliness score by age group, with each person drawn as a jittered, transparent point (n = 3,083).

The combined display shows three things that bars of means would hide. The medians differ by one point (5 in the two younger groups and 6 in the two older groups). Every group includes people across the whole range of the scale, from 3 to 9. And the groups differ in size, from 1,079 people aged 16 to 29 to 366 aged 65 and older, which the density of the points makes visible.

Boxplots compare a numeric variable across the groups of a categorical one. When both variables are numeric, the display is a scatterplot, which the next part extends with a smoother.

Two Numeric Variables and a Smoother

A scatterplot maps two numeric variables to x and y. Because the loneliness score takes only seven values, the points are jittered vertically (height = 0.2) and drawn with transparency. The layer geom_smooth() adds a curve that summarizes the average of y across the values of x, with a grey 95% confidence band. With method = "loess", the curve is a loess smoother (locally estimated scatterplot smoothing), which fits many small weighted regressions along the axis and joins them, so it can bend. With method = "lm", the curve is a straight least-squares line. When it draws the curve, ggplot2 prints the message `geom_smooth()` using formula = 'y ~ x', which only reports the formula it used and is not an error. Figure 6.15 adds a loess smoother to a scatterplot of loneliness against age.

Scatterplot of age from 16 to 100 against loneliness score, with jittered transparent points and a blue loess curve with a grey band. The curve starts near 6.4 at age 16, falls to about 5.3 near age 30, rises to about 6.0 between 50 and 60, and declines after 70, where the band becomes wide.
Figure 6.15. Loneliness score by age with a loess smoother and its 95% confidence band (n = 3,083). Points are jittered vertically.

The curve tells a story that a straight line would miss. Predicted loneliness is highest among the youngest respondents (about 6.4 at age 16), falls to about 5.3 around age 30, rises to about 6.0 between ages 50 and 60, and declines after 70 (about 5.4 at age 80). The band is narrow where most people are (ages 20 to 40) and widens at both ends, where the curve rests on few people: 48 respondents are under 20 and 29 are 80 or older. The ends of a smoother deserve the least trust, and a single respondent aged 100 sits at the far right. A smoother is a description of the data. It suggests that age might enter a later model as a curve or in groups, and it gives no reason, by itself, for the pattern.

The smoother summarizes one relationship across the whole briefing file. The next part asks whether a pattern differs between groups by drawing one panel for each group.

Comparing Across Groups with Facets

Facets draw the same plot in one panel per group. facet_wrap(~ age_group) takes one categorical variable and wraps its panels into a grid, which suits a single grouping variable with several levels. All panels share the same axes by default, so heights and positions can be compared across panels. Figure 6.16 applies facet_wrap() to the histogram of loneliness, with one panel for each age group.

Four histograms of loneliness score in a two by two grid, one for each age group; the 16 to 29 and 30 to 44 panels have tall bars at 5 and 6, and the two older panels have shorter bars spread more evenly.
Figure 6.16. Loneliness score by age group with facet_wrap(~ age_group) (n = 3,083). The panels share both axes.

The shared vertical axis shows both the shape and the size of each group. The two younger groups, with more than 1,000 people each, have tall bars at 5 and 6. The two older groups have shorter, flatter histograms with more people at both ends of the scale, and they hold most of the ceiling pile-up seen in Section 2: in the 45 to 64 group, 111 people (19.2%) scored 9, as many as scored 6 (112), and in the 65 and older group the most common score is 3, the lowest possible (85 people, 23.2%). The older groups are therefore more divided, with more people at the least lonely and the most lonely ends of the scale. If the groups' shapes matter more than their sizes, facet_wrap(~ age_group, scales = "free_y") gives each panel its own vertical axis, at the cost of making heights incomparable across panels.

facet_grid(gender ~ age_group) takes two variables, the first for the rows and the second for the columns, and draws one panel for every combination. Figure 6.17 shows social support against loneliness, with a straight line in each panel.

A grid of twelve scatterplots, three rows for woman, man and non-binary and four columns for the age groups, each with a blue line sloping downward. Lines in the older age groups are steeper. The non-binary row has few points and wide grey bands.
Figure 6.17. Loneliness score against social support by gender (rows) and age group (columns), with a least-squares line and 95% confidence band in each panel (n = 3,083).

The lines slope downward in all twelve panels, so higher support goes with lower loneliness in every combination of gender and age group. The lines look steeper in the two older age groups: within the age groups as a whole, the slopes are about −0.40 and −0.54 points of loneliness per point of support in the two younger groups and about −0.76 and −0.84 in the two older groups. The table of counts printed in Activity 6.5 is the necessary companion to the grid. The non-binary panels hold between 6 and 26 people, so their lines and bands are very uncertain, and any difference they appear to show would need a much larger sample to assess. An exploratory figure should show small groups, and the text should say how small they are.

Table 6.5 gives these counts, the number of people behind each panel of Figure 6.17.

Table 6.5. Number of people in each panel of Figure 6.17, by gender and age group.

Gender16 to 2930 to 4445 to 6465 and older
Woman490477362251
Man564557204109
Non-binary2526126

Facets give each group its own panel. Groups can also share one panel and be told apart by colour, and the next part turns to colour, palettes and themes.

Colour, Palettes and Themes

Mapping a variable to colour (for points and lines) or fill (for bars and boxes) shows groups within one panel. Colour and facets suit different situations. Colour lets the reader compare groups directly on the same axes, and it works well for two or three groups. Facets give each group its own panel, and they work better when there are many groups, when points overlap heavily, or when one group is much smaller than the others and would be hidden. The two can also be combined, with colour for one grouping variable and facets for another.

The colours themselves need care. About 8% of men and 0.5% of women of northern European ancestry have some form of red-green colour-vision deficiency, so a palette that relies on telling red from green fails a sizeable part of any audience. The Okabe-Ito palette (Okabe & Ito, 2008) is a set of eight colours chosen to remain distinguishable for people with the common forms of colour-vision deficiency. Its first three colours after black, orange (#E69F00), sky blue (#56B4E9) and bluish green (#009E73), are used below with scale_fill_manual(). The viridis palettes, available in ggplot2 as scale_fill_viridis_d() for categories, are another safe choice, and they also remain readable when printed in greyscale. Figure 6.18 maps fill to gender in boxplots of loneliness by age group, with the three Okabe-Ito colours and a minimal theme.

Boxplots of loneliness score for four age groups, with three boxes per group filled orange for women, sky blue for men and bluish green for non-binary respondents, on a white minimal theme with a legend titled Gender.
Figure 6.18. Loneliness score by age group and gender, with fill mapped to gender, three Okabe-Ito colours and theme_minimal(base_size = 12) (n = 3,083).

The theme controls the appearance of everything except the data. The default theme, theme_grey(), uses a grey panel with white grid lines. theme_minimal() removes the grey panel, theme_bw() adds a black border, and theme_classic() keeps only the axis lines. The argument base_size sets the size of all text in points, and 11 or 12 suits most journal figures. Labels are set with labs(), which also names the legend through the aesthetic it describes, as in fill = "Gender".

Once a figure has its colours, theme and labels, it is saved to a file at the size at which it will be printed or shown, which the next part describes.

Saving a Figure with ggsave()

The function ggsave() writes a ggplot2 plot to a file. Its first argument is the file name, and the extension sets the format: .png for most uses, .pdf or .svg for figures that must stay sharp at any size, and .tiff when a journal asks for it. The second argument is the plot object, here p, and without it ggsave() saves the last plot shown. The arguments width and height are in inches by default, and dpi sets the resolution, with 300 dots per inch the usual requirement for print. A figure of 7 by 4.5 inches at 300 dpi is 2,100 by 1,350 pixels. Text size is relative to the physical size of the figure, so a plot saved at a large width with a small base_size will have text that is too small when the figure is shrunk to fit a page.

Activity 6.5 draws the plots in Figure 6.15, Figure 6.16, Figure 6.17 and Figure 6.18, prints the counts in Table 6.5, and saves the last plot with ggsave().

R Activity 6.5: smoothers, facets, colour and saving

This activity continues from Activity 6.4, with ggplot2 loaded and the briefing file in memory. It draws a scatterplot with a loess smoother, two faceted plots and a boxplot coloured by gender, and saves the last plot to a file.

set.seed(2021)
ggplot(brief, aes(x = age, y = loneliness)) +
  geom_jitter(height = 0.2, alpha = 0.1) +
  geom_smooth(method = "loess") +             # a smooth curve through the average score
  labs(x = "Age (years)", y = "Loneliness score (3 to 9)")
Console output
`geom_smooth()` using formula = 'y ~ x'

The message reports the formula that geom_smooth() used, a curve of y on x, and is not an error. The blue curve is the loess smoother and the grey band is its 95% confidence band.

ggplot(brief, aes(x = loneliness)) +
  geom_histogram(binwidth = 1, fill = "#0B7B6B", colour = "white") +
  facet_wrap(~ age_group) +                   # one panel for each age group
  labs(x = "Loneliness score (3 to 9)", y = "Number of people")

facet_wrap(~ age_group) draws one histogram per age group, with shared axes.

set.seed(2021)
ggplot(brief, aes(x = support, y = loneliness)) +
  geom_jitter(height = 0.2, alpha = 0.2) +
  geom_smooth(method = "lm") +                # a straight line in each panel
  facet_grid(gender ~ age_group) +            # rows: gender; columns: age group
  labs(x = "Social support (1 to 7)", y = "Loneliness score (3 to 9)")
Console output
`geom_smooth()` using formula = 'y ~ x'

The grid has three rows (gender) and four columns (age group), and method = "lm" draws a straight line in each panel.

table(brief$gender, brief$age_group)         # people in each panel
Console output
16 to 29 30 to 44 45 to 64 65 and older Woman 490 477 362 251 Man 564 557 204 109 Non-binary 25 26 12 6

The table gives the number of people behind each panel of the grid. Six non-binary respondents are aged 65 and older.

p <- ggplot(brief, aes(x = age_group, y = loneliness, fill = gender)) +
  geom_boxplot() +
  scale_fill_manual(values = c("#E69F00", "#56B4E9", "#009E73")) +  # Okabe-Ito colours
  labs(x = "Age group", y = "Loneliness score (3 to 9)", fill = "Gender") +
  theme_minimal(base_size = 12)               # a plain theme with larger text
p                                             # show the plot
ggsave("loneliness_age_gender.png", p, width = 7, height = 4.5, dpi = 300)

Storing the plot as p lets the same object be shown on screen and saved. ggsave() writes loneliness_age_gender.png, 7 by 4.5 inches at 300 dots per inch (2,100 by 1,350 pixels), to the working directory, and prints nothing.

R Reflect on what you just ran

Use the questions below to interpret the output you produced. Look at your console output and plots before answering.

1. Describe the shape of the loess curve of loneliness across age in two or three sentences. Why should the ends of the curve be trusted least?

Model answerThe curve is highest among the youngest respondents (about 6.4 at age 16), dips to about 5.3 around age 30, rises to about 6.0 between 50 and 60, and declines after 70. The ends deserve the least trust because few people are there: 48 respondents are under 20 and 29 are 80 or older, so the curve rests on little data and its confidence band is wide. The shape suggests that a straight line in age would not describe loneliness well.

2. Using the faceted grid and the table of counts, compare the lines in the panels for women and men with the panels for non-binary respondents. What should the briefing say about the non-binary panels?

Model answerIn the panels for women and men, the lines slope downward in every age group and are steeper in the two older groups, with narrow bands because hundreds of people are in most panels (for example, 490 women and 564 men aged 16 to 29). The non-binary panels hold only 25, 26, 12 and 6 people, so their lines are very uncertain, with wide bands. The briefing should show these panels but state how few people they contain and avoid drawing any conclusion about differences for non-binary respondents from them.

3. Explain why the code uses scale_fill_manual(values = c("#E69F00", "#56B4E9", "#009E73")), and state what the width, height and dpi arguments of ggsave() control.

Model answerThe three colours come from the Okabe-Ito palette, which was designed so that people with common forms of colour-vision deficiency, such as red-green deficiency, can still tell the groups apart; the default ggplot2 colours include a red and a green that some readers cannot distinguish. In ggsave(), width and height set the physical size of the figure in inches (7 by 4.5), and dpi sets the resolution in dots per inch (300, the usual print standard), so the saved file is 2,100 by 1,350 pixels.
Saved.

The card below links again to the narrated walkthrough introduced in Section 2.

Learn to do this in R

Narrated R walkthrough: Data Visualization in Base R and ggplot2

The walkthrough builds the ggplot2 plots in this section layer by layer, so that the effect of each line of code can be seen before the next is added.

Open the walkthrough

This section has described the grammar on which ggplot2 is built and used it to draw the briefing plots for one variable, for a numeric variable across groups and for two numeric variables, to compare groups with facets and colour, and to save a figure. The knowledge check and the reflection that follow review these steps, and Section 4 adds displays for many variables and combines the briefing plots into one figure.

Knowledge check: this section

1. In the layered grammar of graphics, what is an aesthetic mapping?

An aesthetic mapping, written inside aes(), connects a variable to a visual property such as x, y, colour or fill. The marks are geoms, and fonts and backgrounds belong to the theme.

2. A data frame has one row per age group and a column of mean scores. Which geom draws one bar per group with the height of its mean?

geom_col() draws bars whose heights are supplied in the data. geom_bar() counts rows, so it would give every group in this table a bar of height 1.

3. Why is outlier.shape = NA used in geom_boxplot() when geom_jitter() is added?

Without it, the boxplot draws its own points for values beyond the whiskers, and the jitter layer draws the same people again. The data are unchanged; only the duplicate marks are hidden.

4. A plot must compare a score across 12 regions, and the points overlap heavily. Which approach is usually clearer?

Twelve colours are hard to tell apart, especially with overlapping points. Facets give each region a panel on shared axes, which keeps the comparison readable.

5. Which ggsave() call saves the plot p as a 7 by 4.5 inch PNG file at a resolution suitable for print?

The file name comes first and its extension sets the format, the plot object comes second, width and height are in inches, and 300 dots per inch is the usual print standard.

✎ Reflection

This section introduced ggplot2 and the layered grammar of graphics, in which a plot is built from data, aesthetic mappings (variables linked to visual properties inside aes()), geoms (marks such as boxes, points and lines), scales, facets (one panel per group) and a theme. It also described two ways to show groups: mapping a variable to colour or fill within one panel, which suits two or three groups compared directly, and facets, which suit many groups, heavily overlapping points or very unequal group sizes. Colours should come from a palette that people with colour-vision deficiency can read, such as the Okabe-Ito palette. A colleague's draft figure shows depressive symptom scores (0 to 6) against age for respondents in five regions, all in one panel, with each region in a different colour from a red-to-green palette and about 3,000 overlapping solid points. Describe three changes you would make to the figure, write the ggplot2 code for your improved version (assume a data frame d with the variables age, depression and region), and explain what each change improves.

Model answerFirst, I would give each region its own panel with facet_wrap(~ region), because five overlapping colours in one panel with 3,000 points make it impossible to see any one region, while facets on shared axes keep the regions comparable. Second, I would deal with overplotting: depression takes only seven whole values, so I would use geom_jitter(height = 0.2, alpha = 0.15) so that dense areas appear darker and individual values are visible, and add geom_smooth(method = "loess") to summarize the average score across age in each panel. Third, with region now shown by the panels, colour is no longer needed for region, which removes the red-to-green palette that readers with red-green colour-vision deficiency cannot read; if colour were still needed, I would use the Okabe-Ito colours. I would also add clear labels and a plain theme. The code would be:
ggplot(d, aes(x = age, y = depression)) +
  geom_jitter(height = 0.2, alpha = 0.15) +
  geom_smooth(method = "loess") +
  facet_wrap(~ region) +
  labs(x = "Age (years)", y = "Depressive symptoms, PHQ-2 (0 to 6)") +
  theme_minimal()

Finally, I would print a table of the number of people in each region beside the figure, so that a panel with few people is not over-interpreted.
✓ Reflection saved!
● Complete the quiz and reflection to continue.
Section 4 of 4

Many Variables and Combined Figures

⏱ Estimated time: 50 minutes
Lesson 6 · Section 4

Many Variables and Combined Figures

An overview of many variables, and one figure that brings the briefing together.

Correlation

Pearson's r for every pair of variables

Loneliness withrDirection
Depressive symptoms (PHQ-2)0.47Positive
Anxiety symptoms (GAD-2)0.43Positive
Social support−0.40Negative
Life satisfaction−0.37Negative
Age0.07Close to zero

The strongest correlation in the matrix is between depression and anxiety (r = 0.59).

Heatmap

A correlation matrix drawn as colours

Correlation heatmap of six variables, red for positive and blue for negative correlations, with the value printed on each square.
Pearson correlations among six briefing variables (n = 3,083).
Interpretation

What a correlation heatmap cannot show

Cause

A correlation fits several causal stories, including reverse causation and common causes.

Curves

Pearson's r measures straight-line association, so a curved pattern can give r near zero.

Shape

One number per pair hides outliers, clusters and the number of people involved.

Marginal histograms

A relationship and two distributions in one display

Jittered scatterplot of social support against loneliness with a red downward line, a histogram of support along the top with spikes at whole numbers, and a histogram of loneliness along the right side.
Drawn with ggMarginal(p_scatter, type = "histogram").
Combining panels

The briefing as one figure

Four-panel figure: A histogram of loneliness, B boxplots by age group, C scatterplot of support and loneliness with a red line, D correlation heatmap.
(p_hist + p_box) / (p_scatter + p_heat) + plot_annotation(tag_levels = "A")
Publication figures

Conventions for a manuscript figure

  • Each axis has a label in words, with units or the range of the scale.
  • Plain labels replace R variable names such as life_sat.
  • Panels carry letters, and legends say what colours and marks mean.
  • Colours remain readable with colour-vision deficiency and in greyscale.
  • Text is legible at the printed size, and the file is saved at about 300 dpi.
Captions

What a figure caption states

ElementExample from the briefing figure
What the figure showsLoneliness and social connection
Data source and yearCanadian Social Connection Survey, 2021
People included3,083 respondents with complete data
Measures and rangesUCLA 3-item scale (3 to 9); support (1 to 7)
What the marks meanBoxes, whiskers, jitter, line and 95% band
Lesson summary

From a first look to a finished figure

StepTool
Look before modellingPlots of every variable and relationship
Choose the displayThe question and the variable types
Check the data quicklyBase R: hist(), boxplot(), plot(), pairs()
Build the figureggplot2: data, aesthetics, geoms, facets, themes
Combine and captionpatchwork, ggsave() and a caption that stands alone

Introduction and Overview

The working group's third question asks how loneliness goes together with social support and mental health, which involves several variables at once. One scatterplot per pair would answer it slowly, because six variables have fifteen pairs. This section shows how to summarize many pairs in one display with a correlation matrix and a heatmap, and it explains the limits of that summary. It then adds marginal histograms to a scatterplot, combines the briefing plots into one multi-panel figure, and sets out the conventions of a publication figure and its caption. The finished figure is the kind of exploratory figure that a research paper based on survey data often includes.

Learning Objectives

  • Compute a correlation matrix with cor() and display it as a heatmap in ggplot2.
  • Explain why a correlation heatmap shows association and cannot identify causes, curved relationships or the shape of the data.
  • Build a scatterplot with marginal histograms using ggExtra::ggMarginal().
  • Combine several ggplot2 plots into one labelled figure with the patchwork package and save it with ggsave().
  • Apply the conventions of a publication figure and write a caption that states what the figure shows.

The section begins with the correlation matrix, the numeric summary from which the heatmap is drawn.

Looking at Many Variables at Once

The Pearson correlation coefficient, written r, measures how closely two numeric variables follow a straight line. It ranges from −1 (all points on a line sloping downward) through 0 (no straight-line relationship) to +1 (all points on a line sloping upward). Its sign gives the direction and its size gives the strength. Cohen (1988) suggested that values of about 0.1, 0.3 and 0.5 be called small, medium and large in the behavioural sciences, and these labels are a rough guide only, since what counts as a strong association depends on the field and the question.

The function cor(), given a data frame of numeric columns, returns a correlation matrix with one row and one column per variable. Each cell holds the correlation between its row variable and its column variable. The diagonal holds the correlation of each variable with itself, which is always 1, and the matrix is symmetric, so the values above the diagonal repeat those below it. Six variables give 36 cells and 15 distinct correlations.

Table 6.6 gives the correlation matrix for the six numeric briefing variables.

Table 6.6. Pearson correlations among the six numeric variables in the briefing file (n = 3,083).

lonelinesssupportdepressionanxietylife_satage
loneliness1.00−0.400.470.43−0.370.07
support−0.401.00−0.38−0.300.35−0.04
depression0.47−0.381.000.59−0.41−0.07
anxiety0.43−0.300.591.00−0.33−0.06
life_sat−0.370.35−0.41−0.331.000.00
age0.07−0.04−0.07−0.060.001.00

In the briefing file, loneliness has moderate positive correlations with depressive symptoms (0.47) and anxiety symptoms (0.43), and moderate negative correlations with social support (−0.40) and life satisfaction (−0.37). Its correlation with age is close to zero (0.07). The strongest correlation in the matrix is between depressive and anxiety symptoms (0.59), two measures that often move together.

Two practical points apply to cor(). First, it returns NA for any pair with a missing value unless told how to handle them, for example with use = "complete.obs". The briefing file has no missing values, so this does not arise here. Second, Pearson's r assumes numeric variables and is sensitive to outliers and to skew. For ordinal scores or skewed variables, Spearman's rank correlation, computed with cor(..., method = "spearman"), uses the ranks of the values and is less affected by extreme values. For loneliness and social support, Spearman's correlation is −0.36, close to the Pearson value of −0.40.

Activity 6.6 computes Table 6.6 with cor(), reshapes it into one row per pair of variables, and draws a heatmap of the correlations.

R Activity 6.6: a correlation matrix and a heatmap

This activity starts from the briefing file with ggplot2 loaded. If you are starting a new R session, run the loading and preparation code from Section 2 and library(ggplot2) first (both are at the top of the answer key).

vars <- c("loneliness", "support", "depression", "anxiety", "life_sat", "age")
cor_mat <- round(cor(brief[, vars]), 2)     # Pearson correlations, rounded
cor_mat
Console output
loneliness support depression anxiety life_sat age loneliness 1.00 -0.40 0.47 0.43 -0.37 0.07 support -0.40 1.00 -0.38 -0.30 0.35 -0.04 depression 0.47 -0.38 1.00 0.59 -0.41 -0.07 anxiety 0.43 -0.30 0.59 1.00 -0.33 -0.06 life_sat -0.37 0.35 -0.41 -0.33 1.00 0.00 age 0.07 -0.04 -0.07 -0.06 0.00 1.00

brief[, vars] keeps the six numeric columns named in vars, cor() computes every pairwise Pearson correlation, and round(..., 2) keeps two decimal places. The diagonal is 1.00 and the matrix is symmetric.

cor_long <- as.data.frame(as.table(cor_mat))   # one row for each pair
names(cor_long) <- c("var1", "var2", "r")
head(cor_long)
Console output
var1 var2 r 1 loneliness loneliness 1.00 2 support loneliness -0.40 3 depression loneliness 0.47 4 anxiety loneliness 0.43 5 life_sat loneliness -0.37 6 age loneliness 0.07

The long data frame has 36 rows, one for each cell of the matrix. head() prints the first six, which are the correlations of each variable with loneliness.

ggplot(cor_long, aes(x = var1, y = var2, fill = r)) +
  geom_tile(colour = "white") +                     # one coloured square per pair
  geom_text(aes(label = r), size = 3.5) +           # print the value on each square
  scale_fill_gradient2(low = "#2166AC", mid = "white", high = "#B2182B",
                       limits = c(-1, 1)) +         # blue negative, red positive
  labs(x = NULL, y = NULL, fill = "r") +
  theme_minimal()

This block draws the heatmap and prints nothing. scale_fill_gradient2() sets a diverging scale with white at zero, and limits = c(-1, 1) fixes the ends of the scale at the possible range of r.

R Reflect on what you just ran

Use the questions below to interpret the output you produced. Look at your console output and plots before answering.

1. Which variable has the strongest correlation with loneliness, and what is its value? Describe this correlation in one sentence using the words direction and strength.

Model answerDepressive symptoms (PHQ-2) has the strongest correlation with loneliness, r = 0.47. The association is positive and moderate in strength: respondents with higher loneliness scores tend to report more depressive symptoms, although many individuals depart from that trend.

2. The correlation between age and loneliness is 0.07. A member of the working group concludes that age is unrelated to loneliness. Using what you saw in Section 3, explain why this conclusion is too strong.

Model answerPearson's r measures only straight-line association. The loess curve in Section 3 showed that loneliness changes with age in a curved way: it is high among the youngest respondents, dips around age 30, rises to about 6 in the fifties and falls after 70. These rises and falls largely cancel out in a single straight-line summary, so r is close to zero even though age and loneliness are related. The briefing should show the curve and avoid saying that age is unrelated to loneliness.

3. Another member reads the heatmap and says that it shows social support reduces loneliness. Write a two-sentence reply that the analyst could give.

Model answerThe heatmap shows that social support and loneliness are negatively correlated (r = −0.40), meaning that people with more support tend to report less loneliness, but a correlation from a survey collected at one time cannot show which comes first or rule out other explanations. Loneliness could reduce the support people perceive, or other factors such as health or life circumstances could affect both, so a causal claim would need a design and an analysis built to test it.
Saved.

The heatmap at the end of Activity 6.6 displays the same matrix in colour, and the next part explains how it is built and how it is read.

Correlation Heatmaps

A matrix of 36 numbers is slow to read. A correlation heatmap draws each cell as a coloured square, so that the pattern can be seen at a glance. Drawing it in ggplot2 takes two steps. The matrix is first turned into a long data frame with one row per pair: as.table() treats the matrix as a table, and as.data.frame() turns that table into three columns, here renamed var1, var2 and r. Then geom_tile() draws one square per row, with var1 mapped to x, var2 to y and r to fill, and geom_text() prints the value on each square. Figure 6.19 shows the heatmap for the six briefing variables.

Correlation heatmap of loneliness, support, depression, anxiety, life satisfaction and age. The diagonal is dark red (1). Loneliness, depression and anxiety are pink to red with each other (0.43 to 0.59) and blue with support and life satisfaction (minus 0.30 to minus 0.41). Age is nearly white in every square.
Figure 6.19. Pearson correlations among six briefing variables (n = 3,083). Red squares are positive and blue squares negative, with white at zero, and each square shows its value.

The colour scale is a diverging scale, set with scale_fill_gradient2(): blue for negative values, white at zero and red for positive values, with limits = c(-1, 1) so that the colours keep the same meaning in any heatmap. A diverging scale suits correlations because zero is a meaningful midpoint. Blue and red remain distinguishable for people with red-green colour-vision deficiency, which a red-to-green scale would not.

The heatmap shows a clear structure. Loneliness, depression and anxiety form a block of positive correlations, and each of the three is negatively correlated with social support and life satisfaction, which are positively correlated with each other (0.35). Age is close to white in every square. For the briefing, this panel gives the working group an overview of the third question in one picture.

The overview has limits, which the following subsection sets out.

What a Heatmap Shows and What It Cannot

A correlation heatmap shows the strength and direction of straight-line association between each pair of variables. Three common readings go beyond what it shows.

The accordion below explains each of these readings and why the heatmap cannot support it.

A correlation does not identify a cause

The correlation of 0.47 between loneliness and depressive symptoms is consistent with several causal stories. Loneliness might increase depressive symptoms. Depression might lead people to withdraw from others and feel lonelier, which is reverse causation. Other factors, such as poor physical health, bereavement or low income, might increase both, which is confounding by a common cause. All three stories, and combinations of them, would produce the same square on the heatmap. The CSCS 2021 data were collected at one time point, so the order of events cannot be established from them. Separating these explanations requires study design, causal diagrams of the kind drawn in earlier courses, and models that adjust for confounders, and even then a cross-sectional survey supports causal claims only weakly. A heatmap is therefore a map of where to look, and the briefing should describe it as showing which measures go together.

A correlation near zero does not mean no relationship

Pearson's r measures how well a straight line describes the data. The correlation between age and loneliness is 0.07, yet the loess curve in Section 3 showed a clear pattern: loneliness dips around age 30, rises into the fifties and falls after 70. Rises and falls in different parts of the age range cancel out in a single straight-line summary. A near-white square means only that no straight line fits well, and a scatterplot is needed to see whether a curve does.

A correlation hides the shape and size of the data

Anscombe's quartet in Section 1 showed four datasets with the same correlation of 0.82 and completely different shapes, including a curve, an outlier and a relationship created by a single point. A heatmap reduces each pair to one number, so it hides outliers, clusters, floors and ceilings, and it says nothing about how many people contributed. Strong or surprising squares should be followed up with a scatterplot of that pair.

A heatmap therefore points to the pairs of variables that deserve a closer look. The next part shows one way to take that closer look: a scatterplot of one pair with the distribution of each variable along its margins.

Scatterplots with Marginal Histograms

A scatterplot shows how two variables go together, and a histogram shows how one variable is distributed. A scatterplot with marginal histograms shows both at once: the scatterplot sits in the middle, with a histogram of the horizontal variable along the top and a histogram of the vertical variable along the right side. The function ggMarginal() from the ggExtra package adds the margins to a finished ggplot2 scatterplot. Its argument type can be "histogram", "density" or "boxplot". Figure 6.20 shows loneliness against social support with a histogram of each variable along its margin.

Scatterplot of social support from 1 to 7 against loneliness from 3 to 9, with jittered grey points and a red line sloping down from about 8 to about 4.3. Along the top, a histogram of support is skewed to the left, with most values between 4 and 6 and tall spikes at 4, 5 and 7. Along the right, a histogram of loneliness has separate bars at each whole score.
Figure 6.20. Loneliness score against social support with marginal histograms (n = 3,083). Points are jittered vertically; the red line is a least-squares line with its 95% confidence band.

The marginal histogram of social support shows a distribution skewed to the left: most people have scores between 4 and 6, and a long, thin tail runs toward low support. It also shows spikes at whole-number scores. In the briefing file, 170 people have a score of exactly 5, 136 exactly 4 and 117 exactly 7. The support score is the mean of twelve items, so a whole-number score is what a person who gives the same answer to every item receives, and a score of 7 requires the highest answer on all twelve. Such answers can reflect genuine views or a respondent moving quickly through the questions (sometimes called straight-lining), and the plot is a reason to look more closely before analyzing the score. The marginal histogram of loneliness shows the seven separate values of that score.

The object that ggMarginal() returns is a finished drawing, and it is not a ggplot2 object, so further layers cannot be added to it and patchwork cannot place it in a grid without extra steps. It can be saved with ggsave(). The combined briefing figure (Figure 6.21) therefore uses the scatterplot without its margins.

The marginal histograms reveal spikes in the support score at whole-number values. The next part places the scatterplot beside three other briefing plots in a single figure made with patchwork.

Combining Panels with patchwork

A multi-panel figure lets a reader see several related displays together, with consistent styling and one caption. The patchwork package combines ggplot2 plots with arithmetic-like operators. A plus sign (+) or a vertical bar (|) places plots side by side, a forward slash (/) stacks them, and parentheses group them. The expression (p_hist + p_box) / (p_scatter + p_heat) therefore gives two plots in the top row and two in the bottom row. The function plot_annotation() adds a title, a subtitle or a caption to the whole figure, and its argument tag_levels = "A" labels the panels A, B, C and D in reading order, which the caption can then refer to. The function plot_layout() controls the number of columns, the relative widths and heights of panels, and whether legends are collected in one place. Figure 6.21 shows the draft briefing figure that this expression produces from the histogram, the boxplots, the scatterplot and the heatmap.

A four-panel figure titled Loneliness and social connection, CSCS 2021, n = 3,083. Panel A: histogram of loneliness scores. Panel B: boxplots of loneliness by age group. Panel C: jittered scatterplot of social support and loneliness with a red line. Panel D: correlation heatmap of six variables.
Figure 6.21. The draft briefing figure produced in Activity 6.7 with patchwork. The caption for this figure is given in Worked Example 6.1 later in this section.

The figure is saved with ggsave("briefing_figure.png", figure, width = 11, height = 8.5, dpi = 300). A four-panel figure needs a larger size than a single plot, so that each panel keeps legible text, and 11 by 8.5 inches suits a full page turned sideways or a slide. When the figure is placed in a narrower space, such as one column of a journal, it is better to save it at that width and adjust the text size with base_size than to shrink a large image.

Activity 6.7 draws the scatterplot with marginal histograms in Figure 6.20, and then assembles and saves the four-panel briefing figure in Figure 6.21.

R Activity 6.7: marginal histograms and the combined briefing figure

This activity continues from Activity 6.6, with the briefing file, ggplot2 and the cor_long data frame in memory. It installs and loads two more packages (installation is needed once per computer), draws a scatterplot with marginal histograms, and assembles and saves the four-panel briefing figure.

# install.packages(c("ggExtra", "patchwork"))   # run once, if not yet installed
library(ggExtra)     # ggMarginal() adds histograms to the margins of a plot
library(patchwork)   # + and / combine ggplot2 plots into one figure
set.seed(2021)
p_scatter <- ggplot(brief, aes(x = support, y = loneliness)) +
  geom_jitter(height = 0.2, width = 0, alpha = 0.15) +
  geom_smooth(method = "lm", colour = "#CC0033") +
  labs(x = "Social support (1 to 7)", y = "Loneliness score (3 to 9)")
ggMarginal(p_scatter, type = "histogram")    # a histogram on each axis
Console output
`geom_smooth()` using formula = 'y ~ x' `geom_smooth()` using formula = 'y ~ x'

The message appears twice because ggMarginal() builds the scatterplot twice while it adds the margins; it is not an error. The plot shows the scatterplot with a histogram of support along the top and a histogram of loneliness along the right side.

p_hist <- ggplot(brief, aes(x = loneliness)) +
  geom_histogram(binwidth = 1, fill = "#0B7B6B", colour = "white") +
  labs(x = "Loneliness score (3 to 9)", y = "Number of people")
p_box <- ggplot(brief, aes(x = age_group, y = loneliness)) +
  geom_boxplot(fill = "#E6F3F0") +
  labs(x = "Age group", y = "Loneliness score (3 to 9)")
p_heat <- ggplot(cor_long, aes(x = var1, y = var2, fill = r)) +
  geom_tile(colour = "white") +
  geom_text(aes(label = r), size = 3) +
  scale_fill_gradient2(low = "#2166AC", mid = "white", high = "#B2182B",
                       limits = c(-1, 1)) +
  labs(x = NULL, y = NULL, fill = "r")

These lines store three more plots without drawing them. Each is a complete ggplot2 object that patchwork can arrange.

figure <- (p_hist + p_box) / (p_scatter + p_heat) +   # + side by side, / stacked
  plot_annotation(tag_levels = "A",                     # label the panels A to D
    title = "Loneliness and social connection, CSCS 2021 (n = 3,083)")
figure
ggsave("briefing_figure.png", figure, width = 11, height = 8.5, dpi = 300)
Console output
`geom_smooth()` using formula = 'y ~ x' `geom_smooth()` using formula = 'y ~ x'

Typing figure draws the combined figure, and ggsave() draws it again to save it, so the message from the scatterplot panel appears twice. The file briefing_figure.png is 11 by 8.5 inches at 300 dots per inch.

R Reflect on what you just ran

Use the questions below to interpret the output you produced. Look at your console output and plots before answering.

1. Describe the marginal histogram of social support in two sentences. What do the spikes at whole-number scores suggest, and why might they matter?

Model answerSocial support is skewed to the left: most people score between about 4 and 6, and a long, thin tail runs toward low support. The spikes at whole numbers such as 4, 5 and 7 (136, 170 and 117 people in the briefing file) are where respondents who gave the same answer to all twelve items land, which may reflect genuine views or quick, uniform answering; it is worth checking how many people answered this way and whether results change without them.

2. Explain in words how the expression (p_hist + p_box) / (p_scatter + p_heat) arranges the four plots, and what plot_annotation(tag_levels = "A") adds.

Model answerThe plus sign places two plots side by side, and the parentheses group each pair, so p_hist + p_box is the top row and p_scatter + p_heat is the bottom row. The forward slash stacks the first group above the second, giving a two-by-two grid with the histogram top left, the boxplot top right, the scatterplot bottom left and the heatmap bottom right. plot_annotation(tag_levels = "A") labels the panels A, B, C and D in reading order and adds the overall title, so the caption can refer to each panel by its letter.

3. List two changes that would bring the draft figure closer to the conventions of a publication figure, and write the code for one of them.

Model answerPanel D shows R variable names such as life_sat, which should be replaced by plain labels, and a plain theme would remove the grey backgrounds. For the theme, the combined figure can be given one theme for every panel with figure & theme_minimal(), where & applies the theme to all panels. For the labels, the names can be recoded before the heatmap is drawn, for example levels(cor_long$var1)[levels(cor_long$var1) == "life_sat"] <- "Life satisfaction" and the same for var2. The legend title could also read "Pearson r" with labs(fill = "Pearson r").
Saved.

The draft figure still shows R variable names in panel D. The next part sets out the conventions that a publication figure follows.

Conventions of a Publication Figure

An exploratory figure in a manuscript has a different audience from the analyst's quick looks. Readers see it once, often without the text nearby, and they should be able to understand it without access to the code. Tufte (1983) argued that a graphic should spend its ink on the data and remove decoration that carries no information, which he called chartjunk. Table 6.7 lists the conventions most journals expect and checks the draft briefing figure against them.

Table 6.7. Conventions of a publication figure, checked against the draft briefing figure (Figure 6.21).

ConventionWhat it meansDraft briefing figure
Axis labels in words, with units or ranges"Loneliness score (3 to 9)" in place of lonelinessMet in panels A to C
Plain labels for variables"Life satisfaction" in place of life_satNot yet met in panel D, which shows R names
Panel lettersA, B, C and D, referred to in the captionMet with tag_levels = "A"
Legends that name their scaleThe fill legend in panel D is labelled rMet, though "Pearson r" would be clearer
Readable coloursColour-blind-safe palettes that also work in greyscaleMet (teal, a red line, a blue-white-red scale)
Legible text at printed sizeText of about 8 points or more when printedMet at 11 by 8.5 inches
ResolutionAbout 300 dots per inch for printMet with dpi = 300
Sample sizeThe number of people shown, in the title or captionMet in the title (n = 3,083)
No chartjunkNo three-dimensional effects, shadows or decorationMet; a plain theme would lighten it further

A figure that meets the conventions in Table 6.7 still needs a caption, which the next part describes.

Writing a Figure Caption

A caption makes a figure stand on its own. Most journals place the caption below the figure, beginning with the figure number and a short title, followed by sentences that tell the reader what is shown. A good caption states what the figure displays and for whom: the data source and year, the number of people included and the reason others were excluded, the measures and their ranges, and the meaning of every panel and mark, such as what the boxes, whiskers, lines and bands represent and whether points were jittered. The caption describes the figure; the interpretation, such as which differences matter, belongs in the results text that refers to it.

Worked Example 6.1 applies these rules to the briefing. It lists the changes made to the draft in Figure 6.21, gives the full caption of the finished figure, and shows the three sentences of text that accompany the figure in the briefing.

📋 Worked Example 6.1: The finished briefing figure and its caption

The analyst replaces the R names in panel D with plain labels (loneliness, social support, depression, anxiety, life satisfaction and age), applies theme_minimal() to all four panels, and saves the figure at 11 by 8.5 inches and 300 dpi. The caption reads as follows.

Figure 1. Loneliness and social connection among respondents to the 2021 wave of the Canadian Social Connection Survey (n = 3,083 respondents with complete data on all variables shown; 962 of 4,045 respondents were excluded for missing data). (A) Distribution of scores on the three-item UCLA Loneliness Scale (range 3 to 9; higher scores indicate more loneliness). (B) Loneliness score by age group; boxes show the median and interquartile range, whiskers extend to the most extreme values within 1.5 times the interquartile range, and points show values beyond the whiskers. (C) Loneliness score against perceived social support (Multidimensional Scale of Perceived Social Support, range 1 to 7); points are jittered vertically by up to 0.2 and drawn with transparency, and the line is a least-squares line with its 95% confidence band. (D) Pearson correlations among loneliness, social support, depressive symptoms (PHQ-2, 0 to 6), anxiety symptoms (GAD-2, 0 to 6), life satisfaction (1 to 10) and age.

In the one-page briefing, three sentences of text accompany the figure. Loneliness scores were spread across the whole scale, with the most common scores at 5 and 6 and a group of respondents at the maximum score. Median loneliness was one point higher among respondents aged 45 and older than among younger respondents, although scores varied widely within every age group. Loneliness was moderately correlated with lower social support, more depressive and anxiety symptoms and lower life satisfaction; these are associations in a survey collected at one time and do not show which factors cause which.

The caption and the three sentences of text complete the briefing. The last part of this section places a figure of this kind within the results of a research paper.

From Analysis to Manuscript

An exploratory figure is one part of the results of a research paper. STROBE, the reporting guideline for cohort, case-control and cross-sectional studies (von Elm et al., 2007), lists what the results section should report. Item 13 asks for the number of people at each stage of the study and the reasons they were not included, which the caption of the briefing figure gives (3,083 of 4,045 respondents, with 962 excluded for missing data). Item 14 asks for the characteristics of the participants and the number with missing data for each variable, usually in a Table 1, and item 15 asks for summaries of the outcome, such as the distribution in panel A. Item 16 asks for the main results: unadjusted and adjusted estimates with their confidence intervals, and a statement of which confounders were adjusted for and why. These estimates are usually set out in a results table with one column for the crude estimates and one for the adjusted estimates, as Lesson 4, Section 1 showed for a logistic model.

The figure and the tables then support the rest of the paper. The abstract reports the main estimate with its confidence interval in one or two sentences. The discussion interprets the results, compares them with earlier studies and states the limitations; for the CSCS, these include the cross-sectional design, which limits causal claims, and the online recruitment, which limits generalization to the Canadian population. For optional reading, HSCI 207 Lesson 12 (Writing Up Research) treats results tables, the discussion and limitations in more detail.

This section has summarized many variables with a correlation matrix and a heatmap, set out the limits of that summary, added marginal histograms to a scatterplot, combined the briefing plots into one figure, and described the conventions of a publication figure and its caption. The knowledge check and the reflection that follow review these ideas, and the final page of the lesson gathers the key takeaways, a closing reflection and the final assessment.

Knowledge check: this section

1. A correlation matrix of five variables is printed with cor(). How many distinct correlations between different variables does it contain?

Five variables give 5 × 5 = 25 cells. Five are on the diagonal (each variable with itself), and the remaining 20 form mirror-image pairs, so there are 20 / 2 = 10 distinct correlations.

2. In a correlation heatmap, the square for loneliness and depressive symptoms shows r = 0.47. Which statement is supported by the heatmap alone?

A correlation describes how two measures go together. It cannot show the direction of cause or rule out common causes, so the causal statements and the intervention claim go beyond what the heatmap shows.

3. The Pearson correlation between age and loneliness is 0.07, but a loess curve shows loneliness falling and rising across age. What explains this?

Rises and falls in different parts of the age range cancel out in a straight-line summary. A scatterplot or smoother shows the curve that r misses.

4. Which patchwork expression places plot a above plot b?

The forward slash stacks plots vertically. The plus sign and the vertical bar place plots side by side.

5. Which of the following belongs in a figure caption for an exploratory figure?

A caption states what the figure shows, the data source and sample size, and the meaning of each panel and mark, so that the figure stands on its own. Interpretation belongs in the results text, and code belongs in the analysis script.

✎ Reflection

This section combined several displays into one figure and set out the conventions of a publication figure. A caption should state what the figure shows, the data source and year, the number of people included and why others were excluded, the measures and their ranges, and what each panel and mark means (for example, what boxes, whiskers, lines and confidence bands represent and whether points were jittered), and it should leave interpretation to the results text. Suppose a two-panel figure from a community survey of 1,850 adults (2,200 were surveyed, and 350 were excluded for missing data) shows, in panel A, a histogram of sleep duration in hours per night and, in panel B, boxplots of a perceived stress score (range 0 to 40, higher is more stress) for three groups of sleep duration (under 6 hours, 6 to 8 hours, more than 8 hours), with the individual points jittered horizontally. Write a complete caption for this figure, then write one sentence for the results text that interprets panel B without making a causal claim, assuming that the median stress score was 22 in the shortest-sleep group and 16 in the other two groups.

Model answerCaption. Figure 2. Sleep duration and perceived stress among adults in a community survey (n = 1,850 respondents with complete data; 350 of 2,200 respondents were excluded for missing data). (A) Distribution of self-reported sleep duration in hours per night. (B) Perceived stress score (range 0 to 40; higher scores indicate more stress) by sleep duration group (under 6 hours, 6 to 8 hours and more than 8 hours per night); boxes show the median and interquartile range, whiskers extend to the most extreme values within 1.5 times the interquartile range, and each point is one respondent, jittered horizontally to reduce overlap.
Results sentence. Respondents who reported sleeping under 6 hours per night had a higher median perceived stress score (22) than those who slept 6 to 8 hours or more than 8 hours (16 in both groups), although scores overlapped considerably across the three groups (Figure 2B).
The caption describes what is shown and how to read it, and the results sentence states the association without saying that short sleep causes stress, since the data cannot establish the direction of the relationship.
✓ Reflection saved!
● Complete the quiz and reflection to continue.
Final Assessment

Lesson 6: Final Assessment

15 questions • 100% required to pass

Bringing It All Together

This lesson followed one request from start to finish: a one-page visual briefing on social connection for a regional loneliness working group, drawn from the 2021 wave of the Canadian Social Connection Survey. Section 1 set out why the analyst looks before modelling. Anscombe's quartet showed four datasets with the same means, standard deviations, correlation and fitted line but completely different patterns, and a decision guide matched the display to the number and types of the variables in a question. A gallery of poor displays showed what a truncated axis, a bar chart of means, a crowded pie chart and an overplotted scatterplot conceal.

Section 2 prepared the briefing file in base R (3,083 people with complete data on nine variables) and used quick base R plots to check it. A histogram revealed a spike of 116 people at the top value of household size, where larger totals had been recorded as 20, and heaping at 40 hours in reported hours of work, problems that the numeric summaries did not show. Section 3 rebuilt the plots with ggplot2, using the layered grammar of data, aesthetics, geoms, scales, facets and themes. A loess smoother showed that loneliness dips around age 30 and rises to a second high point in the fifties, facets compared the relationship between social support and loneliness across gender and age groups, and the Okabe-Ito palette kept the colours readable. Section 4 summarized many variables with a correlation heatmap, explained why a heatmap shows association and cannot show cause, curves or shape, added marginal histograms to a scatterplot, and combined four panels into one figure with a caption.

The same habits apply to every analysis in the rest of the course: each variable is plotted before it is summarized, each relationship is plotted before it is modelled, and each figure carries labels and a caption that let it stand on its own. The next lesson, Measurement and Psychometrics, takes up a question that this lesson left open: whether the scale scores plotted here, such as the loneliness and social support scores, measure what they are meant to measure, and how reliably they do so.

Key Takeaways from this lesson

  • Exploratory analysis looks at the data before modelling, because very different data can share the same summary statistics.
  • The display follows from the question and the variable types: a histogram or bar chart for one variable, a scatterplot or boxplots for two, and a heatmap, scatterplot matrix or facets for many.
  • A bar chart's axis starts at zero, distributions are shown alongside or in place of means, pie charts are kept to a few categories, and overplotting is reduced with jitter and transparency.
  • Base R draws quick diagnostic plots with hist(), barplot(), boxplot(), plot() and pairs(), arranges them with par(mfrow = ...), and saves them with png() and dev.off().
  • Histograms reveal skew, ceilings, heaping, spikes at a top-coded value and impossible values, and each finding is checked against the codebook, handled and reported.
  • A ggplot2 plot is built from data, aesthetic mappings and geoms, with scales, facets and a theme added as layers, and it is saved at a chosen size and resolution with ggsave().
  • Groups are shown with colour for a few groups compared directly, or with facets for many groups, overlapping points or unequal group sizes, using a colour-blind-safe palette.
  • A correlation heatmap summarizes straight-line association among many variables, and it cannot identify causes, detect curved relationships or show the shape of the data.
  • patchwork combines panels into one labelled figure, and a caption states what the figure shows, the data source, the number of people and the meaning of every mark.

The final assessment covers all four sections. All 15 questions must be answered correctly (100%), and the final reflection completed, to finish the lesson.

Reflection

A research question asks whether daily time spent on social media is associated with loneliness among CSCS 2021 respondents. In the data, social media time per day is recorded in six ordered categories (from "Less than 10 minutes per day" to "More than 3 hours per day"); 18 respondents chose "Presented but no response" and 597 have no answer. Loneliness is the three-item UCLA score (whole numbers from 3 to 9, higher is lonelier), and the analyst also has age group (four categories) and gender (woman, man, non-binary; only about 2% of respondents are non-binary). This lesson recommended plotting each variable before summarizing it (a histogram for a numeric variable, a bar chart for a categorical one), checking for skew, ceilings, heaping, spikes at a top value and impossible values, choosing displays by variable type (boxplots with jittered points for a numeric variable across groups, facets or colour to compare groups, a correlation heatmap for many numeric variables), and finishing with a multi-panel figure made with ggplot2 and patchwork, saved with ggsave() and described by a caption that states the data source, the number of people and the meaning of every mark. Describe the exploratory figure you would build for this question: name each panel, the display and the main ggplot2 functions it uses, and what it is meant to show. Then state two data checks you would make before drawing it, and list the elements your caption would include.

Model answerI would build a three-panel figure. Panel A would be a histogram of loneliness with geom_histogram(binwidth = 1), so that each whole score has its own bar and any pile-up at the ceiling of 9 is visible. Panel B would be a bar chart of social media time with geom_bar(), with the six categories kept in their natural order by setting the factor levels in base R first, to show how respondents are spread across the categories. Panel C, the main panel, would show loneliness across the six social media categories with geom_boxplot(outlier.shape = NA) and geom_jitter(width = 0.2, height = 0.2, alpha = 0.1), so that the medians, the spread and the number of people in each category are all visible, and I would add facet_wrap(~ age_group) to see whether the pattern differs by age. I would report gender in a table and leave it out of the facets, because the non-binary panels would hold very few people. I would combine the panels with patchwork as (p_a + p_b) / p_c + plot_annotation(tag_levels = "A"), apply theme_minimal(), and save with ggsave(..., width = 10, height = 8, dpi = 300). Before drawing, I would run table() on social media time, set the 18 "Presented but no response" answers to missing with a bracket recode, and confirm that the six categories appear in their natural order; I would also check that loneliness lies only between 3 and 9 and count how many people na.omit() removes, since 597 people already lack a social media answer. The caption would give the survey and year (CSCS 2021), the number of respondents included and the number excluded for missing data, the measures and their ranges (UCLA 3-item loneliness score, 3 to 9; social media time in six categories), what the boxes, whiskers and jittered points show, and what each panel letter and facet represents.

Minimum 20 characters required.

✓ Reflection saved

Final Knowledge Assessment

Final assessment: the 15 questions

1. Which activity is an example of exploratory analysis?

Exploration asks what the data look like, including whether any values are impossible. The other three options test or estimate a quantity that was set in advance, which is confirmatory work.

2. In one of Anscombe's datasets, ten points share the same x value and one point lies far to the right. The correlation is 0.82. What does this dataset show?

Within the ten points that share an x value there is no relationship to estimate, so the lone point determines the slope and the correlation. Only a plot reveals this.

3. A briefing must compare the distribution of life satisfaction (scored 1 to 10) across four income groups. Which display suits this best?

A numeric variable across groups calls for side-by-side boxplots, and showing the points adds the spread and the group sizes. Bars of means hide the spread, and one combined histogram cannot show group differences.

4. A plot must show whether the percentage of respondents in each self-rated mental health category differs between three provinces of very different sizes. Which display suits this best?

Both variables are categorical. Percentages within each province make the provinces comparable despite their different sizes, whereas raw counts would mainly show which province is largest.

5. A histogram of a symptom score that runs from 0 to 6 shows a taller bar at 6 than at 5. What is the most likely description of this pattern?

A rise at the highest possible score means that a group of respondents gave the most extreme answers and the scale cannot record anything higher, which is a ceiling.

6. Reported cigarettes smoked per day pile up at 10, 20 and 40, with few answers at 9, 11, 19 or 21. What is this pattern called?

Answers that cluster on round or standard numbers (here, half a pack, one pack and two packs) show heaping, which Lesson 2 called digit preference. It usually reflects rounding.

7. Which base R code draws boxplots of loneliness for each gender in the data frame brief?

The formula loneliness ~ gender asks boxplot() for one box of loneliness per gender. barplot(table(...)) counts people in each gender, and hist() does not take a grouping formula.

8. When jitter() or geom_jitter() is used in a scatterplot, what changes?

Jitter adds a small random shift to where each point is drawn so that tied values spread out. The data, and any statistic computed from them, are unchanged.

9. In ggplot(brief, aes(x = support, y = loneliness, colour = gender)) + geom_point(alpha = 0.3), which property is set to a fixed value for every point?

alpha = 0.3 is given outside aes(), so every point has the same transparency. Position and colour are inside aes(), so they are mapped to variables and change from point to point.

10. A data frame has one row per survey respondent and a column for self-rated mental health. Which geom draws a bar chart of the number of respondents in each category?

geom_bar() counts the rows in each category. geom_col() needs the bar heights to be supplied in a column, so it suits a table of computed values.

11. What does facet_grid(gender ~ age_group) produce?

In facet_grid(rows ~ columns), the variable before the tilde defines the rows and the one after it defines the columns, giving one panel per combination.

12. Why would an analyst choose the Okabe-Ito palette for a figure that shows three groups?

The palette was chosen so that people with common forms of colour-vision deficiency, such as red-green deficiency, can tell its colours apart. A legend is still needed.

13. In the briefing file, the Pearson correlation between social support and loneliness is −0.40. Which statement is the best interpretation?

A negative correlation of moderate size means that higher values of one variable tend to go with lower values of the other. It is not a slope, a percentage of people or a proportion of variance explained, and it does not establish cause.

14. Which patchwork expression places plots a and b side by side in a top row, with plot c below them across the full width?

The vertical bar places a and b side by side, the parentheses group them as one row, and the forward slash stacks c beneath that row.

15. Which sentence belongs in the caption of an exploratory figure?

A caption explains what the figure shows and how to read its marks. Causal claims and judgements about importance belong in the results and discussion, and details of effort do not belong in the paper.

✦ Before submitting: pass every section knowledge check (100%) and complete every reflection.

🏆 Congratulations!

The next lesson, Measurement and Psychometrics, asks whether the scale scores plotted in this lesson, such as the loneliness and social support scores, measure what they are meant to measure, and how reliable they are.

You have successfully completed this lesson: Exploratory Data Analysis and Visualization.

Your responses have been downloaded automatically.