Designing Against Bias: Validity and Confounding in Study Protocols
Fundamental Epidemiological Concepts and Approaches
Learning objectives for this lesson:
- Recall the three threats to internal validity from HSCI 230 Lessons 7 to 11, and distinguish a confounder from a mediator and a collider in a described study.
- Compute the odds ratio of the sampling fractions to judge the direction and size of selection bias, and correct an observed odds ratio for assumed sampling fractions.
- Specify protocol decisions on the sampling frame, response maximization and retention that keep selection from depending jointly on exposure and outcome.
- Compute the table that non-differential misclassification produces from sensitivity and specificity, and back-correct an observed table.
- Plan a validation sub-study, and inflate a sample size for the loss of power caused by outcome misclassification.
- Choose among restriction, matching, stratification and adjustment for each confounder, and select an adjustment set from a causal diagram with the change-in-estimate check as support.
- Compute and interpret a Mantel-Haenszel odds ratio, and distinguish confounding from effect modification and from non-collapsibility.
- Plan a sensitivity analysis for unmeasured confounding with external adjustment and the E-value.
- Draft the threats to validity and mitigation subsection of a study protocol, stating for each threat its mechanism, likely direction, prevention, assessment and residual risk.
This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University based on Dohoo, I. R., Martin, S. W., & Stryhn, H. (2012). Methods in Epidemiologic Research. VER Inc.
Glossary: Key Terms, People & Concepts
📚 Reference page, available throughout the lesson
This glossary collects the key concepts, methods and people in this lesson. Use it as a reference while you work through the material, or as a review before assessments. Type in the search box to filter entries.
Validity in the Protocol and Selection Bias
Where this lesson starts
HSCI 230 taught the vocabulary of validity. Its Lessons 7 to 11 defined selection bias, information bias and confounding, named their common forms, and showed how each appears on a causal diagram. In a course on designing research, the vocabulary leads to two practical questions: which decisions in a study protocol prevent each bias, and how the protocol can show, with numbers, how large the bias that remains could be. A protocol often gathers the answers in a subsection titled "threats to validity and mitigation", and this lesson supplies the tools that subsection uses.
Selection bias is the first threat a protocol plans for. Its direction and size can be computed from sampling fractions, it can be bounded with assumed fractions when the true ones are unknown, and it is kept small by decisions about the sampling frame, response and retention. A one-screen recap of all three threats, and a short diagnostic check on the difference between a confounder, a mediator and a collider, come first.
Learning Objectives for this section
- Recall the three threats to internal validity and locate each in HSCI 230 Lessons 7 to 11.
- Distinguish a confounder from a mediator and from a collider in a described study.
- Compute the odds ratio of the sampling fractions, state the direction of the selection bias it implies, and correct an observed odds ratio for assumed sampling fractions.
- Specify protocol decisions on the sampling frame, response maximization and retention that keep selection from depending jointly on exposure and outcome.
Recap: validity and the three threats
A study is internally valid when the association it reports, apart from random error, equals the association in the source population, the population that actually produced the participants. It is externally valid when that association also holds in a target population beyond the source. Bias is a systematic difference between the estimate and the truth, and a larger sample does not remove it. Three threats account for most bias in observational research, and HSCI 230 treated each of them in depth. Table 7.1 is the whole recap. Where a definition is unfamiliar, the HSCI 230 lesson named in the third column is the place to review it.
Table 7.1. The threats to internal validity, how each arises, where HSCI 230 teaches it, and what this lesson adds.
| Threat | How it arises | Where HSCI 230 teaches it | What this lesson adds |
|---|---|---|---|
| Selection bias | Entry into the study, or remaining in it, depends on both exposure and outcome. | Lesson 8 (selection mechanisms, transportability, and missing data from unit non-response and loss to follow-up), Lesson 7 (colliders) and Lesson 11 Section 2 (missing data mechanisms) | Sampling fractions, bounds, and decisions on the sampling frame, response and retention |
| Information bias | Exposure, outcome or covariates are measured with error. | Lesson 9 (misclassification, recall, observer and detection bias), Lesson 7 (measurement) and Lesson 11 Section 2 (item missingness) | The misclassified table, back-correction, validation sub-studies and sample size inflation |
| Confounding | The groups compared differ in other causes of the outcome. | Lesson 11 (confounding, Simpson's paradox) and Lesson 7 (causal diagrams, overadjustment) | Restriction, matching, adjustment sets, Mantel-Haenszel stratification and sensitivity analysis |
| Design-specific and temporal biases | The timing of exposure, outcome and follow-up is mishandled. | Lesson 10 (for example, immortal time and lead time) | Cross-referenced where a protocol has follow-up; Lesson 8 of this course takes up time-to-event data |
The pre-course survey for this course found that most students could define these threats, and that about one student in four still described a variable on the causal pathway as a confounder. The check below targets that confusion. It is ungraded and takes two or three minutes.
Confounder, mediator or collider?
Each item describes a third variable in a study. Choose its role. Feedback appears as soon as you choose. A confounder is a common cause of the exposure and the outcome, a mediator lies on the path from the exposure to the outcome, and a collider is a common effect of the exposure and the outcome (or of their causes).
1. A survey of university students studies loneliness (the exposure) and depressive symptoms (the outcome). Loneliness lowers the social support a student perceives, and low perceived support raises depressive symptoms.
2. In the same survey, financial strain makes students less able to take part in social activities, which raises loneliness, and it also raises depressive symptoms through worry and insecurity. Financial strain is present before either.
3. Lonely students are less likely to complete the survey, and so are students with depressive symptoms. Only students who complete the survey are analyzed.
4. A cohort study examines years of night-shift work (the exposure) and type 2 diabetes (the outcome). Night-shift work raises body mass index, and a higher body mass index raises the risk of diabetes. Body mass index is measured after the shift-work history.
The four items turn on the direction of the arrows in a causal diagram. A confounder is a common cause of the exposure and the outcome and is controlled in the protocol, a mediator lies on the path between them and is left out of the adjustment set for a total effect, and a collider is a common effect, such as completing the survey in the third item, on which an analysis of completers conditions. The next part sets out the protocol decisions that address each threat.
Bias is decided when the protocol is written
Most of the decisions that determine whether a study is biased are made before any data are collected. The sampling frame and the recruitment method decide who can enter the study. The instruments, their mode of administration and the timing of measurement decide how much misclassification there will be. The list of variables to be measured decides which confounders can ever be controlled, because an analysis can adjust only for what was recorded. Restriction and matching are design choices, and stratification and adjustment, although carried out in the analysis, must be planned in the protocol so that the necessary variables are collected and the analysis is not chosen after the results are seen. Table 7.2 pairs each threat with the protocol decisions that prevent it and the analysis steps that check or correct it.
Table 7.2. Protocol decisions and analysis steps for each threat to validity.
| Threat | Decided in the protocol by | Checked or corrected in the analysis by |
|---|---|---|
| Selection bias | The sampling frame, the sampling strategy, the recruitment message, contact and reminder schedules, incentives, and retention procedures | Comparing responders with the frame, weighting, and bounding the bias with assumed sampling fractions |
| Information bias | The choice of instrument, its validity evidence, the mode and timing of measurement, blinding of assessors, and a validation sub-study | Back-correction from sensitivity and specificity, regression calibration, and a sample size that allows for misclassification |
| Confounding | The causal diagram, the list of confounders to measure, restriction, and matching | Stratification, adjustment, the change-in-estimate check, and sensitivity analysis for unmeasured confounders |
The protocol therefore carries a plan for each threat, often written in a subsection titled threats to validity and mitigation. Each section of this lesson ends with the decisions it contributes to that subsection, and Section 4 shows a complete example.
The first threat in such a plan is selection bias. The next part shows how its direction and size can be computed from the sampling fractions of the source population.
Quantifying selection bias with sampling fractions
HSCI 230 described selection bias as conditioning on a common effect. When both the exposure and the outcome influence who takes part, restricting the analysis to participants creates an association between them, or distorts one that exists. Figure 7.1 shows the structure in general and in the study that serves as this lesson's running example, the Campus Connection Study, a worked example that this lesson shares with Lesson 8. That study is a survey of loneliness and depressive symptoms among students aged 18 to 29 at five public post-secondary institutions in British Columbia.
The diagram shows that selection bias is possible. Sampling fractions show how large it is and in which direction it runs. Imagine the source population arranged in the usual two-by-two table, with exposure in the columns and disease in the rows. Upper-case letters denote counts in the source population and lower-case letters denote the counts that end up in the study. Table 7.3 sets out the source population in this form.
Table 7.3. The source population as a two-by-two table of exposure and disease, with counts in upper-case letters.
| Source population | Exposed | Unexposed | Total |
|---|---|---|---|
| Cases | A1 | A0 | M1 |
| Non-cases | B1 | B0 | M0 |
| Total | N1 | N0 | N |
Each cell has a sampling fraction, the proportion of that cell of the source population that entered the study and stayed in it to the analysis. Equation 7.1 defines the four fractions, each as the study count divided by the source count for one cell.
sf21 = b1/B1 sf22 = b0/B0Eq 7.1
Equation 7.2 combines the four fractions into a single quantity, the odds ratio of the sampling fractions (ORsf). The observed odds ratio equals the source-population odds ratio multiplied by the odds ratio of the sampling fractions:
Three consequences follow. When ORsf equals 1, the observed odds ratio is unbiased, even when the four fractions differ from each other. When ORsf is above 1, the observed odds ratio is larger than the true one, which for a harmful exposure is a bias away from the null value of 1. When ORsf is below 1, the observed odds ratio is smaller, which for a harmful exposure is a bias toward the null. The same information can be expressed as sampling odds: the selection odds of cases to non-cases among the exposed (sf11 ÷ sf21) and among the unexposed (sf12 ÷ sf22). There is no selection bias in the odds ratio when these two odds are equal. The result applies exactly to the odds ratio. For a risk ratio it holds approximately when the outcome is rare.
HSCI 341 Lesson 2 Section 3 called selection that depends on the outcome the formal criterion for selection bias, and that criterion holds for a descriptive estimate such as a prevalence, which is biased whenever cases and non-cases take part at different rates. For an odds ratio the criterion is an odds ratio of the sampling fractions different from 1, which Worked Example 7.1 computes: selection that depends on the outcome biases the odds ratio only when it depends on the outcome differently in the exposed and the unexposed.
A low response rate is therefore not, by itself, evidence of selection bias. A meta-analysis of studies that measured both found only a weak relationship between survey response rates and the size of non-response bias (Groves & Peytcheva, 2008). What matters is whether response depends on the exposure and the outcome jointly, which is what ORsf measures.
Worked Example 7.1 applies Equation 7.2 to non-response in a source population of 10,000 people, first when non-response depends on exposure alone and then when it also depends on disease. Calculator 7.1, which follows it, applies Equation 7.2 to any four sampling fractions and an observed odds ratio; its presets include the daycare study of Worked Example 7.2 and the two Campus Connection cases of Worked Example 7.3.
Worked Example 7.1: Non-Response That Depends on Exposure and Outcome
A source population of 10,000 people contains 1,000 exposed people, of whom 250 develop the disease (a risk of 25%), and 9,000 unexposed people, of whom 1,080 develop it (12%). The true risk ratio is 0.25 ÷ 0.12 = 2.08 and the true odds ratio is 2.44.
If 30% of the exposed and 10% of the unexposed fail to respond, and non-response is unrelated to the disease, every risk within an exposure group is unchanged and so are the risk ratio and odds ratio. The fractions are 0.70 for both exposed cells and 0.90 for both unexposed cells, and ORsf = (0.70 × 0.90) ÷ (0.90 × 0.70) = 1.
Suppose instead that, within each exposure group, the risk of disease among non-responders is twice the risk among responders. Among the exposed, the 700 responders then have a risk of 250 ÷ (700 + 2 × 300) = 0.192, and among the unexposed the 8,100 responders have a risk of 1,080 ÷ (8,100 + 2 × 900) = 0.109. The four sampling fractions are 0.538 for exposed cases, 0.754 for exposed non-cases, 0.818 for unexposed cases and 0.911 for unexposed non-cases.
- ORsf = (0.538 × 0.911) ÷ (0.818 × 0.754) = 0.80.
- The observed odds ratio is 2.44 × 0.80 = 1.94, and the observed risk ratio is 0.192 ÷ 0.109 = 1.76, against a true value of 2.08.
Both measures are biased toward the null, because the people missing from the study are disproportionately exposed cases.
Bounding the bias when the fractions are unknown
The true sampling fractions are almost never known, because the source population is not observed cell by cell. The formula is still useful as a sensitivity analysis. The investigator assigns plausible fractions to the four cells, computes ORsf, and divides the observed odds ratio by it to see what the estimate would have been without the selection. Assigning one set of values is a deterministic analysis; drawing the fractions from distributions and repeating the calculation many times is a probabilistic one, which Section 4 describes as part of quantitative bias analysis. Worked Example 7.2 applies a deterministic analysis to a case-control study of daycare attendance and childhood respiratory disease, and the daycare preset of Calculator 7.1 reproduces it.
Worked Example 7.2: Bounding Selection Bias in a Daycare Study
A case-control study of childhood respiratory disease and regular daycare attendance observed an odds ratio of 2.33 (95% CI 1.04 to 5.19). Suppose cases were easier to recruit than controls, and that exposed controls were especially hard to recruit, with assumed sampling fractions of 0.5 for exposed cases, 0.6 for unexposed cases, 0.05 for exposed controls and 0.1 for unexposed controls.
- ORsf = (0.5 × 0.1) ÷ (0.6 × 0.05) = 0.05 ÷ 0.03 = 1.67.
- Corrected odds ratio = 2.33 ÷ 1.67 = 1.40, which is 40% lower than the observed value.
Under these assumptions the association would be considerably weaker than reported. A protocol for such a study would describe how controls are to be recruited so that their participation does not depend on daycare use, and would plan this bounding calculation for the report.
Selection bias can be corrected fully in only two situations. In the first, the factors that drive selection are measured, precede both exposure and outcome, and have a known distribution in the source population, so that they can be handled like confounders, for example by weighting. In the second, the investigator can identify and measure a variable that carries the whole dependence of selection on exposure and outcome (sometimes called a bias breaker) and can obtain unbiased estimates of its distribution in the source population. Both situations require information collected by design, which is why the protocol decisions below include what to record as well as how to recruit.
Protocol decisions that prevent selection bias
The aim of the design is to keep ORsf close to 1: selection may depend on exposure or on outcome, but it should not depend on their combination. The decisions fall into four groups.
The four tabs below take the groups in turn: the sampling frame, response, retention, and what to record. The Response tab opens with Box 7.1, which recalls the strategies for raising response from HSCI 341 Lesson 3.
The sampling frame is the list from which participants are drawn, and it defines who can enter the study. A protocol states the frame, how completely it covers the source population, and how people are selected from it. A frame that is independent of the exposure and the outcome, such as a registrar's enrolment list or a provincial health insurance registry, allows probability sampling, and each person's chance of selection is known. Recruitment through advertisements, social media or clinics produces volunteers whose decision to take part may depend on the study topic, and therefore on the exposure and the outcome. In a case-control study the frame for controls must represent the population that produced the cases, so that controls reflect the exposure distribution of that population. Using incident cases, and drawing controls from the same base as the cases, also keeps selection from depending on survival with the disease.
Box 7.1: Recall: Strategies That Raise Response
HSCI 341 Lesson 3, Section 4 (Response Rates and Data Coding), listed the strategies that raise response, many of them tested in the randomized trials that Edwards et al. (2009) reviewed: repeat contact with non-responders, small incentives, short questionnaires, advance notice and personalized invitations. That lesson stated that low response increases the risk of selection bias. The odds ratio of the sampling fractions refines the statement: low response leaves room for bias, and the odds ratio is biased only when response depends jointly on the exposure and the outcome.
Retrieval question. Which strategy did Lesson 3 single out as the most consistently effective way to raise response?
Repeat contact with non-responders, through reminders and a second delivery of the questionnaire, was the most consistently effective strategy.
Response maximization reduces the room for selection to depend on exposure and outcome. The tailored design method organizes the steps above into a contact schedule (Dillman, Smyth & Christian, 2014). The protocol also controls the recruitment message. An invitation that describes a survey of student wellbeing is less likely to attract or repel students according to their loneliness than one that announces a study of loneliness and depression. The protocol states the expected response rate and the reasoning behind it, because the sample size calculation depends on it, and it states how many contacts each person will receive. It also states the denominator of the response rate it will report, usually completed responses among the eligible people sampled.
In a cohort study, loss to follow-up that depends on both exposure and outcome produces the same bias as differential non-response at entry. Retention procedures are designed in advance: several contact details collected at baseline, including an alternative contact person where ethics approval allows it; a schedule of reminders across modes; low participant burden at each wave; and incentives for completing follow-up, which a systematic review of population-based cohorts associated with higher retention (Booker, Harding & Benzeval, 2011). Equal follow-up procedures for the exposed and the unexposed keep retention from depending on exposure through the study's own actions. A cross-sectional survey has no follow-up, but a multi-page online questionnaire has its own form of attrition in the respondents who stop partway, and the protocol states how partial completions will be handled.
To assess selection afterwards, the protocol records the information that the analysis will need. This includes a flow of participants from the frame through eligibility, invitation, response and completion (the format that STROBE reporting asks for), the reasons for refusal or ineligibility where they are known, and characteristics available for everyone on the frame, so that responders can be compared with non-responders. In a cohort study it includes the baseline characteristics of those later lost, compared with those retained. These data support weighting by the inverse of the probability of response, which HSCI 341 Lesson 2 introduced for design weights, and they inform the sampling fractions assumed in a bounding analysis.
Worked Example 7.3 applies the four groups of decisions to the Campus Connection Study and then uses Equation 7.2 to bound the selection bias that its expected response rate of 15% leaves room for.
Worked Example 7.3: Selection Decisions in the Campus Connection Study
The Campus Connection Study samples students aged 18 to 29 from a frame of about 63,800 students at five institutions, one in each regional health authority, using stratified random sampling with institutions as strata. Invitations are sent by each registrar to institutional email addresses. Recruitment through social media was considered and rejected, because the decision to answer a post about loneliness would depend on both loneliness and depressive symptoms, which makes participation the collider shown in Figure 7.1. The protocol plans 12,200 invitations, on the assumption that 95% of those invited are eligible, 15% respond and 81% of respondents complete the questionnaire, which gives about 1,400 analyzable responses.
A response rate of 15% leaves a great deal of room for selection. The protocol's response decisions are a neutral invitation that describes a survey of student health and connection, three reminders, and a short questionnaire. Its recording decisions are aggregate frame counts by institution, age group and gender from each registrar, so that respondents can be compared with the frame and weighted, and a count of partial completions.
The protocol can then bound the remaining bias. Suppose the probability of completing the survey is 0.18 for students who are neither lonely nor depressed, 0.15 for students who are lonely only, and 0.15 for students with depressive symptoms only. If lonely students with depressive symptoms respond with probability 0.125, the two reductions combine multiplicatively and ORsf = (0.125 × 0.18) ÷ (0.15 × 0.15) = 1.00, so the odds ratio is unbiased despite unequal response. If they respond with probability 0.10, ORsf = (0.10 × 0.18) ÷ (0.15 × 0.15) = 0.80, and an observed odds ratio of 2.0 would correspond to 2.0 ÷ 0.80 = 2.5 without the selection. Calculator 7.1 contains both cases as presets.
Section 1 has shown that selection biases an odds ratio only when participation depends jointly on the exposure and the outcome, that the odds ratio of the sampling fractions measures the size and direction of that bias, and that a protocol keeps it small through its sampling frame, recruitment message, contact schedule, retention procedures and records of the people who do not take part. The key takeaways below summarize these points. The reflection that follows applies the sampling fractions to a study that recruits through social media, and the knowledge check closes the section.
Key Takeaways
- HSCI 230 Lessons 7 to 11 define the three threats to internal validity; this lesson concerns the protocol decisions that prevent them and the calculations that size them.
- A confounder is a common cause of exposure and outcome, a mediator lies between them, and a collider is a common effect, and only the first belongs in an adjustment set for a total effect.
- The observed odds ratio equals the true odds ratio multiplied by the odds ratio of the sampling fractions, (sf11 × sf22) ÷ (sf12 × sf21), so selection biases the odds ratio only when that quantity differs from 1.
- Assumed sampling fractions turn a worry about selection into a bounded estimate, and the protocol should plan that calculation.
- The sampling frame, the recruitment message, contact schedules, incentives, retention procedures and the recording of frame characteristics are the protocol decisions that keep selection from depending jointly on exposure and outcome.
Reflection
A protocol proposes to estimate the association between heavy social media use (the exposure) and poor sleep quality (the outcome) among undergraduates at one university. Participants would be recruited through posts on the university's social media accounts that invite students to take part in a study of social media and sleep. Suppose the probability of taking part is 0.06 for heavy users with poor sleep, 0.04 for heavy users with good sleep, 0.02 for light users with poor sleep and 0.03 for light users with good sleep. The odds ratio of the sampling fractions is (fraction for exposed cases × fraction for unexposed non-cases) ÷ (fraction for unexposed cases × fraction for exposed non-cases), and the observed odds ratio equals the true odds ratio multiplied by it. (1) Compute the odds ratio of the sampling fractions and state the direction of the bias it implies. (2) If the study observed an odds ratio of 2.4, what odds ratio would it have observed without this selection? (3) Propose two changes to the sampling frame or the recruitment method that would reduce the problem, and one piece of information the protocol should record so that selection can be assessed.
Minimum 20 characters required.
1. A study's assumed sampling fractions are 0.5 for exposed cases, 0.6 for unexposed cases, 0.05 for exposed non-cases and 0.1 for unexposed non-cases. What is the odds ratio of the sampling fractions, and what does it imply?
2. A survey achieves a response rate of 15%. Which statement about bias in its odds ratio is correct?
3. In the Campus Connection Study, recruitment through social media posts about loneliness was rejected mainly because:
4. Non-response related to both exposure and outcome gives an odds ratio of the sampling fractions of 0.80, and the study observes an odds ratio of 1.94. Which value estimates the odds ratio that would have been observed without the selection?
5. A cohort protocol expects 25% loss to follow-up over two years. Which plan does most to limit selection bias from that loss?
Misclassification and Validation Sub-Studies
From the test to the study
HSCI 341 Lesson 5 described a measurement in terms of its sensitivity (Se), the probability that a person who truly has a characteristic is classified as having it, and its specificity (Sp), the probability that a person who truly lacks it is classified as lacking it. HSCI 230 Lesson 9 showed the consequence for a study: misclassification that is the same in the groups being compared (non-differential) tends to pull a measure of association toward the null, and misclassification that differs between them (differential) can push it either way. Both statements can be made quantitative, and each calculation corresponds to a protocol decision: the table a study would observe, the correction of an observed table, the validation sub-study that supplies the sensitivity and specificity a correction needs, and a sample size large enough that misclassification does not leave the study underpowered.
Learning Objectives for this section
- Compute the observed two-by-two table that non-differential exposure misclassification produces, and explain why the odds ratio moves toward the null.
- Explain how outcome sensitivity and specificity affect a risk ratio, and why a protocol favours a specific outcome definition.
- Back-correct an observed table from assumed sensitivity and specificity, and explain why the result depends on validation data that transfer to the study population.
- Plan a validation sub-study, including its reference standard, sampling and size, and describe regression calibration.
- Inflate a sample size for the loss of power caused by outcome misclassification.
Computing the misclassified table
The first calculation shows what non-differential misclassification of a binary exposure does to the two-by-two table a study observes. Equation 7.3 states the condition that defines non-differential error.
Misclassification of exposure is non-differential when the sensitivity and specificity of the exposure measurement are the same among cases and non-cases:
Each observed cell is then a mixture of correctly classified people and people who belong in the other exposure column. Among the cases, the observed exposed count contains the truly exposed cases who were correctly classified and the truly unexposed cases who were wrongly classified as exposed. The same mixing happens among the non-cases. Table 7.4 gives each observed count as a mixture of two true counts, weighted by SeE and SpE.
Table 7.4. Observed counts under non-differential exposure misclassification, in terms of the true counts, SeE and SpE.
| True count | Observed count |
|---|---|
| Exposed cases, a1 | a1′ = SeE × a1 + (1 − SpE) × a0 |
| Unexposed cases, a0 | a0′ = (1 − SeE) × a1 + SpE × a0 |
| Exposed non-cases, b1 | b1′ = SeE × b1 + (1 − SpE) × b0 |
| Unexposed non-cases, b0 | b0′ = (1 − SeE) × b1 + SpE × b0 |
Because the mixing blends the exposed and unexposed columns in the same proportions for cases and non-cases, the two columns become more alike and the odds ratio moves toward 1, provided SeE + SpE is greater than 1. Misclassifying exposure moves people between the exposure columns and leaves the case and non-case totals unchanged, which is the property that makes back-correction possible. Error rates that look modest, of 10% to 20%, can attenuate an association considerably. Worked Example 7.4 shows the size of that attenuation in a study whose true table is known, and Calculator 7.2 reproduces the example, with a second preset for better measurement, so that other values of SeE and SpE can be tried.
Worked Example 7.4: The Table a Study Would Observe
A study has 90 exposed and 70 unexposed cases, and 210 exposed and 630 unexposed non-cases, so the true odds ratio is (90 × 630) ÷ (70 × 210) = 3.86. Exposure is measured with SeE = 0.80 and SpE = 0.90 in both groups.
- Observed exposed cases: a1′ = 0.80 × 90 + 0.10 × 70 = 79.
- Observed unexposed cases: a0′ = 0.20 × 90 + 0.90 × 70 = 81.
- Observed exposed non-cases: b1′ = 0.80 × 210 + 0.10 × 630 = 231.
- Observed unexposed non-cases: b0′ = 0.20 × 210 + 0.90 × 630 = 609.
- Observed odds ratio = (79 × 609) ÷ (81 × 231) = 2.57.
Measurement that misses one exposed person in five and wrongly labels one unexposed person in ten has reduced the odds ratio from 3.86 to 2.57. The case total (79 + 81 = 160) and the non-case total (231 + 609 = 840) are unchanged.
Misclassification of the outcome
Outcome misclassification has an asymmetry that matters for design. Suppose the true risk of disease is 0.10 among the exposed and 0.05 among the unexposed, a risk ratio of 2.0. If the outcome measure misses some cases (sensitivity 0.80) but never labels a healthy person as a case (specificity 1), both observed risks are multiplied by 0.80, giving 0.08 and 0.04, and the risk ratio stays at 2.0. If instead the measure finds every case (sensitivity 1) but wrongly labels 5% of healthy people as cases (specificity 0.95), the observed risks become 0.145 and 0.0975, and the risk ratio falls to 1.49. False positives add roughly the same number of spurious cases to each group, which dilutes a ratio, and the dilution is greatest when the outcome is rare. A protocol therefore favours a highly specific outcome definition, for example a confirmed diagnosis or a validated case-finding algorithm, even at some cost in sensitivity.
In a case-control study the same logic applies to the case definition: false-positive cases bias the odds ratio toward the null, so diagnoses are confirmed before enrolment. Sensitivity and specificity estimated in the general population cannot be used directly to correct a case-control study, because sampling cases and controls at different rates changes the proportions of true and false positives among those enrolled.
Differential misclassification
When the accuracy of the exposure measurement differs between cases and non-cases, or the accuracy of the outcome measurement differs between exposed and unexposed people, the misclassification is differential:
Equation 7.4 states the condition for differential exposure misclassification. The resulting bias can run in either direction and its size cannot be predicted without knowing the error rates in each group. Recall bias, in which cases search their memories for exposures more thoroughly than controls, is the familiar example, and HSCI 230 Lesson 9 described observer and detection bias as others. Differential error cannot be corrected by a single pair of sensitivity and specificity values, so the protocol concentrates on preventing it. Figure 7.2 contrasts the two kinds of misclassification, and Table 7.5 lists the protocol decisions that prevent differential error.
Table 7.5. Protocol decisions that prevent differential misclassification.
| Protocol decision | What it prevents |
|---|---|
| Measure the exposure before the outcome occurs, or take it from records made at the time | Recall that depends on outcome status |
| Use the same instrument, mode, wording and order of questions for every participant | Accuracy that differs by group because the measurement differs |
| Keep interviewers and outcome assessors unaware of exposure or case status, and train them to a written protocol | Observer bias and differential probing |
| Apply the same schedule of outcome assessment to exposed and unexposed participants | Detection bias from closer monitoring of one group |
| Describe the study to participants in neutral terms | Reporting shaped by knowledge of the hypothesis |
The second decision in Table 7.5 holds the mode of administration constant. Box 7.2 gives background on modes of administration and on mode effects, and explains why a protocol uses one mode for every participant.
Box 7.2: Background: Mode of Administration and Mode Effects
The mode of administration is the way questions reach a respondent: on the web, on paper, by telephone or in person. HSCI 341 Lesson 3, Section 1 (Planning a Questionnaire), compared these methods by response rate, cost and the potential for interviewer bias. A mode effect is a difference in survey results that arises from the mode by which the data were collected, and it has two components. A selection effect concerns who answers by each mode: people who respond online can differ from people who respond by telephone in age, health and the outcomes under study, so a comparison across modes can mix different groups of people. A measurement effect concerns how the mode changes the answers of the same person: self-completed modes, on paper or online, tend to produce greater disclosure of sensitive behaviour, such as substance use or symptoms of depression, than interviewer-administered modes, in which the wish to give a socially acceptable answer is stronger.
The measurement effect is the reason a protocol uses one mode for every participant. When a study must offer more than one mode, it records the mode of each response so that results can be compared by mode in the analysis. HSCI 207 Lesson 8, Section 2.1 (Mode Effects), treats the topic in more detail and is optional reading.
Non-differential misclassification of a binary exposure moves the odds ratio toward the null by an amount that Table 7.4 allows the analyst to compute, while differential misclassification has to be prevented by design. The calculation in Table 7.4 can also be run in reverse, which the next part shows.
Back-correcting an observed table
When the sensitivity and specificity of the exposure measurement are known, or can be assumed, the true table can be recovered from the observed one. Because misclassifying exposure leaves the non-case total m0 unchanged, the true count of exposed non-cases is:
Equation 7.5 recovers the exposed non-cases. The same reasoning applied to the case total m1 gives Equation 7.6 for the exposed cases.
The unexposed counts follow by subtraction from the totals. Applied to the observed table of Worked Example 7.4 with the true values SeE = 0.80 and SpE = 0.90, the formulas return the original 90, 70, 210 and 630, and the odds ratio of 3.86. Applied to the same observed table with SeE = 0.90 and SpE = 0.95, they return an odds ratio of 3.03. The corrected estimate depends heavily on the assumed error rates, and for some combinations the formulas give a zero or negative cell, which shows that the assumed values are incompatible with the data. For these reasons the correction is reported over a range of plausible values, and those values come from a validation study in a population like the study population. Calculator 7.3 contains both versions of the example.
A back-corrected table is only as accurate as the sensitivity and specificity used to compute it. The next part describes how a protocol plans the validation sub-study that supplies those values.
Planning a validation sub-study
A validation sub-study measures a subset of participants twice: once with the study's instrument and once with a reference standard that is accepted as correct or close to it. It supplies the sensitivity and specificity (or, for a continuous measure, the relationship between measured and true values) that a correction needs. Two directions of estimation should be kept apart. A validation estimate runs from truth to measurement, for example the probability of being classified as exposed among those who are truly exposed, which is the sensitivity. A correction runs from measurement to truth, for example the probability of being truly exposed among those classified as exposed, which is a predictive value and depends on prevalence as well as on the error rates. Equation 7.7 sets the two directions side by side.
A protocol plans the sub-study in five steps. The accordion below sets out each step, from the choice of a reference standard to the use of the results.
The reference standard is the best feasible measurement of the same construct: a clinical or semi-structured diagnostic interview for a mental health screen, pharmacy dispensing records for self-reported medication use, a biomarker for self-reported smoking, or a chart review for an administrative case definition. It should be applied by people who do not know the result of the study instrument, and close enough in time that the true state has not changed.
An internal sub-study validates the measure in a subset of the study's own participants. An external estimate is taken from a published validation in another population. External estimates cost nothing but assume transportability, that is, that sensitivity and specificity are the same in the study population, which can fail when age, language, setting or the prevalence of related conditions differ. Mental health diagnoses identified from physician billing records, for example, tend to have lower sensitivity than chronic physical conditions, because many people with these conditions never generate a billing record for them. An internal sub-study removes the transportability assumption at the cost of extra measurement.
A simple random subsample of participants gives sensitivity and specificity directly. A subsample stratified by the study measurement, for example every screen-positive participant and a random tenth of the screen-negatives, is more efficient when the condition is uncommon, but the estimates must then be weighted by the inverse of the sampling fractions, or they will be distorted by verification bias. When misclassification may be differential, the sub-study samples cases and non-cases (or exposed and unexposed participants) separately, so that sensitivity and specificity can be estimated in each group. This is the two-stage design described in HSCI 341 Lesson 2.
Sensitivity is estimated only among people who truly have the condition, so the size of the sub-study is driven by how many of them it will contain (Buderer, 1996). To estimate an expected sensitivity Se within a margin d with 95% confidence, the sub-study needs Z2 × Se × (1 − Se) ÷ d2 true positives, with Z = 1.96, and the number of participants to verify is that number divided by the expected prevalence of the condition among them. The same calculation with Sp and (1 − prevalence) gives the number needed for specificity, and the protocol takes the larger of the two.
The protocol states which estimate will be corrected, by which method (back-correction of the table for a categorical variable, regression calibration for a continuous one), and how the uncertainty in the validation estimates will be carried into the corrected result, for example by repeating the correction across the confidence limits of sensitivity and specificity. It also states what will be reported if the sub-study cannot be completed, usually a correction across a range of published values.
Worked Example 7.5 applies the five steps to the PHQ-2 in the Campus Connection Study, including the sample size calculation of the fourth step.
Worked Example 7.5: Validating the Outcome Measure in the Campus Connection Study
The Campus Connection Study measures depressive symptoms with the two-item Patient Health Questionnaire (PHQ-2). Against semi-structured diagnostic interviews, a pooled analysis found a sensitivity of 0.72 and a specificity of 0.85 at the standard cut-off (Levis et al., 2020). Those values come from many populations, and their transportability to students aged 18 to 29 at five British Columbia institutions is an assumption. An internal sub-study would invite a random subsample of respondents to a semi-structured interview by video within two weeks of the survey, conducted by trained interviewers who do not see the PHQ-2 responses.
To estimate a sensitivity near 0.72 within ±0.15, the sub-study needs 1.962 × 0.72 × 0.28 ÷ 0.152 = 34.4 participants with depression. If the protocol assumes, for planning, that 15% of interviewed respondents meet the interview criteria, it needs 34.4 ÷ 0.15 = 230 interviews (rounded up). Those 230 interviews would contain about 196 participants without depression, enough to estimate a specificity near 0.85 within about ±0.05. Narrowing the margin for sensitivity to ±0.10 would require about 517 interviews, which shows why the precision of sensitivity drives the cost when the condition is uncommon. The protocol states the planning prevalence, the margin it can afford, and the analysis that will use the results.
Regression calibration in brief
For a continuous exposure measured with error, such as dietary intake from a food-frequency questionnaire, the analogue of back-correction is regression calibration (Rosner, Willett & Spiegelman, 1989). In the validation sub-study, the reference measurement is regressed on the study measurement (and on any other covariates in the main model). The fitted equation is then used to predict a calibrated exposure value for every participant in the main study, and the outcome model is fitted with the calibrated values in place of the measured ones. The slope moves away from the attenuated value toward the true association, and its standard error must be adjusted for the uncertainty in the calibration step, for example by bootstrapping. The method assumes that the measurement error is non-differential and that the calibration equation transfers from the sub-study to the whole sample. A protocol that plans regression calibration therefore plans an internal validation sub-study with both measurements, and HSCI 410 provides the regression skills to carry it out.
Validation and correction address the bias that misclassification causes. Misclassification also reduces power, and the last part of this section turns to the sample size a protocol needs when its outcome measure is imperfect.
Inflating the sample size for misclassification
Sample size formulas assume that the proportions entered into them are the proportions the study will observe. With an imperfect outcome measure, the observed proportions are:
p2′ = Se × p2 + (1 − Sp) × (1 − p2)Eq 7.8
With non-differential error the difference p1′ − p2′ is smaller than p1 − p2, so a study sized with the true proportions would be underpowered. Box 7.3 recalls how HSCI 341 Lesson 2 computed the sample size for comparing two proportions.
Box 7.3: Recall: Sample Size for Comparing Two Proportions
HSCI 341 Lesson 2, Section 5 (Planning Sample Size), computed the sample size for comparing two proportions from four inputs: the Type I error (α, usually 0.05, the chance of declaring a difference that does not exist), the power (1 − β, usually 0.80, the chance of detecting a difference of the planned size), and the two proportions expected in the groups being compared. The closer the two proportions, the larger the sample needed. HSCI 230 Lesson 11, Section 2, described the same two errors from the reader's side, when a published study is appraised.
Retrieval question. An underpowered study finds no statistically significant difference, and its confidence interval runs from a small protective effect to a large harmful one. What can it conclude?
It cannot conclude that there is no difference. The interval shows that the data are compatible with no effect and with effects large enough to matter, so the question remains open, and a Type II error is a real possibility.
Misclassification enters through the last of those inputs, so the protocol computes the observed proportions first and enters them into the two-proportion formula. Worked Example 7.6 applies Equation 7.8 to a cohort comparison and to the Campus Connection Study, and Calculator 7.4 reproduces both cases, with a third preset for a perfect test.
Worked Example 7.6: Two Sample Size Inflations
A cohort comparison. A study is planned to detect a difference between true outcome proportions of 0.15 and 0.10 with 95% confidence and 80% power, which requires 685 participants per group. If the outcome test has Se = 0.85 and Sp = 0.95, the observed proportions are 0.85 × 0.15 + 0.05 × 0.85 = 0.170 and 0.85 × 0.10 + 0.05 × 0.90 = 0.130, and the requirement rises to 1,249 per group, 1.8 times as many.
The Campus Connection Study. The protocol's planning values are positive-screen proportions of 0.34 among lonely students and 0.26 among others, which need 514 per group. Those values came from studies that used screening questionnaires, so they already describe screen-positive status. Suppose instead that the protocol defined its outcome as depression itself and treated 0.34 and 0.26 as true proportions, measured with the PHQ-2 (Se = 0.72, Sp = 0.85). The observed proportions would be 0.344 and 0.298, the observed prevalence ratio would fall from 1.31 to 1.15, and the requirement would rise to 1,643 per group, 3.2 times as many. With a specificity of 0.95 the requirement would be 1,025 per group. The example shows why a protocol defines its outcome as the construct its instrument measures, and why specificity matters as much as sensitivity when the outcome is common.
Two further inflation factors appear in protocols. The first is the design effect from unequal weighting. When respondents carry weights that vary, because strata were sampled at different rates or because the weights were adjusted for non-response as in the Campus Connection Study, the effective sample size is smaller than the number of respondents. Kish's approximation gives this design effect as 1 + CV², where CV is the coefficient of variation of the weights, and the required sample size is multiplied by it. The second is the variance inflation factor, 1 ÷ (1 − R²), where R² is the proportion of the variance of the exposure explained by the other covariates in a regression model. It is the factor in the multicollinearity formula of HSCI 230 Lesson 11 Section 2, and it applies only when a regression-adjusted analysis, taught in HSCI 410, is planned: adjustment for covariates that are correlated with the exposure inflates the variance of the exposure coefficient by this factor, so the sample size is multiplied by it.
Section 2 has shown that misclassification affects a study in two ways. It biases the estimate, which the protocol limits by measuring in the same way in every group and assesses with back-correction and a validation sub-study, and it reduces power, which the protocol offsets by sizing the study with the observed proportions. The key takeaways below summarize these calculations. The reflection that follows applies the sample size inflation and the planning of a validation sub-study to a screen for generalized anxiety, and the knowledge check closes the section.
Key Takeaways
- Non-differential misclassification of a binary exposure mixes the exposure columns in the same proportions for cases and non-cases, which moves the odds ratio toward 1 (on average) and leaves the case and non-case totals unchanged.
- For a risk ratio, imperfect outcome sensitivity with perfect specificity leaves the ratio unbiased, while false positives dilute it, so protocols favour specific outcome definitions.
- Differential misclassification can bias in either direction, and the protocol prevents it through the timing, standardization and blinding of measurement.
- Back-correction recovers the true table from assumed sensitivity and specificity, and the result is only as good as the validation data behind those values.
- A validation sub-study needs a reference standard, a sampling plan, a size driven by the number of true positives, and a stated use for its results.
- Misclassification brings the observed proportions closer together, so the sample size is calculated from the observed proportions p′.
Reflection
A protocol will compare the proportion of students with a positive screen for generalized anxiety between students who live in residence and students who live with family. The planning values are true proportions of 0.30 and 0.20. The screening scale has a sensitivity of 0.80 and a specificity of 0.85 against a diagnostic interview, and its errors are expected to be the same in both groups. The observed proportion is p′ = Se × p + (1 − Sp) × (1 − p). With perfect classification, 293 students per group would give 80% power at the 5% significance level; with the observed proportions, the same two-proportion formula requires 797 per group. (1) Compute the two observed proportions and explain why the required sample size increases. (2) The investigators consider an internal validation sub-study with a diagnostic interview as the reference standard. The number of participants with the condition needed to estimate sensitivity within a margin d with 95% confidence is 1.96² × Se × (1 − Se) ÷ d². How many interviews are needed to estimate a sensitivity of about 0.80 within ±0.10 if 25% of interviewed students are expected to have the condition? (3) Name two other features of the validation sub-study that the protocol should specify.
Minimum 20 characters required.
1. A study's true table has 90 exposed and 70 unexposed cases. Exposure is measured with a sensitivity of 0.80 and a specificity of 0.90 among cases and non-cases alike. How many cases will be observed as exposed?
2. Why does non-differential misclassification of a binary exposure move the odds ratio toward 1 (when Se + Sp exceeds 1)?
3. True risks of disease are 0.10 among the exposed and 0.05 among the unexposed. The outcome measure detects 80% of true cases and never labels a healthy person as a case. What risk ratio will the study observe?
4. A validation sub-study aims to estimate a sensitivity of about 0.80 within ±0.10 with 95% confidence. Using n = Z² × Se × (1 − Se) ÷ d², how many participants who truly have the condition must it include?
5. An outcome test with a sensitivity of 0.85 and a specificity of 0.95 will be used to compare true proportions of 0.15 and 0.10. What happens to the planned sample size?
Controlling Confounding by Design and Stratification
Confounding as a design problem
HSCI 230 Lessons 7 and 11 defined confounding, showed it on causal diagrams, and illustrated Simpson's paradox, in which an association reverses when a third variable is taken into account. In a protocol, the control of confounding is a sequence of decisions. The investigator first decides which variables are confounders, using a causal diagram. For each one, the protocol then states whether it will be controlled by restriction or matching at the design stage, or measured and controlled by stratification or adjustment in the analysis. The stratified analysis that this lesson carries out by hand, the Mantel-Haenszel estimator, completes the sequence, together with two phenomena that can be mistaken for confounding in a stratified table: effect modification and the non-collapsibility of the odds ratio.
Learning Objectives for this section
- Define a confounder as a common cause of the exposure and the outcome, apply the classical criteria that follow from that definition, and keep mediators and colliders out of an adjustment set.
- Choose among restriction, matching, stratification and adjustment for each confounder, and state the costs of each choice.
- Select an adjustment set from a causal diagram, and use the change-in-estimate check on the log scale as a supporting step.
- Compute a Mantel-Haenszel odds ratio and a homogeneity test, and distinguish confounding from effect modification and from non-collapsibility.
Confounders in brief
A confounder is a common cause of the exposure and the outcome, or a proxy for such a cause, that is not on the causal pathway from the exposure to the outcome. This is the definition the diagnostic check in Section 1 used. The three classical conditions follow from it. A confounder is associated with the outcome independently of the exposure, because it causes the outcome or is a proxy for a cause. It is associated with the exposure in the source population, because it causes the exposure or shares a cause with it. And it is not caused by the exposure, so it does not lie on the path from exposure to outcome, and it is not a consequence of the outcome. A mediator satisfies the first two conditions and fails the third. The three conditions are necessary but not sufficient. A variable measured before the exposure that is caused by two unmeasured factors, one that also causes the exposure and one that also causes the outcome, meets all three, yet it is a collider, and adjusting for it opens a non-causal path between the exposure and the outcome (M-bias). The causal diagram, which shows the direction of each arrow, settles the question. Two further distinctions affect what a protocol states. A population confounder is a variable known from previous research to confound the association in the target population, and the protocol plans to control it whatever the study data show. A sample confounder appears to confound in the study data only, and the protocol states in advance how such variables will be judged, so that the adjustment set is not built from every variable that happens to change the estimate.
Worked Example 7.7 introduces a study of carriage of Streptococcus pneumoniae and childhood respiratory disease in which a measured variable, RSV infection, meets these conditions.
Worked Example 7.7: STREP, Childhood Respiratory Disease and RSV
A study of 400 children examines the association between carriage of Streptococcus pneumoniae (STREP) and childhood respiratory disease (CRD). Respiratory syncytial virus (RSV) infection raises the risk of CRD and is more common among STREP carriers. The same data are used throughout this section and in Section 4.
| Children | CRD | STREP+ | STREP− | Odds ratio |
|---|---|---|---|---|
| All | Yes | 70 | 30 | 5.44 (crude) |
| No | 90 | 210 | ||
| RSV+ | Yes | 52 | 11 | 3.55 |
| No | 36 | 27 | ||
| RSV− | Yes | 18 | 19 | 3.21 |
| No | 54 | 183 |
Among children without CRD, RSV is present in 36 of 90 STREP carriers (40%) and in 27 of 210 non-carriers (13%), so RSV is associated with the exposure. Within each RSV stratum the odds ratio lies between 3.2 and 3.6, well below the crude value of 5.44, so part of the crude association reflects RSV.
In the STREP example, RSV is associated with both the exposure and the outcome, and the crude odds ratio overstates the association found within each RSV stratum. Once a protocol has identified such a confounder, it has to decide how to control it, and the next part sets out the four approaches available.
Four ways to control a confounder
A protocol chooses, for each confounder, one of four approaches or a combination of them. The first two act on who is studied; the last two act on how the data are analyzed and require that the confounder be measured well. Table 7.6 compares the four approaches by how each works, what it costs and what the protocol states about it.
Table 7.6. Four approaches to controlling a confounder.
| Approach | How it works | What it costs | What the protocol states |
|---|---|---|---|
| Restriction | Only one level of the confounder is eligible, so it cannot vary. | Smaller eligible population; results apply only to that level. | The eligibility criterion and the reason for it |
| Matching | Comparison groups are made alike on the confounder at recruitment. | Recruitment effort; in case-control studies, a matched analysis is required and the matched variable cannot be studied. | The matching variables, the ratio, and the matched analysis |
| Stratification | The association is estimated within levels of the confounder and pooled. | Strata thin out quickly when several confounders are combined. | The strata, the pooled estimator and the homogeneity test |
| Regression adjustment | The confounder enters a model for the outcome. | Model assumptions; taught in HSCI 410. | The adjustment set and the model, named in advance |
Restriction
Restriction prevents confounding by admitting only one level of the confounder. Restricting a study of cervical cancer to women removes confounding by gender completely. Restricting a study to people aged 60 to 69 reduces confounding by age without removing it, because age still varies within the band. When the confounder is binary, admitting the low-risk level is usually preferred, because the exposure effect is then easier to detect and effect modification by the restricted variable is avoided. In the STREP example, restricting to children without RSV gives an odds ratio of 3.21. That estimate is free of confounding by RSV and describes only children without RSV, which is the cost of restriction: the protocol trades the breadth of its conclusions for their internal validity.
Matching
Matching is the second of the two design-stage approaches in Table 7.6. Box 7.4 recalls how HSCI 230 introduced matching and overmatching.
Box 7.4: Recall: Matching and Overmatching
HSCI 230 Lesson 4, Section 4 (Comparability, Analysis and Reporting), introduced matching as a way to make controls comparable to cases, and noted that a matched design requires a matched analysis. Its glossary defined overmatching as matching on a variable on the causal pathway, an effect of the exposure, or a variable unrelated to confounding, which reduces efficiency or biases the estimate toward the null. The paragraphs below develop both ideas by design. The Section 3 reflection includes a retrieval question on overmatching.
Matching makes the distribution of a confounder the same in the groups being compared. Its consequences differ by design. The three tabs below describe matching in cohort studies, in case-control studies and in matched pairs, and the third tab gives the matched odds ratio and McNemar's test in Equation 7.9.
In a cohort study, each exposed person is matched to one or more unexposed people with the same value of the confounder. Matching happens before the outcome occurs, so it makes the exposure independent of the matched variable in the study and removes confounding by it. The matched variable can still affect the outcome, and it does so equally in both exposure groups, so the crude comparison of the matched cohort is unconfounded by it. A stratified or matched analysis can still improve precision.
In a case-control study, controls are matched to cases, after the outcome has occurred. Matching makes the controls resemble the cases on the matched variable, and because that variable is associated with the exposure, it also makes the controls' exposure distribution resemble the cases'. This is a selection bias, usually toward the null, and it is removed only by a matched or stratified analysis that includes the matched variable. The protocol therefore states the matched analysis together with the matching. Matching on a variable that is associated with the exposure but is not a confounder, or on a consequence of the exposure, is overmatching: it costs precision, or introduces bias, without removing any confounding. The protocol matches only on variables that the causal diagram identifies as confounders.
Frequency matching makes the overall distribution of the confounder the same in the two groups and is analyzed by stratification, which allows effect modification by the matched variable to be examined. Pair matching links each case to one or more specific controls and is analyzed by a matched-pair method. With one control per case, only the discordant pairs carry information, because a pair in which both members are exposed, or both unexposed, says nothing about a difference in exposure between cases and controls. Equation 7.9 computes the matched odds ratio and McNemar's test from the counts of the two kinds of discordant pair.
For example, if 40 pairs have an exposed case and an unexposed control and 16 pairs have the reverse, the matched odds ratio is 40 ÷ 16 = 2.5 and McNemar's χ² = (40 − 16)² ÷ 56 = 10.3 on one degree of freedom. Little precision is gained beyond about four controls per case.
Restriction and matching act on who is studied, and both depend on knowing which variables are confounders, since matching on a variable that is associated with the exposure but is not a confounder is overmatching. The causal diagram supplies that knowledge, and the next part shows how it determines the adjustment set.
Choosing the adjustment set from a causal diagram
A causal diagram, or directed acyclic graph (DAG), records the investigator's assumptions about what causes what. HSCI 230 Lesson 7 taught how to draw one and how conditioning on a collider or a mediator biases an estimate. For a protocol, the diagram has a specific job: it determines the adjustment set, the variables that must be measured and controlled to estimate the effect of the exposure, using the backdoor criterion (Greenland, Pearl & Robins, 1999).
- Draw the diagram from subject-matter knowledge before seeing the study data, including unmeasured variables.
- Set aside every arrow that leaves the exposure, and every variable that the exposure causes.
- List the remaining paths that connect the exposure to the outcome. Each one that begins with an arrow into the exposure is a backdoor path.
- Choose a set of measured variables that blocks every backdoor path, without including colliders on those paths (unless the path is also blocked by another variable in the set) and without including descendants of the exposure.
Figure 7.3 presents the backdoor criterion as a story in seven scenes, from the causal front door to the closing of the backdoor by conditioning on the confounder. Figure 7.4 then applies the four steps to a simplified causal diagram for the Campus Connection Study.
Watch the front door (causal) and back door (confounded) of a DAG, then close the backdoor by conditioning. Next ▶ advances scenes.
Figure 7.3. A seven-scene visualization of Pearl's backdoor criterion: the causal front door from exposure to outcome, the spurious backdoor through a confounder, and conditioning on the confounder as the door closing.
In this simplified diagram there are two backdoor paths from loneliness to depressive symptoms, one through financial strain and one through gender, and the adjustment set {financial strain, gender} blocks both. Perceived social support is caused by loneliness, so it is excluded: adjusting for it would estimate only the part of the effect that does not pass through support. Completing the survey is a collider; it cannot be adjusted away, and Section 1 dealt with it as selection bias. Living arrangement affects loneliness and, in this diagram, affects depressive symptoms only through loneliness, so it opens no backdoor path and is not needed for control of confounding. If the investigators believed that living arrangement also affected depressive symptoms directly, they would draw that arrow and add the variable to the set. The protocol's real diagram contains more variables than this one. Its value lies in making each of these judgements visible to reviewers before the data are collected.
Change-in-estimate as a supporting check
The change-in-estimate approach compares the crude odds ratio (ORc) with the odds ratio adjusted for a candidate variable (ORa). By the common convention, a change of more than 10 percent is taken as evidence that the variable confounds the association, and a protocol that uses another threshold names it in advance as a choice. Three rules keep the check honest. The crude value is the baseline. For ratio measures the change is computed on the log scale, (ln ORc − ln ORa) ÷ ln ORc, because a halving and a doubling are changes of the same size on that scale. For a logistic model this is the percentage change in the exposure coefficient, because the coefficient is the log odds ratio, and it is the form that HSCI 410 Lesson 3 uses. And the check is applied only to variables that the diagram allows to be confounders, because adjusting for a mediator or a collider also changes the estimate. In the STREP example the Mantel-Haenszel estimate in Worked Example 7.8 is 3.36, and the change from the crude value is (ln 5.44 − ln 3.36) ÷ ln 5.44 = 28% on the log scale, so RSV meets the criterion and the diagram supports it as a confounder. A protocol names the diagram as the primary basis for the adjustment set, states the threshold it will use for variables whose role is uncertain, and fixes both before the analysis.
The change-in-estimate check needs an adjusted estimate to compare with the crude one. For one or two categorical confounders that estimate usually comes from a stratified analysis, which the next part carries out by hand.
Stratified analysis with the Mantel-Haenszel estimator
The Mantel-Haenszel procedure stratifies the data by the levels of the confounder, computes the odds ratio within each stratum, and pools the strata into one adjusted estimate with a test of whether that estimate differs from 1 (Mantel & Haenszel, 1959). In practice it is accompanied by a test of whether the stratum-specific values are similar. It is the analysis that a protocol usually plans when it has one or two categorical confounders. Equation 7.10 gives the odds ratio within one stratum, Equation 7.11 pools the strata into the Mantel-Haenszel estimate, Equation 7.12 tests whether the stratum-specific values are similar, and Equation 7.13 tests whether the pooled estimate differs from 1.
Worked Example 7.8 applies Equation 7.11 to the two RSV strata of Worked Example 7.7. Calculator 7.5 reproduces the calculation, reports the homogeneity test of Equation 7.12, and has a second preset in which the strata differ.
Worked Example 7.8: The Mantel-Haenszel Odds Ratio for STREP and CRD
The RSV+ stratum has 126 children and the RSV− stratum has 274.
- Numerator: (52 × 27) ÷ 126 + (18 × 183) ÷ 274 = 11.143 + 12.022 = 23.165.
- Denominator: (11 × 36) ÷ 126 + (19 × 54) ÷ 274 = 3.143 + 3.745 = 6.887.
- ORMH = 23.165 ÷ 6.887 = 3.36.
The stratum odds ratios (3.55 and 3.21) are similar, and the homogeneity test in Calculator 7.5 gives no evidence that they differ, so a single pooled value is appropriate. The pooled estimate is 28% lower than the crude value on the log scale, so RSV confounds the crude association, and about 3.4 is the better estimate of the association between STREP and CRD.
Interactive 7.1 simulates a study of 2,000 people in which a third variable, C, raises the risk of the outcome. Its sliders change the strength of confounding and of effect modification, so that the crude, stratum-specific and Mantel-Haenszel odds ratios can be compared as each changes.
📊 Interactive 7.1: Mantel-Haenszel Stratified Analysis
Crude (unstratified) 2×2
| Y+ | Y− | Total | |
|---|---|---|---|
| E+ | 281 | 536 | 817 |
| E− | 198 | 985 | 1183 |
Stratum 1: C+
| Y+ | Y− | |
|---|---|---|
| E+ | 239 | 278 |
| E− | 145 | 338 |
Stratum 2: C−
| Y+ | Y− | |
|---|---|---|
| E+ | 42 | 258 |
| E− | 53 | 647 |
Crude vs. stratum-specific vs. MH-adjusted OR
Effect modification
Effect modification is present when the association between exposure and outcome differs across levels of a third variable. It is a property of the causal system to be described, while confounding is a bias to be removed. Whether it is present depends on the scale of measurement. The three cards below define interaction on the additive scale and on the multiplicative scale, and the two terms for a joint effect that is greater or smaller than the sum or the product of the separate effects.
On the multiplicative scale, effect modification appears in a stratified analysis as stratum-specific odds ratios that differ. Box 7.5 states what a protocol does in that case.
Box 7.5: When the strata differ
When the stratum-specific odds ratios differ more than chance would explain, a single Mantel-Haenszel summary hides the difference, and the stratum-specific estimates are reported instead. A protocol names in advance the variables it will examine as effect modifiers and the reason for each, because a homogeneity test on many unplanned variables will find apparent modification by chance, and because tests of homogeneity have low power, so a study designed to detect modification needs a larger sample than one designed to detect a main effect.
Non-collapsibility of the odds ratio
The odds ratio has a property that can mimic confounding. Even when a third variable is unrelated to the exposure, so that it cannot confound, the crude odds ratio can differ from the stratum-specific odds ratios whenever that variable affects the outcome and the outcome is common (Greenland, Robins & Pearl, 1999). Suppose 100 exposed and 100 unexposed people are studied in one stratum, with 80 and 50 cases, and 90 exposed and 90 unexposed in a second stratum, with 30 and 10 cases. The exposure is equally common in both strata, so the stratifying variable is unrelated to it. The odds ratio is 4.0 in each stratum and the Mantel-Haenszel estimate is 4.0, but the crude odds ratio from the combined table (110 of 190 exposed and 60 of 190 unexposed are cases) is 2.98. The change-in-estimate rule would flag a confounder where there is none. Two checks separate the explanations. The first asks whether the third variable is associated with the exposure, which it must be to confound (here it is not). The second uses the risk ratio, which is collapsible: without confounding, its crude value is a weighted average of the stratum-specific values. Here the stratum risk ratios are 1.60 and 3.00 and the crude risk ratio is 1.83, which lies between them. When the outcome is rare the odds ratio approximates the risk ratio and non-collapsibility is negligible.
Section 3 has set out the decisions a protocol makes about confounding. The protocol identifies confounders with a causal diagram, controls each one by restriction, matching, stratification or adjustment, and separates confounding from effect modification and from the non-collapsibility of the odds ratio. The key takeaways below summarize these decisions. The reflection that follows asks for the adjustment set in a cohort study of first-year students, and the knowledge check closes the section.
Key Takeaways
- A confounder is a common cause of the exposure and the outcome that is not caused by the exposure; the classical criteria follow from that definition but are not sufficient, and mediators and colliders stay out of the adjustment set.
- Restriction and matching control confounding by design, at the cost of breadth and recruitment effort, and matching in a case-control study requires a matched or stratified analysis.
- The causal diagram, drawn before the data are seen, determines the adjustment set through the backdoor criterion, and the change-in-estimate check on the log scale supports it.
- The Mantel-Haenszel estimator pools homogeneous strata into one adjusted odds ratio; in the STREP example it moves the estimate from 5.44 to 3.36.
- Heterogeneous strata indicate effect modification, which is reported stratum by stratum, and a change in the odds ratio without an association between the third variable and the exposure indicates non-collapsibility.
Reflection
A cohort protocol will estimate the effect of regular physical activity during the first term of university (the exposure) on a positive depression screen at the end of the first year (the outcome) among first-year students. The investigators' causal diagram contains these assumptions. A chronic health condition reduces physical activity and raises the risk of depressive symptoms. Depressive symptoms at university entry reduce physical activity and predict depressive symptoms at the end of the year. Sleep quality at mid-year is improved by physical activity and lowers the risk of depressive symptoms. Varsity athlete status raises physical activity and has no other path to the outcome. Completing the end-of-year survey is more likely among active students and less likely among students with depressive symptoms. A confounder is a common cause of the exposure and the outcome, a mediator lies on the path from the exposure to the outcome, and a collider is a common effect. (1) Identify the adjustment set and justify each variable in it. (2) Identify the variables that should stay out of the adjustment set, and explain why for each. (3) For one confounder, decide whether you would control it by restriction or matching at the design stage or by stratification in the analysis, and explain how you would use the change-in-estimate check. (4) Recall overmatching from HSCI 230 Lesson 4: a colleague suggests that a case-control version of this study match each case to a control on sleep quality at mid-year. Explain why that would be overmatching and what it would do to the estimate.
Minimum 20 characters required.
1. In a survey of loneliness and depressive symptoms, perceived social support is caused by loneliness and affects depressive symptoms. If the protocol adds it to the adjustment set, the adjusted estimate will:
2. In a case-control study, controls are matched to cases on age, and age is associated with the exposure. Which statement is correct?
3. In the STREP example, the RSV+ stratum (126 children) has 52 exposed cases, 11 unexposed cases, 36 exposed non-cases and 27 unexposed non-cases, and the RSV− stratum (274 children) has 18, 19, 54 and 183. What is the Mantel-Haenszel odds ratio?
4. The crude odds ratio is 5.44 and the Mantel-Haenszel odds ratio is 3.36. What is the change on the log scale, and what does it suggest?
5. Two strata each give an odds ratio of 4.0, the exposure is equally common in both strata, the outcome is common, and the crude odds ratio is 2.98. What is the best explanation for the difference?
Unmeasured Confounding and the Threats-to-Validity Plan
From measured to unmeasured confounders
Sections 1 to 3 dealt with biases that a protocol can prevent or correct with information it collects. Some confounders cannot be measured at all, and three computations show how much such a confounder could change a result: external adjustment, a planned sensitivity analysis and the E-value. A short orientation to the model-based methods that later courses teach comes first. The threats to validity and mitigation subsection of a protocol, with a complete example for the Campus Connection Study, then brings the four sections together.
Learning Objectives for this section
- Describe in outline what multivariable regression, standardization, propensity scores, marginal structural models, difference-in-differences and instrumental variables do, and where each is taught.
- Apply external adjustment with a bias factor to estimate what an association would be if an unmeasured confounder had been measured.
- Plan a sensitivity analysis for a protocol, including its bias parameters, their sources and the way the results will be reported, and compute and interpret an E-value.
- Write a threats to validity and mitigation subsection that states, for each threat, its mechanism, likely direction, prevention, assessment and residual risk.
Methods that scale to many confounders: an orientation
Stratification works well for one or two confounders. With five confounders of three levels each there are 243 strata, many of them nearly empty, and the Mantel-Haenszel estimator loses most of its information. Model-based methods summarize the confounders more efficiently. Later courses teach them. A protocol for a study with one or two confounders often plans a stratified analysis and names the model-based method that a fuller analysis would use. Table 7.7 summarizes what each method does and where it is taught, and it completes, at the level of orientation, the confounding-control methods that HSCI 341 Lesson 1 Section 4 previewed: propensity scores, regression and difference-in-differences.
Table 7.7. Model-based methods for many confounders: what each does, its key assumption and where it is taught.
| Method | What it does | Key assumption | Where it is taught |
|---|---|---|---|
| Multivariable regression | Estimates the exposure effect with the confounders held constant in a model for the outcome. | No unmeasured confounding, and a correctly specified model | HSCI 410 Lesson 3 (linear and logistic regression) |
| Standardization | Applies stratum-specific risks to a chosen standard population to give one summary that is valid even when the strata differ. | No unmeasured confounding | HSCI 341 Lesson 4 (standardized rates and ratios) |
| Propensity scores | Summarize the confounders as the probability of exposure, then match, stratify or weight on that probability. | No unmeasured confounding, and overlap between exposed and unexposed people | Introduced in HSCI 230 Lesson 3 Section 1; taught in HSCI 826 Lesson 7 (graduate course) |
| Marginal structural models | Recall from HSCI 230 Lesson 11 Section 1, where the CD4 count example showed the problem they solve: each person is weighted by the inverse of the probability of the exposure received, to handle confounders that change over time and are affected by earlier exposure. | No unmeasured confounding at any time point | Introduced in HSCI 230 Lesson 11; HSCI 826 Lesson 7 Section 3.7 teaches inverse probability weighting for a treatment at one time |
| Difference-in-differences | Compares the change over time in an exposed group with the change in an unexposed group, which removes differences between the groups that are stable over time. | Parallel trends: without the exposure, the two groups would have changed by the same amount | HSCI 826 Lesson 7 (graduate course) |
| Instrumental variables | Use a variable that affects the exposure, and affects the outcome only through the exposure, to estimate the effect without measuring the confounders. | A valid instrument, which is rare in observational data | HSCI 826 Lesson 8 (graduate course) |
In the STREP example, logistic regression of CRD on STREP and RSV gives an odds ratio of 3.35, and stratification on the propensity score gives 3.36, identical to the Mantel-Haenszel estimate because with a single binary confounder the propensity score takes only two values. The methods differ in how they handle many covariates. Except for instrumental variables, and difference-in-differences for confounders that are stable over time, they share one limitation: they adjust only for what was measured.
A protocol therefore also needs a way to reason about the confounders it could not measure, which the next part takes up.
When a confounder was never measured
Every protocol leaves some confounders unmeasured, because they were unknown, too costly or too sensitive to ask about, or because the data source does not record them. The protocol cannot remove this residual confounding, but it can name the likely candidates, state the direction of the bias each would cause, and plan a calculation that shows how strong a confounder would have to be to change the conclusion.
External adjustment
When a confounder U was not measured in the study but its prevalence and its association with the disease are known from other sources, the observed estimate can be divided by a bias factor (Schlesselman, 1978):
The formula is exact for risk ratios when the confounder is binary and its risk ratio with the disease is the same among the exposed and the unexposed; it is approximate for odds ratios, and the approximation is best when the disease is rare. In a case-control study the two prevalences describe the source population, which the controls represent. Worked Example 7.9 applies Equation 7.14 to the STREP data as if RSV had not been recorded, and Calculator 7.6 reproduces the example, with a second preset for a strong confounder that is far more common among the exposed.
Worked Example 7.9: External Adjustment for STREP and CRD
Suppose RSV had not been recorded, so that only the crude odds ratio of 5.44 was available. External studies of children without respiratory disease suggest that about 40% of STREP carriers and 13% of non-carriers have RSV (p1 = 0.40, p2 = 0.13), and that RSV multiplies the odds of CRD by about 4.2 (ORUD = 4.2).
- Bias factor: B = [0.40 × 3.2 + 1] ÷ [0.13 × 3.2 + 1] = 2.28 ÷ 1.416 = 1.61.
- Externally adjusted odds ratio = 5.44 ÷ 1.61 = 3.38.
The result is close to the Mantel-Haenszel estimate of 3.36 obtained in Section 3 when RSV was measured. The agreement depends on the accuracy of the external information, and because CRD is common in these data the odds-ratio approximation is less exact than it would be for a rare disease.
External adjustment gives one corrected estimate for one set of outside values, and the agreement in Worked Example 7.9 depends on the accuracy of those values. A protocol therefore also plans for the possibility that they are wrong, which is the subject of the next part.
Planning a sensitivity analysis
When no single set of external values is trustworthy, the protocol plans a sensitivity analysis that varies the bias parameters over plausible ranges and reports how the estimate changes. The same logic applies to all three threats. For selection bias the parameters are the sampling fractions of Section 1, for misclassification they are the sensitivity and specificity of Section 2, and for an unmeasured confounder they are the prevalences and the strength used in external adjustment. Quantitative bias analysis is the general name for these methods (Lash, Fox & Fink, 2009). It ranges from simple analyses with one set of fixed values, through grids of fixed values, to probabilistic analyses that draw the parameters from probability distributions and repeat the correction many times. HSCI 230 Lesson 12 interprets sensitivity analyses that published studies report; a protocol plans one in five steps.
- Name each threat to be analyzed and the direction in which it would move the estimate.
- Define its bias parameters.
- State where plausible values will come from: published studies, national surveys, pilot data or the study's own validation sub-study.
- Choose the form of the analysis: a grid of fixed values (deterministic) or distributions (probabilistic).
- State how the results will be reported, including the values at which the conclusion would change.
Worked Example 7.10 applies these steps to an unmeasured confounder of the STREP and CRD association, with a grid of fixed values for the prevalence and the strength that enter the bias factor of Equation 7.14.
Worked Example 7.10: A Planned Grid for an Unmeasured Confounder
Suppose a confounder U of the STREP and CRD association could not be measured. The grid varies the prevalence of U among STREP carriers (p1), with a prevalence of 0.10 among non-carriers, and the strength of the U and CRD association (ORUD). Each cell divides the observed 5.44 by the bias factor for that combination.
| ORUD | p1 = 0.4 | p1 = 0.6 | p1 = 0.8 |
|---|---|---|---|
| 2.0 | 4.27 | 3.74 | 3.32 |
| 5.0 | 2.93 | 2.24 | 1.81 |
| 10.0 | 2.25 | 1.62 | 1.26 |
A moderately strong confounder (ORUD = 5, present in 40% of carriers and 10% of non-carriers) leaves the estimate near 2.9. Only a strong confounder that is also far more common among carriers (ORUD = 10, 80% against 10%) brings it down to about 1.3, and no cell of the grid removes the association. The protocol would state this grid in advance, together with the sources for the values chosen.
The E-value
VanderWeele and Ding (2017) summarized this kind of analysis in a single number. The E-value is the minimum strength of association, on the risk-ratio scale, that an unmeasured confounder would need to have with both the exposure and the outcome, beyond the measured confounders, to explain away an observed association completely. Equation 7.15 gives the E-value for a risk ratio above 1.
For a risk ratio of 2.0, E = 2.0 + √(2.0 × 1.0) = 3.41, so a confounder would need a risk-ratio association of at least 3.4 with both the exposure and the outcome to explain the result away. An odds ratio or hazard ratio can be treated as a risk ratio when the outcome is rare; when the outcome is common (above about 15%), the square root of the odds ratio approximates the risk ratio (VanderWeele & Ding, 2017). In a protocol the E-value is a planning tool. The investigator computes it for the effect the study is designed to detect, compares it with the strength that the named unmeasured confounders could plausibly have, and states in advance that it will be reported for the estimate and for the confidence limit closer to 1. Worked Example 7.11 carries out this planning step for the Campus Connection Study, and Calculator 7.7 reproduces it, with presets for two other estimates.
Worked Example 7.11: The E-value at the Planning Stage
The Campus Connection Study is sized to detect a prevalence ratio of 0.34 ÷ 0.26 = 1.31. A prevalence ratio is a ratio of proportions, so the risk-ratio formula applies: E = 1.31 + √(1.31 × 0.31) = 1.95. If the lower confidence limit were 1.10, its E-value would be 1.10 + √(1.10 × 0.10) = 1.43. Suppose the protocol identifies a history of depression before university entry as a confounder that the questionnaire cannot measure well. If such a history could plausibly be associated with both current loneliness and current depressive symptoms by risk ratios of 2 or more, then a result of the planned size could be explained away by it. The protocol says so in advance, which tells reviewers that the study's contribution will be a well-measured, well-sampled association whose causal interpretation remains open, and it may prompt the investigators to add a retrospective question on earlier depression.
External adjustment, a planned grid and the E-value give a protocol three ways to show how much an unmeasured confounder could change a result. The last part of this lesson places these results, with those of Sections 1 to 3, in the protocol subsection on threats to validity and mitigation.
Writing the threats to validity and mitigation subsection
The subsection brings the four sections of this lesson together for one study. It usually sits in the methods of a protocol, after the analysis plan and sample size. A reviewer reading it should be able to see that each threat has been identified for this particular study, that its likely direction has been reasoned out, and that the protocol contains specific steps to prevent it and to measure what remains. A strong subsection is specific to the study, quantitative where the tools of this lesson allow, consistent with the causal diagram, sampling plan and sample size elsewhere in the protocol, and candid about the threats that cannot be removed. Table 7.8 lists what the subsection states for each threat and the tools from this lesson that supply it.
Table 7.8. The elements of a threats to validity and mitigation subsection and the tools from this lesson that supply them.
| For each threat, state | Tools from this lesson |
|---|---|
| The mechanism in this study: who or what could produce the bias | The causal diagram; the recap of HSCI 230 Lessons 7 to 11 |
| The likely direction, and its size where it can be estimated | Sampling fractions; misclassified tables; the bias factor |
| The design steps that prevent it | Frame, response and retention decisions; instrument choice, timing and blinding; restriction and matching |
| The analysis steps that assess or correct it | Comparison with the frame and weighting; back-correction; Mantel-Haenszel stratification and the homogeneity test |
| The residual risk and the planned sensitivity analysis | Bounding with assumed sampling fractions; correction over a range of sensitivity and specificity; external adjustment; the E-value |
| For missing data, the expected unit and item missingness, the assumed mechanism and the handling method | Response and retention decisions (Section 1); the missing data mechanisms and complete-case analysis or multiple imputation (HSCI 230 Lesson 11 Section 2) |
The subsection is usually organized by threat (selection, information, confounding, and any design-specific or temporal threat), in about half a page to a page. It does not need to repeat the full sample size or analysis plan; it refers to them. Missing data cut across the threats: unit non-response and loss to follow-up are selection problems, while item missingness affects information and the measured confounders, so the subsection states the missingness it expects, the mechanism it assumes and the method it will use. Box 7.6 recalls the missing data mechanisms that HSCI 230 Lesson 11 described.
Box 7.6: Recall: Missing Data Mechanisms
HSCI 230 Lesson 11, Section 2, classified missing data by mechanism. Data are missing completely at random when the chance of being missing is unrelated to any data, and complete-case analysis is then unbiased but loses power. They are missing at random when missingness depends only on observed variables, and multiple imputation that uses those variables is valid. They are missing not at random when missingness depends on the unobserved value itself, which no standard method fixes, so a sensitivity analysis is required.
Retrieval question. Younger respondents skip the financial strain items more often than older respondents, and age is recorded for everyone. Which mechanism does this describe, and which handling method fits it?
Missing at random, provided that within each age group missingness does not depend on the level of financial strain itself. Multiple imputation that includes age is then valid, while a complete-case analysis would under-represent younger respondents.
Worked Example 7.12 gives the complete subsection for the Campus Connection Study. It draws on the selection decisions and bound of Worked Example 7.3, the size of the validation sub-study in Worked Example 7.5 and the E-value of Worked Example 7.11, and it adds a plan for missing data.
Worked Example 7.12: Threats to Validity and Mitigation in the Campus Connection Study
Selection bias. Participation depends on responding to an emailed invitation, and the planned response rate of 15% leaves room for selection. Selection would bias the association only if response depended jointly on loneliness and depressive symptoms, for example if lonely students with depressive symptoms were especially unlikely to respond, which would bias the estimate toward the null. The study prevents recruitment through self-selection by sampling at random from registrar frames at five institutions, and it reduces differential response with a neutral invitation describing a survey of student health and connection, three reminders and a short questionnaire. Respondents will be compared with registrar counts by institution, age group and gender, and the design weights will be adjusted for non-response. A bounding analysis will report the odds ratio under sampling fractions of 0.18, 0.15, 0.15 and 0.10 for the four combinations of loneliness and depressive symptoms; under these values an observed odds ratio of 2.0 corresponds to 2.5.
Information bias. Loneliness is measured with the six-item De Jong Gierveld scale and depressive symptoms with the PHQ-2, administered online in the same order to every respondent, with the PHQ-2 placed before the loneliness items. The outcome is defined as a positive PHQ-2 screen, the construct that the planning values describe. Against diagnostic interviews the PHQ-2 has a pooled sensitivity of 0.72 and specificity of 0.85 (Levis et al., 2020), so associations with depression itself would be attenuated if errors are non-differential. A secondary analysis will correct the prevalence ratio for outcome misclassification across a range of published sensitivity and specificity values. An internal validation sub-study, which would need about 230 diagnostic interviews, is beyond the study's resources, and the transportability of published values to students aged 18 to 29 is stated as a limitation.
Confounding. The causal diagram presented earlier in the protocol identifies financial strain, gender and the other covariates in the adjustment set as confounders, measured with Statistics Canada standard items. Perceived social support is a mediator and is excluded from the adjustment set for the total effect. The primary analysis stratifies by financial strain with the Mantel-Haenszel estimator and reports the homogeneity test; regression adjustment for the full set, which HSCI 410 teaches, is named as a secondary analysis. A history of depression before university entry is a plausible unmeasured confounder that would bias the estimate away from the null. The E-value will be reported for the estimate and its lower confidence limit; for the planning prevalence ratio of 1.31 it is 1.95, which a strong prior-depression effect could reach, and the report will say so.
Temporality. Loneliness and depressive symptoms are measured at the same time, so the survey cannot show that loneliness precedes depressive symptoms, and depressive symptoms may themselves increase loneliness. The estimand is described as a prevalence ratio for a cross-sectional association, and the reverse pathway is stated as a limitation.
Missing data. Unit non-response is treated under selection bias. The sample size assumes that 81% of respondents complete the questionnaire, so about one respondent in five is expected to stop partway; partial completions will be counted and compared with complete responses on the items they answered. Among completers, item missingness is expected to be highest for sensitive items such as financial strain. The protocol assumes that item missingness is at random given the observed variables (institution, age group, gender and the other questionnaire items), uses multiple imputation for missing confounder values in the primary analysis, and reports a complete-case analysis as a sensitivity analysis.
The note that follows summarizes how a subsection like Worked Example 7.12 is drafted.
The subsection is drafted from material already in the protocol: the causal diagram, the sampling plan and sample size, and the instruments chosen. Each threat is described with the five elements in Table 7.8. Where this lesson gives a calculation (an odds ratio of the sampling fractions, a sample size inflated for misclassification, an E-value for the effect the study is designed to detect), the subsection reports the result. The final reflection of this lesson applies the five elements to a cohort study of first-year students, and its model answer shows a second complete example for a different design.
Section 4 has shown how a protocol reasons about confounders it cannot measure, and how the threats to validity and mitigation subsection gathers the tools of all four sections into one part of the protocol. The key takeaways below summarize these points. The reflection that follows plans a sensitivity analysis for an unmeasured confounder with the bias factor and the E-value, and the knowledge check closes the section.
Key Takeaways
- Regression, standardization, propensity scores and marginal structural models adjust only for measured confounders; instrumental variables are the exception, and valid instruments are rare.
- External adjustment divides an observed estimate by a bias factor built from the confounder's prevalence among the exposed and the unexposed and its association with the disease.
- A planned sensitivity analysis names the threat, its bias parameters and their sources, and states how the results will be reported, for selection bias and misclassification as well as for unmeasured confounding.
- The E-value, computed at the planning stage for the effect a study is designed to detect, tells reviewers how strong an unmeasured confounder would have to be to explain that effect away.
- The threats to validity and mitigation subsection states, for each threat in a particular study, its mechanism, likely direction, prevention, assessment and residual risk.
Reflection
A cohort protocol anticipates a risk ratio of about 1.5 (95% confidence interval 1.2 to 1.9) for the association between living alone in the first year of university and a positive depression screen at the end of the year. Adverse childhood experiences cannot be measured. The investigators expect them to be present in about 40% of students who live alone and 25% of those who do not, and to double the risk of a positive screen (a risk ratio of 2.0). The bias factor is B = [p₁ × (RRUD − 1) + 1] ÷ [p₂ × (RRUD − 1) + 1], where p₁ and p₂ are the prevalences among the exposed and the unexposed and RRUD is the confounder's association with the outcome; the adjusted estimate is the observed estimate divided by B. The E-value for a risk ratio above 1 is RR + √[RR × (RR − 1)]. (1) Compute the bias factor and the externally adjusted risk ratio. (2) Compute the E-values for the point estimate and for the lower confidence limit, and interpret them. (3) Describe how the protocol should plan and report this sensitivity analysis. (4) Recall from HSCI 230 Lesson 3 Section 1: what does a propensity score summarize, and why could matching on it not remove confounding by adverse childhood experiences here?
Minimum 20 characters required.
1. An observed odds ratio of 5.44 is externally adjusted for a confounder with a prevalence of 0.40 among the exposed and 0.13 among the unexposed and an odds ratio of 4.2 with the disease. What is the adjusted odds ratio?
2. What is the E-value for an observed risk ratio of 1.5?
3. A study reports a risk ratio of 1.40 with a 95% confidence interval of 0.95 to 2.06. What is the E-value for the confidence interval?
4. Which method can, when its key assumption holds, estimate an effect without measuring the confounders?
5. Which sentence best fits the threats to validity and mitigation subsection of a protocol for a student survey of loneliness and depressive symptoms?
Final Assessment
Bringing It All Together
This lesson treated validity as a set of decisions made when a protocol is written. HSCI 230 had already defined selection bias, information bias and confounding; this lesson added the calculations that size each one and the design steps that prevent it. Section 1 showed that selection biases an odds ratio by the factor ORsf, so that what matters is whether participation depends jointly on exposure and outcome, and it turned the sampling frame, the recruitment message, contact schedules and retention procedures into protocol decisions. Section 2 computed the table that misclassification produces, recovered the true table from assumed sensitivity and specificity, planned a validation sub-study to supply those values, and inflated a sample size so that misclassification does not leave a study underpowered.
Section 3 framed the control of confounding as a choice among restriction, matching, stratification and adjustment, made for each confounder that the causal diagram identifies, and it carried out the Mantel-Haenszel analysis that protocols with one or two categorical confounders plan, together with the checks that separate confounding from effect modification and from non-collapsibility. Section 4 oriented the model-based methods taught in later courses, computed external adjustment and the E-value for confounders that cannot be measured, and brought everything together, including the handling of missing data, in the threats to validity and mitigation subsection of a protocol. The STREP and childhood respiratory disease example ran through the confounding sections, and the table below collects its results.
How the adjustment methods compare: STREP and CRD, adjusting for RSV
| Method | Odds ratio | Key feature |
|---|---|---|
| Crude (unadjusted) | 5.44 | No control for RSV |
| Restriction (children without RSV) | 3.21 | Removes RSV by design; describes only children without RSV |
| Mantel-Haenszel stratification | 3.36 (95% CI 1.96 to 5.76) | Pools two homogeneous strata (3.55 and 3.21) |
| Propensity-score stratification | 3.36 | With one binary confounder, identical to stratifying on RSV |
| Logistic regression (STREP and RSV) | 3.35 (95% CI 1.96 to 5.73) | Extends to many confounders (HSCI 410) |
| External adjustment (RSV unmeasured) | 3.38 | Uses external prevalence and strength of RSV |
The adjusted methods agree on an odds ratio of about 3.2 to 3.4, against a crude value of 5.44. Methods that share the assumption of no unmeasured confounding can agree and still miss a confounder that none of them measured, which is why the protocol also plans a sensitivity analysis and reports the E-value.
Key Takeaways from this lesson
- HSCI 230 Lessons 7 to 11 define the threats to validity; a protocol prevents them by design and plans the calculations that size what remains.
- A confounder is a common cause of exposure and outcome, a mediator lies between them and a collider is a common effect, and only confounders belong in the adjustment set for a total effect.
- Selection multiplies the odds ratio by the odds ratio of the sampling fractions, so a low response rate biases the estimate only when response depends jointly on exposure and outcome.
- The sampling frame, a neutral recruitment message, contact schedules, incentives, retention procedures and the recording of frame characteristics are the protocol's defences against selection bias.
- Non-differential misclassification of a binary variable mixes the groups being compared and moves the estimate toward the null on average, and differential misclassification is prevented through the timing, standardization and blinding of measurement.
- Back-correction and regression calibration need sensitivity and specificity, or a calibration equation, from a validation sub-study whose design, size and use the protocol states.
- Sample sizes are calculated from the proportions the study will observe, which misclassification brings closer together.
- Restriction, matching, stratification and adjustment are chosen for each confounder in the protocol, with the causal diagram as the basis and the change-in-estimate check on the log scale as support.
- The Mantel-Haenszel estimator pools homogeneous strata; heterogeneous strata indicate effect modification, and a change without an association between the third variable and the exposure indicates non-collapsibility.
- External adjustment, a planned sensitivity analysis and the E-value show how strong an unmeasured confounder would need to be, and the threats to validity and mitigation subsection records all of these decisions for one study.
Core Concepts Reviewed
Section 1: The recap of HSCI 230 Lessons 7 to 11; the diagnostic check on confounders, mediators and colliders; sampling fractions, the odds ratio of the sampling fractions and sampling odds; the non-response and daycare worked examples; bias breakers; and protocol decisions on the sampling frame, response, retention and what to record, illustrated with the Campus Connection Study.
Section 2: The misclassified two-by-two table; outcome sensitivity and specificity and the risk ratio; differential misclassification and the decisions that prevent it; back-correction; the five steps of a validation sub-study, including its size; regression calibration in brief; and sample size inflation for an imperfect outcome measure.
Section 3: The definition of a confounder and the classical criteria; population and sample confounders; the STREP and CRD example; restriction, matching, overmatching and McNemar's test; the backdoor criterion and the Campus Connection diagram; the change-in-estimate check; the Mantel-Haenszel estimator, its homogeneity and overall tests; effect modification; and non-collapsibility.
Section 4: An orientation to regression, standardization, propensity scores, marginal structural models and instrumental variables; external adjustment with a bias factor; planning a sensitivity analysis; the E-value at the planning stage; and the threats to validity and mitigation subsection with the Campus Connection example.
The final reflection applies the whole lesson to one study: a first draft of the threats to validity and mitigation subsection for a cohort of first-year students.
Reflection
Draft the threats to validity and mitigation subsection, in about 300 to 450 words, for this study: a cohort of 900 first-year students recruited at September orientation, with weekly participation in a campus club (yes or no, self-reported in October) as the exposure and a positive PHQ-2 depression screen in April as the outcome. About 30% of participants are expected to be lost by April, the PHQ-2 has a sensitivity of 0.72 and a specificity of 0.85 for depression, and the investigators can measure age, gender, living arrangement, financial strain and a PHQ-2 score at orientation. For each threat (selection bias, information bias, confounding, and any temporal or design-specific threat), state the mechanism in this study; its likely direction and, where you can, its size; the design steps that prevent it; the analysis steps that assess or correct it; and the residual risk with the sensitivity analysis the study will report. The E-value for a risk ratio below 1 is computed from its reciprocal, R, as R + √[R × (R − 1)].
Minimum 20 characters required.
Final Knowledge Assessment
Complete the following 15-question assessment, which draws on all four sections of the lesson. A score of 100% is required to complete the lesson. You may retake the assessment as many times as needed.
Question 1: In a cohort study of physical activity and depressive symptoms one year later, sleep quality at six months is improved by activity and lowers depressive symptoms. What is its role, and how is it handled?
Question 2: Response probabilities in a study are 0.30 for exposed cases, 0.20 for unexposed cases, 0.40 for exposed non-cases and 0.40 for unexposed non-cases. What is the odds ratio of the sampling fractions, and what does it imply?
Question 3: Which protocol decision most directly prevents differential misclassification of exposure in a case-control study?
Question 4: When an observed table is back-corrected with assumed values of sensitivity and specificity, one corrected cell comes out negative. What does this indicate?
Question 5: A protocol will correct for exposure misclassification with sensitivity and specificity taken from a validation study conducted in another country. What is the main assumption?
Question 6: Using p′ = Se × p + (1 − Sp) × (1 − p), what observed proportion does a test with a sensitivity of 0.72 and a specificity of 0.85 give for a true proportion of 0.26?
Question 7: Restricting the STREP study to children without RSV gives an odds ratio of 3.21. Which statement is correct?
Question 8: In a 1:1 pair-matched case-control study, 40 pairs have an exposed case and an unexposed control, 16 have an unexposed case and an exposed control, 25 have both members exposed and 19 have neither exposed. What is the matched odds ratio?
Question 9: A protocol's causal diagram shows that financial strain causes both the exposure and the outcome, and that completing the survey is caused by both the exposure and the outcome. Which decision follows?
Question 10: The stratum-specific odds ratios in a Mantel-Haenszel analysis are 1.2 and 4.8, and the homogeneity test gives p < 0.001. What should the planned report give?
Question 11: Why does a protocol state its causal diagram and its change-in-estimate threshold before the analysis?
Question 12: A cohort study finds a risk ratio of 2.0 with a 95% confidence interval of 1.5 to 2.7. What are the E-values for the estimate and for the confidence interval?
Question 13: For one or two categorical confounders, which analysis does a protocol usually name as its primary method?
Question 14: A planned sensitivity analysis divides an observed odds ratio of 5.44 by a bias factor, with a confounder prevalence of 0.10 among the unexposed. With a confounder-disease odds ratio of 5 and a prevalence of 0.40 among the exposed, what is the adjusted value?
Question 15: Which statement belongs in the threats to validity and mitigation subsection for a cross-sectional survey of loneliness and depressive symptoms?