Case-Control
Studies
Evaluating Epidemiological Research
Learning objectives for this lesson:
- Describe the major design features of risk-based and rate-based case-control studies
- Identify hypotheses and population types consistent with each design
- Differentiate between primary-base and secondary-base case-control studies
- Elaborate the principles used to select and define the case series
- Explain the principal features for selecting controls in open and closed populations
- Design and implement a valid case-control study to meet specific objectives
- Describe the key features of six hybrid observational designs (case-crossover, self-controlled case-series, case-case, case-case-control, case-only, and case-cohort), including the logic of using cases as their own controls across time and the choice of referent periods
- Describe two-stage sampling designs, explain when they enhance the efficiency of cross-sectional, cohort, and case-control studies, and design the basic sampling strategy for a two-stage case-control study
- Apply hybrid design concepts to research questions involving rare exposures, transient triggers, expensive covariates, or surveillance data without a usable control group
This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University.
Glossary: Key Terms, People & Concepts
📚 Reference page, available throughout the lesson
This glossary collects the key concepts, people, and ideas you will meet in this lesson. Use it as a reference while you work through the material, or as a review before assessments. Type in the search box to filter entries.
Introduction & The Study Base
⏱ Estimated reading time: 15 minutes
Introduction and Overview
Modern teaching of case-control design follows the “study base” framework laid out by Vandenbroucke and Pearce (2012). An earlier lesson introduced the three sampling approaches that organize observational analytic studies: cross-sectional (sample without regard to disease), case-control (sample on the disease), and cohort (sample on the exposure). It walked through the cross-sectional design in detail. This lesson does the same work for case-control studies. The seven content sections proceed from the most general design choices to the most specific, and then outward to the design’s variants: this section sets up the basic logic and the concept of the study base; a later section covers how cases are identified and how controls are selected; a later section distinguishes the two main flavors (risk-based and rate-based) and shows what the odds ratio actually estimates under each; a later section closes the loop on comparability, analysis, and reporting; and the last three sections introduce hybrid observational designs, which adapt case-control sampling to transient exposures, to surveillance data without a usable control group, and to cohorts in which measurement is expensive.
Two ideas from an earlier lesson carry over directly. First, the cross-sectional limit of measuring prevalence rather than incidence is one of the things case-control studies are designed to overcome. Second, the unified-approach discipline (think experiment first, fix design before seeing data, project forward to alternative results) applies just as much here as it did to cross-sectional designs, arguably more, because the choice of cases and controls creates more opportunities for things to go wrong.
Learning Objectives
- Describe the fundamental logic of the case-control study design.
- Distinguish between primary-base and secondary-base case-control studies.
- Explain the concept of nested case-control studies.
- Identify when case-control designs are performed prospectively vs. retrospectively.
What Is a Case-Control Study?
The basis of the case-control study design is to select individuals who have newly developed the disease or outcome of interest (the cases) and, as a comparison, individuals who have not developed the disease at the time of selection (the controls). We then contrast the frequency of exposure factors in the cases with the frequency of exposure factors in the controls.
Walk through the 1950 Doll & Hill (1950) case-control study scene by scene. Next ▶ advances at your pace.
A 7-scene reenactment of the first major case-control study: the rising lung-cancer ward, the backward-looking design, case and control interviews, the 2×2 table populating live, the tilting scale, and the moment OR ≈ 14 lands in print.
Key Distinction
A case-control study is not a comparison between a set of cases and a set of ‘healthy’ subjects. It is a comparison between a set of cases and a set of non-case subjects (people who have not developed the specific disease but may have other diseases) whose exposure to the factors of interest reflects the exposure in the source population.
The controls would have been included as cases if they had developed the outcome (disease) of interest. Most frequently, individual people are the units of interest, but the design also applies to aggregates of individuals.
Figure, The logic of case-control design: select cases and controls from the same source population, then compare their exposure histories.
Usually, case-control studies are performed retrospectively since the outcome (usually disease) has occurred when the study begins. However, it is possible to conduct case-control studies prospectively; in these, the cases have not yet developed until after the study begins, so the cases are enrolled as they occur over time.
The diagram makes the logic look simple, but it conceals a hard question: which source population? In case-control studies that question has a name and three standard answers, and the rest of the lesson essentially turns on getting it right.
The Study Base
The study base is the population from which the cases and (possibly) the controls are obtained. The nature of the study base determines how controls should be selected. The three flip cards below introduce the standard typology, primary base, secondary base, and the special case of nested designs. Click each one and notice that they differ in how directly the source population can be enumerated, which in turn determines how easy it is to draw a valid control sample.
The three study-base types make more sense when you can see them in published studies. The four examples below recur throughout the rest of the lesson; expand each one to see how the design choices were made and which combination of features (primary vs. secondary base; risk-based vs. rate-based; nested or not) the investigators chose. We will refer back to these examples by number in later sections.
Key Examples
Dorgan et al (2010) used serum samples from a secondary-base case-control study. A total of 6,915 women who were free of cancer donated blood between 1977–1989. Of the 6,720 women in extended follow-up, 1,751 were identified as deceased. For each of the 117 potential cases, 2 potential controls were matched on age (±2 years), date (±1 year), and menstrual cycle day (±2 days). This is a risk-based sampling strategy. Conditional logistic regression was used to evaluate the association.
Dore et al (2004) conducted a rate-based study in Alberta, British Columbia, and Saskatchewan, Canada (Dec 1999–Nov 2000). Eligible cases had diarrheal illness with S. Typhimurium from stool samples. Controls were matched 1:1 on age and province of residence, randomly selected from provincial health registries. Cases and controls were interviewed by telephone using a pre-tested, standardised questionnaire covering demographics, health history, medication use, travel history, and animal contact.
Magura et al (2008) used a risk-based, secondary-base case-control design. Cases were men newly diagnosed with prostate cancer at Meritcare hospital between 2004–2006. Controls were identified from the primary-care database of the same hospital: men without cancer, aged 50–74, who had annual physicals and lipid profiles within a year. Exclusion criteria included other cancers and non-Caucasian race. The authors used a widely accepted definition of hypercholesterolemia (total cholesterol >5.17 mmol/l) and estimated odds ratios using multiple logistic regression.
Rodrigo et al (2011) conducted a community-based (primary-base), nested, rate-based case-control study within a larger randomised controlled trial in South Australia. 300 households maintained weekly health diaries. The outcome (highly credible gastroenteritis (HCG)) was defined as 2+ loose stools, 2+ vomiting episodes, or combinations with abdominal pain/nausea in 24 hours. Controls were matched to cases by study week. Logistic regression was used, allowing for familial clustering and repeated observations.
Key Takeaways
- Case-control studies select subjects based on disease status and look backward at exposure.
- The study base can be a primary base (enumerable population) or secondary base (clinic/registry).
- Nested designs allow estimation of disease frequency by exposure, a unique advantage.
- Controls should represent the exposure experience of the source population that gave rise to the cases.
Now that you can see the design's full logic and a handful of working examples, the worked example below brings the analysis back to the 2×2 contingency table you met at the end of an earlier lesson. The same structure shows up; but because we sampled on case status this time, the only valid summary measure is the odds ratio.
First, what are “odds”?
The odds of an event are the number of times it happens divided by the number of times it does not. Take the smoking table in the worked example just below. Among the 50 cases, 45 were smokers and 5 were not, so the odds of being a smoker among cases are 45 to 5, which is 9. Among the 150 controls, 60 were smokers and 90 were not, so the odds there are 60 to 90, or about 0.67. The odds ratio divides one by the other: 9 ÷ 0.67 ≈ 13.5. The cross-product shortcut, (a × d) / (b × c), reaches the same value in one step, which is why it is the formula you will see.
Worked example: the odds ratio from a 2×2 case-control table
A hypothetical study recruits 50 lung-cancer cases and 150 controls and records whether each person smoked.
| Status | Smoker | Nonsmoker | Total |
|---|---|---|---|
| Cases (lung cancer) | a = 45 | b = 5 | 50 |
| Controls | c = 60 | d = 90 | 150 |
The cross-product odds ratio is (45 × 90) / (5 × 60) = 4,050 / 300 = 13.5, so cases had about 13.5 times the odds of being smokers compared with controls. A 95% confidence interval is computed on the log scale using the Woolf method, in which the standard error of ln(OR) is √(1/a + 1/b + 1/c + 1/d) = √(1/45 + 1/5 + 1/60 + 1/90) = 0.50. Exponentiating ln(13.5) ± 1.96 × 0.50 gives a 95% CI of 5.07 to 35.97.
The interval excludes 1, so chance alone is an implausible explanation for an association of this size. The interval is also wide, and the formula shows why: the smallest cell, the 5 nonsmoking cases, contributes 0.20 of the 0.25 total under the square root, so the precision of an odds ratio is set by its sparsest cell. Because the study sampled on case status, the proportion of smokers in each row reflects the sampling fractions, and a risk ratio or risk difference cannot be estimated from this table without outside information on the source population. The odds ratio is the appropriate summary measure.
The reflection below is a personal application of the primary/secondary distinction. After working through it and the knowledge check, a later section takes the same design logic one level deeper: how do you actually identify cases, and how do you choose controls so that the comparison is fair?
Reflection
Reflection
Consider a disease that is of interest to you. Would a primary-base or secondary-base case-control study be more feasible? What would be the advantages and trade-offs of each approach for your specific research question?
Minimum 20 characters required.
1. In a case-control study, what do we compare between cases and controls?
2. What distinguishes a primary study base from a secondary study base?
3. What unique advantage does a nested case-control study provide?
4. Case-control studies are most commonly performed:
The Case Series & Principles of Control Selection
⏱ Estimated reading time: 15 minutes
Introduction and Overview
An earlier section set up the design at the level of the source population. This section steps inside it. The two halves of a case-control study (the case series and the control group) each carry their own design decisions, and historically far more case-control studies have been ruined by control selection than by anything else. We start with the case series, where the choices are mostly about definition and ascertainment, then turn to the harder problem of choosing controls.
Learning Objectives
- Describe the key elements in selecting and defining the case series.
- Discuss the importance of diagnostic criteria and case ascertainment.
- Articulate the four major principles of control selection.
- Compare different sources of controls and their strengths and limitations.
The Case Series (Section 9.3)
Key elements in selecting the case series include: specifying the disease (including diagnostic criteria), identifying the source(s) of the cases, deciding whether only incident or both incident and prevalent cases are to be included, and estimating the required number of cases and total sample size.
Incident vs. Prevalent Cases
There is virtually unanimous agreement that, when possible, only incident cases should be used. There are specific circumstances where prevalent cases may be justified, but this would be the exception, not the rule. Usually, only the first occurrence of the outcome in each study subject is included (Examples 9.1 and 9.3); however, multiple occurrences of the same disease can be included (Example 9.4).
Where Do Cases Come From?
The primary/secondary distinction we met in an earlier section also shapes how cases themselves are identified. The two tabs below revisit each option with the case series specifically in view; notice how the source choice creates a very different downstream problem for control selection.
Primary-base cases come from a specific registry that contains virtually all cases for a defined population (e.g., provincial or state disease registries). Sampling or taking a census of cases directly from the primary source population avoids a number of potential selection biases, but may be more difficult to implement and more costly.
Primary-base designs are moderately common because provincial or state records allow complete enumeration of people and their health events.
Secondary-base cases are obtained from a physician’s clinic, one or more hospitals, or registries. A major challenge is to conceptualise the actual source population from which the cases arose. A common solution is to select controls from records at the same source (e.g., the same hospital; see Example 9.3).
Every effort should be made to obtain complete case ascertainment. In secondary-base studies, the set of cases from a tertiary care facility could become increasingly different from cases in the broader source population.
Diagnostic Criteria
The diagnostic criteria for a subject to become a case should include specific, well-defined manifestational (i.e., clinical) signs where appropriate and, when possible, clearly documented diagnostic criteria (e.g., laboratory test results) that can be applied to all study subjects in a uniform manner. In some instances, it might be desirable to subdivide the case series into subgroups based on differences in disease characteristics.
Diagnostic criteria settle who counts as a case. The harder question (the one that makes or breaks a case-control study) is who counts as an appropriate comparison. That is the rest of this section.
Principles of Control Selection (Section 9.4)
The selection of appropriate controls is often one of the most difficult aspects of a case-control design. The key guideline is that controls should be representative of the exposure experience in the population which gave rise to the cases.
The Four Major Principles
Wacholder, McLaughlin, Silverman, & Mandel (1992a; 1992b; 1992c) provide the classic discussions of control selection. The four principles below act together (not independently) to ensure that the controls' exposure experience really does mirror the population that gave rise to the cases.
Sources of Controls
The four principles tell you what valid controls look like in the abstract. In practice, every choice of control source trades a different strength against a different bias. The table below catalogues the six most common sources; the column to read most carefully is the third one, because the limitation of each source is exactly the kind of bias that source most often produces.
| Source | Strengths | Limitations |
|---|---|---|
| Population controls | Representative of source population | Low response rates; recall bias; less motivated |
| Hospital controls | Accessible; cooperative; similar recall ability | Exposure may be related to hospitalisation |
| Friend controls | Similar recall; willing to participate | Over-matching; biased estimates (Bunin et al, 2011) |
| Neighbourhood controls | Similar socioeconomic background | If neighbourhood related to exposure, causes bias |
| Random digit dialling (RDD) | Population-representative sampling | Business vs. home phone issues; declining response rates |
| Partner controls | Shared environment; cooperative | Age-sex distribution differs; over-matching on exposures |
Key Takeaways
- Incident cases are strongly preferred over prevalent cases.
- Cases can come from primary bases (registries) or secondary bases (clinics/hospitals).
- Controls must represent the exposure experience of the source population.
- The four key principles: same study base, closed/open population rules, and temporal eligibility.
The reflection below asks you to make a real control-selection decision and defend it against the biases the table above just named. Once you have done that and the knowledge check, a later section turns to a parallel choice: should the design be risk-based or rate-based, and how does that decision change what the odds ratio actually estimates?
Reflection
Reflection
Imagine you are studying whether a specific dietary factor is associated with colorectal cancer. You plan to recruit cases from a hospital. What type of control group would you select (hospital, population, friend, etc.) and why? What biases might arise from your choice?
Minimum 20 characters required.
1. In case-control studies, which type of cases should preferably be used?
2. The key guideline for valid control selection is that controls should be:
3. What is a major limitation of using hospital controls?
4. Using friend controls in a case-control study can lead to:
Controls in Risk-Based & Rate-Based Designs
⏱ Estimated reading time: 15 minutes
Introduction and Overview
Earlier sections settled how the case series and the control group are assembled. The remaining design choice is how time enters the picture, specifically, whether the controls are people who survived the whole study period without becoming cases (a closed-population, risk-based design) or people sampled from the population at the moment each case occurs (an open-population, rate-based design). That choice changes what the odds ratio you eventually compute actually means. The two halves of this section unpack each design in turn, ending in matched 2×2 tables and the equations that connect them to risk and rate ratios.
Learning Objectives
- Describe the data layout and sampling approach for risk-based case-control studies.
- Derive and interpret the odds ratio (OR) in a risk-based design (Eq 9.1).
- Describe the data layout and incidence density sampling for rate-based case-control studies.
- Explain why the OR estimates the risk ratio in risk-based designs and the rate ratio in rate-based designs.
Risk-Based Case-Control Designs (Section 9.5)
The traditional approach to case-control studies has been risk-based (cumulative incidence) design. Controls are selected from among the people that did not become cases by the end of the study period. A subject can be selected as a control only once.
Design Requirements
This design is appropriate if the population is closed and is most informative if the risk period for the outcome has ended before subject selection begins. It fits situations such as outbreaks from infectious or toxic agents where the risk period is short and essentially all cases have occurred within the defined study period.
2×2 Table: Risk-Based Case-Control Design
The closed-source population can be categorised with respect to exposure and outcome (upper-case = population, lower case = sample):
| Exposed | Non-exposed | Total | |
|---|---|---|---|
| Cases | a1 | a0 | m1 |
| Controls (Non-cases) | b1 | b0 | m0 |
The cases (M1) are those that arose during the study period, while the controls (M0) are those that remained free of the outcome. Usually all or most cases are included (sampling fraction sf among cases approaches 1). We select controls independently of exposure status so that the sampling fractions in the two exposure groups should be equal:
The measure of association in risk-based designs is the odds ratio (OR):
What Does the OR Estimate?
The OR is a valid measure of association in its own right. It also estimates the ratio of risks (RR) if the outcome is relatively infrequent (e.g., <5%) in the source population. Whether the OR approximates the RR or rate ratio depends on the study design and assumptions about the source population (Knol, Vandenbroucke, Scott, & Egger, 2008).
Here is the intuition for why rarity matters. In the closed source population, the risk of disease among the exposed is A1/(A1 + B1), where A1 is the number of exposed people who become cases and B1 the number of exposed people who do not, while the odds of disease among the exposed is A1/B1. When the disease is rare, very few of the exposed become cases, so A1 is tiny next to B1; that makes A1 + B1 almost equal to B1, and so the risk almost equals the odds. The same holds among the unexposed (A0 and B0), so the odds ratio and the risk ratio nearly coincide. When the disease is common, A1 is no longer negligible and the two measures pull apart, with the odds ratio landing farther from 1 than the risk ratio. Because controls are sampled independently of exposure, the sample odds ratio in Eq 9.1 estimates this population odds ratio, which is why the rare-disease assumption is needed before it can be read as a risk ratio.
The risk-based design works beautifully when the population stays put for the whole study window, an outbreak investigation, a closed cohort with short follow-up. Most of the populations epidemiology actually studies do not behave that way: people enter, leave, age, and accumulate exposure over time. For those populations the case-control design has to be rebuilt around person-time rather than head counts.
Rate-Based Case-Control Designs (Section 9.6)
Because the populations we study are often open, the case-control designs for these populations should use a rate-based approach (incidence density sampling), which ensures that the time-at-risk is taken into account when control subjects are selected.
2×2 Table: Rate-Based Case-Control Design
| Exposed | Non-exposed | Total | |
|---|---|---|---|
| Cases | A1 | A0 | M1 |
| Person-time at risk | T1 | T0 | T |
Recall that in a cohort study, the two rates of interest would be:
In a rate-based case-control study, we select controls using a sampling rate (sr) that is equal in exposed and non-exposed populations:
Therefore, the ratio of exposed to unexposed controls equals the ratio of the cumulative exposed and unexposed subject times:
This means the OR from the case-control data estimates the incidence rate ratio (IR) in the source population:
Key Advantage of Rate-Based Design
In this design, the OR estimates the IR (from a cohort study) and no assumption about rarity of outcome is necessary for a valid estimate. This is a major advantage over risk-based designs where the rare disease assumption is needed for the OR to approximate the RR.
Equations 9.2–9.5 describe the relationship between the case-control sample and the underlying source-population rates. The practical question is how to draw the control sample so those equations actually hold. The answer is a specific sampling rule.
Incidence Density Sampling
The most common method of obtaining controls is by selecting a specified number of non-cases from the risk set, matched time-wise to the occurrence of each case. This is called incidence density sampling. At each time a subject develops the outcome, we choose b controls from the non-case subjects that exist in the source population at that point. Key features:
- We do not need to know the time-at-risk for potential controls.
- We do not need to assume the population is stable.
- The number of controls per case can vary.
- Subjects initially identified as controls can subsequently become cases.
- Controls can subsequently become cases (and vice versa in rate-based designs).
A third scheme, case-cohort (case-base) sampling, draws the controls as a random subcohort of the source population at baseline, regardless of who later becomes a case. A later section of this lesson covers it, together with the other hybrid designs that adapt case-control sampling.
Key Takeaways
- Risk-based designs use closed populations; the OR estimates the RR when the outcome is rare (Eq 9.1).
- Rate-based designs use open populations and incidence density sampling (Eqs 9.2–9.5).
- In rate-based designs, the OR directly estimates the IR with no rarity assumption needed.
- Incidence density sampling matches controls to cases by time of occurrence.
The reflection below pulls earlier sections together: it asks you to make the design-flavor choice explicitly and trace its consequences for interpretation. A later section then turns to the four practical questions that any case-control investigator has to answer once the design is chosen, how many controls, whether to use multiple control groups, how to assess exposure, how to keep cases and controls comparable, and how to report the result honestly.
Reflection
Reflection
Why is the distinction between risk-based and rate-based case-control designs important for interpreting the odds ratio? In what situations would you recommend a rate-based design over a risk-based design, and how would this affect control selection?
Minimum 20 characters required.
1. In a risk-based case-control study, controls are selected from:
2. The odds ratio in a risk-based case-control study estimates the risk ratio when:
3. What is the key advantage of the rate-based OR over the risk-based OR?
4. In incidence density sampling, at each time a case occurs we select controls from:
Comparability, Analysis & Reporting
⏱ Estimated reading time: 15 minutes
Introduction and Overview
Earlier sections walked through the major design choices: study base, case ascertainment, control selection principles, and the risk-based/rate-based split. By the time you get here, the design is essentially set. This section is about the practical implementation choices that follow, how many controls, whether to use more than one control group, how to assess exposure, how to maintain comparability, how to analyse the resulting data, and how to report it. Each of these is a place where a sound design can still be undermined.
Learning Objectives
- Discuss the number of controls per case and the use of multiple control groups.
- Describe exposure and covariate assessment in case-control studies.
- Explain the three approaches to keeping cases and controls comparable.
- Describe the analysis of case-control data and STROBE reporting guidelines.
Number of Controls per Case (Section 9.8)
Most studies use a 1:1 case-control ratio; however, other than being statistically efficient, there is nothing magical about this ratio. If the information on covariates and exposure is already recorded (i.e., exposure data is ‘free’), one might use all qualifying non-cases as controls to avoid sampling issues.
Practical Guidelines
When the number of cases is small, the precision of association measures can be improved by selecting more than one control per case. There are formal approaches for deciding the optimal number (Schlesselman, 1974), but usually the benefit of increasing the number of controls per case is small; often 3–4 controls per case is the practical maximum.
Number of Control Groups (Section 9.9)
Beyond the question of how many controls per case is the related question of how many control groups. Some researchers use multiple control groups to balance a perceived bias with one specific control group (Examples 9.5 and 9.6). However, this should be clearly defined, as it adds complexity and can be difficult to interpret if the different control groups produce different results. The two examples below show the strategy in practice; in both cases the second control group functioned mainly as a robustness check on the first.
Example 9.5, Secondary-Base Study with Population Controls
Abubakar et al (2007) studied Crohn’s disease risk factors from 9 hospitals in England using both hospital-derived and community controls. The a priori design was matched with 104 cases. For community controls, 2 general practitioners per Crohn’s patient were randomly selected, matched by age (±1 year) and gender. The authors noted that the choice of control group had little impact on their results.
Example 9.6, Primary-Care and Population-Based Controls
Brenner et al (2010) evaluated lung cancer risk factors in never-smokers in Toronto. They used both population-based controls (randomly sampled from property tax files, n=425) and hospital-based controls (from a family medicine clinic, n=523). Unconditional logistic regression models were used. A separate analysis based on 156 non-smoking cases with 466 non-smoking controls confirmed the main findings.
Once you have settled how many controls and how many control groups, the next implementation question is how exposure and covariates are actually measured, and, especially in retrospective designs, how to keep that measurement from being shaped by case status itself.
Exposure & Covariate Assessment (Section 9.10)
Most case-control studies are retrospective, so a concise, workable definition of ‘exposure’ (and also of confounders) is needed when implementing the study design. When ascertaining exposure status and information on confounders, it is preferable to obtain the greatest accuracy possible using the same process for both cases and controls.
General Rules for Exposure Assessment
When possible, have data collectors blinded to case status. As a general rule, the exposure status of cases should be the exposure category that existed at the time of outcome occurrence. For controls, their exposure status reflects their exposure situation at the time of their selection.
Keeping Cases and Controls Comparable (Section 9.11)
Accurate exposure measurement is necessary but not sufficient. Even with perfect measurement, a confounded comparison gives biased answers. The three flip cards below describe the three standard tools for preventing that, restriction, matching, and analytic control. They are not interchangeable, and a well-designed case-control study often uses two or three of them in combination.
With design choices and comparability tools in place, what remains is the analysis itself, and, just as importantly, what the resulting odds ratio means under each combination of design and sampling decision we have made so far.
Analysis of Case-Control Data (Section 9.12)
The data format and analysis for both risk-based and rate-based designs proceeds in a similar manner. In a 2×2 table:
| Exposed | Non-exposed | Total | |
|---|---|---|---|
| Cases | a1 | a0 | m1 |
| Controls | b1 | b0 | m0 |
Remember that we cannot directly estimate disease frequency (unless the study is nested) because the m1:m0 ratio was fixed by the sampling design. Chapter 6 outlines the analysis including hypothesis testing, estimating the odds ratio, and developing confidence intervals.
The three tabs below summarize what the OR estimates under each of the design combinations we have built up across this lesson. The pattern to read away with: the same number, computed from the same 2×2 table, is interpreted differently depending on the sampling decisions you made before any data were collected.
With risk-based designs and sampling of controls at the end of the follow-up period, the odds ratio estimates the risk ratio if the frequency of disease in the source population is low (e.g., below 10%), and censoring is unrelated to exposure.
If concurrent sampling (incidence density sampling) is used, the odds ratio estimates the rate ratio in both closed and open populations. For validity, stability of exposure is needed in the closed population but not in the open population.
When controls are selected from an open population without concurrent sampling of controls, the odds ratio estimates the rate ratio only if the population is stable, otherwise it is just the odds ratio. If matching is used to select controls but is ignored in the analysis, the impact depends on the extent of exposure changes during the study period (Knol, Vandenbroucke, Scott, & Egger, 2008).
The final piece of the implementation puzzle is making the design transparent to the next reader. The STROBE statement, which we previewed in an earlier lesson and then introduced in an earlier lesson, has a case-control extension that names the items most likely to go missing in a write-up.
Reporting Guidelines (Section 9.13)
Vandenbroucke et al (2007) described the key elements of case-control studies that should be reported (STROBE). The complete listing is in Table 7.3; items specific to case-control studies are included in Table 9.1, expanded in the accordion below.
Methods:
- Item 6a: Give the eligibility criteria, and the sources and methods of case ascertainment and control selection. Give the rationale for the choice of cases and controls.
- Item 6b: For matched studies, give matching criteria and the number of controls per case.
- Item 12: If applicable, explain how matching of cases and controls was addressed.
Results:
- Item 15: Report numbers in each exposure category, or summary measures of exposure.
Appraising a Case-Control Study: A Worked CASP Example
The principles of this lesson become a reader's tool when they are applied to a published study. Appraisal checklists put the same questions in a fixed order. The CASP case-control checklist (Critical Appraisal Skills Programme, 2024) has 11 items: whether the study addressed a focused issue with an appropriate method (items 1 and 2), whether the cases and controls were recruited acceptably (3 and 4), whether exposure was measured accurately (5), whether the groups were comparable and confounding was taken into account (6), the size and precision of the effect (7 and 8), whether the results are believable (9), and whether they apply to the population of interest and fit other evidence (10 and 11). The worked example below applies the items that carry most of the weight in a case-control study, the study base, exposure measurement and confounding, to a fictional excerpt.
Fictional excerpt: night-shift work and type 2 diabetes
Investigators studied whether working night shifts is associated with type 2 diabetes. Cases were 240 adults with newly diagnosed type 2 diabetes referred to the diabetes clinics of two hospitals in one city between 2023 and 2025. Controls were 480 adults without diabetes attending the fracture clinics of the same hospitals over the same period, frequency-matched to cases on age group and sex. A research nurse, who knew whether each participant was a case or a control, interviewed everyone about their work history; night-shift work was defined as at least three night shifts a month for a year or more before diagnosis (or before the interview, for controls). Night-shift work was reported by 74 of 240 cases (31%) and 91 of 480 controls (19%), a crude odds ratio of 1.9. The odds ratio adjusted for age, sex and family history of diabetes was 1.8 (95% CI 1.2 to 2.6). Body mass index and household income were not recorded.
Open each item to compare your own answer with a model answer.
Cases: largely yes. They are incident cases with a stated diagnosis and period, which avoids the survival problems of prevalent cases. The study base is secondary: it is the population that would be referred to these two diabetes clinics, so the reader has to ask who reaches them.
Controls: a concern. Controls should represent the exposure distribution of that same study base. Fracture-clinic patients would probably have been referred to the same hospitals had they developed diabetes, which helps, but their fractures may be related to the exposure: shift work is common in manual and health-care jobs, where injuries and fall-related fractures may differ from other occupations. If night-shift work is more common among fracture patients than in the study base, the odds ratio is biased toward 1; if it is less common, the odds ratio is biased away from 1.
Partly. The definition of night-shift work is explicit and was applied to the same period before diagnosis or interview, which is good practice. The interviewer, however, knew each participant's status, and newly diagnosed cases may search their history for causes more thoroughly than controls. Both can make reporting of night shifts more complete among cases, a differential misclassification that would tend to inflate the odds ratio. Blinding the interviewer, or checking a sample of answers against employer or payroll records, would have reduced the concern.
Partly. Matching on age group and sex, and including both in the analysis, deals with those variables, and family history is adjusted for. Socioeconomic position is the main gap: it plausibly influences both the chance of working night shifts and the risk of diabetes, and household income was not recorded, so residual confounding is likely. Body mass index needs more thought. Weight before the start of shift work is a possible confounder, but weight gained because of shift work lies on the causal pathway, and adjusting for it would remove part of the effect being estimated.
The adjusted odds ratio of 1.8 (95% CI 1.2 to 2.6) says that cases had about 1.8 times the odds of having worked night shifts compared with controls, and the interval excludes 1. Because the cases are incident and the controls were sampled over the period in which the cases arose, the odds ratio can be read as an estimate of the incidence rate ratio, provided the controls represent the study base. Whether to believe it depends on the three concerns above: the direction of the control-selection bias is uncertain, interviewer and recall bias would push the estimate up, and residual confounding by socioeconomic position would probably push it up as well. The study suggests an association that deserves a better-designed test; it does not establish one.
Items 1 and 2 are satisfied here (a focused question, and a design suited to an outcome with a long induction period), and items 10 and 11 call for the wider literature, which Lesson 2 showed how to find and appraise. The JBI critical appraisal checklist for case-control studies (Moola et al., 2020) covers the same ground in ten items, including the comparability of cases and controls, matching, measurement of exposure in the same way for both groups, and strategies for dealing with confounding.
Key Takeaways
- 3–4 controls per case is usually the practical maximum for improving precision.
- Multiple control groups add complexity; the general experience is that more than one control group has limited value.
- Exposure assessment should use the same process for cases and controls, with blinding when possible.
- Comparability is achieved through exclusion, matching, or analytic control (multivariable techniques).
- What the OR estimates (RR or IR) depends on the study design and sampling approach.
- An appraisal checklist such as CASP turns these principles into questions about a published study, with most of the weight on the study base, exposure measurement and confounding.
The reflection below is the section's payoff, a short colleague-asks-you-a-question prompt that requires you to use everything in the lesson so far to give a careful answer. Once you have worked through it and the knowledge check, the lesson turns to hybrid designs, which adapt the case-control machinery you have now assembled to problems the standard design handles poorly.
Reflection
Reflection
A colleague presents a case-control study with an odds ratio of 2.5 and asks: “Does this mean exposed people have 2.5 times the risk?” How would you respond? Consider the study design (risk-based vs. rate-based), the rarity of the outcome, and what the OR actually estimates under different conditions.
Minimum 20 characters required.
1. What is the practical maximum number of controls per case in most case-control studies?
2. What approach to preventing confounding is ‘most often relied upon’ in case-control studies?
3. When concurrent (incidence density) sampling is used, the OR estimates:
4. According to STROBE guidelines for case-control studies, which of the following should be reported?
Hybrid Observational Designs: Case-Crossover and Self-Controlled Case-Series
⏱ Estimated reading time: 18 minutes
Introduction and Overview
Earlier sections of this lesson assembled the standard case-control study: a study base, a case series, a control series sampled from that base, an odds ratio whose meaning depends on how the controls were sampled, and a matched analysis when controls are matched to cases. This section and the two that follow introduce hybrid designs, which rework that machinery to handle problems the standard design handles poorly, such as rare or expensive exposures, transient triggers (Suissa, 1995), and surveillance data without obvious controls. The three sections move from time-based case-only designs (this section: case-crossover and self-controlled case-series), through designs that compare one case series with another (the next section), to case-cohort and two-stage sampling designs that subsample from larger cohorts to make biomarker-heavy studies affordable (the final content section).
Learning Objectives
- Understand why hybrid study designs were developed and how they fit alongside the traditional cohort, case-control, and cross-sectional designs.
- Describe the design logic of case-crossover studies and identify when they are appropriate.
- Distinguish between unidirectional and bidirectional referent selection strategies.
- Describe the self-controlled case-series design and recognise the contexts in which it is most useful.
What Are Hybrid Observational Designs?
An earlier lesson introduced the three classic sampling approaches of observational analytic studies: cross-sectional, case-control, and cohort. This lesson has worked through the case-control design in detail. Hybrid designs are variants of these classic designs that have been developed to address particular methodological challenges such as expensive covariates, rare exposures, and transient triggers, or surveillance data where traditional control selection is problematic. The qualifier observational matters because implementation science uses “hybrid” for a different family, the effectiveness-implementation hybrid designs, which HSCI 826 Lesson 9 teaches.
This section and the two that follow cover six hybrid designs plus one important sampling strategy. Four of the hybrid designs use only cases (no separate control group), while two use a control series. The two-stage sampling design, by contrast, is a strategy that can be layered onto any of the traditional designs to enhance efficiency.
Why a Family of Hybrid Designs?
Each hybrid design solves a specific problem. Case-crossover studies eliminate the difficulty of choosing controls for transient exposures. Case-cohort studies allow one comparison group to support the study of multiple outcomes. Case-only studies allow inferences about gene-environment interactions when a control group is impractical. Two-stage designs let researchers spend money on detailed measurement only where it matters most. Knowing the “problem” each design was created to solve makes it much easier to remember when to use it.
The Six Hybrid Designs at a Glance
Click any card to see a brief description of the design and its key feature.
Case-Crossover Studies
The case-crossover study is the observational analogue of the experimental crossover design. Each case serves as its own control by contrasting exposure during a defined time window before the event with exposure during one or more comparison time windows.
Maclure (1991) introduced the design to answer the “why now” question, in contrast to the “why me” question answered by traditional case-control studies. By using the same person as both case and control, the design automatically controls for all time-invariant confounders, including ones the investigator never measured or even thought of. Maclure and Mittleman (2000) review a decade of applications.
When Is a Case-Crossover Design Appropriate?
Three Conditions Must Hold
1. The exposure must be transient. Stable exposures (such as smoking status or chronic medication use) cannot be evaluated because they would be present in all time windows.
2. The outcome must be acute. The event must happen close in time to the exposure if a causal relationship exists. Diseases with long induction periods are unsuitable.
3. The exposure must not be affected by the outcome. If experiencing the event changes future exposure (e.g., a heart attack alters subsequent activity), bidirectional control selection is problematic.
Defining the Risk Period and Control Period
Two design choices drive the validity of a case-crossover study: the length of the risk period (sometimes called the case-risk window) and the strategy for selecting control periods (sometimes called referent periods).
The risk period is the time during which the exposure, if causal, would have produced the event. Choosing a risk period that is too long increases the chance of detecting spurious associations; too short, and real associations may be missed. For physical exertion and myocardial infarction the risk window might be a few hours; for mobile phone use and motor vehicle crashes it might be five minutes; for air pollution effects on respiratory hospitalisations it is typically one day.
Figure 11.1. A symmetric bidirectional case-crossover design. One control window is selected before the event and one after, balancing potential time trends in exposure.
Strategies for Selecting Control Periods
Unidirectional (Backward) Referent Selection
Control periods are chosen only from time before the event. This was the original case-crossover approach. It is the appropriate choice when the event itself alters future exposure, for example when a leg injury changes subsequent training distance, or food poisoning alters what someone eats afterward.
Limitation: If exposure prevalence changes over time (a long-term trend), comparing only earlier control periods with the case-risk period can produce biased estimates.
Symmetric Bidirectional Referent Selection
Control periods are selected both before and after the case event, often equally spaced. The intent is that, if exposure is trending, the higher and lower exposure values from the two flanking control periods will roughly cancel out. This is now the most widely used approach.
Limitation: Cases that occur very early or very late in the study period may have only one control period feasible. Bidirectional selection is only valid if the event itself does not affect future exposure.
Time-Stratified Referent Selection
Janes, Sheppard, & Lumley (2005) proposed this method when shared-exposure data (such as daily air pollution measurements) are available across the entire observation period. The study period is stratified a priori (e.g., by month). When a case occurs, say on a Wednesday in July, all the other Wednesdays in July serve as control periods. This effectively matches on day-of-week and month and avoids the need to specify a single lag time.
Advantage: Eliminates the controversy over how to choose the spacing between case and control periods. Naturally accommodates shared exposure data.
Example 11.1: Weather Events and Waterborne Disease Outbreaks
Thomas and colleagues (2006) studied 92 waterborne disease outbreaks in Canada between 1975 and 2001. They hypothesised that extreme rainfall and warm spring conditions might trigger outbreaks. For each outbreak, the six weeks immediately before onset served as the case-risk period. The 27-year period was stratified into six time windows, and within each non-case window a six-week control period was selected, matched to the case on month, day, and ecozone. Conditional logistic regression identified warmer temperatures and extreme rainfall as plausible contributors.
Notice how the design eliminates the need to find “control communities” that did not have an outbreak, a perennial difficulty in waterborne disease epidemiology. Each outbreak community is its own control.
Example 11.2: Salmonella Outbreak in Long-Term Care
Haegebaert and colleagues (2003) used a case-crossover design within a foodborne Salmonella outbreak that affected mostly residents of chronic-care institutions. Food exposures during the three days before illness onset were compared with food exposures during a control period three days long, ending two days before the case-risk period. Because the illness itself would change subsequent food intake, only earlier (unidirectional) control periods were used. Mantel-Haenszel matched-pair odds ratios were calculated for each meat product. Notice that the design avoided the difficult problem of selecting institutionalised “controls” whose food intake would otherwise have to be matched.
Analysis of Case-Crossover Data
Because each case is matched to one or more control periods within the same individual, the data are analysed as if from a matched case-control study. With one control period per case, the data fit a 2×2 table and McNemar's test applies. Intuitively, this comparison ignores occasions when exposure was the same in both windows and asks, among the occasions where it differed, whether exposure fell in the risk window more often than in the control window. With multiple control periods, conditional logistic regression is the standard approach, and the exponentiated coefficient represents the change in odds of the event associated with a one-unit short-term increase in exposure. This is the matched analysis described in an earlier section, applied to matched sets that each consist of one person observed in several time windows.
When daily exposure data are available for the entire observation period (the “shared exposure” setting common in air pollution studies), the data can equivalently be analysed as a Poisson time series. The two analytical frameworks are mathematically linked when time-stratified referents are used.
Self-Controlled Case-Series Studies
The self-controlled case-series design (often shortened to “case-series” in this literature, but distinct both from the descriptive case series of clinical reports and from the case series of a case-control study described in an earlier section) was developed by Farrington (1995) largely for vaccine safety research. It is a close cousin of the case-crossover design but generalises the comparison from discrete control periods to all of an individual's observation time outside the risk window. Whitaker, Farrington, Spiessens, & Musonda (2006) provide an accessible tutorial.
The Logic of the Design
For each individual who has experienced the outcome of interest, an observation period is defined, namely a calendar window during which exposure history and event occurrence are tracked. Within that observation period, one or more risk periods are designated based on the biology of the exposure (e.g., 6–35 days after vaccination for febrile conditions). All remaining time within the observation period constitutes the control period.
The analysis compares the rate of events during risk time with the rate during control time, after adjusting for the duration of each. As with the case-crossover design, this is a within-person comparison: every time-invariant characteristic of the case (genetics, sex, baseline health) is automatically controlled by design. Age and season can be adjusted for analytically because they vary across the observation period.
Figure 11.2. The observation period for a single case is partitioned into risk periods (after each exposure) and control periods (everything else). The number of events and the duration of each period type drive the relative incidence estimate.
Key Assumptions
If a febrile reaction after a first vaccine dose causes parents to skip the booster, the exposure pattern is no longer independent of outcome. One way to deal with this is to ignore post-event exposures (i.e., consider only the first vaccination). Whitaker and colleagues note that the bias from violating this assumption is often small in practice, but it should be considered explicitly.
If the outcome is death (which clearly ends observation) or a serious illness that prompts withdrawal from the study, the design's assumptions are violated. The standard self-controlled case-series is designed for outcomes that occur and resolve, allowing observation to continue.
Multiple recurrences of the outcome can be included as long as they are conditionally independent given exposure. If they are not (e.g., one event makes another more likely), only first events should be analysed.
If the observation period or risk window does not cover the full duration over which exposure can affect the outcome, any resulting estimate of relative incidence is biased toward the null. Sample size formulae are given in Whitaker, Hocine, & Farrington (2009).
Analysis
The standard analytic tool is a conditional Poisson regression model, where the outcome is the count of events in each risk and control time interval and the logarithm of the duration of each interval is the offset. The parameter of interest is the relative incidence, that is, the rate during the risk period relative to the rate during the control period. Poisson regression models counts of events in relation to the time over which they could occur; a later lesson develops it in the cohort setting, and here it is enough to read the relative incidence as a rate ratio computed within each person.
Example 11.3: Falls and Antihypertensive Medication
Gribbin and colleagues (2011) used UK primary-care databases to study whether starting an antihypertensive medication transiently increased the risk of falls in adults aged 60 and older. They identified 9,862 falls between 2003 and 2006. For each patient, episodes of continuous medication exposure of up to 60 days were defined. After each prescription, the exposure period was further subdivided into day 0, days 1–21, and days 22–60. All remaining person-time was the unexposed baseline. Poisson regression yielded incidence rate ratios for each post-exposure period, allowing the temporal pattern of risk after initiation to be characterised.
Why is this question well-suited to a self-controlled case-series? Because the comparison is within-person, all the patient-level confounders that complicate fall risk (frailty, polypharmacy, comorbidity, age) are automatically controlled.
Key Takeaways
- Hybrid designs are variants of the classic observational designs developed to address specific methodological challenges.
- Case-crossover studies use each case as its own control by comparing exposure during a risk period with exposure during one or more control periods at other times. They are best for transient exposures and acute outcomes.
- Three referent-selection strategies for case-crossover studies are unidirectional, symmetric bidirectional, and time-stratified. The choice depends on whether the event itself alters subsequent exposure and whether time trends in exposure are likely.
- Case-crossover data are typically analysed by conditional logistic regression; with shared daily exposure data, an equivalent Poisson time-series approach is available.
- The self-controlled case-series partitions each case's observation period into risk windows (defined by exposure timing) and control time (everything else). Conditional Poisson regression yields the relative incidence.
- Both designs automatically control all time-invariant confounders, including those the investigator has not measured.
The reflection below asks you to choose between the two self-controlled designs for a real public-health question and to name the assumption you would find hardest to defend. After the reflection and the knowledge check, the next section moves from comparisons within one person across time to comparisons between different kinds of cases.
Reflection
Reflection
A city health department wants to know whether days of heavy wildfire smoke trigger asthma emergency visits. State whether you would use a case-crossover or a self-controlled case-series design, which referent or risk-window strategy you would choose, and which assumption of the design would be hardest to defend.
Minimum 20 characters required.
1. The case-crossover design is most appropriate for studying:
2. A unidirectional (backward-only) referent selection strategy is preferred when:
3. The chief advantage shared by both case-crossover and self-controlled case-series designs is that they:
4. The self-controlled case-series design is most commonly used to study:
Hybrid Observational Designs: Case-Case, Case-Case-Control, and Case-Only Studies
⏱ Estimated reading time: 17 minutes
Introduction and Overview
The previous section covered designs that use each case as their own control across time. This section turns to designs that compare different kinds of cases to one another, useful when traditional controls are unavailable, when subtypes of disease have different aetiologies, or when surveillance data only contain ill people. Each design in this section sacrifices something in exchange for not needing a healthy control group.
Learning Objectives
- Describe the design and applications of case-case studies.
- Distinguish case-case studies from case-case-control studies.
- Explain the logic of case-only studies for evaluating gene-environment interactions.
- Recognise the assumptions and limitations of each design.
Case-Case Studies
The case-case design is a variant of the case-control design in which the comparison group consists of cases of a different disease subtype drawn from the same surveillance system. McCarthy and Giesecke (1999) proposed it as an efficient way to identify risk factors that distinguish closely related etiological subgroups using routine surveillance data.
For example, the cases might be people infected with Salmonella Typhimurium, while the “controls” might be people infected with Salmonella Heidelberg. Both groups have salmonellosis (both are cases), but the design seeks to identify exposures that distinguish one serotype from the other.
When Is the Case-Case Design Useful?
Two Common Settings
1. Identifying differential risk factors for related endemic diseases. When all subjects who appear in a surveillance system have undergone similar selection (e.g., they all sought medical care, all had stool cultured), comparing them with one another minimises selection bias and recall bias. Comparing them to community controls who never had any salmonellosis would be far more vulnerable to these biases.
2. Distinguishing outbreak cases from sporadic cases of the same organism. In an outbreak investigation, the cases are people whose isolates match the outbreak strain. The “controls” are sporadic cases of the same serotype during the same time window. The exposures that differentiate them point to the outbreak vehicle.
Strengths and Limitations
Comparable selection experience. Because both groups appear in the same surveillance system, both have passed through similar diagnostic and reporting filters. Selection bias is minimised.
Comparable recall experience. Both groups have had a similar clinical experience (an episode of gastrointestinal illness). Their motivation to recall recent food exposures is similar, reducing differential recall bias.
Efficient use of surveillance data. No new control recruitment is required; the comparison group is already in the database.
Cannot identify shared risk factors. Exposures that cause both serotypes equally (such as eating any contaminated food) will not be detected because they are present in both groups.
Surveillance limitations. Wilson and colleagues (2008) note tendencies for selection bias (only severe cases reported), information bias (data collected by people who know the diagnosis), confounding (limited covariate information), and lack of detail on exposure.
The OR is not a true risk measure. Because the “controls” are not drawn from the underlying source population, the odds ratio reflects the relative difference in exposure between two case subtypes, not the absolute risk of either disease.
If the analysis identifies poultry consumption as a stronger risk factor for S. Typhimurium than for S. Heidelberg, this does not tell us that eating poultry causes Typhimurium in absolute terms. Rather, it tells us that poultry consumption is more strongly associated with the Typhimurium subgroup than with the Heidelberg subgroup, which is useful for tracing distinct food sources or transmission routes.
To get an absolute risk estimate, a traditional case-control study with population-based controls would still be required. Case-case findings are best treated as hypothesis-generating about subtype-specific exposures.
Example 11.4: Two Campylobacter Species
Gillespie and colleagues (2002) used population-based surveillance data from England and Wales to compare the exposure histories of people with Campylobacter coli infection (the much rarer species) with those of people with Campylobacter jejuni infection. Standard structured questionnaires from the surveillance system provided the exposure data. Backward stepwise logistic regression identified differential risk factors and tested for interaction. The authors emphasised that exposures common to both species would not be detected by this design; only those that distinguish the two species could emerge.
If you wanted to know which exposures were common to both species, what design would you use instead?
Example 11.5: A Salmonella Outbreak in Germany
Krumkamp and colleagues (2008) investigated a 2003 outbreak of Salmonella 1,4,[5],12:i:- in a German district. Ten outbreak cases were compared with 97 sporadic cases of other Salmonella serotypes that occurred in the same area during the same year. Telephone interviews collected exposure histories. Fisher's exact tests and odds ratios identified meat sold from a single butcher shop as the only significant risk factor, a finding that would have been very difficult to achieve with a traditional community-based control group.
Analysis
Case-case data are analysed by the same techniques as the risk-based case-control studies described in an earlier section, typically logistic regression. The exponentiated coefficient is interpreted as the relative odds of exposure between the two case subtypes, not as a risk ratio.
Case-Case-Control Studies
The case-case-control design (Kaye, Harris, Samore, & Carmeli, 2005) was developed to overcome a specific limitation of traditional case-control studies in the context of antimicrobial resistance. The original example was vancomycin-resistant Enterococcus (VRE) versus vancomycin-susceptible Enterococcus (VSE).
The Problem the Design Solves
Suppose you want to identify risk factors for VRE infection. A traditional case-control approach would compare VRE cases with non-infected controls. But many of the exposures associated with VRE (prior antibiotic use, prolonged hospitalisation, ICU stay) are also strong risk factors for VSE, and indeed for any hospital-acquired infection. So a traditional design tells you what causes hospital infection in general, not what specifically drives the resistant phenotype.
An alternative is a case-case design comparing VRE with VSE. But Kaye and colleagues argue against this: VRE often emerges from external sources (transmission of an already-resistant strain) rather than from within-patient evolution of a susceptible strain. So contrasting VRE directly with VSE conflates “risk of acquiring a resistant strain” with “risk of selection pressure on a susceptible one.”
The Case-Case-Control Solution
The design uses two case series (resistant and susceptible) and one control series (people without infection from the same source population). Two separate logistic regression models are fitted, one comparing each case series with the controls. The risk factors are then sorted into three categories.
Figure 11.3. In the case-case-control design, two case series are each compared separately with the same control series. Comparison of the two resulting models identifies which risk factors are unique to the resistant phenotype.
Interpreting the Three Variable Categories
These are risk factors unique to the resistant phenotype. They are the variables most useful for understanding what drives resistance specifically. In the original VRE example, exposure to vancomycin itself or to a roommate carrying VRE might fall in this category.
These are risk factors unique to the susceptible phenotype. They tell us what predisposes to acquiring the susceptible strain in particular (perhaps community sources for the susceptible organism but not for the resistant one).
These are risk factors for the target organism in general, regardless of resistance status. Hospitalisation, prior antibiotic exposure, indwelling catheters, and severity of illness typically appear here. They are real risk factors but they do not distinguish resistance from susceptibility.
Design Considerations
- First positive culture per patient. Only the first positive culture should be included to avoid double-counting; for nosocomial infection studies, restrict to cultures taken >48 hours after admission.
- Source population matters. Controls should come from the same source population as the cases. For nosocomial infections, controls should be other patients hospitalised >48 hours, ideally with documented negative cultures for both phenotypes.
- Confounding control. Because there is only a single control series, restricted sampling and matching are difficult. Confounding is typically handled through multivariable unconditional logistic regression.
Example 11.6: MRSA Colonisation in an ICU
Melo and Fortaleza (2009) investigated risk factors for nasopharyngeal colonisation with methicillin-resistant Staphylococcus aureus (MRSA) in an ICU. They enrolled 122 patients who had been screened weekly for S. aureus colonisation. The two case series were patients colonised with MRSA and patients colonised with methicillin-susceptible S. aureus (MSSA). Controls were patients in whom no colonisation was detected during their ICU stay. Comparing the resulting two models revealed which exposures were specifically associated with the resistant phenotype rather than with general susceptibility to S. aureus colonisation.
Case-Only Studies
The case-only design uses only cases, with no observed control group recruited. The expected exposure distribution in the hypothetical “control population” is derived from theoretical or external sources. The design originated in genetic epidemiology, where the population frequency of common alleles can often be specified from external reference data (Khoury & Flanders, 1996).
Key Concept: Effect Modification and Interaction
Effect modification is variation in the association between an exposure and an outcome across levels of a third variable, and it is always stated on a scale: additive (comparing risk differences) or multiplicative (comparing risk or odds ratios), because an association can be modified on one scale and not on the other. Statistical interaction is its representation by a product term in a model, and the case-only design estimates interaction on the multiplicative scale. Lesson 6, Section 3 (Inferential Errors and Sources of Ecologic Bias) develops the concept for group-level variables.
What the Design Can and Cannot Estimate
Important Restriction
The case-only design cannot estimate main effects; it cannot tell you whether a gene or an environmental exposure independently raises the risk of disease. What it can estimate is interaction between two factors among cases, provided the two factors are independent of each other in the source population.
The intuition is this: among cases, if a genetic risk factor and an environmental exposure are independent in the source population but appear together more often than expected by chance, that excess co-occurrence is evidence of statistical interaction on the multiplicative scale. If the gene and exposure were also causally associated in the source population (not independent), this signal would be confounded.
Required Assumptions
- Independence in the source population. The exposure and the proposed effect modifier (often a gene, but can be sex, age, or another stable trait) must be independent in the population from which cases arose. For a heritable polymorphism not influenced by the environmental exposure, this is biologically plausible.
- Stable, well-defined effect modifier. Genetic variants, sex, race, and age are common choices because they don't change over time and can be measured reliably.
- The disease must be rare. Like many odds-ratio-based estimators, the case-only interaction estimate approximates the true interaction parameter most closely when the outcome is rare in the source population.
Recent Extensions Beyond Genetics
The design has been extended to study how non-genetic stable characteristics modify the effects of time-varying exposures. Armstrong (2003) and Schwartz (2005) used case-only designs to ask whether sex, race, age, or socioeconomic class modify the effect of extreme weather on mortality. Because age, sex, and socioeconomic class can reasonably be considered independent of daily weather exposures, the case-only approach yields a valid interaction estimate.
Analytic Logic
Suppose we are interested in whether sex modifies the effect of an extreme heat day on mortality. A case-only logistic regression takes the form:
The logic feels strange at first because we appear to be modelling a covariate (sex) as a function of an exposure. But this is mathematically equivalent to a Poisson model of mortality count as a function of heat, sex, and a heat×sex interaction term, where the case-only regression coefficient is the interaction term. The trick is that we never need a control group at all; the “control” expectation is built into the assumed independence of sex and heat in the source population.
Example 11.7: Effect Modifiers of Mortality from Temperature Extremes
Schwartz (2005) investigated whether sex, non-white race, or age over 85 modified the effect of extreme temperatures on mortality in Wayne County, Michigan. Weather data identified excessively hot and cold days. Demographic data on people who died came from medical records. Separate models were fitted for heat and for cold, and one-day and three-day average temperature exposures were both examined. All three covariates emerged as effect modifiers. Notice that no control group of survivors was needed; the inference depended on whether the demographic profile of cases differed between extreme-weather days and other days.
Why is this design especially appealing for studying mortality? Because building a comparable control group of “people who did not die” is conceptually awkward when daily death registry data are already complete.
Example 11.8: Heat Waves and Hospital Admissions in New South Wales
Khalaj and colleagues (2010) used a case-only design to identify which underlying medical conditions raised the risk of hospital admission during heat waves across five regions of New South Wales, Australia. Daily admission records and weather data covered the warm months of 1998–2006. The analysis fitted logistic regression models with each primary diagnosis as the “outcome” and an extreme-heat indicator as the predictor. Sine and cosine terms were included to control for season, since otherwise season could confound the interaction (some chronic conditions have stronger seasonal patterns than others).
Key Takeaways
- Case-case studies compare two related disease subtypes drawn from the same surveillance system, identifying differential risk factors while minimising selection and recall bias. The OR reflects relative differences in exposure between subtypes, not a true risk measure.
- Case-case-control studies use two case series (e.g., resistant and susceptible) compared separately with one control series, sorting risk factors into category A (unique to resistance), B (unique to susceptibility), and C (shared by the organism in general).
- Case-only studies use only cases and rely on external knowledge of the exposure distribution in “controls.” They estimate interaction between an exposure and an effect modifier, but not main effects, and require independence between the two factors in the source population.
- All three designs use cases-only or two-case data structures because constructing a satisfactory traditional control group would be impractical, biased, or uninformative for the specific research question.
The reflection below asks you to apply a case-only comparison design to routine surveillance data and to say what its odds ratio can and cannot tell you. After the reflection and the knowledge check, the next section turns to designs that subsample a large cohort so that expensive measurement becomes affordable.
Reflection
Reflection
A provincial surveillance system holds standard exposure questionnaires for every laboratory-confirmed case of two related enteric infections but no healthy controls. Which case-only comparison design would you use to find exposures that distinguish the two infections, what would its odds ratio mean, and what could the design not tell you?
Minimum 20 characters required.
1. A case-case study comparing Salmonella Typhimurium with Salmonella Heidelberg cases:
2. In a case-case-control study of vancomycin-resistant Enterococcus, a Category C variable is one that:
3. The case-only design's most important limitation is that it:
4. A case-only study asks whether a common genetic polymorphism modifies the effect of a dietary exposure on colorectal cancer. Its interaction estimate is valid only if:
Hybrid Observational Designs: Case-Cohort and Two-Stage Sampling
⏱ Estimated reading time: 18 minutes
Introduction and Overview
The two previous sections were about designs that use only cases, or that compare one case series with another. This section returns to designs that use cohorts as their backbone but subsample from them to make expensive measurements feasible. Both case-cohort and two-stage designs let you mount essentially a cohort study while only paying for biomarker measurements on a fraction of the participants, the kind of design that makes large biobanks practical. A later lesson treats cohort studies in full; here a cohort simply means a defined group of people followed forward in time from a common starting point.
Learning Objectives
- Describe the structure and rationale of the case-cohort design.
- Distinguish risk-based from rate-based case-cohort analyses.
- Explain why a single subcohort can support investigation of multiple outcomes.
- Describe the logic of two-stage sampling and identify when it is most efficient.
- Design a basic two-stage sampling strategy for a case-control study.
Case-Cohort Studies
The case-cohort design, introduced by Prentice (1986), combines features of cohort and case-control studies. From a defined source cohort, the investigator draws a random sample called the subcohort at the start of follow-up. Detailed exposure and covariate data are obtained on the subcohort. As follow-up proceeds, all incident cases that arise from the full source cohort, whether or not they happen to fall within the subcohort, are also studied.
The design has the same advantages as a full cohort study (clear temporal ordering, multiple outcomes, direct disease frequency estimates) but achieves them with much smaller measurement costs because expensive covariate or biomarker assays are performed only on the subcohort plus the cases, not on the entire source cohort.
Case-cohort sampling is a third way of drawing a comparison series from a study base, beside the cumulative and incidence density schemes described in an earlier section. The comparison series is drawn once, at baseline, without regard to who later becomes a case, so a subcohort member who later develops the disease belongs to both series. It also differs from the nested case-control study introduced at the start of this lesson, in which the controls are a subsample of the cohort members who have not become cases.
Figure 11.4. The case-cohort layout. Detailed exposure and covariate data are needed only on the random subcohort plus the cases. Most of the full source cohort never requires expensive measurement.
The Big Win: Multiple Outcomes from a Single Subcohort
Why Researchers Love Case-Cohort Designs
One subcohort can serve as the comparison group for multiple disease outcomes. If researchers are interested in cardiovascular disease, several cancers, and diabetes within the same large cohort, they need only one set of expensive biomarker measurements on the subcohort. Each outcome study then adds detailed measurements on its own cases. By contrast, a nested case-control study requires fresh control selection for each outcome, and a full cohort analysis would require measuring everyone for everything.
Risk-Based vs. Rate-Based Designs
Risk-Based (Closed-Cohort) Case-Cohort
Suitable when the source cohort is closed (a fixed group followed for a defined period) and exposures are stable over follow-up. The subcohort is sampled by simple or stratified random sampling at the start of follow-up. Cases arising outside the subcohort during follow-up are added.
Analysis: Combine the two case groups (those in and those outside the subcohort) and analyse the data in the familiar 2×2 case-control format using logistic regression. The odds ratio approximates the relative risk when the disease is rare.
Example: Matsuda and colleagues (2011) studied placental abruption and placenta previa among 5,036 of 242,715 births in Japan, using multivariable logistic regression with the subcohort plus all cases.
Rate-Based (Open-Cohort) Case-Cohort
Suitable when the source cohort is open (entries and exits possible during follow-up) or when exposures change over time. At the moment a case occurs, eligible members of the subcohort are those who have not yet experienced the outcome. Their current exposure status (which may have been updated through repeated surveys or stored serial samples) is recorded.
Analysis: A weighted Cox proportional-hazards model is the standard approach. Weights account for the sampling fraction; for example, if the subcohort represents 20% of the source cohort, controls are typically up-weighted by 5. Three Cox-weighting schemes have been proposed historically; Prentice's method most closely reproduces the estimates from a full-cohort analysis.
Cox regression (the proportional-hazards model) relates the exposure and covariates to the rate at which events occur over follow-up time, and its exponentiated coefficient is a hazard ratio, a ratio of instantaneous event rates. A later lesson develops Cox and Poisson regression in the cohort setting; here the point is that the weights let the subcohort stand in for the full cohort at each event time.
Example: Agalliu and colleagues (2011) followed a subcohort of 1,979 men for prostate cancer risk, with exposure and supplement use updated through repeated surveys.
Practical Considerations
- Eligibility for the subcohort. Members must be willing to provide health history, lifestyle data, and (often) biological samples. Stratified sampling can ensure that the subcohort's covariate profile matches anticipated cases (e.g., over-sampling young adults if young-adult disease is the focus).
- Stored specimens. Serially stored tissue or blood samples allow detection of exposure changes over time and support post-hoc biomarker assays as new hypotheses emerge.
- Sampling adjustments for non-response. If 20% are sampled but only 80% of those agree to participate, the weighting should reflect the actual participation, not the original sampling probability.
- Robust standard errors are recommended for case-cohort analyses to account for the sampling variability.
- Clustering of cases. If cases tend to be diagnosed at the same clinic, marginal models with adjusted variances or frailty models should be used to account for within-cluster correlation.
Example 11.9: Drinking Water Quality and Stomach Cancer
Auvinen and colleagues (2005) studied radon and other radionuclides in drinking water and the risk of stomach cancer in a Finnish population of over 144,000 people who drew their water from drilled wells between 1967 and 1980. An initial subcohort of 4,590 was sampled with stratification by age and sex. Many of these did not actually meet the long-term-exposure criterion, leaving an effective subcohort of 371 long-term users. Stomach cancer cases (n=107) were identified through the cancer registry. Water samples were collected blindly with respect to case status and analysed for radionuclides. A proportional-hazards model accounted for how long each subject had been exposed to each level of radon. All hazard ratios were below 1, suggesting a protective association, a surprising finding that illustrates how case-cohort designs can efficiently support investigation of unusual exposures using stored samples and registry data.
Two-Stage Sampling Designs
A two-stage (or two-phase) sampling design is a strategy that can be layered on top of any traditional design, whether cohort, case-control, or cross-sectional. The first stage collects readily available, inexpensive data on a large group. The second stage collects more detailed (and usually more expensive) data on a strategically selected subsample.
The same logic appears at the group level in ecological research, where a two-phase design supplements aggregate data with individual-level data on a sample from each area; a later lesson on ecological studies describes that version.
Why Two-Stage Designs Make Sense
The Core Problem
Imagine you want to study whether occupational solvent exposure increases birth defect risk. Hospital records can give you basic information on hundreds of thousands of pregnancies cheaply, but a detailed occupational exposure assessment requires a one-hour interview at $200 per participant. Spending $200 on every pregnancy is unaffordable. A two-stage design lets you do the cheap step on everyone and the expensive step only where the information is most valuable.
Three Common Use Cases
The first stage uses an inexpensive surrogate exposure measure (e.g., job title from a registry). The second stage performs a detailed work-up (e.g., personal interview about specific solvent contacts, dose, duration) on a subsample. This is the most common application.
If the inexpensive first-stage measure has known measurement error, the second stage applies a near-gold-standard measurement to a subsample. The relationship between the two measures (the “measurement model”) is estimated, and inferences from the full first-stage data set can be corrected for the measurement error. McNamee (2002, 2005) describes optimal designs for this purpose. A later lesson on information bias discusses measurement error and validation sub-studies in more detail.
When key covariate data are missing for many subjects, instead of assuming missingness at random, the missing-data subjects can be the explicit target of the second-stage data collection. This concentrates resources on filling specific gaps rather than dropping incomplete records.
How to Sample at Stage 2
The key design question in any two-stage study is: how should we choose whom to include in the expensive second stage? The optimal answer depends on the design.
| Stage 1 Design | Recommended Stage 2 Sampling | Rationale |
|---|---|---|
| Cohort | Fixed numbers of exposed and unexposed | Balanced sampling on exposure ensures precision in the exposure–outcome estimate. |
| Case-control | Fixed numbers of cases and controls | Balanced sampling on disease ensures precision; oversampling cases is efficient when the disease is rare. |
| Either, when surrogate exposure is available | Approximately equal numbers from each of the four exposure× | Optimal efficiency: extracts the most information from a fixed second-stage budget by ensuring that small cells (rare combinations) are not under-represented. |
A Worked Two-Stage Case-Control Sampling Strategy
Suppose your stage 1 data come from a hospital registry of 50,000 pregnancies. From the registry you can identify 2,000 birth-defect cases and 48,000 non-cases. A crude (and possibly mismeasured) exposure indicator, namely whether the mother held a job classified as “industrial”, is available for everyone.
Stage 1 cross-classification might look like this:
| Industrial job (surrogate exposed) | Other job (surrogate unexposed) | Total | |
|---|---|---|---|
| Cases | 120 | 1,880 | 2,000 |
| Non-cases | 1,500 | 46,500 | 48,000 |
For stage 2, balanced sampling across the four cells is the most efficient strategy. Suppose your budget allows 400 detailed interviews:
- Industrial-exposed cases (120 available): Take all 120.
- Other-exposure cases (1,880 available): Sample 100.
- Industrial-exposed non-cases (1,500 available): Sample 100.
- Other-exposure non-cases (46,500 available): Sample 80.
Because we have oversampled the small cells, weighting must be applied in analysis to recover correct association estimates (inverse probability weighting, in which each interviewed person is weighted by the reciprocal of the sampling fraction in their cell); this is what the Cain & Breslow (1988) and later Flanders & Greenland (1991) methodologies handle. Hanley and colleagues (2005) provide worked examples of the adjusted odds ratio and its variance.
Example 11.10: A Two-Stage Case-Control Study of Childhood Asthma
Martel and colleagues (2009) used a two-stage design with three linked Quebec administrative health databases. Stage 1 was a nested case-control study within a cohort of pregnant women and their children: 5,226 asthmatic children (cases) and 20 non-asthmatic children per case were selected using density sampling matched to time of case occurrence. Covariate data from the administrative databases were used at this stage. Stage 2 was a mailed questionnaire to a subsample of mothers, balanced across the cells of the first-stage exposure–outcome cross-table to overrepresent small cells. Conditional logistic regression was used at stage 1; unconditional logistic regression with sample-fraction weighting at stage 2. Final corrected estimates were obtained by combining the stages.
Notice how the design exploits the cheap administrative data to identify cases and screen on rough covariates, while reserving the expensive questionnaire for the subset where new information will most improve the estimate.
Practical Pitfalls
- Budget allocation between stages. Hanley and colleagues (2005) note that tools for optimal allocation have not advanced much in the past two decades; in practice, simulation can help determine the relative number of stage 1 and stage 2 subjects to maximise precision under a fixed budget.
- Variance estimation. The variance of the final estimate depends on both stages. Naive variances that ignore stage 2 sampling will be too small. Hanley and colleagues provide details for dichotomous covariates; software exists for more complex situations.
- Sampling fractions must be known. Weighting requires knowing exactly what proportion of each stage 1 cell was sampled at stage 2. If non-response or other losses change those fractions, the realised (not planned) fractions should be used.
Key Takeaways
- Case-cohort studies sample a random subcohort at the start of follow-up and add all incident cases. Detailed measurement is needed only on the subcohort plus the cases, not the full source cohort.
- A single subcohort can serve as the comparison group for many outcomes, making the design particularly attractive for large prospective studies with stored biological specimens.
- Risk-based case-cohort analyses combine subcohort and outside cases in a logistic regression. Rate-based analyses use weighted Cox models, with weights reflecting the inverse sampling probability.
- Two-stage sampling lets investigators pay for cheap, low-quality data on everyone and high-quality data only on a strategically chosen subsample.
- For two-stage case-control studies, the most efficient stage 2 sampling allocates approximately equal numbers across the four cells of the stage 1 exposure×disease table.
- Two-stage analyses must use weighting to recover unbiased estimates and must use variance formulae that account for both stages of sampling.
The reflection below asks you to match one of the hybrid designs to a question of your own and to name the assumptions you would have to defend. After the reflection and the knowledge check, the lesson moves to its final assessment, which draws on all seven sections.
Reflection
Reflection
Consider a research question of interest to you. It might involve a transient environmental trigger of an acute event, an outbreak you would like to characterise, a gene-environment interaction, or a long cohort follow-up where biomarker measurement is expensive. The hybrid designs are the case-crossover study (each case serves as their own control, comparing exposure just before the event with exposure at other times), the self-controlled case series (event rates are compared within individuals between exposed and unexposed periods), the case-case study (cases of one subtype are compared with cases of another), the case-only design (interaction is estimated from cases alone, assuming the two factors are independent in the population), the case-cohort study (a random subcohort sampled at baseline is compared with all cases), and the two-stage design (cheap data are collected on everyone and expensive data on a chosen subset). Which hybrid design would you choose, and why? What assumptions would you need to defend, and what limitations would you have to acknowledge in your discussion?
Minimum 20 characters required.
1. The key efficiency advantage of a case-cohort design over a full cohort study is that:
2. In a rate-based case-cohort analysis of an open cohort, the standard analytic approach is:
3. For a two-stage case-control study with a binary surrogate exposure measure available at stage 1, the most statistically efficient stage 2 sampling strategy is to:
4. A two-stage sampling design is most useful when:
Final Review & Assessment
⏱ Estimated time: 20 minutes
Bringing It All Together
This lesson worked from the inside out. An earlier section fixed the conceptual core, cases and controls must arise from a single, well-defined study base, and getting that right is what separates a clean primary-base design from a shaky secondary-base one. Earlier sections then translated that core into mechanics: how to define and recruit a case series, the four principles of control selection, and the choice between risk-based sampling (which gives you the OR as an approximation of the risk ratio) and rate-based / incidence density sampling (which gives you the OR directly as a rate ratio).
An earlier section closed the loop by moving from design to execution and reporting, choosing the number of controls, deciding when to use multiple control groups, ensuring comparability through exclusion, matching, and analytic control, and then communicating the whole package transparently using STROBE. Read end-to-end, the lesson is a single argument: case-control studies are valuable precisely because they are efficient, but every efficiency has a price, and the design choices you make have to be explicit, defensible, and reported.
The last three sections extended the design into its hybrid variants. Case-crossover and self-controlled case-series designs use each case as their own control across time; case-case and case-case-control designs substitute or add a second case series for the traditional control group; the case-only design estimates interaction without any control series; case-cohort sampling lets one random subcohort serve several outcomes; and two-stage sampling layers cheap measurement on everyone with expensive measurement on a strategically chosen subsample. Each hybrid responds to a specific limitation of the standard design, such as recall bias, between-person confounding, costly biomarker measurement, or the lack of a usable control group, by changing the comparison group or the sampling rule, and each carries an assumption that has to be stated and defended in the same way.
The final reflection asks you to put that argument to work by sketching a brief case-control proposal of your own. The 24-question assessment then checks the conceptual content of all seven sections directly. From here, a later lesson turns the design around to follow exposed and unexposed people forward in time (cohort studies) and a lot of the vocabulary you just built (study base, comparability, sampling logic) will travel with you.
Key Takeaways from this lesson
- A valid case-control study begins with a clearly specified study base; primary-base designs make it explicit, secondary-base designs reconstruct it.
- Cases must satisfy a stable diagnostic definition; whether you use incident or prevalent cases changes what your odds ratio means.
- Controls must be sampled from the same source population as the cases, independently of exposure, the four principles of control selection are non-negotiable.
- Risk-based sampling and rate-based (incidence density) sampling answer different questions; the OR estimates a risk ratio in one and a rate ratio in the other.
- Comparability is engineered, not assumed (through exclusion, matching, and analytic control) and matched designs require matched analyses.
- Transparent reporting using STROBE is the bridge between a defensible design and a study other people can appraise, replicate, or extend.
- Case-crossover studies and self-controlled case-series use each case as their own control across time, which suits transient exposures and acute outcomes; the choice of control periods depends on whether the event alters later exposure and whether exposure trends over time.
- Case-case, case-case-control, case-only, case-cohort, and two-stage designs each replace, reuse, or subsample a comparison group to gain efficiency, and each rests on an assumption (a shared surveillance source, independence of the two factors, a subcohort drawn without regard to outcome, known sampling fractions) that must be defended.
Reflection
Design a brief case-control study proposal for a health question of your choice. Specify: (1) the research question, (2) whether you would use a primary or secondary study base and why, (3) how you would define and identify cases, (4) how you would select controls and from what source, (5) whether a risk-based or rate-based design is more appropriate, and (6) how you would ensure comparability.
Minimum 20 characters required.
Final Knowledge Assessment
This assessment covers all sections of this lesson. You must score 100% to complete the lesson. Review the feedback after each attempt.
1. The fundamental logic of a case-control study is to:
2. A case-control study is NOT a comparison between cases and:
3. A secondary study base refers to a source population that is:
4. A unique advantage of a nested case-control study is that it can:
5. Why are incident cases preferred over prevalent cases?
6. According to the principles of control selection, controls should:
7. The odds ratio in a risk-based case-control study (Eq 9.1) is calculated as:
8. In rate-based case-control designs, the OR estimates the incidence rate ratio because:
9. What is incidence density sampling?
10. What is the practical maximum number of controls per case before benefits diminish?
11. Why might using friend controls lead to biased estimates?
12. In Example 9.3, the secondary-base case-control study of prostate cancer used controls from:
13. When ascertaining exposure in case-control studies, what is recommended?
14. The general experience regarding multiple control groups is that:
15. According to STROBE, which item is specific to case-control study reporting?
16. The case-crossover design controls for time-invariant confounders by:
17. Investigators propose a self-controlled case-series to study whether a new medication triggers sudden cardiac death. Which assumption of the design does this outcome violate most directly?
18. A case-crossover study of daily air pollution and asthma admissions takes, as control periods for a case admitted on a Wednesday in July, all the other Wednesdays in that July. This referent strategy is called:
19. A case-case study comparing Campylobacter coli with Campylobacter jejuni would NOT be useful for identifying:
20. In a case-case-control study of vancomycin-resistant Enterococcus (VRE), prior vancomycin use is a significant risk factor in the model comparing VRE cases with uninfected controls, but not in the model comparing vancomycin-susceptible cases with the same controls. Vancomycin use therefore belongs to:
21. In Schwartz’s (2005) case-only study of temperature extremes in Wayne County, Michigan, no control group of survivors was recruited. The evidence that age over 85 modified the effect of extreme heat on mortality came from:
22. A major attractive feature of the case-cohort design is that:
23. How does the comparison series in a case-cohort study differ from the controls in a nested case-control study that uses incidence density sampling?
24. In the worked two-stage example in this lesson, all 120 surrogate-exposed cases but only 80 of the 46,500 surrogate-unexposed non-cases were interviewed at stage 2. How should the analysis treat these interviewed subjects?