HSCI 341 · Lesson 5

Screening &
Diagnostic Tests

Fundamental Epidemiological Concepts and Approaches

Learning objectives for this lesson:

  • Define accuracy and precision as they relate to test characteristics
  • Interpret measures of precision for quantitative tests and calculate kappa for categorical tests
  • Define sensitivity and specificity, and calculate their estimates and confidence intervals
  • Define predictive values and explain the factors that influence them
  • Choose appropriate cutpoints using ROC curves and likelihood ratios
  • Use multiple tests and interpret results in series or parallel

This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University based on Dohoo, I. R., Martin, S. W., & Stryhn, H. (2012). Methods in Epidemiologic Research. VER Inc.

Reference

Glossary: Key Terms, People & Concepts

📚 Reference page, available throughout the lesson

This glossary collects the key concepts, people, and ideas you will meet in this lesson. Use it as a reference while you work through the material, or as a review before assessments. Type in the search box to filter entries.

Test Performance Concepts
Screening Test A test applied to asymptomatic individuals to identify those at higher risk of disease so they can undergo further diagnostic evaluation. Goal is early detection in apparently healthy people.
Diagnostic Test A test used in symptomatic individuals or those with positive screens to confirm or rule out disease. Generally more invasive, expensive, and accurate than screening tests.
Gold (Reference) Standard The best available test or set of criteria used to define true disease status when evaluating a new test. The benchmark against which sensitivity and specificity are measured.
Accuracy How close a measurement is to the true value. In test evaluation, the proportion of all results (positive and negative) that are correct.
Precision The reproducibility or repeatability of measurements: how close repeated measurements are to one another, regardless of accuracy.
Sensitivity (True Positive Rate) The probability that a test correctly identifies a person with the disease: P(test+ | disease+). High sensitivity is needed to rule out disease (SnNout).
Specificity (True Negative Rate) The probability that a test correctly identifies a person without the disease: P(test− | disease−). High specificity is needed to rule in disease (SpPin).
Positive Predictive Value (PPV) Among those who test positive, the proportion who actually have the disease: P(disease+ | test+). Strongly dependent on the prevalence in the tested population.
Negative Predictive Value (NPV) Among those who test negative, the proportion who truly do not have the disease: P(disease− | test−).
Prevalence Threshold The point along the prevalence axis below which the predictive value of a test deteriorates rapidly. A reminder that a test’s usefulness depends on the population it is applied to.
False Positive A test result that incorrectly indicates disease in someone who is disease-free. Tied to specificity (1 − specificity = false positive rate).
False Negative A test result that incorrectly indicates absence of disease in someone who actually has it. Tied to sensitivity (1 − sensitivity = false negative rate).
Cutpoint (Threshold) The numeric value of a continuous test that separates “positive” from “negative.” Lowering the cutpoint typically raises sensitivity at the expense of specificity.
Methods & Measures
ROC Curve Receiver Operating Characteristic curve: a plot of sensitivity (true positive rate) vs. 1 − specificity (false positive rate) across all possible cutpoints, used to compare tests and choose thresholds.
Area Under the Curve (AUC) A summary of overall test discrimination from the ROC curve, ranging from 0.5 (no better than chance) to 1.0 (perfect). Equivalent to the c-statistic.
Positive Likelihood Ratio (LR+) Sensitivity / (1 − specificity). Indicates how much a positive test increases the odds of disease. Values > 10 strongly rule in disease (Deeks & Altman, 2004).
Negative Likelihood Ratio (LR−) (1 − sensitivity) / specificity. Indicates how much a negative test decreases the odds of disease. Values < 0.1 strongly rule out disease.
Cohen’s Kappa A chance-corrected measure of agreement between two raters or tests on categorical outcomes, κ = (po − pe) ÷ (1 − pe). Ranges from −1 to 1, with 0 for chance agreement, 1 for perfect agreement and negative values when agreement is below chance.
Series Testing Sequential testing in which a positive on one test triggers a second test; disease is declared only if both are positive. Increases overall specificity, decreases overall sensitivity.
Parallel Testing Two tests done simultaneously; disease is declared if either is positive. Increases overall sensitivity, decreases overall specificity.
Screening Programme Concepts
Lead-Time Bias Apparent improvement in survival that arises only because screening detects disease earlier in its course, even when actual time of death is unchanged.
Length Bias Tendency for screening to preferentially detect slow-progressing (longer pre-clinical phase) cases, making screened cases appear to have better outcomes than non-screened cases.
Overdiagnosis Detection of disease that would never have caused symptoms or harm in the patient’s lifetime. Inflates apparent screening benefit and exposes patients to unnecessary treatment (Welch & Black, 2010; Brodersen et al., 2018).
Wilson & Jungner Criteria Ten classic criteria (Wilson & Jungner, 1968, WHO) for evaluating whether a screening programme is appropriate, covering the disease, the test, treatment availability, costs, and ethics. See also Wikipedia: Screening (medicine) and the modern revisit by Andermann et al. (2008).
No matching entries. Try a different search term.
Section 1

Introduction & Test Attributes

⏱ Estimated reading time: 12 minutes

Lesson 5 · HSCI 341

From Populations to Individual Results

Given a single test result, what does it actually tell you about this person?

Section 1 of 4

Introduction & Test Attributes

What a test is, the difference between analytic and diagnostic properties, and how we quantify precision and agreement.

What is a test?

Screening vs. diagnostic tests

Screening test

Applied to apparently healthy populations. Goal: early detection before symptoms appear.

Diagnostic test

Applied to individuals already suspected of disease. Goal: confirm or rule out a condition.

Evaluation principles are the same for both. The clinical context differs.

A critical distinction

Analytic vs. diagnostic sensitivity and specificity

Analytic sensitivity

The lowest concentration of a compound the test can detect. A laboratory property.

Diagnostic sensitivity

The probability that a truly diseased person tests positive. An epidemiological property, covered next.

Analytic specificity: cross-reactivity with other compounds. Diagnostic specificity: correct negative rate. Different quantities entirely.

Test quality

Accuracy and precision

Accurate & precise

Mean near truth; low scatter. The ideal.

Precise, not accurate

Consistent results, but all biased from the truth.

Accurate, not precise

Mean near truth; high variability.

Neither

High scatter and a biased mean.

Coefficient of variation = σ / μ. Lower values indicate greater precision.

Categorical agreement

Cohen's kappa (κ)

Kappa (Cohen, 1960)
\[ \color{#0B7B6B}{\kappa} = \frac{\color{#C2410C}{p_o} - \color{#6D28D9}{p_e}}{1 - \color{#6D28D9}{p_e}} \]
κ chance-corrected agreementpₒ observed agreementpₑ agreement expected by chance

κ ≤ 0: Poor  |  0.01–0.20: Slight  |  0.21–0.40: Fair

0.41–0.60: Moderate  |  0.61–0.80: Substantial  |  0.81–1.00: Almost perfect  (Landis & Koch, 1977)

Weighted kappa

For ordinal scales: near-misses (4 vs. 5) penalised less than large discrepancies (1 vs. 5).

Carry forward

The thread into the next section

  • Screening applies to healthy populations; diagnostic testing to suspected cases. Evaluation principles are shared.
  • Analytic and diagnostic sensitivity are different quantities. Keep them separate.
  • Accuracy, precision, and agreement each ask a version of: how much can you trust the result?

Introduction and Overview

An earlier lesson covered measures of disease frequency in populations. This lesson takes the same probabilistic vocabulary and applies it at the level of a single test administered to a single person. Whether you're evaluating a new screening assay, interpreting a clinical result, or designing a surveillance algorithm, the same four-cell 2×2 logic appears: a test result that's either positive or negative, against a true disease state that's either present or absent (Sackett & Haynes, 2002). The four content sections build up from the basic attributes of a test (this section), through sensitivity and specificity (a later section), to the predictive values that depend on disease prevalence (a later section), and finally to ROC curves and likelihood ratios for tests with continuous output (a later section).

Learning Objectives

  • Distinguish between screening tests and diagnostic tests.
  • Define analytic sensitivity and specificity of a test.
  • Explain the difference between accuracy and precision.
  • Describe measures of agreement, including Cohen’s kappa and weighted kappa.

The section begins by defining a test and distinguishing screening tests from diagnostic tests. It then turns to the attributes of a test itself: its analytic sensitivity and specificity, its accuracy and precision, and the agreement between two tests or raters.

What Is a Test?

A test is any device or procedure designed to detect or quantify a sign, substance, tissue change, or body response in an individual. Tests can also be applied at the household or other levels of aggregation. In epidemiology, the term “test” extends broadly to include clinical signs, history-taking questions, survey items, and post-mortem findings.

Box 5.1 sets out why the characteristics of a test matter, both for decisions about individual patients and for the quality of research data.

Box 5.1: Why Evaluate Tests?

In a decision-making context (e.g., clinical diagnosis), the selection of an appropriate test should alter your assessment of the probability that a disease exists, and guide subsequent actions (further testing, treatment, quarantine). In a research context, understanding test characteristics is essential for knowing how they affect data quality.

Screening vs. Diagnostic Tests

Tests are applied for two broad purposes. The two cards below describe screening tests, which are applied to apparently healthy people, and diagnostic tests, which are applied to people already considered abnormal, with examples of each.

Click each card to learn more:

Screening TestsClick to learn more
Diagnostic TestsClick to learn more

Despite their different uses, the principles of evaluation and interpretation are the same for both screening and diagnostic tests.

The rest of the lesson therefore treats the two kinds of test together, beginning with the attributes of the test itself.

Attributes of the Test Per Se

Some properties of a test describe it as a measurement and can be assessed without knowing the true disease status of the people tested. This part covers analytic sensitivity and specificity, accuracy and precision, and the measures used to quantify precision and agreement for quantitative and for categorical results.

Analytic Sensitivity and Specificity

The analytic sensitivity of an assay refers to the lowest concentration of a chemical compound the test can detect. The analytic specificity refers to the capacity of a test to react to only one chemical compound. These are distinctly different from diagnostic (epidemiologic) sensitivity and specificity, which are discussed in a later section.

Accuracy and Precision

The laboratory accuracy of a test relates to its ability to give a true measure of the substance of interest. To be accurate, a test need not always be close to the true value, but if repeat tests are run, the resulting average should be close to the true value.

The precision of a test relates to how consistent the results are. If a test always gives the same value for a sample (regardless of whether it is the correct value), it is said to be precise.

Figure 5.1 illustrates the two properties with four targets, in which the bullseye represents the true value and the dots represent repeated test results.

Accurate & Precise Inaccurate but Precise Accurate but Imprecise Inaccurate & Imprecise

Figure 5.1. Laboratory accuracy and precision. The bullseye represents the true value.

Precision and Agreement

Repeatability refers to variability obtained from repeated testing of the same sample within the same laboratory. Reproducibility refers to variability from testing the same sample in different laboratories. Agreement refers to how well two different tests (or raters) agree when applied to the same sample.

Measuring Precision: Quantitative Tests

When a test reports a number, such as a glucose concentration or a systolic blood pressure, its precision is judged by how closely repeated or paired results agree. Three measures are in common use, and the accordion below describes each of them.

Common measures for quantifying variability between pairs of test results include:

Coefficient of Variation (CV)▼

The CV is computed as CV = σ / μ, where σ is the standard deviation among test results on the same sample and μ is the mean. A lower CV indicates greater precision.

Concordance Correlation Coefficient (CCC)▼

The CCC (Lin, 1989) compares two sets of test results and better reflects agreement than a Pearson correlation. It is computed from three parameters: the location-shift (how far data are from the equality line), the scale-shift (difference in slopes), and the Pearson r. A CCC of 1 indicates perfect agreement.

Limits of Agreement (Bland-Altman Plot)▼

A Bland-Altman plot (Bland & Altman, 1986) plots the differences between paired test results against their mean value. The mean difference (μd) and limits of agreement (μd ± 1.96σd) are shown. This reveals systematic bias and whether disagreement varies with the magnitude of the measurement.

Two of these measures have simple formulae. The tabs below give each formula with a worked example and a calculator.

The first tab gives the coefficient of variation (Equation 5.1). Worked Example 5.1 applies it to five repeated measurements of one serum sample, and Calculator 5.1 reproduces that example so that the effect of a larger spread on the coefficient can be seen. The second tab gives the limits of agreement (Equation 5.2), which Worked Example 5.2 applies to two blood pressure monitors and Calculator 5.2 reproduces.

Coefficient of variation
\[ {\color{#0B7B6B}{CV}} = \dfrac{{\color{#C2410C}{\sigma}}}{{\color{#6D28D9}{\mu}}} \]Eq 5.1
the coefficient of variation equals the standard deviation of repeated results on the same sample divided by the mean of those results.

Worked Example 5.1: Coefficient of Variation

A laboratory runs the same serum sample through a glucose assay five times and obtains 102, 98, 105, 95, and 100 mg/dL. The mean of the five results is μ = 100 mg/dL, and their sample standard deviation (dividing by n − 1 = 4) is σ = 3.81 mg/dL.

CV = 3.81 / 100 = 0.038, or 3.8%. Because the CV divides by the mean, it has no units, which allows the precision of assays measured on different scales to be compared.

Limits of agreement (Bland-Altman)
\[ {\color{#0B7B6B}{\text{limits of agreement}}} = {\color{#C2410C}{\mu_d}} \pm 1.96\,{\color{#6D28D9}{\sigma_d}} \]Eq 5.2
the limits of agreement run from the mean difference between paired results minus 1.96 times the standard deviation of those differences to the mean difference plus the same amount.

Worked Example 5.2: Limits of Agreement

Two automated blood pressure monitors are applied to the same 50 patients. The difference in systolic pressure (monitor A minus monitor B) has a mean of μd = 2.4 mmHg and a standard deviation of σd = 5.0 mmHg.

  • Lower limit = 2.4 − 1.96 × 5.0 = 2.4 − 9.8 = −7.4 mmHg.
  • Upper limit = 2.4 + 1.96 × 5.0 = 2.4 + 9.8 = 12.2 mmHg.

For about 95% of patients, monitor A is expected to read between 7.4 mmHg lower and 12.2 mmHg higher than monitor B. The mean difference of 2.4 mmHg indicates a small systematic bias, and whether a range of this width is acceptable is a clinical judgement.

The coefficient of variation and the limits of agreement both describe tests that report a number. Many tests instead place a result in a category, such as positive or negative, and agreement between categorical results needs a measure of its own, which the next subsection introduces.

Measuring Agreement for Categorical Tests: Kappa (κ)

Cohen’s kappa is the usual measure of agreement between two categorical classifications of the same items. Box 5.2 gives the background: why some agreement is expected by chance alone, the formula for kappa (Equation 5.3), and the conventional labels for its values in Table 5.1. Worked Example 5.3 then applies Equation 5.3 to two radiologists who read the same 100 chest X-rays, and Calculator 5.3 reproduces the example so that the effect of changing the cell counts on kappa can be seen.

Box 5.2: Background: Agreement Beyond Chance

When test results are categorical (dichotomous or ordinal), Cohen’s kappa (κ) measures agreement beyond what would be expected by chance alone (Cohen, 1960). Two raters, or two tests, agree on some results by chance: if each calls about half of the results positive, they would agree on about half of them even by guessing. The chance-expected agreement, pe, is computed from how often each rater uses each category, and kappa compares the observed agreement, po, with it:

Cohen’s kappa
\[ {\color{#0B7B6B}{\kappa}} = \dfrac{{\color{#C2410C}{p_o}} - {\color{#6D28D9}{p_e}}}{1 - {\color{#6D28D9}{p_e}}} \]Eq 5.3
kappa equals the observed agreement minus the agreement expected by chance, divided by one minus the chance agreement.

Kappa ranges from −1 to 1. A value of 0 means agreement no better than chance, 1 means perfect agreement, and negative values mean agreement below chance, which can happen when two tests tend to pull in opposite directions. The labels in Table 5.1 come from Landis and Koch (1977). They are conventions for describing kappa values, and other authors draw the bands differently, so a report gives the value itself as well as its label. HSCI 241 Lesson 7, Section 3 (Screening Studies: Stages, Dual Screening and Agreement), applies kappa to agreement between two people screening studies for a systematic review and is optional reading.

Table 5.1. Conventional labels for values of kappa (Landis and Koch, 1977).

κ ValueInterpretation
≤ 0Poor agreement
0.01 – 0.20Slight agreement
0.21 – 0.40Fair agreement
0.41 – 0.60Moderate agreement
0.61 – 0.80Substantial agreement
0.81 – 1.00Almost perfect agreement

Worked Example 5.3: Cohen’s Kappa

Two radiologists independently read the same 100 chest X-rays and classify each one as positive or negative for pneumonia.

Radiologist B positiveRadiologist B negativeTotal
Radiologist A positive401050
Radiologist A negative54550
Total4555100
  • Observed agreement: po = (40 + 45) / 100 = 0.85.
  • Agreement expected by chance: radiologist A calls 50% of the films positive and radiologist B calls 45% positive, so pe = (0.50 × 0.45) + (0.50 × 0.55) = 0.225 + 0.275 = 0.50.
  • κ = (0.85 − 0.50) / (1 − 0.50) = 0.35 / 0.50 = 0.70.

The radiologists agree on 85% of the films, but agreement on half of the films (pe = 0.50) would be expected by chance alone. Kappa rescales the agreement beyond chance so that it ranges from −1 to 1, with 0 for chance agreement, 1 for perfect agreement and negative values when agreement is below chance; a value of 0.70 falls in the substantial band of Table 5.1.

How Prevalence and Bias Affect Kappa

Kappa depends on more than how well two raters agree. Because pe is computed from how often each rater uses each category, kappa also depends on the prevalence of the condition among the items rated and on any difference between the raters in how often they call a result positive (Feinstein & Cicchetti, 1990; Byrt, Bishop & Carlin, 1993).

Box 5.3 describes each of these two influences, and the example that follows it shows the prevalence effect with numbers.

Box 5.3: Factors Affecting Kappa

Prevalence: When the condition is very common or very rare, most of the agreement falls in one cell, pe is close to 1, and a little disagreement produces a low κ. Two tests will therefore have a higher κ when prevalence is moderate (~0.5) than when it is very high or very low, even when they make the same errors. This is the paradox of high observed agreement with a low kappa.

Bias: If one test consistently produces more positive results than the other, κ will be affected. Use McNemar’s χ² test to check whether the two tests classify the same proportion as positive before evaluating agreement.

A small example shows the prevalence effect. Two readers classify 100 films; reader A calls 3 positive and reader B calls 2 positive, and they never call the same film positive. They agree on the 95 films that both call negative, so po = 0.95. Chance agreement is pe = (0.03 × 0.02) + (0.97 × 0.98) = 0.951, so κ = (0.95 − 0.951) ÷ (1 − 0.951) = −0.02. Agreement of 95% is no better than chance here, because almost all of it is agreement on negatives that chance alone would produce. Byrt and colleagues (1993) therefore recommended reporting a prevalence index and a bias index alongside kappa, together with the prevalence-adjusted bias-adjusted kappa, PABAK = 2po − 1, which shows what kappa would be if both categories were equally common and the raters used them equally often. For the two readers, PABAK = 2 × 0.95 − 1 = 0.90, and the gap between 0.90 and −0.02 shows how much of the low kappa comes from the rarity of positives.

Weighted Kappa

For tests measured on an ordinal scale, a weighted kappa accounts for partial agreement. Pairs of test results that are close (e.g., scores of 4 and 5) receive more credit than pairs that are far apart (e.g., scores of 1 and 5). Each cell of the k × k agreement table, where k is the number of categories, receives a disagreement weight that grows with the distance between the two categories assigned, i and j: linear weights are |i − j| ÷ (k − 1), and quadratic weights are (i − j)² ÷ (k − 1)². Weighted kappa is 1 minus the ratio of the weighted observed disagreement to the weighted disagreement expected by chance, κw = 1 − Σwijoij ÷ Σwijeij, where oij and eij are the observed and chance-expected proportions in each cell. With quadratic weights it is closely related to the intraclass correlation coefficient (Fleiss & Cohen, 1973). The choice of weights changes the value, so a report states which weights were used. Weighted kappa gives a better reflection of agreement for ordinal data than the unweighted version, which treats every disagreement as equally serious.

This section has distinguished screening tests from diagnostic tests and described the attributes of a test itself: analytic sensitivity and specificity, accuracy and precision, and the coefficient of variation, the limits of agreement and kappa as measures of precision and agreement. None of these attributes shows how often a test classifies a person’s disease status correctly, which requires comparison with a gold standard and is the subject of the next section. The key takeaways and the knowledge check below review the material of this section.

Key Takeaways

  • A test is any procedure designed to detect or quantify a sign, substance, or response.
  • Screening tests are applied to healthy populations; diagnostic tests are applied to individuals suspected of disease.
  • Accuracy measures closeness to the true value; precision measures consistency of results.
  • Cohen’s kappa quantifies agreement beyond chance for categorical tests; weighted kappa extends this to ordinal scales.
  • Prevalence and bias both affect kappa values.
Knowledge Check: this section

1. A test that always gives the same result for a sample, but the result is consistently wrong, is best described as:

Precision relates to consistency of results. If the test always gives the same value, it is precise. However, if that value is wrong, it is inaccurate. This corresponds to the “inaccurate but precise” target pattern.

2. Cohen’s kappa measures:

Kappa measures the extent of agreement between two sets of categorical test results (or raters) beyond what would be expected by chance alone.

3. Which statement about screening and diagnostic tests is correct?

Screening tests are applied to healthy populations to detect disease early, while diagnostic tests are used to confirm disease in individuals already suspected of being ill. Despite different uses, the principles of evaluation are the same.

✦ Pass the knowledge check with 100% to continue

Section 2

Sensitivity & Specificity

⏱ Estimated reading time: 15 minutes

Section 2 of 4

Sensitivity & Specificity

The 2×2 table, gold standards, and the test properties that travel wherever the test goes.

Reference standard

The gold standard

A gold standard is the reference procedure assumed to be perfectly accurate: all cases classified correctly, no misclassification.

In practice few true gold standards exist. Biological variability and measurement limits mean test evaluation often relies on composite or imperfect references.

The core tool

The 2×2 contingency table

Disease + (D+) Disease − (D−)
Test + (T+) a true positive c false positive
Test − (T−) b false negative d true negative
The key measures

Sensitivity and specificity

Sensitivity
\[ \color{#0B7B6B}{Se} = \frac{\color{#C2410C}{a}}{\color{#C2410C}{a} + \color{#6D28D9}{b}} = \frac{\text{TP}}{\text{all D+}} \]
Se sensitivity (true positive rate)a true positivesb false negatives
Specificity
\[ \color{#0B7B6B}{Sp} = \frac{\color{#C2410C}{d}}{\color{#6D28D9}{c} + \color{#C2410C}{d}} = \frac{\text{TN}}{\text{all D}-} \]
Sp specificity (true negative rate)d true negativesc false positives

SnNOut

High Se: a Negative result rules Out disease.

SpPIn

High Sp: a Positive result rules In disease.

Prevalence effects

True vs. apparent prevalence

Apparent prevalence (Eq 5.6)
\[ \color{#0B7B6B}{AP} = \color{#C2410C}{P} \cdot \color{#6D28D9}{Se} + (1 - \color{#C2410C}{P})(1 - \color{#1D4ED8}{Sp}) \]
AP apparent (test-measured) prevalenceP true prevalenceSe sensitivitySp specificity
Rogan-Gladen correction (1978)
\[ \color{#0B7B6B}{\hat{P}} = \frac{\color{#C2410C}{AP} + \color{#1D4ED8}{Sp} - 1}{\color{#6D28D9}{Se} + \color{#1D4ED8}{Sp} - 1} \]
P̂ corrected true prevalenceAP apparent prevalenceSe sensitivitySp specificity
Carry forward

What sensitivity and specificity cannot tell you

  • Se and Sp are test properties: prevalence does not enter their calculation.
  • They answer the question from the test's viewpoint, not the patient's.
  • To answer “does this positive result mean disease?” you also need disease prevalence. A later section does exactly that.

Introduction and Overview

An earlier section named the attributes a test should have in the abstract. This section turns to the two quantitative properties that capture most of what we care about: sensitivity (the test's ability to find disease that is truly present) and specificity (its ability to correctly say “no” when disease is truly absent). Both are properties of the test itself, not of the population to which it is applied; that distinction becomes essential at the predictive values in a later section. Se and Sp can still differ between settings when the mix of mild and severe disease differs (the spectrum effect), so values estimated in one setting, such as a hospital study, may not hold in a screening population.

Figure 5.2 shows the 2×2 test table on which the section is built and the direction in which each measure is read from it.

The 2x2 test table: sensitivity is read down the diseased column, specificity down the non-diseased columnDisease +Disease −Test +Test −atrue positivecfalse positivebfalse negativedtrue negativeSe = a / (a + b)read down the Disease + columnSp = d / (c + d)read down the Disease − columnPredictive values, by contrast, read across the test-result rows.
Figure 5.2. The 2×2 test table. Sensitivity and specificity are computed down the disease-status columns (so they are properties of the test); predictive values are computed across the test-result rows (so they depend on prevalence).

Learning Objectives

  • Explain the concept of a gold standard and its role in test evaluation.
  • Calculate sensitivity, specificity, false positive fraction, and false negative fraction from a 2×2 table.
  • Distinguish between true prevalence and apparent prevalence.
  • Estimate true prevalence from apparent prevalence using the Rogan-Gladen formula.
  • Apply sensitivity, specificity, predictive value and the Rogan-Gladen correction to the validation of an administrative case definition.

The section first defines the gold standard against which a test is judged, then sets out the 2×2 table and the measures computed from it, and finally shows how sensitivity and specificity link the true prevalence of a condition to the prevalence that a test reports.

The Gold Standard

Sensitivity and specificity are estimated by comparing the results of a test with a reference that is taken to give the true disease status. Box 5.4 recalls how this kind of comparison was classified when questionnaires were validated in Lesson 3, and the paragraphs after it define the gold standard and its limits.

Box 5.4: Recall: Criterion Validity

HSCI 341 Lesson 3, Section 3 (Wording, Structure and Pre-Testing), listed four approaches to validating a questionnaire: comparison with a gold standard, comparison with established methods, repeated administration and method comparison. Comparison with a gold standard gives criterion validity evidence, and sensitivity and specificity, the subject of this section, are that evidence for a categorical measure. Repeated administration to the same respondents is a test-retest design, so it gives reliability evidence: it shows that answers are repeatable, and a repeatable instrument can still measure the wrong thing.

Retrieval question. Classify each of the four Lesson 3 approaches as reliability evidence or validity evidence.

Show the answer▼

Comparison with a gold standard gives criterion validity evidence. Comparison with a well-validated established instrument gives validity evidence (concurrent or convergent), which is only as strong as the established instrument. Repeated administration gives test-retest reliability evidence. Method comparison, which checks whether mail and telephone administration give comparable results, shows agreement between modes, which is reliability (equivalence) evidence unless one mode is treated as the reference.

A gold standard (GS) is a test or procedure that is absolutely accurate: it diagnoses all cases of a specific disease and misdiagnoses none. In reality, very few true gold standards exist. Much of the error in test evaluation is due to biological variability: people do not immediately become “diseased” upon exposure, and the timescale for crossing a detectable threshold varies from person to person.

Box 5.5 notes the approaches available when no true gold standard exists.

Box 5.5: Important Caveat

When no true gold standard exists, alternative approaches for estimating sensitivity and specificity are needed, including the use of results from several different tests, repeated testing of selected samples, and latent class models (discussed in Section 5.7 of the textbook). Latent class models estimate sensitivity and specificity by treating true disease status as an unobserved category, inferred from the pattern of results when several imperfect tests are applied to the same people.

With a gold standard, or the best available reference, in place, each person tested can be classified twice: once by the test and once by the reference. The next part arranges these two classifications in a 2×2 table.

The 2×2 Contingency Table

The concepts of sensitivity and specificity are most easily understood through a 2×2 contingency table comparing disease status to test results:

Table 5.2. The 2×2 table of disease status by test result, with the notation for its cells and totals.

Test Positive (T+)Test Negative (T−)Total
Disease Positive (D+)a (true positive)b (false negative)m1
Disease Negative (D−)c (false positive)d (true negative)m0
Totaln1n0n

In Table 5.2 the rows give the true disease status and the columns give the test result. Cells a and d hold the correct results and cells b and c the errors; m1 and m0 are the numbers with and without the disease, and n1 and n0 are the numbers who test positive and negative.

Key Measures from the 2×2 Table

Four measures are computed within the disease-status groups of Table 5.2. The cards below define sensitivity, specificity, the false positive fraction and the false negative fraction, and give the mnemonics SnNOut and SpPIn for the use of sensitive and specific tests.

Click each card to explore:

Sensitivity (Se)Click to explore
Specificity (Sp)Click to explore
False Positive FractionClick to explore
False Negative FractionClick to explore

Equation 5.4 and Equation 5.5 give sensitivity and specificity in terms of the cells of Table 5.2. Worked Example 5.4 applies both to a study of 188 stool samples tested for norovirus with an EIA, and Calculator 5.4 reproduces the example so that the effect of changing the cell counts on both measures can be seen.

Sensitivity
\[ {\color{#0B7B6B}{Se}} = \dfrac{{\color{#C2410C}{a}}}{{\color{#C2410C}{a}} + {\color{#6D28D9}{b}}} \]Eq 5.4
sensitivity equals the true positives divided by everyone who has the disease (the true positives plus the false negatives); the false negative fraction is 1 − Se.
Specificity
\[ {\color{#0B7B6B}{Sp}} = \dfrac{{\color{#BE185D}{d}}}{{\color{#1D4ED8}{c}} + {\color{#BE185D}{d}}} \]Eq 5.5
specificity equals the true negatives divided by everyone who does not have the disease (the false positives plus the true negatives); the false positive fraction is 1 − Sp.

Worked Example 5.4: Norovirus EIA Data

From a study of 188 stool samples tested with an EIA against a gold standard:

GS+ (D+)GS− (D−)Total
T+71374
T−11103114
Total82106188
  • Se = 71/82 = 86.6% (95% CI: 77.3%, 93.1%)
  • Sp = 103/106 = 97.2% (95% CI: 92.0%, 99.4%)
  • FNF = 1 − 0.866 = 13.4%
  • FPF = 1 − 0.972 = 2.8%

Sensitivity and specificity describe how a test performs within the diseased and the non-diseased groups. When a test is applied to a whole population, these two quantities and the prevalence together determine the proportion that tests positive, which is the subject of the next part.

True and Apparent Prevalence

A survey that uses an imperfect test counts the people who test positive, which is generally a different number from the people who have the condition. This part distinguishes true prevalence from apparent prevalence and shows how each can be calculated from the other. Box 5.6 first recalls the definition of prevalence and its relation to incidence and duration from Lesson 4.

Box 5.6: Recall: Prevalence

Prevalence is the proportion of a population that has a condition at a point in time (point prevalence) or during a period (period prevalence). HSCI 341 Lesson 4, Section 3 (Prevalence, Mortality and Other Frequency Measures), showed that it reflects both how often new cases arise and how long they last: in a steady state, P ÷ (1 − P) = I × D, so P ≈ I × D when the condition is uncommon. HSCI 230 Lesson 3, Section 4, called this the prevalence-duration confound, because factors that prolong a condition can look like causes of it in a cross-sectional study.

Retrieval question. A condition has an incidence rate of 2 per 1,000 person-years and an average duration of 10 years. What is its approximate prevalence in a steady state?

Show the answer▼

I × D = 0.002 × 10 = 0.02, so the prevalence odds are 0.02 and the prevalence is 0.02 ÷ 1.02 = 0.0196, about 2%.

The true prevalence (P) is the actual proportion of the population that has the disease. In Worked Example 5.4, P = 82/188 = 43.6%.

The apparent prevalence (AP) is the proportion that tests positive, which includes both true positives and false positives. In Worked Example 5.4, AP = 74/188 = 39.4%.

Equation 5.6 expresses the apparent prevalence in terms of the true prevalence and the sensitivity and specificity of the test. Worked Example 5.5 applies it to the norovirus study and recovers the apparent prevalence of 39.4% counted in Worked Example 5.4, and Calculator 5.5 reproduces the example.

Apparent prevalence
\[ {\color{#0B7B6B}{AP}} = {\color{#C2410C}{P}}\,{\color{#6D28D9}{Se}} + (1 - {\color{#C2410C}{P}})(1 - {\color{#1D4ED8}{Sp}}) \]Eq 5.6
apparent prevalence equals true prevalence times sensitivity, plus the non-diseased fraction times one minus specificity (the false positives).

Worked Example 5.5: Apparent Prevalence from the Formula

In the norovirus study, the true prevalence is P = 82/188 = 0.4362, and the test has Se = 0.8659 and Sp = 0.9717.

AP = (0.4362 × 0.8659) + (1 − 0.4362) × (1 − 0.9717) = 0.3777 + (0.5638 × 0.0283) = 0.3777 + 0.0160 = 0.3937, or 39.4%.

The formula reproduces the 74/188 = 39.4% counted directly from the table. The first term is the share of the population that is diseased and correctly detected (the true positives), and the second is the share that is healthy but tests positive (the false positives).

Estimating True Prevalence from Apparent Prevalence

If the Se and Sp of a test are known, the true prevalence can be estimated from the apparent prevalence using the Rogan-Gladen formula (Rogan & Gladen, 1978):

True prevalence from apparent prevalence
\[ {\color{#0B7B6B}{P}} = \dfrac{{\color{#C2410C}{AP}} + {\color{#1D4ED8}{Sp}} - 1}{{\color{#6D28D9}{Se}} + {\color{#1D4ED8}{Sp}} - 1} \]Eq 5.7
true prevalence equals apparent prevalence plus specificity minus one, divided by sensitivity plus specificity minus one.

Equation 5.7 solves Equation 5.6 for the true prevalence. Worked Example 5.6 applies it to an apparent prevalence of 0.150 measured with a test whose sensitivity is 0.363 and specificity 0.876, and Calculator 5.6 reproduces the example so that the effect of different values of Se and Sp on the corrected estimate can be seen.

Worked Example 5.6: True Prevalence from the Rogan-Gladen Formula

If AP = 0.150, Se = 0.363, and Sp = 0.876, then:

P = (0.150 + 0.876 − 1) / (0.363 + 0.876 − 1) = 0.026 / 0.239 = 0.109 (10.9%)

Note: Some combinations of Se, Sp, and AP can produce estimates of P outside the range 0–1, indicating that the Se and Sp estimates may not be applicable to the population being studied.

Validating an Administrative Case Definition

The measures of this section apply to any rule that classifies people as cases or non-cases, including a case definition applied to health administrative records. Box 5.7 recalls the Canadian Chronic Disease Surveillance System from Lesson 2, and Worked Example 5.7 validates a diabetes case definition against chart review and then applies the Rogan-Gladen formula (Equation 5.7) to a provincial prevalence estimate.

Box 5.7: Recall: The Canadian Chronic Disease Surveillance System

HSCI 341 Lesson 2, Section 1 (Surveillance Systems and Canadian Data Sources), introduced the Canadian Chronic Disease Surveillance System (CCDSS), which applies validated case definitions to health administrative records to estimate the prevalence and incidence of chronic conditions such as diabetes, hypertension and dementia. A case definition of this kind works as a test: each person is classified as a case or a non-case from their records, and the classification can be checked against a reference standard such as chart review.

Retrieval question. Which two kinds of administrative record does the CCDSS draw on?

Show the answer▼

Physician billing claims and hospital discharge abstracts.

Worked Example 5.7: Validating an Administrative Case Definition

A diabetes case definition used with Canadian administrative data counts a person as a case if they have one hospital discharge abstract, or two physician claims within two years, with a diabetes diagnosis. To validate it, a study abstracts the charts of 1,000 adults and compares the case definition with diabetes as documented in the chart, the reference standard. The counts are hypothetical, chosen for teaching.

Chart: diabetesChart: no diabetesTotal
Case definition met8618104
Case definition not met14882896
Total1009001,000
  • Se = 86/100 = 0.86
  • Sp = 882/900 = 0.98
  • PV+ = 86/104 = 0.83

The definition misses 14% of people with diabetes and wrongly includes 2% of people without it, and 17% of the people it counts as cases do not have diabetes according to the chart. The predictive value depends on the prevalence of diabetes in the validation sample (10% here), so it would be lower where diabetes is less common, as the next section shows.

These values change a prevalence estimate. Suppose the case definition, applied to a province's administrative data, gives an apparent prevalence of 8.0%. The Rogan-Gladen formula gives P = (0.080 + 0.98 − 1) ÷ (0.86 + 0.98 − 1) = 0.060 ÷ 0.84 = 0.071, or 7.1%. The administrative estimate is too high by about 0.9 percentage points, because the false positives among the 92.9% of people without diabetes (0.02 × 0.929, or 1.9 percentage points) outnumber the cases the definition misses (0.14 × 0.071, or 1.0 percentage point). A lower specificity would widen the gap, which is why validation studies of administrative case definitions report specificity and predictive value as well as sensitivity.

This section has defined the gold standard, computed sensitivity and specificity from the 2×2 table, and used them to move between true and apparent prevalence. Worked Example 5.7 also computed a predictive value, which depends on the prevalence in the population tested; the next section develops predictive values in full. The reflection, key takeaways and knowledge check below review the material of this section.

Reflection

A new rapid test for influenza has a sensitivity of 75% and a specificity of 98%. In a population where the true prevalence of influenza is 5%, calculate the apparent prevalence using the formula AP = P × Se + (1 − P) × (1 − Sp). What does this tell you about relying solely on test results to estimate disease burden?

Model answerAP = 0.05×0.75 + 0.95×(1−0.98) = 0.0375 + 0.019 = 0.057 (5.7%). The apparent prevalence (5.7%) is close to but biased upward from the true prevalence (5%): the false-positive rate of 2% applied to the 95% non-diseased population is more numerous than the 25% false negatives among the 5% diseased. The implication: using raw test results without correction systematically misestimates disease burden, with the direction of bias depending on the relative magnitudes of (1−Sp) and Se. For surveillance reporting you must correct for known test performance: P = (AP − (1 − Sp)) / (Se + Sp − 1). Routine surveillance dashboards that report ‘positivity rate’ as if it were prevalence are mathematically misleading whenever Se and Sp are imperfect.

Minimum 20 characters required.

✓ Reflection saved

Key Takeaways

  • A gold standard is the reference test assumed to be perfectly accurate; in practice, few truly exist.
  • Sensitivity = probability of testing positive given disease; specificity = probability of testing negative given no disease.
  • High Se is important for ruling out disease (SnNOut); high Sp is important for confirming disease (SpPIn).
  • Apparent prevalence differs from true prevalence due to test imperfections.
  • The Rogan-Gladen formula estimates true prevalence from apparent prevalence when Se and Sp are known.
Knowledge Check: this section

1. In a 2×2 table, the false negative fraction (FNF) is calculated as:

The false negative fraction is the proportion of truly diseased individuals that test negative. Since Se = a/(a+b), the FNF = b/(a+b) = 1 − Se.

2. If a test has Se = 90% and Sp = 95%, and the true prevalence is 10%, what is the apparent prevalence?

AP = P × Se + (1 − P) × (1 − Sp) = 0.10 × 0.90 + 0.90 × 0.05 = 0.09 + 0.045 = 0.135 or 13.5%.

3. A highly specific test is most useful for:

A highly specific test has few false positives, so a positive result strongly suggests the individual truly has the disease (SpPIn: Specificity, Positive result, Rules In).

✦ Pass the knowledge check with 100% and complete the reflection to continue

Section 3

Predictive Values

⏱ Estimated reading time: 12 minutes

Section 3 of 4

Predictive Values

From the test’s perspective to the patient’s: what a result actually means in a given population.

The patient-centred measures

Positive and negative predictive value

PV+ (Eq 5.8)
\[ \color{#0B7B6B}{PV^+} = \frac{\color{#C2410C}{P} \cdot \color{#6D28D9}{Se}}{\color{#C2410C}{P} \cdot \color{#6D28D9}{Se} + (1-\color{#C2410C}{P})(1-\color{#1D4ED8}{Sp})} \]
PV+ positive predictive valueP prevalence (pre-test probability)Se sensitivitySp specificity
PV− (Eq 5.9)
\[ \color{#0B7B6B}{PV^-} = \frac{(1-\color{#C2410C}{P}) \cdot \color{#1D4ED8}{Sp}}{(1-\color{#C2410C}{P}) \cdot \color{#1D4ED8}{Sp} + \color{#C2410C}{P}(1-\color{#6D28D9}{Se})} \]
PV− negative predictive valueP prevalence (pre-test probability)Se sensitivitySp specificity
The prevalence effect

The same test, very different answers

Prevalence PV+ PV−
50%96.9%87.9%
5%61.9%99.3%
0.1%3.0%~100%

Se = 86.6%, Sp = 97.2% held constant. Only prevalence changes.

A public health scenario

HIV universal screening: the arithmetic

PV+ calculation at P = 0.3%
\[ \color{#0B7B6B}{PV^+} = \frac{\color{#C2410C}{0.003} \times \color{#6D28D9}{0.995}}{\color{#C2410C}{0.003} \times \color{#6D28D9}{0.995} + \color{#1D4ED8}{0.997} \times \color{#BE185D}{0.002}} = 60\% \]
0.003 prevalence P0.995 sensitivity Se0.997 1 − P (non-diseased)0.002 1 − Sp (false positive rate)

Se = 99.5%, Sp = 99.8%: an excellent test. Yet 40% of positives are false alarms at low prevalence. Confirmatory testing is arithmetically necessary.

Practical strategies

Increasing PV+ in low-prevalence settings

Target high-risk groups

Higher local prevalence raises PV+ without changing the test itself.

More specific confirmation

Apply a high-Sp test to initial positives. Overall false-positive rate falls sharply.

Series vs. parallel

Series: higher overall Sp, lower Se. Parallel: higher Se, lower Sp. Match strategy to stakes.

Carry forward

What to take into the next section

  • PV+ and PV− are not portable across populations with different prevalence.
  • Prevalence does not enter sensitivity and specificity, so they transfer more readily.
  • A later section asks what happens when we stop forcing continuous results into a binary yes/no.

Introduction and Overview

An earlier section covered sensitivity and specificity, which are properties of the test itself. This section introduces the predictive values: what an individual person should believe about their disease status given the test result. Importantly, predictive values depend on disease prevalence in the population being tested, which is why the same test can be useful in one setting and useless in another. This is the most clinically important section in the lesson.

Learning Objectives

  • Define predictive value positive (PV+) and predictive value negative (PV−).
  • Calculate PV+ and PV− from a 2×2 table and using Bayesian formulas.
  • Explain how prevalence affects predictive values.
  • Describe strategies for increasing the predictive value of a positive test.

The section defines the two predictive values and computes them from the 2×2 table and from Bayesian formulas, shows how strongly they depend on prevalence, and ends with strategies for raising the predictive value of a positive test.

What Are Predictive Values?

While Se and Sp are characteristics of the test, predictive values tell us how useful the test is for individuals of unknown disease status. Once we decide to use a test, we want to know the probability that the individual has or does not have the disease, given the test result.

Figure 5.3 is an interactive story that follows 1,000 people through a test with 95% sensitivity and 95% specificity at a prevalence of 1%, and shows how the positive predictive value emerges from the counts of true and false positives.

▸ INTERACTIVE STORY: 1000 PIXEL PEOPLE Open full screen ↗

Watch a 95-95 test scan 1,000 people and see PPV emerge from the math. Next ▶ advances scenes.

Figure 5.3. A 6-scene Bayesian-reasoning visualization: a population of 1,000 with 1% prevalence, a 95%-sensitive 95%-specific test scanning across, the four buckets (TP/FP/FN/TN) populating in real time, and the surprising PPV that follows.

The two tabs below define the predictive values. The first tab gives the positive predictive value (Equation 5.8), which Worked Example 5.8 applies to the norovirus study and Calculator 5.7 reproduces. The second tab gives the negative predictive value (Equation 5.9), which Worked Example 5.9 applies to the same study and Calculator 5.8 reproduces.

Predictive Value Positive (PV+)

The PV+ is the probability that an individual who tests positive actually has the disease: p(D+|T+) = a / n1.

Positive predictive value
\[ {\color{#0B7B6B}{PV^+}} = \dfrac{{\color{#C2410C}{p(D^+)}}\,{\color{#6D28D9}{Se}}}{{\color{#C2410C}{p(D^+)}}\,{\color{#6D28D9}{Se}} + {\color{#1D4ED8}{p(D^-)}}(1 - {\color{#BE185D}{Sp}})} \]Eq 5.8
the positive predictive value is the true positives, prevalence times sensitivity, divided by all positives (true positives plus false positives, the non-diseased fraction times one minus specificity).

In the norovirus example: PV+ = 71/74 = 95.9% (95% CI: 88.6%, 99.2%)

Worked Example 5.8: PV+ from the Formula

The same answer follows from the formula, using the norovirus study’s prevalence p(D+) = 0.436, Se = 0.866, and Sp = 0.972, so that p(D−) = 1 − 0.436 = 0.564.

PV+ = (0.436 × 0.866) / [(0.436 × 0.866) + 0.564 × (1 − 0.972)] = 0.3776 / (0.3776 + 0.0158) = 0.3776 / 0.3934 = 0.960, or 96.0%.

The numerator is the share of all people tested who are true positives, and the denominator is the share who test positive for any reason. The small difference from 95.9% reflects rounding of the inputs.

Predictive Value Negative (PV−)

The PV− is the probability that an individual who tests negative truly does not have the disease: p(D−|T−) = d / n0.

Negative predictive value
\[ {\color{#0B7B6B}{PV^-}} = \dfrac{{\color{#C2410C}{p(D^-)}}\,{\color{#6D28D9}{Sp}}}{{\color{#C2410C}{p(D^-)}}\,{\color{#6D28D9}{Sp}} + {\color{#1D4ED8}{p(D^+)}}(1 - {\color{#BE185D}{Se}})} \]Eq 5.9
the negative predictive value is the true negatives, the non-diseased fraction times specificity, divided by all negatives (true negatives plus false negatives, the diseased fraction times one minus sensitivity).

In the norovirus example: PV− = 103/114 = 90.4% (95% CI: 83.4%, 95.1%)

Worked Example 5.9: PV− from the Formula

With p(D−) = 0.564, Sp = 0.972, p(D+) = 0.436, and Se = 0.866:

PV− = (0.564 × 0.972) / [(0.564 × 0.972) + 0.436 × (1 − 0.866)] = 0.5482 / (0.5482 + 0.0584) = 0.5482 / 0.6066 = 0.904, or 90.4%.

The numerator is the share of all people tested who are true negatives, and the denominator is the share who test negative for any reason, including the diseased people the test misses.

In the norovirus study both predictive values were high, at 95.9% for a positive result and 90.4% for a negative result. Both were computed at the prevalence in that study, 43.6%, and the next part shows how they change when the same test is used where the disease is less common.

Effect of Prevalence on Predictive Values

Predictive values depend heavily on the prevalence of disease in the population being tested (an application of Bayes's theorem). This is why PV+ and PV− are not good measures of a test’s intrinsic performance; they vary from population to population.

Box 5.8 holds the sensitivity and specificity of the norovirus EIA fixed and recomputes both predictive values at three levels of prevalence, which Table 5.3 sets out.

Box 5.8: Dramatic Impact of Prevalence

Using Se = 86.6% and Sp = 97.2% from the norovirus example, observe how PV+ and PV− change as prevalence drops:

Table 5.3. Predictive values of the norovirus EIA (Se = 86.6%, Sp = 97.2%) at three levels of prevalence.

Prevalence (%)PV+ (%)PV− (%)
5096.987.9
561.999.3
0.13.0100.0

As you can see, when prevalence drops to 0.1%, the PV+ falls to just 3%, meaning 97% of positive results are false positives. Meanwhile, the PV− approaches 100%. This is a fundamental challenge in screening low-prevalence populations.

Worked Example 5.10 reaches the same conclusion by counting people. It follows 10,000 people through a test with 99% sensitivity and 95% specificity at a prevalence of 1%.

Worked Example 5.10: Predictive Value in Natural Frequencies

The formula can feel abstract, so it helps to walk a whole group of people through the test and simply count. Imagine screening 10,000 people for a disease with a prevalence of 1%, using a strong test with Se = 99% and Sp = 95%.

GroupPeople
Have the disease (1% of 10,000)100
Diseased and test positive (true positives, 99% of 100)99
Diseased and test negative (false negatives)1
Do not have the disease (99% of 10,000)9,900
Healthy and test positive (false positives, 5% of 9,900)495
Healthy and test negative (true negatives)9,405

Now read the positive predictive value straight off the counts. A total of 99 + 495 = 594 people test positive, but only 99 of them truly have the disease, so PV+ = 99 / 594 = 16.7%. About five of every six positive results are false alarms, even though the test is correct 99% of the time in the sick and 95% of the time in the healthy. The reason is arithmetic, not a flaw in the test: the 9,900 healthy people are so numerous that their small 5% error rate yields more false positives (495) than there are true cases in the whole group (100).

Interactive 5.1 brings the cutoff, the prevalence and the overlap between the test scores of healthy and diseased people together in one simulator. It shows that moving the cutoff trades sensitivity against specificity, while lowering the prevalence leaves both unchanged and lowers the positive predictive value.

🧪 Interactive 5.1: Sensitivity, Specificity, PPV & the Cutoff

This simulator shows the test scores of diseased and healthy people as two overlapping curves, with a cutoff that labels every score to its right as test positive, and fills in the resulting 2×2 table for 10,000 people tested. It is meant to show you that sensitivity and specificity trade off against each other as the cutoff moves, while the predictive values also depend on how common the disease is. As you drag the cutoff line to the right (or move the Cutoff slider), notice that specificity rises and sensitivity falls; as you lower the disease prevalence, sensitivity and specificity stay fixed while PPV (PV+) falls and NPV (PV−) rises. Moving the two mean scores further apart reduces the overlap, so that one cutoff can give high sensitivity and high specificity together.
Distribution of test scores

Drag the dashed cutoff line. Right of the line = test positive.

cutoffTest score02468101214HealthyDiseased
2×2 confusion matrix (per 10,000 tested)
D+D−Total
T+195420203974
T−4659806026
Total2000800010,000
Sensitivity
97.7%
Specificity
74.8%
PPV
49.2%
NPV
99.2%
Presets:
Move the cutoff: see Sn and Sp trade off. Then move prevalence: see PPV/NPV swing while Sn/Sp stay fixed.

Strategies to Increase PV+

Because a low prevalence produces many false positives, the positive predictive value can be raised by changing the population tested or the way tests are used. The three cards below describe targeting high-risk groups, increasing specificity and using more than one test.

Click each card to explore:

Target High-Risk GroupsClick to explore
Increase SpecificityClick to explore
Use Multiple TestsClick to explore

Worked Example 5.11 applies Equation 5.8 to a proposal for universal HIV screening and shows why positive results in a low-prevalence population need confirmation.

Worked Example 5.11: Scenario: Universal HIV Screening

A country considers implementing universal HIV screening using a rapid test with Se = 99.5% and Sp = 99.8%. The national HIV prevalence is 0.3%.

PV+ = (0.003 × 0.995) / [(0.003 × 0.995) + (0.997 × 0.002)] = 0.002985 / (0.002985 + 0.001994) = 60.0%

Even with an excellent test (99.5% Se, 99.8% Sp), 40% of positive results in this low-prevalence population would be false positives. This is why confirmatory testing is essential.

This section has shown that the predictive values of a test depend on the prevalence of the condition as well as on its sensitivity and specificity, so that in a low-prevalence population even a highly specific test can produce mostly false positive results. The reflection, key takeaways and knowledge check below review this material, and the next section turns to tests whose results lie on a continuous scale.

Reflection

Consider a screening programme for a rare genetic condition affecting 1 in 10,000 newborns. The test has Se = 99% and Sp = 99.9%. Calculate the positive predictive value using PV+ = (P × Se) / [P × Se + (1 − P) × (1 − Sp)], where P is the prevalence, and discuss the implications of the result for clinical decision-making. What strategies would you recommend to improve the programme?

Model answerAt P = 1/10,000 = 0.0001 with Se = 0.99 and Sp = 0.999: PPV = (0.0001×0.99) / (0.0001×0.99 + 0.9999×0.001) = 0.000099 / 0.001099 ≈ 0.09 (9%). Even with an extraordinarily specific test, 91% of positive screens are false alarms. Implications: every positive screen must be followed by a confirmatory test (different assay or repeat with different conditions), genetic counselling, and family-history workup; never act on the first positive alone. Programme-improvement strategies: (a) tighten screening criteria, restricting to higher-prevalence subgroups (family history) when feasible; (b) add a second-stage confirmatory test (e.g., DNA sequencing after the rapid immunoassay) before any treatment decision; (c) combine multiple markers in a panel to multiply specificity; (d) improve specificity at the cost of sensitivity if the disease is treatable late as well as early.

Minimum 20 characters required.

✓ Reflection saved

Key Takeaways

  • PV+ is the probability of disease given a positive test; PV− is the probability of no disease given a negative test.
  • Predictive values are driven by both test characteristics (Se, Sp) and the prevalence of disease.
  • In low-prevalence populations, even highly specific tests can produce mostly false positive results.
  • Strategies to increase PV+ include targeting high-risk groups, increasing Sp, and using multiple tests in series.
Knowledge Check: this section

1. As the prevalence of a disease decreases, what happens to PV+ (assuming Se and Sp stay constant)?

When prevalence decreases, there are proportionally more non-diseased individuals who can produce false positives, driving PV+ down. PV− tends to increase as prevalence drops.

2. PV+ is best described as:

PV+ = p(D+|T+), the probability that an individual who tests positive actually has the disease. This is distinct from sensitivity, which is p(T+|D+).

3. Which strategy would NOT help increase PV+?

Lowering the cutpoint increases sensitivity but decreases specificity, leading to more false positives and a lower PV+. The other strategies all help increase PV+.

✦ Pass the knowledge check with 100% and complete the reflection to continue

Section 4

Cutpoints, ROC Curves & Likelihood Ratios

⏱ Estimated reading time: 15 minutes

Section 4 of 4

Cutpoints, ROC Curves & Likelihood Ratios

What to do when the test gives a number, and how to use it to update the probability of disease.

The overlap problem

Every cutpoint is a trade-off

Healthy Diseased Cutpoint False Negatives False Positives

Shift the cutpoint right: more false negatives, fewer false positives. Shift left: the reverse.

ROC curves

Plotting performance across all cutpoints

1 − Specificity (FPF) Sensitivity chance ROC curve AUC = area under curve
Area under the curve

Summarising discriminatory ability

AUC interpretation (Hanley & McNeil, 1982)
\[ \color{#0B7B6B}{AUC} = P(\color{#C2410C}{\text{score}_{D+}} > \color{#6D28D9}{\text{score}_{D-}}) \]
AUC area under the ROC curvescore (D+) test value of a diseased personscore (D−) test value of a healthy person

0.50: Chance  |  0.50–0.70: Poor  |  0.70–0.80: Acceptable

0.80–0.90: Excellent  |  >0.90: Outstanding

Equivalent to the Mann–Whitney U statistic. Non-parametric confidence intervals available.

Likelihood ratios

Combining Se and Sp into one update factor

LR+ (Eq 5.10)
\[ \color{#0B7B6B}{LR^+} = \frac{\color{#C2410C}{Se}}{1 - \color{#6D28D9}{Sp}} \]
LR+ positive likelihood ratioSe sensitivity (true positive rate)Sp specificity
LR− (Eq 5.11)
\[ \color{#0B7B6B}{LR^-} = \frac{1 - \color{#C2410C}{Se}}{\color{#6D28D9}{Sp}} \]
LR− negative likelihood ratioSe sensitivitySp specificity (true negative rate)

Category-specific LR: \( LR_{cat} = P(\text{result} \mid D+) \;/\; P(\text{result} \mid D-) \). Grades evidence by result magnitude, not just whether it crossed a threshold.

Updating probability

Pre-test to post-test: three steps

Step 1

Pre-test odds = P / (1 − P)

Step 2

Post-test odds = pre-test odds × LR

Step 3

Post-test P = odds / (1 + odds)

Example: pre-test P = 2%, LR = 25.95 → post-test P = 35%. Multiplying the pre-test odds by the LR raises them about 26-fold.

Carry forward

The arc of the lesson

  • ROC curves visualise the Se/Sp trade-off; AUC summarises it.
  • LR+ and LR− combine Se and Sp into a single update factor for each test result.
  • Category-specific LRs grade evidence continuously, not just above or below a threshold.
  • The final review and assessment are just below.

Introduction and Overview

Earlier sections treated tests as if they were strictly binary, either positive or negative. In practice, most tests produce a continuous result (a blood pressure reading, an antibody titre, a probability score) that gets dichotomized at a chosen cutpoint. This section makes the cutpoint visible and shows how to choose it well: ROC curves trade sensitivity against specificity at every possible cutpoint, and likelihood ratios let a clinician update probability of disease without doing any of that arithmetic by hand.

Learning Objectives

  • Explain the trade-off between sensitivity and specificity when choosing a cutpoint.
  • Describe receiver operating characteristic (ROC) curves and the area under the curve (AUC).
  • Define and calculate likelihood ratios for positive and negative test results.
  • Apply likelihood ratios to update pre-test probability to post-test probability.
  • Apply the Wilson and Jungner criteria, and their revision by Andermann and colleagues, to decide whether a test should become a screening program.

The section explains the overlap between the test values of healthy and diseased people, uses ROC curves to describe all possible cutpoints at once, introduces likelihood ratios for moving from a pre-test to a post-test probability, and ends by asking when an accurate test should become a screening program.

Interpreting Continuous Test Results

Many tests produce results on a continuous or semi-quantitative scale (e.g., blood urea nitrogen levels, optical density values, enzyme activity). To classify individuals as positive or negative, we select a cutpoint (also called a cut-off or threshold) to determine what level indicates a positive test result.

Box 5.9 explains why every cutpoint produces errors when the test values of the two groups overlap, and Figure 5.4 shows the overlap and a cutpoint that divides it.

Box 5.9: The Overlap Problem

In reality, the distributions of test values for healthy and diseased individuals often overlap. Whatever cutpoint we choose will result in both false positive and false negative results. Raising the cutpoint increases Sp (fewer false positives) but decreases Se (more false negatives). Lowering the cutpoint has the opposite effect.

Healthy Diseased Cutpoint False Negatives False Positives Test Value

Figure 5.4. Overlap between healthy and diseased distributions. Moving the cutpoint left or right trades off sensitivity for specificity.

Each cutpoint therefore gives its own pair of sensitivity and specificity values. The next part describes a single graph that displays all of these pairs together.

Receiver Operating Characteristic (ROC) Curves

A ROC curve plots the Se (y-axis) against the false positive fraction (1 − Sp) (x-axis) computed at a number of different cutpoints (Hanley & McNeil, 1982; see also Wikipedia: ROC curve). This graphical tool helps select the optimum cutpoint and evaluate overall test performance.

The accordion below explains how to read a ROC curve, how to choose an optimal cutpoint from it, and how parametric and non-parametric ROC curves differ.

Interpreting the ROC Curve▼

The 45° diagonal line represents a test with no discriminating ability (no better than chance). The closer the ROC curve gets to the top-left corner, the better the test discriminates between D+ and D− individuals. The top-left corner represents a test with Se = 100% and Sp = 100%.

Choosing the Optimal Cutpoint▼

If sensitivity and specificity are given equal weight, the optimal cutpoint occurs where Se + Sp is at a maximum, which corresponds to the point farthest from the 45° line. This maximised value of Se + Sp − 1 is known as Youden’s J index, and it is the quantity Interactive 5.2 reports as you drag the cutoff. The point closest to the top-left corner is a separate criterion; it agrees with the Youden point when the ROC curve is symmetric and can differ from it when the curve is lopsided. Equal weight on Se and Sp corresponds to equal costs for each false negative and each false positive only when prevalence is 50%. In general, the Youden point minimises the expected cost of errors when the cost of a false negative multiplied by the prevalence equals the cost of a false positive multiplied by (1 − prevalence); at 10% prevalence, for example, it treats one missed case as costing as much as nine false positives. However, if the costs depart from this balance, you might emphasise Se or Sp depending on the clinical context.

Parametric vs. Non-Parametric ROC Curves▼

A non-parametric ROC curve simply plots Se and (1 − Sp) using each observed test value as a cutpoint. A parametric ROC curve provides a smoothed estimate by assuming that the latent variables follow a specified distribution (usually binormal). Both approaches can generate 95% confidence intervals.

Area Under the Curve (AUC)

The AUC summarises the overall discriminatory ability of the test across all cutpoints. It can be interpreted as the probability that a randomly selected D+ individual has a greater test value than a randomly selected D− individual, equivalent to the Mann–Whitney U statistic (Hanley & McNeil, 1982).

Table 5.4 gives conventional labels for ranges of the AUC.

Table 5.4. Conventional interpretation of the area under the ROC curve (AUC).

AUC ValueInterpretation
0.50No discrimination (chance alone)
0.50 – 0.70Poor discrimination
0.70 – 0.80Acceptable discrimination
0.80 – 0.90Excellent discrimination
> 0.90Outstanding discrimination

Interactive 5.2 builds a ROC curve from two overlapping score distributions and reports the sensitivity, the false positive fraction, Youden’s J index and the AUC. Dragging the cutoff moves a point along the curve while the AUC stays the same, and increasing the separation of the distributions or reducing their spread bows the curve towards the upper-left corner and raises the AUC.

📊 Interactive 5.2: ROC Curve Builder

This tool builds a receiver operating characteristic (ROC) curve from the same kind of score distributions as the previous tool: each possible cutoff gives one point, plotting sensitivity against 1 − specificity, and the area under the curve (AUC) summarises how well the test separates the two groups. It is meant to show you that the cutoff decides where you sit on the curve, while the overlap between the two groups decides the shape of the curve itself. As you drag the cutoff line or move its slider, notice that the yellow dot travels along the curve while the AUC stays the same; as you increase the distribution separation or reduce the SD, the curve bows towards the upper-left corner and the AUC climbs towards 1.
Test score distributions

Drag the dashed cutoff line.

02468101214test score
ROC curve

Yellow dot = current cutoff. Diagonal = random-chance reference.

0.00.00.20.20.40.40.60.60.80.81.01.01 − Specificity (False Positive Rate)SensitivityAUC = 0.970
Sensitivity
97.7%
1 − Specificity
25.2%
Youden J
0.725
AUC
0.970
Outstanding discrimination (AUC = 0.970). The curve hugs the upper-left; almost any cutoff is good.

The ROC curve and the AUC summarise a test across all of its cutpoints. A clinician who receives a particular result also needs to know how much that result changes the probability of disease, and the next part introduces the measure that answers this question.

Likelihood Ratios

A likelihood ratio (LR) is the ratio of the probability of a given test result among D+ individuals to the probability of that same result among D− individuals (Deeks & Altman, 2004). LRs combine information from both Se and Sp, and allow the determination of post-test odds from pre-test odds via Bayes's theorem.

The three tabs below give the likelihood ratio for a positive result (Equation 5.10), for a negative result (Equation 5.11) and for a particular category of result (Equation 5.12). Worked Example 5.12 applies Equation 5.10 to the norovirus EIA and Calculator 5.9 reproduces it; Worked Example 5.13 and Calculator 5.10 do the same for Equation 5.11; and Worked Example 5.14 applies Equation 5.12 to a serum marker grouped into three categories, which Calculator 5.11 reproduces.

Likelihood Ratio for a Positive Test (LR+)

Positive likelihood ratio
\[ {\color{#0B7B6B}{LR^+}} = \dfrac{\color{#C2410C}{Se}}{1 - {\color{#6D28D9}{Sp}}} \]Eq 5.10
the positive likelihood ratio is sensitivity divided by one minus specificity (the true positive rate over the false positive rate).

An LR+ of a positive test result is the odds of disease given a positive test result divided by the pre-test odds. Higher LR+ values mean a positive test result is more informative for confirming disease.

Worked Example 5.12: LR+

For the norovirus EIA, Se = 71/82 = 0.866 and Sp = 103/106 = 0.9717, so 1 − Sp = 0.0283.

LR+ = 0.866 / 0.0283 = 30.6. A positive EIA result is about 31 times as likely in a person with norovirus as in a person without it.

Likelihood Ratio for a Negative Test (LR−)

Negative likelihood ratio
\[ {\color{#0B7B6B}{LR^-}} = \dfrac{1 - {\color{#C2410C}{Se}}}{\color{#6D28D9}{Sp}} \]Eq 5.11
the negative likelihood ratio is one minus sensitivity divided by specificity (the false negative rate over the true negative rate).

Lower LR− values mean a negative test result is more informative for ruling out disease. An LR− close to 0 is ideal.

Worked Example 5.13: LR−

For the same test, LR− = (1 − 0.866) / 0.9717 = 0.134 / 0.9717 = 0.138. A negative EIA result is about one-seventh as likely in a person with norovirus as in a person without it, so a negative result divides the odds of infection by roughly seven.

Category-Specific LR

Instead of simply classifying results as positive or negative, researchers in diagnostic settings often calculate category-specific LRs based on the actual test value. This uses the actual result rather than just positive/negative, so the strength of evidence is graded by how extreme the value is.

Category-specific likelihood ratio
\[ {\color{#0B7B6B}{LR_{\text{cat}}}} = \dfrac{\color{#C2410C}{P(\text{result}\mid D^+)}}{\color{#6D28D9}{P(\text{result}\mid D^-)}} \]Eq 5.12
a category-specific likelihood ratio is the probability of that result among diseased people divided by its probability among non-diseased people.

Worked Example 5.14: Category-Specific LRs

A serum marker is measured in 100 people with a disease and 200 people without it, and the results are grouped into three categories.

Marker levelD+ (n = 100)D− (n = 200)LRcat
High604(60/100) / (4/200) = 0.60 / 0.02 = 30.0
Intermediate3036(30/100) / (36/200) = 0.30 / 0.18 = 1.67
Low10160(10/100) / (160/200) = 0.10 / 0.80 = 0.125

A high result is 30 times as likely in a person with the disease as in a person without it and strongly supports the diagnosis. An intermediate result changes the odds very little, and a low result reduces them to one-eighth. Collapsing the three categories into a single positive or negative result would discard this gradation.

From Pre-Test to Post-Test Probability

Likelihood ratios allow you to update your assessment of disease probability after receiving a test result:

Worked Example 5.15: Three-Step Process

  1. Convert pre-test probability to pre-test odds: odds = P / (1 − P)
  2. Multiply by the likelihood ratio: post-test odds = pre-test odds × LR
  3. Convert post-test odds back to probability: P = odds / (1 + odds)

Example: Pre-test probability = 2%, test result at a cutpoint where LRcat = 25.95.

  • Pre-test odds = 0.02/0.98 = 0.0204
  • Post-test odds = 0.0204 × 25.95 = 0.5294
  • Post-test probability = 0.5294 / (1 + 0.5294) = 35%

After obtaining the test result, the estimated probability of disease rises from 2% to 35%.

Equation 5.13 writes the three steps of Worked Example 5.15 as formulae, and Calculator 5.12 applies them to any pre-test probability and likelihood ratio.

Pre-test to post-test probability
\[ \begin{aligned} \text{odds}_{\text{pre}} &= \dfrac{{\color{#C2410C}{P_{\text{pre}}}}}{1 - {\color{#C2410C}{P_{\text{pre}}}}} \\[4pt] \text{odds}_{\text{post}} &= \text{odds}_{\text{pre}} \times {\color{#6D28D9}{LR}} \\[4pt] {\color{#0B7B6B}{P_{\text{post}}}} &= \dfrac{\text{odds}_{\text{post}}}{1 + \text{odds}_{\text{post}}} \end{aligned} \]Eq 5.13
convert the pre-test probability to odds, multiply those odds by the likelihood ratio for the result obtained, and convert the resulting odds back into the post-test probability.

Likelihood ratios combine the result of a test with what was known before testing to give the probability of disease for an individual. Whether a test should be offered to a whole population without symptoms is a further question, which the next part addresses.

From a Test to a Screening Program

The earlier parts of this lesson judged a test by its accuracy. Offering a test to a whole population as a screening program raises further questions about the condition, its treatment and the program itself. Box 5.10 recalls newborn screening for phenylketonuria (PKU) from HSCI 130, and the text that follows sets out the principles of Wilson and Jungner (1968) and their revision by Andermann and colleagues (2008).

Box 5.10: Recall: Newborn Screening for PKU

HSCI 130 Lesson 7, Section 4 (Newborn Screening, Precision Medicine, and DTC Testing), presented newborn screening for phenylketonuria (PKU) as public health genetics that works. PKU can be detected from a heel-prick blood spot before any symptoms appear (the Guthrie assay, 1962), and a phenylalanine-restricted diet started early prevents its harm. Every Canadian province now screens newborns for a panel of conditions.

Retrieval question. Which features of PKU made newborn screening for it worthwhile?

Show the answer▼

PKU is serious, it can be detected before symptoms by a simple and acceptable test on a blood spot, and an effective treatment, the restricted diet, works best when it starts early.

The accuracy of a test is one requirement for screening among several. A screening program offers a test to people without symptoms, so it does harm to some of them (false positives, unnecessary follow-up and the treatment of disease that would never have caused illness) as well as good to others. Wilson and Jungner (1968), writing for the World Health Organization, set out ten principles for deciding whether screening for a condition is worthwhile:

  1. The condition sought should be an important health problem.
  2. There should be an accepted treatment for people with recognized disease.
  3. Facilities for diagnosis and treatment should be available.
  4. There should be a recognizable latent or early symptomatic stage.
  5. There should be a suitable test or examination.
  6. The test should be acceptable to the population.
  7. The natural history of the condition, including its development from latent to declared disease, should be adequately understood.
  8. There should be an agreed policy on whom to treat as patients.
  9. The cost of case-finding, including diagnosis and treatment, should be economically balanced in relation to possible expenditure on medical care as a whole.
  10. Case-finding should be a continuing process.

The first eight principles concern the condition, the test and the treatment, and the sensitivity, specificity and predictive values of this lesson are the evidence for the fifth. Andermann and colleagues (2008) reviewed the criteria proposed in the following forty years and synthesized them into a revised set that adds requirements for the program itself: it should respond to a recognized need; define its objectives and its target population at the outset; rest on scientific evidence of effectiveness; integrate education, testing, clinical services and program management; include quality assurance that minimizes the risks of screening; ensure informed choice, confidentiality and respect for autonomy; promote equity and access for the whole target population; plan its evaluation from the outset; and show that its overall benefits outweigh its harms.

Box 5.11 applies both sets of criteria to colorectal cancer screening in Canada, first through the principles that concern the condition, the test and the treatment and then through the program-level criteria.

Box 5.11: Applying the Criteria: Colorectal Cancer Screening in Canada

Colorectal cancer is one of the most commonly diagnosed cancers in Canada and a leading cause of cancer death, so it is an important health problem. Most colorectal cancers develop over years from adenomatous polyps, which gives a recognizable early stage, and polyps can be removed during colonoscopy, which gives an accepted treatment. The fecal immunochemical test (FIT), which detects blood in a stool sample collected at home, is a suitable and acceptable test, and the Canadian Task Force on Preventive Health Care (2016) recommends screening adults aged 50 to 74 at average risk with a fecal test every two years or with flexible sigmoidoscopy every ten years. Organized provincial programs offer the test to people in the target age range and recall them when the next test is due, which makes case-finding a continuing process.

The program-level criteria show where such programs succeed or struggle. A positive FIT leads to colonoscopy, so the facilities criterion depends on colonoscopy capacity: long waits after a positive result reduce the benefit of early detection. Quality assurance covers the follow-up of every positive result and the standards for colonoscopy, whose rare complications include bleeding and perforation. Participation varies across population groups, so the equity criterion requires that programs measure who is screened and reach those who are not. Randomized trials of fecal occult blood testing have shown reductions in colorectal cancer mortality, which supports the judgement that benefits outweigh harms for the target age range.

Screening programs are judged by outcomes that three biases can distort. Lead-time bias makes survival measured from diagnosis look longer in screen-detected cases because the diagnosis is earlier, even when death is not delayed. Length-biased sampling arises because screening at intervals preferentially detects slow-growing disease, which spends longer in the detectable preclinical phase and has a better prognosis. Overdiagnosis is the detection of disease that would never have caused symptoms or death in the person's lifetime. For these reasons a screening program is evaluated by its effect on mortality from the condition in the whole screened population, ideally in a randomized trial. HSCI 230 Lesson 9, Section 2 (Observer and Detection Bias), and HSCI 230 Lesson 10, Section 2 (Immortal Time and Lead-Time Bias), treat these biases in detail.

This section has shown how the choice of cutpoint determines sensitivity and specificity, how the ROC curve and the AUC summarise a test across all cutpoints, how likelihood ratios convert a pre-test probability into a post-test probability, and how the accuracy of a test fits among the wider criteria for a screening program. The reflection, key takeaways and knowledge check below review this material.

Reflection

A disease screening programme uses a test with Se = 92.7% and Sp = 77.4% at a particular cutpoint. Calculate LR+ for this cutpoint using LR+ = Se / (1 − Sp). If the pre-test probability of disease is 10%, convert it to pre-test odds (odds = probability / (1 − probability)), multiply by LR+ to obtain the post-test odds, and convert back to a post-test probability (probability = odds / (1 + odds)). Discuss whether this cutpoint is appropriate for a screening programme where false negatives are very costly.

Model answerLR+ = Se / (1 − Sp) = 0.927 / 0.226 = 4.10. Pre-test probability 10% → pre-test odds = 0.10/0.90 = 0.111. Post-test odds = 0.111 × 4.10 = 0.456. Post-test probability = 0.456/(1+0.456) = 0.313 (31%). For a screening test where false negatives are very costly, this cutpoint is questionable: an LR+ of 4.1 multiplies the disease odds by only about four, raising the probability of disease from 10% to 31%, and the 92.7% Se still misses 7% of true cases. The 77.4% Sp also produces many false positives (each requiring follow-up). Better strategies: (a) move the threshold down to raise Se (accept more FP); (b) use this cutpoint as a first-stage triage with mandatory confirmatory testing on all positives; (c) re-screen at intervals to catch FN at the next round; (d) supplement with a second independent test for parallel screening, which raises overall Se.

Minimum 20 characters required.

✓ Reflection saved

Key Takeaways

  • The choice of cutpoint involves a trade-off between sensitivity and specificity.
  • ROC curves plot Se vs. (1 − Sp) across cutpoints; the AUC summarises overall test performance.
  • An AUC of 0.5 represents chance; values closer to 1.0 indicate better discrimination.
  • LR+ = Se/(1 − Sp); LR− = (1 − Se)/Sp. LRs combine both Se and Sp into a single metric.
  • LRs allow conversion of pre-test probability to post-test probability using a three-step odds-based calculation.
Knowledge Check: this section

1. A ROC curve that perfectly follows the 45° diagonal indicates:

The 45° diagonal represents a test that performs no better than random chance (AUC = 0.5). A good test produces a ROC curve that bows toward the top-left corner.

2. If a test has Se = 90% and Sp = 80%, what is LR+?

LR+ = Se / (1 − Sp) = 0.90 / (1 − 0.80) = 0.90 / 0.20 = 4.5. This means a positive test result is 4.5 times more likely in a diseased individual than in a non-diseased individual.

3. Raising the cutpoint for a continuous test will generally:

Raising the cutpoint means fewer individuals test positive. This reduces false positives (increasing Sp) but increases false negatives (decreasing Se).

✦ Pass the knowledge check with 100% and complete the reflection to continue

Section 5

Final Review & Assessment

⏱ Estimated time: 20 minutes

Bringing It All Together

This lesson built up the toolkit for evaluating tests, from the basic distinction between screening (in healthy populations) and diagnosis (in suspected cases), through sensitivity and specificity, into predictive values, and finally into the more sophisticated machinery of cutpoints, ROC curves, and likelihood ratios. The arc moves from how does the test perform? to what does this result mean for this patient in this setting?

The deepest idea in the lesson is that test performance is never just a property of the test. The same sensitivity and specificity produce very different predictive values when prevalence changes, which is why a screening protocol that works in a high-prevalence clinic can collapse into mostly false positives when applied to the general population. Published diagnostic-accuracy studies themselves are subject to design-related bias that inflates reported performance (Lijmer et al., 1999), motivating the QUADAS-2 quality-assessment tool (Whiting et al., 2011) and STARD 2015 reporting standard (Bossuyt et al., 2015). As you finish the assessment, the takeaways below are the practical companions: keep them in mind whenever someone tells you a test is “accurate.”

Key Takeaways from this lesson

  • Test performance has two layers: accuracy (closeness to truth) and precision (consistency); agreement is quantified with Cohen's kappa.
  • Sensitivity (Se = a/m1) and specificity (Sp = d/m0) are properties of the test: SnNOut for ruling out, SpPIn for ruling in.
  • Predictive values (PV+ and PV−) depend strongly on prevalence: even excellent tests yield mostly false positives in low-prevalence settings.
  • Strategies to raise PV+ include targeting high-risk groups, using more specific confirmatory tests, and testing in series rather than parallel.
  • For continuous tests, the chosen cutpoint is a Se/Sp trade-off; the ROC curve and AUC summarise performance across cutpoints.
  • Likelihood ratios integrate Se and Sp into a single quantity that updates pre-test odds to post-test odds, the most direct way to interpret a single test result.

Reflection

You are advising a public health agency that wants to implement a two-stage screening programme for a disease with a population prevalence of 2%. The first-stage test has Se = 95% and Sp = 90%, and the second-stage (confirmatory) test has Se = 85% and Sp = 99%. When tests are used in series, only people who test positive on the first test receive the second, and a person is classified as positive only if both tests are positive; the combined sensitivity is Se1 × Se2 and the combined specificity is 1 − (1 − Sp1)(1 − Sp2). Positive predictive value is PV+ = (P × Se) / [P × Se + (1 − P) × (1 − Sp)], where P is the prevalence. Discuss how using these tests in series would affect the overall Se, Sp, and PV+ compared to using just the first test alone. What are the practical implications of this approach?

Model answerSeries (two-stage) screening: test 1 first, test 2 only on positives. Overall Se = 0.95×0.85 = 0.808; overall Sp = 1 − (1−0.90)×(1−0.99) = 1 − 0.001 = 0.999; at P = 0.02, PPV = (0.02×0.808)/(0.02×0.808 + 0.98×0.001) = 0.0162/0.0172 ≈ 0.94. Compared to test 1 alone (PPV at P = 0.02 with Se 0.95, Sp 0.90 is 0.162), the two-stage approach dramatically improves PPV (94% vs. 16%) at the cost of reduced overall Se (81% vs. 95%). Trade-offs: fewer false alarms (good for downstream costs, anxiety, and inappropriate treatment) but more missed cases (bad for outcomes when early detection matters). Ethics: a programme must transparently report both Se and PPV and disclose what fraction of true cases will go undetected by the screen. Where missed cases are catastrophic (newborn metabolic disorders), parallel testing or repeat screens at intervals are preferable to series.

Minimum 20 characters required.

✓ Reflection saved

Final Knowledge Assessment

Complete all 15 questions below with 100% accuracy to finish this lesson. You must also complete the reflection above before submitting.

Final Assessment: Screening and Diagnostic Tests

1. The analytic sensitivity of a test refers to:

Analytic sensitivity refers to the lowest concentration the test can detect. This is distinct from diagnostic (epidemiologic) sensitivity, which is the proportion of truly diseased individuals testing positive.

2. A kappa value of 0.55 between two diagnostic tests indicates:

According to the Landis and Koch interpretation scale, a kappa of 0.41–0.60 indicates moderate agreement.

3. In a 2×2 table for test evaluation, cell “c” represents:

In the standard 2×2 table, cell c represents false positives: individuals who do not have the disease (D−) but test positive (T+).

4. If Se = 80% and Sp = 95%, what is the false positive fraction (FPF)?

FPF = 1 − Sp = 1 − 0.95 = 0.05 or 5%. The false positive fraction depends only on specificity.

5. The Rogan-Gladen formula is used to:

The Rogan-Gladen formula estimates true prevalence: P = (AP + Sp − 1) / (Se + Sp − 1), correcting for test imperfections.

6. A screening programme tests 10,000 people for a disease with 1% prevalence using a test with Se = 99% and Sp = 95%. How many false positives would you expect?

Non-diseased individuals = 10,000 × 0.99 = 9,900. False positives = 9,900 × (1 − 0.95) = 9,900 × 0.05 = 495.

7. PV+ depends on which of the following?

PV+ is determined by the formula: PV+ = (P × Se) / [P × Se + (1 − P) × (1 − Sp)]. It depends on all three: Se, Sp, and prevalence.

8. In the context of ROC curves, the area under the curve (AUC) of 0.85 indicates:

An AUC of 0.80–0.90 is generally interpreted as excellent discrimination between diseased and non-diseased individuals.

9. LR+ = Se / (1 − Sp). If a test has Se = 95% and Sp = 90%, what is LR+?

LR+ = 0.95 / (1 − 0.90) = 0.95 / 0.10 = 9.5. A positive result is 9.5 times more likely in a diseased individual.

10. The mnemonic “SnNOut” means:

SnNOut stands for Sensitivity, Negative result, Rules Out. If a highly sensitive test is negative, the individual is very unlikely to have the disease because the test catches almost all true cases.

11. A Bland-Altman plot is used to:

A Bland-Altman (limits of agreement) plot displays the differences between paired measurements against their mean, revealing systematic bias and whether disagreement varies with measurement magnitude.

12. Using tests in series (sequential testing) will generally:

Testing in series requires both tests to be positive to classify as positive. This reduces false positives (increasing Sp and PV+) but may miss some true positives (decreasing overall Se).

13. McNemar’s χ² test is used before evaluating kappa to:

McNemar’s test checks for systematic bias between two tests. If one test produces significantly more positive results than the other, the detailed assessment of agreement could be misleading.

14. To convert pre-test probability to post-test probability using a likelihood ratio, the correct sequence is:

The three-step process is: (1) convert pre-test probability to pre-test odds, (2) multiply pre-test odds by the LR to get post-test odds, (3) convert post-test odds back to post-test probability.

15. Which factor does NOT directly affect the predictive value of a test?

Predictive values are determined by Se, Sp, and prevalence. The coefficient of variation (CV) is a measure of test precision/reproducibility and does not directly enter the PV formula.

✦ Complete the final reflection above before submitting