HSCI 230, Lesson 2

Systematic Reviews and Meta-Analysis

Evaluating Epidemiological Research

Learning objectives for this lesson:

  • Place systematic reviews and meta-analyses within the hierarchy of evidence, and explain why the pyramid is a heuristic, not a verdict
  • Use the DIKW hierarchy (data → information → knowledge → wisdom) to frame the role of evidence synthesis in public-health decision-making
  • Outline the steps of a systematic review, from specifying the question to synthesising results
  • Build a database search from MeSH headings, free-text synonyms and Boolean operators, and document it in a search log
  • Appraise a published systematic review against the seven AMSTAR 2 critical domains
  • Complete the data-extraction process to provide data suitable for meta-analysis
  • Calculate summary estimates of effect and evaluate heterogeneity among study results
  • Choose between fixed- and random-effects models and explain when each is appropriate
  • Present and interpret forest plots and other graphical displays of meta-analysis results
  • Evaluate potential causes of heterogeneity using subgroup analysis, stratification, and meta-regression
  • Evaluate the potential impact of publication bias using funnel plots and related methods
  • Determine if results have been influenced by an individual study (sensitivity analysis)

This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University.

Reference

Glossary: Key Terms, People & Concepts

📚 Reference page, available throughout the lesson

This glossary collects the key concepts, people, and ideas you will meet in this lesson. Use it as a reference while you work through the material, or as a review before assessments. Type in the search box to filter entries.

Key Concepts & Ideas
Hierarchy of Evidence (Evidence Pyramid) A heuristic ranking of study designs by the amount of variation each design typically rules out, from expert opinion and case reports at the base, through cross-sectional, case-control, and cohort studies, to randomised controlled trials, and finally to systematic reviews and meta-analyses at the apex. A useful first sort for "where to look for the strongest available evidence," not a verdict on the quality of any individual study.
DIKW Pyramid A four-tier model from informatics and decision science: Data → Information → Knowledge → Wisdom. Raw data become information when organised; information becomes knowledge when appraised and integrated; knowledge becomes wisdom when exercised in a specific context under uncertainty. A meta-analysis is the highest-order knowledge layer; the wisdom layer is where contextual, value-laden public-health judgement happens.
Systematic Review A structured, protocol-driven review that uses pre-specified, reproducible methods to identify, appraise, and synthesise all studies relevant to a focused research question. Aims to minimise selection and reporting bias in evidence synthesis.
Narrative (Traditional) Review An informal qualitative summary of literature, typically without protocol, comprehensive search, or risk-of-bias appraisal. Useful for orientation but vulnerable to author selection bias.
Meta-Analysis The statistical pooling of effect estimates from multiple studies into a single weighted summary estimate, with quantification of between-study variation.
Protocol & PROSPERO A systematic-review protocol pre-specifies the question, eligibility criteria, search strategy, and analysis plan. PROSPERO is the international prospective register of systematic reviews where protocols are deposited before screening begins.
PICO(S/T) Framework Population, Intervention, Comparator, Outcome (and optionally Study design / Time), the structure used to write a focused, answerable review question about an intervention.
PECO Framework Population, Exposure, Comparator, Outcome, the parallel structure used when a review question concerns an exposure, such as an environmental or behavioural risk factor, in place of an intervention.
Grey Literature Research published outside conventional peer-reviewed journals (theses, conference abstracts, government reports, preprints). Searching grey literature helps mitigate publication bias.
Heterogeneity Variation in true effect estimates across studies beyond what is expected from sampling error. Sources include clinical, methodological, and statistical heterogeneity.
I² Statistic The percentage of total variation across studies attributable to heterogeneity rather than chance (Higgins & Thompson, 2002). Rough benchmarks: 25% low, 50% moderate, 75% high.
Cochran's Q Test A chi-squared test of the null hypothesis of homogeneity across studies. Often underpowered with few studies and overpowered with many; used alongside I² and tau².
Tau² (between-study variance) The estimated variance of the true effect across studies in a random-effects meta-analysis. Larger tau² means greater between-study variation.
Fixed-Effect Model Assumes a single common true effect underlies all studies; differences between studies arise only from sampling error. Inverse-variance weighting gives more weight to larger studies.
Random-Effects Model Assumes the true effect varies across studies according to a distribution. Pooled estimate has a wider confidence interval and weighting is more even across studies (DerSimonian & Laird, 1986).
Publication Bias Systematic distortion of meta-analytic results because studies with statistically significant or favourable findings are more likely to be published, indexed, or available in English.
Risk of Bias The likelihood that features of a study's design or conduct have made its result systematically too large or too small, judged domain by domain with a tool matched to the design (RoB 2 for randomised trials, ROBINS-I for non-randomised studies of interventions). The categories selection, performance, detection, attrition and reporting bias organised the 2011 Cochrane tool (Higgins et al., 2011). Drives sensitivity analyses and GRADE downgrading.
GRADE Grading of Recommendations Assessment, Development and Evaluation, a system for rating the certainty of evidence (high, moderate, low, very low) and the strength of resulting recommendations (Guyatt et al., 2008).
Methods, Measures & Tools
PRISMA 2020 Preferred Reporting Items for Systematic Reviews and Meta-Analyses (Page et al., 2021): a 27-item checklist and a flow diagram with identification, screening and included phases, with a separate column for records found by other methods, for transparent reporting of systematic reviews. The four-phase diagram with an eligibility phase belongs to the 2009 statement.
Forest Plot A graphical display of effect estimates from each study (squares sized by weight) with their confidence intervals, plus the pooled estimate (diamond) and a line of no effect.
Funnel Plot A scatter plot of effect size against standard error (or precision) used to inspect for asymmetry suggestive of publication bias or small-study effects.
Egger's Test & Begg's Test Statistical tests for funnel-plot asymmetry. Egger's (1997) regresses standardised effect on precision; Begg's uses rank correlation. Both are underpowered with few studies.
Trim-and-Fill A non-parametric method that imputes hypothetical missing studies to make a funnel plot symmetric, then re-estimates the pooled effect to gauge sensitivity to publication bias.
RoB 2 (Risk of Bias 2) The Cochrane risk-of-bias tool for randomised trials, organised by signalling questions across five domains (randomisation, deviations, missing data, measurement, selection of result) (Sterne et al., 2019).
ROBINS-I Risk Of Bias In Non-randomised Studies of Interventions, the Cochrane tool for appraising non-randomised intervention studies, adding domains for confounding and participant selection (Sterne et al., 2016).
AMSTAR 2 A 16-item critical appraisal tool for systematic reviews of randomised and non-randomised studies of healthcare interventions (Shea et al., 2017). Seven items are critical domains, and the overall confidence in a review is rated high, moderate, low or critically low according to its critical flaws and non-critical weaknesses.
MeSH (Medical Subject Headings) The controlled vocabulary used to index MEDLINE. Each heading gathers the records on one concept under a single label, lists entry terms (synonyms) that map to it, and sits in a tree of broader and narrower headings; PubMed explodes a heading by default to include its narrower headings. Embase uses a parallel thesaurus, Emtree.
Boolean Operators The operators that combine search terms: OR joins synonyms within one concept and widens a search, AND joins different concepts and narrows it, and NOT removes records containing a term and is used with care. A search log records each string, its date, the number of records and the reason for each change.
Leave-One-Out Sensitivity Analysis A meta-analytic robustness check that recalculates the pooled estimate with each study removed in turn to identify influential studies.
Meta-Regression & Subgroup Analysis Methods for exploring whether study-level covariates (year, dose, region, risk-of-bias rating) explain heterogeneity in effect sizes across studies.
Key People & Organisations
Archie Cochrane (1909–1988) Scottish epidemiologist who argued in Effectiveness and Efficiency (1972) that medical practice should be guided by systematically appraised RCT evidence; namesake of the Cochrane Collaboration.
Iain Chalmers (1943– ) British health-services researcher who founded the Cochrane Collaboration in 1993 and helped establish the modern infrastructure for systematic reviews.
Cochrane (Collaboration) An international not-for-profit network producing the Cochrane Database of Systematic Reviews and stewarding the standard tools (RoB 2, ROBINS-I, RevMan) for evidence synthesis.
No matching entries. Try a different search term.
Section 1 of 5

Hierarchy of Knowledge & Systematic Reviews

⏱ Estimated reading time: 55 minutes
Section 1 of 5

Hierarchy of Knowledge & Systematic Reviews

The evidence pyramid, the DIKW framework, and the seven-step systematic review process.

Why start here

The apex before the rungs below

Almost every public health guideline rests on a systematic review. Understanding how that synthesis works, and where it can mislead, is the lens for reading every study that follows.

This section sets up that lens before the course works down the pyramid.

The evidence pyramid

Ruling out variation, one tier at a time

Systematic Reviews & Meta-Analyses Randomised Controlled Trials Cohort Studies Case-Control Studies Cross-Sectional · Case Series Expert Opinion · Anecdote

Higher tiers rule out more sources of variation, but the hierarchy is a heuristic. Appraise what you find.

The DIKW hierarchy

What the evidence is for

Data

Raw observations: counts, measurements, lab values.

Information

Data given structure, context, and comparison.

Knowledge

Information appraised, integrated, and defensible in print.

Wisdom

Knowledge applied in context, under uncertainty, for a specific decision.

A meta-analysis is the highest knowledge layer. Wisdom is what happens next, and that step is irreducibly human.

Why narrative reviews fail

From subjective summary to documented procedure

Narrative review

Subjective selection. No documented search. All studies weighted equally. Prone to reviewer preconceptions.

Systematic review

Structured, transparent methodology. Reproducible search. Quality appraisal. Pre-specified inclusion criteria.

Antman et al. (1992): expert narrative reviews of myocardial-infarction treatments lagged years behind the cumulative meta-analytic evidence.

Seven steps

The systematic review process (Sargeant et al., 2006)

  1. Specify the question: intervention, population, outcome, comparator
  2. Lay out the protocol: transparent, pre-specified, reproducible
  3. Find all the studies: databases, reference lists, grey literature
  4. Determine relevance: inclusion/exclusion criteria, two independent reviewers
  1. Evaluate quality: RoB 2, ROBINS-I, or equivalent tools
  2. Extract the data: point estimates and precision, standardised template
  3. Synthesise results: qualitatively or quantitatively (meta-analysis)
Quality assurance

PROSPERO and GRADE

PROSPERO

Prospective protocol registration before study selection. Creates a public audit trail. Mirrors the role of ClinicalTrials.gov for primary trials.

GRADE

Rates certainty of evidence per outcome: High, Moderate, Low, or Very Low. Accounts for risk of bias, inconsistency, indirectness, imprecision, and publication bias.

Carry forward

What to take into the next section

  • The evidence pyramid is a heuristic for where to look, not a verdict on individual studies.
  • Systematic reviews replace subjective selection with documented, reproducible procedure.
  • PROSPERO registers the protocol; GRADE rates certainty in the conclusions.
  • A meta-analysis is knowledge, not wisdom. The decision step is still human.

Introduction and Overview

An earlier lesson set up the foundations of epidemiology, its history, its ways of knowing, and the trust scaffolding (research integrity, reproducibility) that lets a body of evidence accumulate. A later section of that lesson made the case that empiricism is one way of knowing among several, and a later section left you with a personal stance on which of those ways you would lean on when the stakes are public-health decisions. This lesson picks up at exactly that seam: once you commit to evidence-based reasoning, how do you tell stronger evidence from weaker? Before we open up the machinery of systematic reviews and meta-analyses, we need a shared way of organising the answer. That organising idea (the hierarchy of knowledge) is where this lesson starts, and it is the lens through which every later lesson in the series will be read.

The Hierarchy of Knowledge

Before reading further, sit with the prompt below for a moment. The rest of the section is easier to follow if you have already put your own intuitions on paper.

💭 Discussion Prompt, Rank the Evidence

Suppose four sources tell you the same intervention works:

  1. A trusted senior clinician's experience over a 30-year career.
  2. A single, well-written case report.
  3. A well-designed cohort study of 4,000 people followed for five years.
  4. A meta-analysis of twelve randomised trials covering 18,000 participants.

Rank them from weakest to strongest evidence for a public-health recommendation. Then ask: what is doing the ranking work, sample size? design? the number of independent studies? Could any single feature of one of the weaker sources overturn your ranking?

Epidemiologists answer this question with a layered model usually called the hierarchy of evidence, or simply the evidence pyramid. The pyramid is not a strict ranking of every study against every other; it is a heuristic for thinking about how much variation a particular design has ruled out. At the base sit forms of evidence that are easy to produce but hard to defend on their own: expert opinion, individual experience, single case reports. Moving up, we add structured comparison (case series → cross-sectional → case-control → cohort). Higher still are interventional designs that randomise the exposure (randomised controlled trials), which under ideal conditions rule out confounding entirely. At the apex sit the designs that synthesise across all of the layers below: systematic reviews and meta-analyses.

The hierarchy of evidence pyramid A six-tier pyramid showing levels of evidence from expert opinion at the base to systematic reviews and meta-analyses at the apex. Systematic Reviews & Meta-Analyses Randomised Controlled Trials Cohort Studies Case-Control Studies Cross-Sectional Studies · Case Series Expert Opinion · Anecdote · Case Reports Synthesis layer , pools across studies Interventional , randomises exposure Observational , structured comparison Unstructured , easy, but hard to defend Strength of inference (more variation ruled out)
Figure 2.1, The traditional hierarchy of evidence. Higher tiers rule out more sources of variation, but the hierarchy is a heuristic, not a verdict on any individual study.

A Complementary Hierarchy: DIKW

The evidence pyramid tells us where to look. A second, complementary hierarchy tells us what the evidence is for. The DIKW pyramid (Data → Information → Knowledge → Wisdom) is widely used in informatics, decision science, and clinical practice:

  • Data are raw observations: counts, measurements, dates, lab values.
  • Information is data given structure: organised, contextualised, compared.
  • Knowledge is information appraised: filtered for quality, integrated across sources, ready to defend in print.
  • Wisdom is knowledge exercised in a specific context, under uncertainty, in service of a particular decision.

A meta-analysis is not the same thing as wisdom. It is the highest-order knowledge layer in the evidence pyramid, the most thoroughly appraised, most heavily aggregated form of knowledge the field knows how to produce. Wisdom is what a clinician, policy-maker, or community partner does with that knowledge in their particular case, and that step is irreducibly human, contextual, and value-laden. An earlier lesson's call for a critical epidemiology that asks who benefits, who is harmed, whose priorities lives at the wisdom layer, not the knowledge layer. Keeping the two hierarchies side-by-side is a useful corrective against treating a forest plot as if it answered every question on its own.

Why Start at the Top of the Pyramid?

This course spends later lessons working through the rungs below the apex: case-control, cohort, ecological, and cross-sectional designs, plus the threats (selection, information, confounding) that compromise each. So why does this lesson begin at the top? Three reasons, each of which shapes the rest of the series:

  1. You need the appraisal lens before you read any single study. When you encounter a primary observational study in later lessons, the first question is not "is this study good?" but "what does the synthesised literature on this question already say, and how does this study fit in?" The synthesis layer is the lens through which every later study should be read.
  2. The hierarchy is a heuristic, not a verdict. A poorly conducted systematic review is weaker evidence than a well-designed cohort study. A well-designed case series of an unprecedented exposure may be the only evidence in existence. The hierarchy tells you where to look for the strongest available evidence; it does not exempt you from appraising what you find. A later section of this lesson will make that point concrete with publication bias, influential studies, and outcome-scale issues that can compromise even a textbook meta-analysis.
  3. Every tool in this lesson exists because the lower tiers can mislead. PROSPERO registration, PRISMA reporting, GRADE certainty ratings, and the Cochrane risk-of-bias appraisals (so named for Archie Cochrane, whose 1972 monograph Effectiveness and Efficiency framed the case for systematically appraised RCT evidence); each is a field-wide response to a specific failure of unstructured evidence accumulation. An earlier lesson's trust ledger and the open-science reforms in its later sections are the wider movement these tools belong to. The hierarchy of evidence is, in a real sense, a history of the lessons epidemiologists learned the hard way.

With the hierarchy in mind, the four content sections of this lesson move from broad to specific. This section sets up systematic reviews as a structured way to identify and appraise all relevant studies; a later section turns to meta-analysis as the quantitative pooling of effect estimates, including the central choice between fixed- and random-effects models; a later section covers the forest plot and heterogeneity analysis; a later section closes with the threats to a meta-analysis that survive correct technique. You will revisit these tools throughout the remainder of this course as you appraise individual observational studies, every appraisal sits inside the question, "and what does the synthesised literature say?"

Learning Objectives

  • Place systematic reviews and meta-analyses within the hierarchy of evidence, and explain why the pyramid is a heuristic rather than a verdict on any single study.
  • Articulate how the DIKW hierarchy (data → information → knowledge → wisdom) frames evidence synthesis in public-health decision-making.
  • Distinguish a narrative review from a systematic review and explain why narrative reviews are unsuitable for guiding policy.
  • Outline the seven steps of a systematic review (Sargeant et al., 2006), from question specification through synthesis.
  • Build a database search for a focused question from MeSH headings, free-text synonyms and Boolean operators, and document it in a search log.
  • Appraise a published systematic review against the seven AMSTAR 2 critical domains and reach an overall confidence rating.
  • Describe the role of the PRISMA reporting checklist and the Cochrane Risk of Bias tools in producing a trustworthy review.
  • Identify the inclusion/exclusion criteria, search strategy, and quality-appraisal steps that make a systematic review reproducible.

16.1 Why Systematic Reviews?

When making decisions about health interventions, we want to use all available information. Unfortunately, the literature is often inconclusive and conflicting, individual studies may produce results ranging from statistically significant to inconsequential, and the variation among results may be greater than expected from chance alone. A classic demonstration of the cost of relying on narrative summaries is Antman et al. (1992), who showed that expert reviews of myocardial-infarction treatments lagged years behind what a cumulative meta-analysis of randomised trials had already established.

There are two fundamental approaches to formally reviewing available data: a narrative review and a systematic review (which may include a meta-analysis).

Narrative Reviews

In a narrative review, each study is considered individually, and the reviewer subjectively assesses the evidence. Narrative reviews have several limitations:

  • They tend to be carried out by subject experts who may bring preconceived opinions, resulting in biased review
  • They often lack a structured methodology for identifying and assessing relevant studies
  • Small but well-designed studies may be omitted if they lack statistical power
  • Inclusion criteria are often not described in adequate detail
  • There is a tendency to weight all studies equally, when they should not all receive equal weight

Narrative reviews should only be used to provide an overview of literature, not to guide treatment or policy decisions.

HSCI 241 Lesson 1, which can be taken before or after this course, describes the other types of review, including scoping and rapid reviews, and the circumstances in which a narrative review is appropriate.

Systematic Reviews

A systematic review uses a structured, transparent methodology to identify, evaluate, and synthesise all relevant studies on a specific question. It minimises bias and provides reproducible results. A systematic review may or may not include a quantitative meta-analysis, depending on the nature and quality of available data.

16.2 Steps of a Systematic Review

A systematic review follows seven steps (Sargeant et al., 2006), summarised here at the level a reader needs in order to judge a published review.

  1. Specify the question. The question comes from a health-policy or clinical objective and is framed with PICO for an intervention or PECO for an exposure.
  2. Lay out the protocol. A written plan for every later step is fixed before screening and ideally registered on PROSPERO (16.2.1).
  3. Find all the studies. The team searches bibliographic databases, reference lists and grey literature, and documents every search (16.2.2).
  4. Determine relevance. Two reviewers independently apply the eligibility criteria to titles and abstracts, then to full texts.
  5. Evaluate study quality. Each included study is assessed for risk of bias with a tool matched to its design. For randomised trials, RoB 2 (Sterne et al., 2019) judges five domains: bias arising from the randomisation process, deviations from intended interventions, missing outcome data, measurement of the outcome, and selection of the reported result. ROBINS-I (Sterne et al., 2016) is the parallel tool for non-randomised studies of interventions. Sequence generation, allocation concealment and blinding were items of the original Cochrane tool (Higgins et al., 2011) that RoB 2 replaced. Risk-of-bias judgements are reported beside each study's result and are used in pre-specified sensitivity analyses that exclude studies at high risk of bias, in subgroup analyses by risk of bias, and in the GRADE risk-of-bias domain; the meta-analysis itself weights studies by inverse variance, and summary quality scores are avoided because different scales give different conclusions (Jüni et al., 1999).
  6. Extract the data. Two reviewers independently record each study's estimate and its precision on a standardised form.
  7. Summarise and synthesise. Results are summarised narratively or pooled in a meta-analysis, the subject of the next three sections.

HSCI 241, which can be taken before or after this course, teaches each step in a lesson of its own.

16.2.1 Two Tools That Underpin a Trustworthy Review

Modern systematic reviews are expected to do more than simply follow the seven steps; they must do so transparently and they must communicate how confident readers should be in the conclusions. Two widely adopted tools have become the de facto standards for these expectations: PROSPERO for protocol registration and GRADE for rating the certainty of the evidence.

📝 Call-Out: PROSPERO, Prospective Protocol Registration

What it is. PROSPERO is the international prospective register of systematic reviews, hosted by the Centre for Reviews and Dissemination at the University of York. Reviewers submit their protocol, research question, eligibility criteria, search strategy, planned analyses, and outcomes, before beginning study selection or data extraction.

Why it matters. Prospective registration mirrors the role of ClinicalTrials.gov for primary trials. It creates a public, time-stamped record of the review’s plan, which:

  • Reduces the risk of outcome-reporting bias and post hoc changes to inclusion criteria
  • Helps avoid duplication of effort across research teams working on the same question
  • Allows readers, peer reviewers, and editors to compare the final review against the planned protocol

PROSPERO registration is now required or strongly encouraged by Cochrane, by many journals (e.g., BMJ, JAMA, Annals of Internal Medicine), and by the PRISMA 2020 reporting guideline (Page et al., 2021; original statement: Moher et al., 2009). Pair the registration with PRISMA at submission and you have a fully transparent audit trail from question to conclusion.

🎯 Call-Out: GRADE, Rating the Certainty of the Evidence

What it is. The Grading of Recommendations, Assessment, Development and Evaluations (GRADE) framework (Guyatt et al., 2008) provides a structured method for rating how much confidence we should have in the body of evidence for each outcome of a systematic review, separately from rating the strength of any clinical recommendation that follows.

How it works. Each outcome starts at a baseline level of certainty based on study design (RCTs start at high; observational studies start at low), and is then downgraded or upgraded based on eight domains:

DomainDirectionWhat it captures
Risk of biasDowngradeMethodological limitations of the included studies
InconsistencyDowngradeUnexplained heterogeneity across studies
IndirectnessDowngradeDifferences in population, intervention, comparator, or outcome from the question of interest
ImprecisionDowngradeWide confidence intervals or few events
Publication biasDowngradeSuspected selective reporting of positive findings
Large effectUpgradeVery large or very consistent observed effects
Dose-responseUpgradePlausible dose-response gradient
Plausible confoundingUpgradeResidual confounding would reduce rather than create the observed effect

The body of evidence is then summarised in one of four certainty levels:

  • High: we are very confident the true effect lies close to the estimate
  • Moderate: the true effect is likely close to the estimate, but could be substantially different
  • Low: the true effect may be substantially different from the estimate
  • Very low: we have very little confidence in the estimate

GRADE certainty ratings are typically presented in a Summary of Findings table alongside the pooled effect estimates, allowing decision-makers to weigh the magnitude of the effect against the trustworthiness of the underlying evidence. GRADE is endorsed by Cochrane, the WHO, NICE (UK), and over 100 other organisations worldwide.

16.2.2 Finding the Studies for an Appraisal

Step 3 is where many reviews succeed or fail, and the same knowledge lets a reader judge whether a published review could have found the relevant studies. Most searching happens in bibliographic databases such as MEDLINE (searched free of charge through PubMed), Embase and CINAHL. A bibliographic database does not hold the articles themselves. It holds a record for each article (title, abstract, authors, journal, date and publication type), and in MEDLINE trained indexers add subject headings that describe what the article is about. A search runs against those records by two complementary routes.

  • Controlled vocabulary. MEDLINE's thesaurus is Medical Subject Headings (MeSH), and Embase uses its own, Emtree. A heading gathers the records on one concept under a single label, whatever words the authors used. Each heading lists entry terms, the synonyms that map to it, and sits in a tree of broader and narrower headings. PubMed explodes a heading by default, so a search on it also retrieves records indexed with any narrower heading.
  • Free text. Words searched in titles and abstracts (the [tiab] field in PubMed) find records that indexers have not yet reached, records indexed before a heading existed, and PubMed records that are never indexed.

The worked example uses the question from this section's reflection: in adults with type 2 diabetes (P), do SGLT2 inhibitors (I), compared with standard care or DPP-4 inhibitors (C), reduce cardiovascular mortality (O)? The first table shows one MeSH heading beside free-text synonyms for the same concept.

MeSH heading: Diabetes Mellitus, Type 2Free-text synonyms for the same concept
Entry terms (a selection): Type 2 Diabetes; Type 2 Diabetes Mellitus; Diabetes Mellitus, Non-Insulin-Dependent; NIDDM; Maturity-Onset Diabetes; Diabetes Mellitus, Adult-Onset"type 2 diabetes"[tiab]
"type II diabetes"[tiab]
T2DM[tiab]
NIDDM[tiab]
Place in the tree: broader heading Diabetes Mellitus; narrower heading Diabetes Mellitus, Lipoatrophic
Explosion: "Diabetes Mellitus, Type 2"[Mesh] also retrieves records indexed only with Diabetes Mellitus, Lipoatrophic

A search usually combines two or three concepts from the question. Here the population and the intervention are searched; the comparator and the outcome are left to screening, because studies describe them in too many ways for a search to capture them reliably.

Concept 1: type 2 diabetes (P)Concept 2: SGLT2 inhibitors (I)
MeSH heading"Diabetes Mellitus, Type 2"[Mesh]"Sodium-Glucose Transporter 2 Inhibitors"[Mesh] (introduced in 2019)
Free text"type 2 diabetes"[tiab], "type II diabetes"[tiab], T2DM[tiab], NIDDM[tiab]"SGLT2 inhibitor*"[tiab], "SGLT-2 inhibitor*"[tiab], gliflozin*[tiab], empagliflozin[tiab], dapagliflozin[tiab], canagliflozin[tiab], ertugliflozin[tiab]
Combine withOR within the columnOR within the column

Combining the terms. Boolean operators join the pieces of a search. OR joins the synonyms within one concept and widens the set; AND joins different concepts and keeps only the records that match both; NOT removes every record that matches a term. NOT is used with care, because it also removes relevant records that happen to match the excluded term. Quotation marks make a phrase search, so "type 2 diabetes" finds the words together and in that order, and an asterisk truncates a word stem, so gliflozin* finds gliflozin and gliflozins. Joining the two columns of the table gives the worked PubMed string.

Worked PubMed search string (Run 2 in the log below)
("Diabetes Mellitus, Type 2"[Mesh] OR "type 2 diabetes"[tiab] OR "type II diabetes"[tiab] OR T2DM[tiab] OR NIDDM[tiab]) AND ("Sodium-Glucose Transporter 2 Inhibitors"[Mesh] OR "SGLT2 inhibitor*"[tiab] OR "SGLT-2 inhibitor*"[tiab] OR gliflozin*[tiab] OR empagliflozin[tiab] OR dapagliflozin[tiab] OR canagliflozin[tiab] OR ertugliflozin[tiab])

Keeping a search log. A reproducible search is documented as it is built. The log below records three runs of the example search in PubMed, with the string, the date, the number of records and the reason for each change. The counts are those PubMed returned on 8 October 2026; the same strings return more records later, as new articles are added.

RunDateSearchRecordsReason for the change
18 Oct 2026"Diabetes Mellitus, Type 2"[Mesh] AND "Sodium-Glucose Transporter 2 Inhibitors"[Mesh]6,637Starting point: the MeSH heading for each concept.
28 Oct 2026The worked string above (MeSH OR free text within each concept, AND between concepts)11,453The SGLT2 heading dates from 2019, so free text was added to reach earlier records, records not yet indexed and records that are never indexed.
38 Oct 2026Run 2 NOT (animals[mh] NOT humans[mh])11,042Removes the 411 records indexed only as animal studies. A first attempt, Run 2 NOT animals[mh], returned only 2,242 records: the heading Animals contains Humans in its tree, so that string also removed the human studies.

A log like this lets anyone rerun the search and see why it changed, and it shows the searcher when a refinement has removed relevant records, as the first NOT attempt did. HSCI 241 Lessons 3 and 4, which can be taken before or after this course, teach systematic searching in full.

16.2.3 Appraising a Published Review

A systematic review can be well reported and still be unreliable, so a reader appraises it before relying on it. AMSTAR 2 (Shea et al., 2017) is a 16-item tool for appraising systematic reviews of randomised and non-randomised studies of healthcare interventions. Seven of its items are critical domains, because a weakness in any of them can undermine a review's conclusions. The overall confidence rating depends mainly on those domains: high (no weakness, or one non-critical weakness), moderate (more than one non-critical weakness and no critical flaw), low (one critical flaw, with or without non-critical weaknesses) and critically low (more than one critical flaw). AMSTAR 2 asks for this overall judgement in place of a summed score.

Worked appraisal: a review of SGLT2 inhibitors (fictional summary)

The summary below describes a fictional review, written for this example so that each domain can be shown; the same questions apply to any published review. The authors registered their protocol on PROSPERO before screening began. They searched MEDLINE, Embase and CENTRAL from inception with MeSH, Emtree and free-text terms, published the full strategies in an appendix, searched two trial registries, checked reference lists, consulted experts, set no language limits, and ran the final search ten months before publication. Their PRISMA flow diagram reports that 41 full texts were excluded, without a list of the studies or the reasons. Two reviewers assessed each of the nine included trials with RoB 2 and reported the judgements by domain; three trials were at high risk of bias from missing outcome data. The trials were pooled with a random-effects model, chosen because they differed in dose and length of follow-up, and heterogeneity was reported with I² and tau². The authors repeated the analysis without the three high-risk trials and discussed the difference. They presented a funnel plot and noted that, with nine trials, a test of asymmetry has little power to detect publication bias.

AMSTAR 2 critical domain (item)JudgementReason
Protocol registered before the review began (2)YesRegistered on PROSPERO before screening.
Adequacy of the literature search (4)YesThree databases with controlled vocabulary and free text, full strategies, registries, reference lists and experts, no language limits, and a recent search.
Justification for excluding individual studies (7)NoThe 41 excluded full texts are neither listed nor explained, so a reader cannot check whether relevant trials were dropped.
Risk of bias from individual studies included in the review (9)YesRoB 2 applied by two reviewers, with judgements reported by domain.
Appropriateness of meta-analytic methods (11)YesA random-effects model justified by differences between the trials, with heterogeneity quantified.
Consideration of risk of bias when interpreting the results (13)YesA sensitivity analysis without the three high-risk trials, discussed in the text.
Assessment of the presence and likely impact of publication bias (15)YesA funnel plot, with the limited power of an asymmetry test acknowledged.

Overall confidence: low. The review has one critical flaw (item 7), so under AMSTAR 2 it may not provide an accurate and comprehensive summary of the available studies, whatever the nine non-critical items show. Before relying on its pooled estimate, a reader would look for the list of excluded studies, or check the trial registries for trials the review might have dropped.

Reflection

Think of a public health question relevant to your interests. How would you specify the question for a systematic review? What databases would you search, and what inclusion/exclusion criteria would you set?

Model answerA strong response converts the topic into a PICO(S) question: Population (e.g., adults with type 2 diabetes), Intervention (SGLT2 inhibitors), Comparator (standard care or DPP-4 inhibitors), Outcome (cardiovascular mortality) and Study type (RCTs and large cohorts). A question about an exposure, such as air pollution and childhood asthma, would use PECO, with Exposure in the second position. Databases should include MEDLINE/PubMed, Embase, CENTRAL (Cochrane), CINAHL for nursing-relevant outcomes, and PsycINFO for behavioural exposures. Grey literature: ClinicalTrials.gov, WHO ICTRP, conference abstracts. Inclusion: study design, year range with justification, language with translation plan, outcome measured by accepted instrument. Exclusion: editorials, animal studies, single-case reports, irrelevant comparators. Document hand-searching of key journals and forward/backward citation tracking. The point is reproducibility: another team should be able to run your search string and arrive at the same hit list.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

1. Which of the following is a limitation of narrative reviews?

Narrative reviews tend to be subjective, with reviewers potentially bringing biased perspectives and selectively including studies that support their opinions. They lack the structured methodology of systematic reviews.

2. What is the first step in conducting a systematic review?

The first step is to specify a clear research question driven by a clinical or health-policy objective. This question guides all subsequent steps of the review, including the search strategy and inclusion criteria.

3. Why should data extraction in a systematic review be carried out by two independent investigators?

Duplicate independent data extraction minimises errors and subjective bias. The two datasets are then compared, and any differences are resolved by discussion, ensuring the accuracy and reliability of extracted data.

4. In a database search, which operator joins synonyms for a single concept, such as T2DM and "type 2 diabetes"?

OR joins synonyms within one concept and widens the set of records. AND joins different concepts and keeps only records that match both, and NOT removes records and is used with care because it can discard relevant studies.

5. Under AMSTAR 2, what overall confidence rating does a review receive if it registered no protocol before it began and did not assess the risk of bias of its included studies?

Both are critical domains (items 2 and 9). One critical flaw gives a rating of low, and more than one critical flaw gives a rating of critically low, which means the review should not be relied on to summarise the available studies.
Section 2 of 5

Meta-Analysis: Data Types & Effect Models

⏱ Estimated reading time: 50 minutes
Section 2 of 5

Meta-Analysis: Data Types & Effect Models

Combining effect estimates; choosing between fixed- and random-effects models.

Definition and purpose

What a meta-analysis does

Glass (1976): the statistical analysis of a large collection of results from individual studies for the purpose of integrating the findings.

Objective 1

Provide an overall estimate of an association or effect, pooled across all included studies.

Objective 2

Explore reasons for variation in the observed effect across studies.

Types of data

Summary, group, and individual patient data

Summary data

Point estimate (RR, OR, MD) plus SE or CI. Most common; extracted from published reports.

Group data

Cell values (2×2 table or group means and SDs). Allows computation of different effect measures.

Individual patient data

Raw outcome values per person. Most flexible for exploring heterogeneity; rarely available.

Fixed-effects model

One true effect for all studies

Fixed-effects model (Eq 28.1)
\[ \color{#0B7B6B}{T_i} = \color{#C2410C}{\theta} + \color{#6D28D9}{\varepsilon_i} \quad \text{where} \quad \color{#6D28D9}{\varepsilon_i} \sim N(0,\, \color{#1D4ED8}{V_i}) \]
Ti study effect θ common true effect εi within-study error Vi within-study variance

\(T_i\) is the observed effect in study \(i\); \(\theta\) is the single common true effect; \(V_i = [\text{SE}(T_i)]^2\) is the within-study variance. Study weights are \(W_i = 1/V_i\) (inverse variance weighting).

Limitation: Assumes one true effect across all populations and settings. Violated in most real-world pools.

Random-effects model

A distribution of true effects

Random-effects model (Eq 28.4)
\[ \color{#0B7B6B}{T_i} = \color{#C2410C}{\theta} + \color{#BE185D}{u_i} + \color{#6D28D9}{\varepsilon_i} \quad \text{where} \quad \color{#BE185D}{u_i} \sim N(0,\,\color{#047857}{\tau^2}),\; \color{#6D28D9}{\varepsilon_i} \sim N(0,\, \color{#1D4ED8}{V_i}) \]
Ti study effect θ average true effect ui study deviation εi within-study error τ² between-study variance Vi within-study variance

\(\tau^2\) is the between-study variance. Weights become \(W_i = 1/(V_i + \tau^2)\), giving smaller studies more weight than under fixed effects.

How study weights redistribute from a fixed-effects to a random-effects model: large studies lose share and small studies gain it. Fixed effects Random effects Large study 52% Medium 30% 12% 6% Large study 36% Medium 29% 21% 14% Adding between-study variance τ² to every denominator shrinks the gap between large and small studies, so small studies pull the pooled estimate more under random effects. Illustrative weights.

Produces a wider confidence interval that accounts for true between-study variation. DerSimonian & Laird (1986) is the foundational estimator for \(\tau^2\).

Carry forward

Choosing between the models

Use fixed effects when…

You have strong grounds to believe all studies estimate one exchangeable true effect. Rare in practice.

Use random effects when…

Studies differ in population, intervention, or setting. The default for most real-world reviews.

For sparse binary data: Mantel-Haenszel or Peto weighting. For different outcome scales: standardised mean differences (Cohen’s d, Hedges’ g).

Introduction and Overview

An earlier section covered how to identify and appraise the studies that go into a review. This section turns to the quantitative half: combining the effect estimates from those studies into a single pooled estimate. The central design choice here (fixed-effects versus random-effects) is not arbitrary; it reflects an underlying assumption about whether all the studies are estimating the same true effect.

Learning Objectives

  • Define meta-analysis and state its two principal objectives (pooled estimate; explore variation).
  • Distinguish summary, group, and individual-patient data, and describe the tradeoffs of each.
  • Compare the assumptions of fixed-effects and random-effects models and state when each is appropriate.
  • Interpret pooled point estimates, confidence intervals, and study weights produced by either model.

16.3 What Is a Meta-Analysis?

A meta-analysis is “the statistical analysis of a large collection of analysis results from individual studies for the purpose of integrating the findings” (Glass, 1976). It is a formal process for combining results from multiple studies and is considered the “gold standard” for providing summary information about health interventions.

Objectives of Meta-Analysis

The objectives are to: (1) provide an overall estimate of an association or effect based on data from multiple studies, and (2) explore reasons for variation in the observed effect across studies. Because it combines data from multiple studies, meta-analysis gains statistical power for detecting effects.

16.3.1 Types of Data in Meta-Analysis

Three types of data can be used in a meta-analysis, each with different capabilities:

Data TypeBinary OutcomeContinuous Outcome
Summary estimatePoint estimate: RR, OR, RD, IR
Precision: SE or CI
Point estimate: mean difference (MD)
Precision: SE or CI
Group dataCell values for treated and control groups (2×2 table)Number, mean, and SD in each group
Individual patient data (IPD)Raw data: outcome value (0 or 1) and individual characteristicsRaw data: outcome value and individual characteristics

Summary data are most commonly used. Group data allow computation of various effect measures. IPD are the most flexible but rarely available; they allow evaluation of study-, group-, and individual-level variables as sources of heterogeneity.

16.4 Fixed- vs. Random-Effects Models

A fundamental decision in any meta-analysis is whether to use a fixed-effects or random-effects model:

Fixed-Effects Model

Assumes the true treatment effect is constant across all studies. Any variation among observed study results is due solely to within-study random variation (sampling error).

Fixed-effects model (Eq 28.1)
\[ \color{#0B7B6B}{T_i} = \color{#C2410C}{\theta} + \color{#6D28D9}{\varepsilon_i} \quad\text{where}\quad \color{#6D28D9}{\varepsilon_i} \sim N(0,\, \color{#1D4ED8}{V_i}) \]
The observed effect in a study equals the single true effect shared by all studies plus within-study error, where that error has variance equal to the within-study variance.

Where Ti is the observed effect from study i, θ is the true overall effect, and Vi = [SE(Ti)]² is the known within-study variance. Weights are computed as Wi = 1/Vi (inverse variance weighting).

Advantage: Does not require estimating between-study variance (τ²).

Limitation: The assumption of a constant effect across all studies is often untenable, and ignoring between-study variation can lead to Type I errors and confidence intervals that are too narrow.

Random-Effects Model

Assumes a distribution of true treatment effects across studies (heterogeneity), with additional variability beyond within-study sampling error.

Random-effects model (Eq 28.4)
\[ \color{#0B7B6B}{T_i} = \color{#C2410C}{\theta} + \color{#BE185D}{u_i} + \color{#6D28D9}{\varepsilon_i} \quad\text{where}\quad \color{#BE185D}{u_i} \sim N(0,\,\color{#047857}{\tau^2}),\; \color{#6D28D9}{\varepsilon_i} \sim N(0,\,\color{#1D4ED8}{V_i}) \]
The observed effect in a study equals an average true effect plus a study-specific deviation plus within-study error. The deviations vary with between-study variance, and the error with within-study variance.

Where ui is the random effect for study i, and τ² is the between-study variance (heterogeneity). Weights become Wi = 1/(Vi + τ²).

Result: Produces a similar point estimate to fixed-effects but with a wider confidence interval (because it accounts for between-study variation). Random-effects models are now more commonly used; the foundational estimator is from DerSimonian & Laird (1986).

Key Distinction

The fixed-effects model asks: “What is the single true effect?” The random-effects model asks: “What is the average of the distribution of true effects?” Random-effects models are generally preferred because the assumption of a constant treatment effect across all studies is rarely justified.

16.4.1 Weighting Methods

The most common weighting procedure is inverse variance weighting, applicable to both continuous and binary outcomes. The name states the recipe: each study's weight is one divided by its variance, so a precise estimate (small variance) earns a large weight and a noisy one earns little, which is why large, well-measured studies dominate the pooled result. For binary outcomes with sparse data, the Mantel-Haenszel procedure (Mantel & Haenszel, 1959) or the Peto method may be preferred. For continuous outcomes, when studies use different scales, standardised mean differences (effect sizes such as Cohen’s d or Hedges’ g) are used.

Reflection

Consider a meta-analysis of 10 studies examining the effect of a drug on blood pressure. Five studies were conducted in elderly populations and five in young adults. Would you expect a fixed-effects or random-effects model to be more appropriate? Why?

Model answerRandom effects is more appropriate. The two sub-populations (elderly vs. young adults) almost certainly differ in baseline BP, comorbidity, polypharmacy, and dose-response, so it is implausible that all 10 studies are estimating one common effect. A fixed-effects model assumes a single true effect and weights only by within-study precision; with biologically distinct subgroups that assumption is wrong, the FE pooled estimate becomes a population-weighted average that doesn't describe either group, and its CI is artificially narrow. A random-effects pool reflects an average true effect plus between-study variance τ². Better still: pre-specify a subgroup analysis (or meta-regression on age) so the heterogeneity is investigated rather than absorbed into a wider CI.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

1. In a fixed-effects model, what is assumed about the true treatment effect?

The fixed-effects model assumes there is one true effect (θ) common to all studies. Any observed variation in study results is attributed to within-study random variation (sampling error) only.

2. What does τ² represent in a random-effects meta-analysis?

τ² represents the between-study variance, the variability in true treatment effects across studies. It quantifies heterogeneity beyond what would be expected from within-study sampling error alone.

3. Which type of data provides the most flexibility for exploring sources of heterogeneity in a meta-analysis?

IPD allow evaluation of study-, group-, and individual-level variables as sources of heterogeneity. Summary data can only evaluate study-level variables, while group data add some flexibility but not at the individual level.
Section 3 of 5

Forest Plots & Heterogeneity

⏱ Estimated reading time: 50 minutes
Section 3 of 5

Forest Plots & Heterogeneity

Reading the key visual output; measuring and explaining variability across studies.

Anatomy of a forest plot

Reading every element

Study A Study B Study C Study D Pooled Null Pooled est. Box area = weight   Line = 95% CI
Cochran's Q

Testing for more variation than chance predicts

Cochran’s Q statistic (Eq 28.7)
\[ \color{#0B7B6B}{Q} = \sum_i \color{#C2410C}{w_i} \left(\color{#6D28D9}{T_i} - \color{#1D4ED8}{\hat{\theta}}\right)^2 \]
Q heterogeneity statistic wi study weight Ti study effect θ̂ pooled estimate

Under no heterogeneity, \(Q \sim \chi^2_{k-1}\). Low power when the number of studies is small. Use a relaxed threshold of \(p = 0.10\) rather than \(0.05\) to avoid missing real heterogeneity.

Higgins I² and τ²

How much of the spread is real?

Higgins I² (Eq 28.8)
\[ \color{#6D28D9}{I^2} = \frac{\color{#0B7B6B}{Q} - (\color{#1D4ED8}{k}-1)}{\color{#0B7B6B}{Q}} \times 100\% \]
I² heterogeneity index Q Cochran’s Q k number of studies

Benchmarks: 25% = low  |  50% = moderate  |  75% = high. Any value above 25% warrants investigation of causes.

\(\tau^2\) is the between-study variance on the same scale as the effect measure, useful for judging practical magnitude.

Explaining heterogeneity

Four diagnostic approaches

Subgroup & stratified analysis

Compare effects across defined study categories. Pre-specify in the protocol to control Type I error.

Galbraith plot

Plots Z-statistic vs. 1/SE. Points outside ±2 units are potential outliers driving heterogeneity.

Meta-regression

Weighted regression of effects on study-level predictors. Most flexible, but still observational.

Ecological caution

Predictors are study-level averages. Multiple comparisons inflate Type I error. Pre-specify analyses.

Carry forward

What to take into the next section

  • The forest plot makes the consistency of the evidence visible at a glance.
  • \(I^2\) above 25% signals real heterogeneity worth explaining, not just absorbing into a wider CI.
  • Subgroup analysis and meta-regression investigate causes; pre-specify them to avoid inflated error rates.
  • A pooled estimate is only as trustworthy as the studies behind it, and the variation among them.

Introduction and Overview

An earlier section produced a single pooled estimate. This section introduces the visual and quantitative tools for inspecting the underlying variation that produced that estimate. The forest plot makes the constituent studies and the pooled result visible at a glance; heterogeneity statistics quantify how much the studies actually disagree with one another. Both are essential before trusting the pooled number.

Learning Objectives

  • Read a forest plot: identify point estimates, confidence intervals, study weights, the null line, and the pooled diamond.
  • Quantify heterogeneity using Cochran's Q, I-squared, and tau-squared, and interpret the conventional thresholds.
  • Distinguish statistical heterogeneity from clinical and methodological heterogeneity.
  • Decide when subgroup analysis or meta-regression is the appropriate response to detected heterogeneity.

16.5 Presentation of Results: The Forest Plot

The forest plot is the most important graphical output of a meta-analysis. It displays the point estimate and confidence interval of the effect observed in each study, along with the summary estimate.

Anatomy of a Forest Plot

Each horizontal line represents one study’s results. The length of the line is the 95% CI. The centre box marks the point estimate, and the area of the box is proportional to the study’s weight. The dashed vertical line shows the overall summary estimate. The diamond at the bottom represents the pooled estimate and its CI. The solid vertical line marks the null value (e.g., 0 for mean difference, 1 for ratio measures).

▸ INTERACTIVE STORY, THE FOREST & THE FUNNEL Open full screen ↗

Watch a forest plot build itself one study at a time, see the PRISMA flow filter thousands of records down to twelve, and watch the pooled diamond emerge. Next ▶ advances scenes.

A 6-scene visualization of the systematic-review pipeline: PICO question, PRISMA filter funnel, forest plot construction one study at a time, heterogeneity check, pooled diamond, and the evidence pyramid.

Reading a Forest Plot

If all study CIs overlap considerably and cluster near the summary estimate, there is little heterogeneity. If CIs are widely scattered and many do not overlap, heterogeneity is substantial. Studies may be ordered by publication year (to detect time trends), quality score, or effect size.

16.6 Heterogeneity

Heterogeneity refers to variability among study results beyond what would be expected from random variation alone. It should always be evaluated in a meta-analysis.

16.6.1 Real vs. Artifactual Heterogeneity

Real HeterogeneityClick to explore
Artifactual HeterogeneityClick to explore

An important distinction is between clinical heterogeneity (real differences between populations, interventions, and settings) and statistical heterogeneity (variation in observed results beyond chance). Clinical heterogeneity is always expected; the key question is whether statistical heterogeneity is also present.

16.6.2 Measuring Heterogeneity: Cochran’s Q and Higgins I²

Cochran’s Q statistic (Eq 28.7)
\[ \color{#0B7B6B}{Q} = \sum_i \color{#C2410C}{w_i}\left(\color{#6D28D9}{T_i} - \color{#1D4ED8}{\hat{\theta}}\right)^2 \]
The heterogeneity statistic is a weighted sum of squared distances between each study effect and the pooled estimate. Larger values indicate more disagreement among studies than chance alone would produce.

Where wi are the study weights, Ti are the study effects, and θ is the pooled estimate. Under the null hypothesis of no heterogeneity, Q follows a χ² distribution with k−1 degrees of freedom. However, the Q test has low power when the number of studies is small, so a non-significant result does not rule out heterogeneity. Consider using a relaxed P-value threshold (e.g., 0.10 instead of 0.05).

Higgins I² (Eq 28.8)
\[ \color{#6D28D9}{I^2} = \frac{\color{#0B7B6B}{Q} - (\color{#1D4ED8}{k}-1)}{\color{#0B7B6B}{Q}} \times 100\% \]
The heterogeneity index rescales the Q statistic by its degrees of freedom (the number of studies minus one) into the percentage of total variation that reflects real between-study differences rather than chance.

I² (Higgins & Thompson, 2002) quantifies the proportion of variance between studies that is due to heterogeneity rather than chance. Benchmarks: 25% = low, 50% = moderate, 75% = high heterogeneity. An evaluation of possible causes should be undertaken whenever I² exceeds 25%.

The formula is more intuitive than it first looks. If every study were really estimating the same effect, Cochran's Q would on average equal its degrees of freedom, the number of studies minus one. Subtracting that quantity from Q removes the disagreement expected from chance alone, and dividing by Q re-expresses what remains as a fraction of the total spread. When the studies agree no more than chance predicts, that fraction falls to zero and I² is reported as 0%.

16.6.3 Evaluating Causes of Heterogeneity

Subgroup Analysis

Identify a specific subgroup of studies defined by a characteristic of interest and examine the effect within that subgroup. However, results should be interpreted with caution, the best estimate for any subgroup is provided by considering all the evidence (Stein’s Paradox) rather than the subgroup data alone. Subgroup analyses should be pre-specified in the review protocol.

Stratified Analysis

Data are stratified by a factor thought to influence the treatment effect, and a separate meta-analysis is carried out in each stratum. The between-strata heterogeneity can be tested using QB = QT − ΣQS. A disadvantage is that individual strata may contain few studies.

Galbraith Plot

A Galbraith plot plots the Z statistic (Ti/SE(Ti)) against the inverse of the SE (1/SE). The slope of the resulting line is the overall fixed-effect estimate, and lines at ±2 units from this line should encompass 95% of observations if there is no significant heterogeneity. Points outside these bounds are potential outliers contributing to heterogeneity.

Meta-Regression

Meta-regression is the most flexible approach: a weighted regression of observed treatment effects against study-level predictors (with inverse variance weights). It extends the random-effects model by adding predictors. Cautions: (1) even with RCTs, meta-regression is observational; (2) multiple comparisons inflate Type I error; (3) ecological fallacy applies since predictors are study-level averages.

Reflection

You conduct a meta-analysis of 20 studies and find I² = 82%. The forest plot shows widely scattered effect sizes. What steps would you take to investigate the causes of this high heterogeneity? Which methods from this section would you prioritise and why?

Model answerWith I² = 82% a single pooled estimate is uninformative; the priority is to explain the heterogeneity, not bury it. Steps: (a) re-examine the forest plot for outliers and clusters; (b) pre-specified subgroup analyses on population (age, sex, severity), intervention (dose, duration, formulation), comparator (active vs. placebo), and methodological quality (low vs. high risk of bias); (c) meta-regression on continuous covariates (mean age, baseline severity, year, follow-up length); (d) leave-one-out sensitivity; (e) re-extract outcomes to check for definitional drift (e.g., "response" defined differently across trials). Prioritise meta-regression on substantive moderators because it both quantifies and explains, and prioritise risk-of-bias subgrouping because high-RoB studies often drive heterogeneity for non-substantive reasons. If no moderator explains it, report each subgroup separately rather than forcing a single pooled number.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

1. In a forest plot, what does the area of the box on each study line represent?

In a forest plot, the area of the box is proportional to the weight assigned to that study in the meta-analysis. Studies with more precise estimates (smaller SEs) receive larger weights and thus larger boxes.

2. A meta-analysis reports I² = 75%. How should this be interpreted?

I² = 75% means that 75% of the observed variance between study results is attributable to real heterogeneity rather than sampling error. This is considered “high” heterogeneity, and an investigation of its causes is warranted.

3. What is meta-regression used for in a meta-analysis?

Meta-regression is a weighted regression of observed treatment effects against study-level predictors. It is the most flexible approach for evaluating whether specific study characteristics (e.g., study design, population, intervention type) explain heterogeneity.
Section 4 of 5

Publication Bias, Influential Studies & Data Issues

⏱ Estimated reading time: 50 minutes
Section 4 of 5

Publication Bias, Influential Studies & Data Issues

Three threats that survive correct technique.

Publication bias

When the pool is a biased sample

Studies with significant or favourable results are more likely to be published. A pool that includes only published studies will overestimate the true effect.

The mechanism

Null results sit in file drawers. Positive results reach journals. The published record is a skewed sample of all research conducted.

The defence

Search grey literature and trial registries upstream. Use funnel plots and Egger's test downstream.

Funnel plots

Detecting asymmetry in the evidence pool

gap here = missing studies? Small SE Large SE Pooled effect

Begg's test: rank correlation. Egger's test: regression. Both have low sensitivity below 20 studies.

Trim-and-fill

Adjusting for publication bias (Duval & Tweedie, 2000)

Step 1: Trim

Remove the most extreme studies on the over-represented side until the funnel is roughly symmetric. Estimate a new pooled effect from the trimmed set.

Step 2: Fill

Replace removed studies and add hypothetical mirror counterparts on the sparse side. Re-estimate the pooled effect.

The difference between original and adjusted estimates quantifies potential impact. Treat it as a sensitivity analysis, not a correction.

Influential studies

Leave-one-out sensitivity analysis

Repeat the meta-analysis \(k\) times, each time omitting one study. Watch how the pooled estimate and \(I^2\) shift.

Worked example from the lesson

25-study pool  •  pooled estimate = −2.121  •  \(I^2\) = 95.6%
Remove refid 218: estimate → −2.011 (5% shift)  •  \(I^2\) → 88.1%
That single study had a documentable and meaningful influence on the result.

Outcome-scale issues

When pooling is inappropriate

Different scales

Continuous outcomes on different instruments require standardised mean differences (Cohen’s d, Hedges’ g), not raw mean differences.

Mixed outcome types

Binary, continuous, and time-to-event data require conversion to a compatible effect measure before pooling.

Duplicate publication

The same patients appearing in multiple papers inflate apparent study count. Check for overlapping cohorts.

Carry forward

What to take into the final assessment

  • Publication bias shifts pooled estimates away from the null. Funnel plots and trim-and-fill are diagnostics, not cures.
  • Leave-one-out analysis identifies whether a single study drives your conclusions.
  • Outcome-scale mismatches make pooling inappropriate regardless of model choice.
  • All three threats are manageable if the protocol anticipates them before extraction begins.

Introduction and Overview

Earlier sections produced and inspected a pooled estimate. This section addresses three threats that can survive correct technique: publication bias (some studies never make it to the literature), influential studies (a single trial driving the entire pooled result), and outcome-scale issues that make it inappropriate to pool studies that look superficially similar. Each of these is a place where a meta-analysis can produce a precise but misleading answer.

Learning Objectives

  • Define publication bias and explain why it biases pooled estimates away from the null.
  • Read a funnel plot, apply Begg's and Egger's tests, and interpret the trim-and-fill adjustment.
  • Identify influential studies via leave-one-out sensitivity analysis and decide how to report their effect.
  • Recognize when outcome-scale issues (binary vs continuous, time-to-event, dose-response) make pooling inappropriate.

16.7 Publication Bias

A critical concern in meta-analysis is publication bias, studies with statistically significant or favourable results are more likely to be published than those with null or unfavourable results. Consequently, published studies may represent a biased subset of all work conducted on a topic.

Why Publication Bias Matters

If the meta-analysis only includes published studies, and published studies tend to overestimate the effect, the summary estimate will be biased away from the null. This can lead to erroneous conclusions about the effectiveness of interventions.

16.7.1 Detecting Publication Bias: The Funnel Plot

A funnel plot displays each study’s SE (or its inverse, 1/SE) plotted against its estimated effect. In the absence of publication bias, the plot should resemble an inverted funnel, symmetric around the summary estimate, with small studies (large SEs) scattered widely at the bottom and large studies (small SEs) clustered near the top.

Interpreting a Funnel Plot

Asymmetry in the funnel plot suggests publication bias. For example, if studies with large effects and large SEs are present, but studies with small or null effects and large SEs are missing (a “gap” on one side), this suggests that null-result studies were not published. However, asymmetry can also arise from other factors, so interpretation should be cautious.

16.7.2 Statistical Tests for Publication Bias

Two commonly used tests evaluate the relationship between study results and their precision:

  • Begg’s test: A rank correlation between effect estimates and their SEs. Simple but low power with few studies.
  • Egger’s test: A linear regression approach that is generally more sensitive at detecting publication bias (Egger et al., 1997).

Neither test is very sensitive when the number of studies is small (<20), and both may produce false positives when there are large treatment effects, few events per trial, or all trials are of similar size. For comprehensive guidance on interpreting funnel-plot asymmetry, see Sterne et al. (2011).

16.7.3 Trim-and-Fill Method

How Trim-and-Fill Works

The trim-and-fill method (Duval and Tweedie, 2000) is a practical approach to assessing and adjusting for publication bias:

  1. “Trim”: Produce a funnel plot and sequentially omit the most extreme studies on one side until the plot is approximately symmetrical.
  2. Determine the centre of the trimmed, symmetrical plot (a new estimate of the treatment effect).
  3. “Fill”: Replace the omitted studies along with their hypothetical “counterparts” on the other side of the centre line.
  4. Redo the meta-analysis including both the original data and the hypothetical studies.

This provides an estimate of what the treatment effect would be if all studies had been published. The difference between the original and adjusted estimates indicates the potential impact of publication bias.

16.8 Influential Studies

It is important to determine whether individual studies have a profound influence on the summary estimate. A study might be much larger than others or have an extreme effect size. To evaluate this, sequentially delete each study from the meta-analysis and observe how the summary estimate changes.

Example: Sensitivity Analysis

In a meta-analysis of 25 studies with a pooled estimate of −2.121, one study (refid 218) was identified as a potential outlier in the Galbraith plot. Removing it changed the estimate to −2.011 (a 5% reduction in magnitude) and reduced I² from 95.6% to 88.1%. While the heterogeneity remained high, the analysis demonstrated that this single study had a meaningful influence on the results.

16.9 Outcome Scales and Data Issues

Published studies vary substantially in how they present data. Several practical issues arise:

Computing SEsClick to explore
Different ScalesClick to explore
Combining OutcomesClick to explore

Reflection

You produce a funnel plot for your meta-analysis and notice asymmetry; there appear to be “missing” studies with small effects and large standard errors. What are the possible explanations for this pattern beyond publication bias? How would you investigate further?

Model answerFunnel asymmetry has several explanations beyond publication bias. (1) Small-study effects: small trials are often run in more selected populations, with greater treatment fidelity and shorter follow-up; they really do produce larger effects, independent of selective publication. (2) Methodological quality differences: smaller studies frequently have higher risk of bias (poor randomisation, less blinding), which inflates effects. (3) True heterogeneity: if effects differ by subgroup and smaller studies happen to be drawn from high-effect subgroups, you get asymmetry without selection. (4) Chance, especially with fewer than 10 trials. Investigation: Egger’s test or its variants for formal asymmetry, trim-and-fill as sensitivity, search trial registries and grey literature for unpublished work, and stratify the funnel plot by risk-of-bias categories. Conclude carefully: asymmetry is a flag, not a diagnosis.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

1. What pattern in a funnel plot suggests publication bias?

An asymmetric funnel plot, where studies with certain characteristics (e.g., small effects and large SEs) appear to be “missing,” suggests publication bias. However, asymmetry can also be caused by other factors such as heterogeneity.

2. What is the purpose of the trim-and-fill method?

The trim-and-fill method creates a symmetrical funnel plot by adding hypothetical “missing” studies, then re-estimates the pooled effect. This provides an adjusted estimate showing what the result might be if all studies had been published.

3. Why is sensitivity analysis (sequentially removing studies) important in meta-analysis?

Sensitivity analysis identifies studies that have a disproportionate influence on the summary estimate. If removing one study substantially changes the result, it warrants careful evaluation of that study’s quality and characteristics.
Section 5 of 5

Final Assessment

📝 15 questions • 100% required to pass

Bringing It All Together

This lesson introduced evidence synthesis: the structured process by which the literature on a question is identified, appraised, and pooled into a single defensible answer. The arc moved from the seven steps of a systematic review, through the central design choice between fixed- and random-effects models, to the diagnostic tools (forest plots, heterogeneity statistics, funnel plots) that show whether the pooled answer should be trusted, and finally to the threats (publication bias, influential studies, outcome-scale issues) that survive correct technique.

The final assessment below asks you to integrate across all four sections: distinguishing systematic from narrative reviews, choosing between fixed- and random-effects, reading a forest plot, computing and interpreting heterogeneity, and recognizing when a precise pooled estimate is nonetheless misleading.

The same checks apply, in miniature, whenever you read a published review: explicit inclusion criteria, a structured and documented search, transparent appraisal, and an honest accounting of what the existing literature can and cannot answer. The search subsection and the AMSTAR 2 worked appraisal in Section 1 give you the tools for that reading.

Key Takeaways from this lesson

  • Systematic reviews are reproducible; they specify the question, search strategy, inclusion criteria, appraisal, and synthesis transparently, in contrast with narrative reviews that depend on the reviewer's preconceptions.
  • Meta-analysis is the quantitative half: a pooled effect estimate, gained statistical power, and a structured way to explore variation across studies.
  • Fixed- vs random-effects is a substantive choice, fixed-effects assumes one true effect, random-effects allows the true effect to vary across studies. Random-effects is the default when there is meaningful clinical or methodological heterogeneity.
  • The forest plot makes the pooled story visible at a glance, and I-squared, Cochran's Q, and tau-squared quantify how much the studies actually disagree.
  • Publication bias, influential studies, and outcome-scale mismatch can produce a precise but misleading pooled estimate, funnel plots, leave-one-out diagnostics, and pre-specified scale checks are essential safeguards.
  • PRISMA is the reporting standard for systematic reviews and meta-analyses; trustworthy reviews follow it explicitly.
  • A reader can check a review's search (MeSH headings, free text, Boolean logic and a dated search log) and appraise the review as a whole against the seven AMSTAR 2 critical domains.

This assessment covers all sections of this lesson. You must answer all 15 questions correctly to complete the lesson. Review the feedback after each attempt.

Reflection

Reflect on the entire process of systematic review and meta-analysis. What do you see as the greatest strength of this approach compared to narrative reviews? What is the most challenging aspect, and how would you address it in your own research?

Model answerThe greatest strength is transparency-enforced reproducibility: a systematic review's protocol, search string, inclusion criteria, extraction template, and risk-of-bias tool make every choice auditable, so disagreement between two reviewers reduces to a disagreement about specific decisions rather than about the bottom line. Meta-analysis then converts that auditable evidence base into a quantitative summary with formal uncertainty, something a narrative review can only describe. The hardest part is the inputs: poor-quality primary studies make even the cleanest synthesis misleading ("garbage in, garbage out"), and heterogeneity is often substantive, not statistical, so a clean I² doesn’t guarantee a clean inference. In my own research I would address this by pre-registering on PROSPERO, applying risk-of-bias tools (RoB 2 for trials, ROBINS-I for observational studies), reporting both fixed and random-effects pools as sensitivity, and committing in advance to not pool when conceptual heterogeneity is too high.

Minimum 20 characters required.

✓ Reflection saved

Final Knowledge Assessment

Final Assessment, the 15 Questions

1. What is the primary purpose of a systematic review?

A systematic review uses a structured, reproducible methodology to comprehensively identify, evaluate, and synthesise all relevant studies on a specific question, minimising bias compared to narrative reviews.

2. Which step of a systematic review involves specifying inclusion and exclusion criteria?

Step 4 of a systematic review involves determining which studies are relevant by specifying inclusion criteria (intervention, population, outcome, study type) and exclusion criteria (language, date, accessibility).

3. In inverse variance weighting, how are weights computed?

Inverse variance weighting assigns Wi = 1/Vi, where Vi = [SE(Ti)]². Studies with more precise estimates (smaller variance) receive larger weights, contributing more to the pooled estimate.

4. What key assumption distinguishes a random-effects model from a fixed-effects model?

The random-effects model assumes a distribution of true effects across studies (heterogeneity), with between-study variance τ². The fixed-effects model assumes one constant true effect with variation only from sampling error.

5. In a forest plot, what does the diamond at the bottom represent?

The diamond at the bottom of a forest plot represents the pooled (summary) estimate. The centre of the diamond is the point estimate, and the width shows the confidence interval.

6. Cochran’s Q statistic tests the null hypothesis that:

Cochran’s Q tests the null hypothesis of no heterogeneity (τ² = 0). Under this assumption, Q follows a χ² distribution with k−1 degrees of freedom. A significant Q indicates evidence of heterogeneity.

7. An I² value of 50% indicates:

I² = 50% means that half the observed variance between study results is due to real heterogeneity rather than sampling error. This is considered “moderate” heterogeneity.

8. Why might the Q test fail to detect heterogeneity even when it exists?

The Q test has relatively low power for detecting heterogeneity when the number of studies is small or when study sizes or SEs vary considerably. A non-significant Q should not be taken as proof of homogeneity.

9. What does an asymmetric funnel plot typically suggest?

Asymmetry in a funnel plot typically suggests publication bias, though it can also result from other factors such as heterogeneity, differences in study quality, or true small-study effects. Interpretation should be cautious.

10. The trim-and-fill method is used to:

Trim-and-fill creates a symmetric funnel plot by adding hypothetical “missing” studies, then re-estimates the pooled effect including these studies. It provides an adjusted estimate that accounts for potential publication bias.

11. In a random-effects model, study weights are computed as:

In a random-effects model, weights incorporate both within-study variance (Vi) and between-study variance (τ²): Wi = 1/(Vi + τ²). This gives relatively more equal weight to all studies compared to fixed-effects models.

12. Which effect size measure expresses the treatment effect relative to within-study variability?

Standardised mean differences (e.g., Cohen’s d, Hedges’ g) express the treatment effect as the mean difference divided by the pooled SD, making it possible to combine studies that measured outcomes on different scales.

13. Why is the choice of effect measure (RR vs. OR vs. RD) important in meta-analysis?

The same data can show no heterogeneity when assessed as RRs but substantial heterogeneity when assessed as ORs or RDs. In general, ratio measures tend to be more stable across studies than difference measures.

14. What is Stein’s Paradox in the context of subgroup analysis?

Stein’s Paradox states that the best estimate of the effect in any particular subgroup is obtained by considering all of the available evidence rather than the data from that specific subgroup alone. This counterintuitive result has important implications for interpreting subgroup analyses.

15. A sensitivity analysis reveals that removing one study changes the pooled estimate from RR = 0.65 to RR = 0.89. What should you conclude?

A large shift in the pooled estimate when one study is removed indicates that study is highly influential. The study should be carefully evaluated for quality, but should not be automatically excluded; it may simply be a larger or more precise study.
✦ Complete every section knowledge check and every reflection before submitting.

Congratulations!

You have successfully completed this lesson: Systematic Reviews and Meta-Analysis. You now understand the steps of a systematic review, how to conduct and interpret a meta-analysis using fixed- and random-effects models, how to assess heterogeneity and publication bias, and how to evaluate the influence of individual studies on pooled estimates.

Earlier lessons set the foundations of this course: the discipline's history, its ways of knowing, its trust scaffolding, and the structured process by which evidence is synthesized across studies. The remaining lessons take you back down to the level of individual studies, the designs that produce the evidence that systematic reviews and meta-analyses pool. A later lesson (Introduction to Observational Studies) sets up the design vocabulary you will use to appraise individual case-control, cohort, and ecological studies in the lessons that follow.