HSCI 241 · Lesson 9

Synthesizing Quantitative Findings

Finding & Synthesizing Health Evidence

Learning objectives for this lesson:

  • Explain what synthesis is and build characteristics tables, cross-tabulations and results tables for the studies included in a review.
  • Define groups for synthesis by population, intervention, comparator, outcome and design, and write decision rules that select one result per study.
  • Describe the three synthesis methods that the Cochrane Handbook accepts when effects are not pooled, and summarize effect estimates with the median, interquartile range and range.
  • Describe the nine items of the SWiM reporting guideline and use them to plan and report a synthesis without meta-analysis.
  • Explain, with reference to statistical power, why vote counting by statistical significance misleads, and carry out vote counting based on the direction of effect with a sign test.
  • Read and construct effect direction plots and harvest plots from tabulated study results.
  • Decide whether pooling is justified for a group of studies by asking five questions, and identify what HSCI 230 Lesson 2 teaches about meta-analytic methods.
  • Plan the synthesis for a rapid scoping review, including its groups, decision rules, results table or cross-tabulation, SWiM-based methods paragraph and pooling decision.

This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University, drawing on the Cochrane Handbook for Systematic Reviews of Interventions and the JBI Manual for Evidence Synthesis.

Lesson 9 · HSCI 241

Synthesizing Quantitative Findings

This lesson moves from tables of extracted results to clear statements about what a body of evidence shows.

About three hours, including the final assessment
Running case

The fictional Cedar Valley evidence review

42included studies
36studies with quantitative results
24controlled studies
6qualitative studies, synthesized in Lesson 10

All study results in this lesson are illustrative and belong to the fictional case.

Overview

Four sections

1. Tables and groups

Make the studies comparable before interpreting them.

2. SWiM

Synthesize and report effects without pooling them.

3. Beyond vote counting

Count by direction and display results with plots.

4. When to pool

Decide whether a meta-analysis is justified.

Where this fits

Synthesis here, computation in HSCI 230

HSCI 230 Lesson 2

The course teaches meta-analytic models, forest plots, heterogeneity and publication bias.

HSCI 241 Lesson 9

This lesson teaches tabulation, synthesis without pooling and the decision to pool.

Lesson goals

What you will be able to do

  • Define groups for synthesis and cross-tabulate the included studies.
  • Write decision rules and build a results table or cross-tabulation for a main outcome.
  • Write a synthesis methods paragraph that addresses the SWiM items.
  • Decide whether pooling would be justified for a main outcome.
Section 1 of 5

Tabulating and Grouping Studies

⏱ Estimated reading time: 35 minutes
Section 1 of 5

Tabulating and Grouping Studies

This section covers the tables, groups and decision rules on which every synthesis rests.

35 minutes
The route

From extraction forms to synthesis

1. Tabulate

Describe and cross-tabulate the studies.

2. Group

Define groups by comparison, outcome and design.

3. Results

Tabulate one result per study in each group.

4. Synthesize

Summarize, count and display, or pool.

The 36 studies with quantitative results follow this route; qualitative findings go to Lesson 10.

Three tables

Each table answers a different question

Characteristics table

What was studied, in whom, and how?

Cross-tabulation

Where is the evidence, and where are the gaps?

Results table

What did the studies in this group find?

Groups for synthesis

Lumping and splitting

Lumping

Broad groups hold more studies but may combine programs that work differently.

Splitting

Narrow groups give specific answers but may hold only one or two studies.

Groups are planned in the protocol, and any later change is reported with its reason.

Cross-tabulation

Where the Cedar Valley evidence sits

12community connector studies, none randomized
14randomized trials, all of group or one-to-one programs
5controlled studies of intergenerational and technology programs
Decision rules

One result per study, chosen in advance

  • A ranked list of scales decides which loneliness measure is used.
  • The time point closest to the end of the program is used first.
  • Trials contribute their analysis of all randomized participants.
  • Every effect is aligned so that a negative value favours the program.
Results table

Group-based programs and loneliness

8randomized trials with 1,236 participants
7 of 8estimates favour the program
2 of 8confidence intervals exclude zero

Results are standardized mean differences, ordered by risk of bias and then by size.

Carry forward

From tables to synthesis

The tables make the studies comparable, and the synthesis methods turn them into conclusions.

Section 2 introduces synthesis without meta-analysis and the SWiM guideline.

Learning Objectives for this section

  • Explain what synthesis means in a review and why tabulating the included studies is the first step of a quantitative synthesis.
  • Distinguish a characteristics table, a cross-tabulation and a results table, and state the question each one answers.
  • Define the groups for a synthesis by population, intervention, comparator, outcome and design, and explain why they should be planned in the protocol.
  • Write decision rules that select one result per study when a study reports several scales, time points or analyses.
  • Build a results table for one comparison and outcome, with effects aligned in one direction and studies ordered by design and risk of bias.

Introduction

In Lesson 8, the Cedar Valley evidence team extracted data, appraised each study's risk of bias and rated the certainty of evidence for loneliness. Those ratings depend on a step that this lesson teaches: synthesis, the process of bringing the results of several studies together to answer the review question. A synthesis says what the studies, taken together, show about a comparison and an outcome, and how consistent they are.

A quantitative synthesis combines numerical results, such as differences in loneliness scores between program and comparison groups; Lesson 10 introduces the synthesis of qualitative findings. Section 1 of this lesson tabulates and groups the studies. Section 2 introduces synthesis without meta-analysis and the SWiM reporting guideline. Section 3 explains why counting studies by statistical significance misleads and presents better displays. Section 4 asks when pooling results in a meta-analysis is justified, a method whose computation HSCI 230 Lesson 2 teaches.

Case: The Cedar Valley evidence review (fictional)

Before launching a community connector (social prescribing) program for older adults, the planning team of the fictional Cedar Valley Health Authority in British Columbia, which serves about 46,000 people aged 65 and older, asked an evidence team (an evidence officer, a university librarian and a student intern) for a rapid scoping review and environmental scan within twelve weeks. The review asks which community-based interventions have been evaluated for reducing loneliness or social isolation among adults aged 65 and older, and with what outcomes.

The searches and screening produced 42 included studies: 14 randomized trials (4 of them cluster-randomized), 10 non-randomized controlled studies, 9 uncontrolled before-and-after studies, 6 qualitative studies and 3 mixed-methods studies. In Lesson 8, no trial reached low risk of bias overall for loneliness, and GRADE rated the evidence as low certainty for group-based programs and for one-to-one befriending or telephone programs, and very low certainty for community connector programs. All study labels, results and other numbers in this lesson beyond these shared figures are illustrative and belong to the fictional case.

This lesson works with the 36 studies that report quantitative results: the 33 quantitative studies and the quantitative strands of the 3 mixed-methods studies. The 6 qualitative studies and the qualitative strands of the mixed-methods studies go to the qualitative evidence synthesis in Lesson 10. The figure shows the route from the extraction forms to the synthesis.

Extraction forms for the 42 included studies (Lesson 8) 36 studies with quantitative results 33 quantitative studies and 3 mixed-methods studies 6 qualitative studies and qualitative strands: Lesson 10 1. Tabulate characteristics and cross-tabulate category by design 2. Group by comparison, outcome and design 3. Tabulate results for each group, ordered by design and risk of bias 4. Synthesize and display (Sections 2 and 3) or pool (Section 4)
The route from extraction to quantitative synthesis in the fictional Cedar Valley review. Studies with qualitative findings follow a separate route to the qualitative evidence synthesis in Lesson 10.

1.1 Why Synthesis Starts with Tables

A common first draft of a review's results describes the included studies one at a time: "Smith and colleagues found that ... Jones and colleagues found that ...". Readers of such a draft must hold every study in mind and do the synthesis themselves, and the writer has no structure for noticing patterns. A structured table solves both problems. It lays the studies side by side on the same columns, so that similarities, differences and gaps become visible before anyone writes a sentence of interpretation.

The guidance on narrative synthesis by Popay and colleagues (2006) describes four elements: developing a theory of how the intervention works, why and for whom; developing a preliminary synthesis; exploring relationships within and between studies; and assessing the robustness of the synthesis. Tabulation is the main tool of the preliminary synthesis, and the Cochrane Handbook chapter on synthesis without meta-analysis (McKenzie and Brennan, 2019) likewise recommends structured tables. A review usually needs three kinds of table.

A characteristics table describes each included study in one row: its design, country and setting, participants, the program and its comparator, the outcomes measured and the time points. It answers the question "What was studied, in whom, and how?" Every included study appears, and Lesson 12 shows how the table is presented.

StudyDesignParticipantsProgram and comparatorLoneliness measure; time point
Trial ARandomized trial96 adults aged 65 to 89 living alone, EnglandWeekly arts group in a seniors' centre for 8 weeks; usual activitiesUCLA Loneliness Scale (20 items); 8 weeks
Trial LRandomized trial160 adults aged 70 and older, the NetherlandsWeekly telephone befriending call for 6 months; waiting listDe Jong Gierveld scale (6 items); 6 months
Study PNon-randomized controlled study420 adults aged 65 and older referred from primary care, ScotlandCommunity connector referral; usual care in matched practicesRevised UCLA Loneliness Scale (20 items); 6 months
Study BA1Before-and-after study140 adults aged 65 and older, OntarioCommunity connector program run by a seniors' services agency; no comparatorThree-item loneliness scale; 12 months

A cross-tabulation counts studies in the cells formed by two characteristics, most often intervention category by design or intervention category by outcome. It answers the question "Where is the evidence, and where are the gaps?" Section 1.3 gives the Cedar Valley version, and the evidence and gap maps in Lesson 10 extend the idea.

A results table lists, for one comparison and one outcome, each study's result in a consistent form, together with the information a reader needs to weigh it: the number of participants analyzed, the measure and time point, and the study's risk of bias. It answers the question "What did the studies in this group find?" Section 1.5 builds one for group-based programs and loneliness.

1.2 Grouping Studies for Synthesis

A review question is usually broader than any single synthesis. A statement that "community-based interventions reduce loneliness" would combine programs that work in different ways and would help no one plan a service. The review therefore divides its studies into groups for synthesis, each defined by a population, intervention, comparator and outcome. The Cochrane Handbook calls this the "PICO for each synthesis" and treats it as part of planning the review (McKenzie et al., 2019). Design can also define groups, since randomized and non-randomized studies are usually synthesized separately.

Grouping involves a trade-off called lumping and splitting. Lumping places broadly similar programs in one group, giving more studies per synthesis but possibly combining programs that differ in ways that matter. Splitting places studies in narrow groups, giving specific answers from as few as one or two studies each. The right level depends on the decision the review supports. The Cedar Valley planning team must choose a kind of program to fund, so the team grouped by the main component of each program.

Populationv

The eligibility criteria already restrict the population to adults aged 65 and older. Few studies selected participants by baseline loneliness, so the team treated that characteristic as a possible source of variation (Section 2) and did not make it a separate group.

Interventionv

The extraction form assigned each study to one of six categories by its main component: group activity; one-to-one befriending or telephone contact; community connector or social prescribing; intergenerational; technology-based; and other. A data dictionary rule classed a connector program that also ran a lunch club as a connector program, because referral defined it.

Comparatorv

Comparators were usual care, usual activities, no program or a waiting list. The protocol set apart any study comparing two programs, because that comparison answers a different question; no controlled study in the review used such a comparator.

Outcomev

Loneliness is the main outcome. The protocol (Lesson 2) grouped other outcomes into domains: social isolation and participation, well-being, mental health, physical health and function, and health service use. Each domain is a separate synthesis, so that results for loneliness are never mixed with results for, say, depressive symptoms.

Designv

Randomized trials, non-randomized controlled studies and uncontrolled before-and-after studies protect against bias to different degrees. The team synthesized each design separately and based conclusions about effects on controlled designs, a choice Section 2 shows how to report.

Groups should be planned in the protocol, because a group defined after seeing the results can be drawn to make a program look better or worse. When plans must change, for example because a category contains no studies, the change is reported with its reason, as the SWiM guideline in Section 2 asks.

1.3 Cross-Tabulating the Cedar Valley Studies

The team's first table counted the 42 included studies by intervention category and design. Five of the six categories contain studies; no study fell into "other".

Intervention categoryRandomized trialsNon-randomized controlledBefore-and-afterMixed methodsQualitativeTotal
Group activity8031214
One-to-one befriending or telephone601018
Community connector or social prescribing0532212
Intergenerational021014
Technology-based031004
Total141093642

The table shows a pattern that matters for the planning team. Community connector programs, the model Cedar Valley intends to launch, have 12 studies, second only to group activities (14), yet none is a randomized trial; all 14 trials evaluated group or one-to-one programs. This explains in advance why Lesson 8 rated the connector evidence as very low certainty. The table also shows that the intergenerational and technology-based categories hold too few controlled studies for a confident synthesis, which the review should state plainly.

Try it: Read the cross-tabulation

Using the table above, answer three questions. (1) How many controlled studies (randomized or non-randomized) evaluated community connector programs? (2) What proportion of the 42 studies evaluated group activities or one-to-one programs? (3) Name one gap the table reveals that a future study in British Columbia could fill. Suggested answers: (1) five, all non-randomized; (2) 22 of 42, or about 52 percent; (3) the review contains no randomized trial of a community connector program, which a staggered introduction of the Cedar Valley program across clinics could begin to address.

1.4 Choosing One Result per Study

Many studies report several results that could fill one cell of a results table: loneliness on two scales, at three time points, by two kinds of analysis. Reviewers who choose among these after seeing them may pick, without meaning to, the result that fits their expectations. The safeguard is a set of decision rules, written before extraction, that select one result per study for each synthesis, as the Cochrane Handbook recommends (McKenzie et al., 2019). The accordion lists the Cedar Valley rules.

Rule 1: more than one loneliness scalev

Use the first scale available in this order: the UCLA Loneliness Scale (any version), the De Jong Gierveld Loneliness Scale, the three-item loneliness scale, and a single direct question. The order was added to the protocol as a dated amendment in week 7, after the charting form was piloted and before the results fields were charted.

Rule 2: several time pointsv

Use the time point closest to the end of the program, and tabulate later follow-up separately, because effects may fade after a program ends.

Rule 3: several analysesv

For randomized trials, use the analysis that includes all randomized participants (the intention-to-treat analysis) when it is reported. For non-randomized studies, use the estimate adjusted for the confounding domains named in the protocol, which Lesson 8 listed as baseline loneliness, depressive symptoms and living alone.

Rule 4: cluster-randomized trialsv

Use estimates that account for clustering. A cluster trial analyzed as if individuals had been randomized has a confidence interval that is too narrow, and the team flags such results.

Rule 5: several program armsv

When a trial compares two programs with one control group, place each arm in its own category and note that the comparisons share control participants, who must not be counted twice in one synthesis.

Rule 6: direction of the scalev

Record every effect so that the same sign means the same thing. On the loneliness scales used here, a higher score means more loneliness, so a negative difference favours the program. For outcomes where a higher score is better, such as social participation, the team reversed the sign so that a negative difference still favours the program, and noted the reversal.

1.5 Building a Results Table

With groups defined and one result chosen per study, the team built a results table for each synthesis. To compare results from different loneliness scales, it expressed each as a standardized mean difference (SMD): the difference in mean scores between program and comparison groups divided by a standard deviation, which puts every scale in standard deviation units. HSCI 230 Lesson 2 shows the calculation. The table gives the results for the 8 trials of group-based programs.

TrialDesignParticipants analyzedScaleProgram lengthSMD (95% CI)Risk of bias (RoB 2)
CCluster-randomized210De Jong Gierveld (6 items)16 weeks−0.18 (−0.51 to 0.15)Some concerns
FRandomized182De Jong Gierveld (6 items)24 weeks−0.15 (−0.44 to 0.14)Some concerns
ERandomized148UCLA (20 items)16 weeks−0.20 (−0.52 to 0.12)Some concerns
BRandomized120Three-item scale12 weeks0.03 (−0.33 to 0.39)Some concerns
ARandomized96UCLA (20 items)8 weeks−0.22 (−0.62 to 0.18)Some concerns
DRandomized64Three-item scale10 weeks−0.55 (−1.05 to −0.05)Some concerns
GCluster-randomized230UCLA (20 items)12 weeks−0.02 (−0.34 to 0.30)High
HRandomized186Three-item scale8 weeks−0.31 (−0.60 to −0.02)High

SMD: standardized mean difference, program minus comparison; values below zero mean less loneliness in the program group. CI: confidence interval, the range of effects reasonably compatible with the trial's data. Trials C and G report estimates adjusted for clustering. The 8 trials analyzed 1,236 participants in total. All results are illustrative.

Three features make this table useful. Every result points in a stated direction, so a reader sees at once that seven of the eight estimates fall below zero. The rows are ordered by risk of bias and then by size, so the two trials at high risk sit together and a reader can check whether they differ from the rest; here one shows almost no difference and the other a small benefit. And the table includes columns, such as program length, that might explain variation (Section 2). A study that reported only "no significant difference" still gets a row, marked "not estimable", so that readers can see how much evidence lacks usable data.

The table does not yet state a conclusion. A reader might be tempted to count the confidence intervals that exclude zero (two of eight) and decide that group programs mostly fail. Section 3 explains why that reading is wrong and what to do instead, and Section 2 first sets out the methods and reporting standard for synthesis without meta-analysis.

Copying the authors' conclusionsClick to explore
Mixing time pointsClick to explore
Dropping studies without numbersClick to explore
Inconsistent directionClick to explore

Reflection

A student review examines community walking programs for adults aged 65 and older and includes seven studies. Four are randomized trials comparing a walking group with usual activities: two measure loneliness with the UCLA Loneliness Scale, one with the De Jong Gierveld scale, and one with both scales. Two of these four trials report loneliness at 12 weeks (the end of the program) and again at 6 months. A fifth randomized trial compares a walking group with a seated exercise class. The last two studies are before-and-after evaluations of walking groups with no comparison group. The student's draft results table has one row for every loneliness result in every study, records each result as "effective" or "not effective" following the authors' abstracts, and leaves out one of the four usual-activities trials because it reported only "no significant difference". Propose the groups for synthesis, write two decision rules for choosing one result per study, and identify three problems in the draft table, with a correction for each.

Model answer

Groups. The main synthesis should contain the four trials comparing walking groups with usual activities, for loneliness. The trial with a seated exercise class as comparator answers a different question (walking compared with another program) and should be reported separately. The two before-and-after studies should be tabulated and described, but conclusions about effect should rest on the controlled trials, because without a comparison group a change over time cannot be attributed to the program.

Decision rules. First, when a study reports more than one loneliness scale, use the UCLA Loneliness Scale, then the De Jong Gierveld scale. Second, use the time point closest to the end of the program (12 weeks) in the main synthesis and tabulate the 6-month results separately.

Problems and corrections. (1) One row per result lets the trial with two scales and the trials with two time points appear more than once; the table should have one row per study for each synthesis, chosen by the decision rules. (2) "Effective" and "not effective" repeat the authors' significance-based interpretation; the table should record each effect estimate with its confidence interval, aligned so that the same sign always favours the program. (3) Omitting the trial with no numbers makes the evidence look more complete; it should keep its row, marked "not estimable", and the synthesis should report that one of four trials lacked usable data.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: What does a cross-tabulation in a review show?

A cross-tabulation counts studies in the cells formed by two characteristics, such as intervention category by design, and shows where evidence is concentrated and where gaps lie. Option c describes a characteristics table and option a describes a results table.

Question 2: Why should decision rules for choosing one result per study be written before extraction begins?

When a study reports several scales, time points or analyses, a choice made after seeing the results can, even unintentionally, favour the result that fits expectations. Rules fixed in advance prevent this. The rules reduce the results to one per study, so option b is the reverse of their effect.

Question 3: What does the community connector row of the Cedar Valley cross-tabulation show?

Community connector programs have 12 studies, second only to group activities (14), made up of 5 non-randomized controlled studies, 3 before-and-after studies, 2 mixed-methods studies and 2 qualitative studies. All 14 randomized trials evaluated group activities or one-to-one programs, which helps explain the very low certainty rating for connector programs.

Question 4: A trial reports loneliness on a scale where a higher score means more loneliness, and social participation on a scale where a higher score means more participation. In a results table where negative values favour the program, what does the Cedar Valley Rule 6 require?

Rule 6 aligns direction so that the same sign always means the same thing: a rise in participation is a benefit, so its sign is reversed to make benefit negative, and the reversal is recorded. Leaving signs as reported would make readers misread the table, and the two outcomes belong to different syntheses in any case.
Section 2 of 5

Synthesis Without Meta-Analysis and the SWiM Guideline

⏱ Estimated reading time: 40 minutes
Section 2 of 5

Synthesis Without Meta-Analysis and the SWiM Guideline

This section covers the reasons for synthesis without pooling, three accepted methods and the nine SWiM items.

40 minutes
Why not pool

Reasons for synthesis without meta-analysis

  • Some studies report results incompletely.
  • Some effect measures cannot be converted to a common metric.
  • A high risk of bias would make a pooled estimate falsely precise.
  • Very different studies would produce a meaningless average.
Three methods

Synthesis without pooling

Summarize estimates

The method describes the size and spread of effects with the median and interquartile range.

Combine P values

The method tests for an effect in at least one study.

Count by direction

The method counts studies favouring each side, whatever their size.

Worked example

Summarizing the group-based trials

−0.19median standardized mean difference
−0.24 to −0.12interquartile range
−0.55 to 0.03range

Negative values mean less loneliness in the program group.

SWiM

Nine reporting items

1 Grouping
2 Metric
3 Methods
4 Prioritization
5 Heterogeneity
6 Certainty
7 Presentation
8 Results
9 Limitations

Items 1 to 7 go in the methods, item 8 in the results and item 9 in the discussion (Campbell et al., 2020).

Variation and writing

Investigate, then state the finding

Program length

The median was −0.22 for 12 weeks or less and −0.18 for longer programs.

Structured statement

Group-based programs may reduce loneliness slightly (low certainty).

Carry forward

From methods to counting and displays

Counting studies is useful only when the count is based on direction of effect.

Section 3 explains the problem with significance counts and introduces effect direction plots and harvest plots.

Learning Objectives for this section

  • Give reasons why a review may synthesize quantitative results without a meta-analysis, and explain why "narrative synthesis" is an incomplete description of a method.
  • Describe the three synthesis methods that the Cochrane Handbook accepts when effects are not pooled, and state the question each answers and the data each needs.
  • Summarize a set of effect estimates with the median, interquartile range and range, and interpret the summary correctly.
  • Describe the nine items of the SWiM reporting guideline and apply them to a synthesis plan.
  • Investigate variation in results without meta-analysis by tabulating results against characteristics planned in advance, and write a structured synthesis paragraph.

Introduction

Many reviews in public health do not combine their results in a meta-analysis, because the studies are too few, too varied or too incompletely reported, or because the review maps evidence. Such reviews still synthesize, and many have described what they did as a narrative synthesis. The term has a defined meaning in the guidance of Popay and colleagues (2006), but it has often been used for any synthesis that is not a meta-analysis, including study-by-study descriptions with no stated method.

Campbell and colleagues (2019) examined a sample of Cochrane reviews that synthesized results without meta-analysis and found that most did not describe their synthesis methods or link their conclusions clearly to the data. The same group then led the development of the Synthesis Without Meta-analysis (SWiM) reporting guideline (Campbell et al., 2020), which sets out nine items to report when a review synthesizes quantitative effects without pooling them.

2.1 When a Review Synthesizes Without Meta-Analysis

The Cochrane Handbook chapter on synthesis without meta-analysis (McKenzie and Brennan, 2019) lists the circumstances that lead reviews to other methods. Effect estimates may be incompletely reported, for example as "no significant difference" with no numbers. Studies may use effect measures that cannot be converted to a common metric. The evidence may be at such high risk of bias that a pooled estimate would give misleading precision. And the studies may differ so much that an average effect would have no clear meaning. Section 4 returns to these conditions.

The Cedar Valley review meets several of them. Its protocol (Lesson 2) planned no meta-analysis, because a twelve-week rapid scoping review aims to map evidence. Its community connector studies are non-randomized and mostly at serious risk of confounding, and several categories contain only a handful of controlled studies. Even so, the planning team asked what the evidence says about loneliness, and the review needs a method that answers that question transparently.

2.2 Three Methods That Synthesize Without Pooling

McKenzie and Brennan (2019) describe three methods that are acceptable when effects are not pooled. Each answers a different question and needs different data. The table summarizes them.

MethodQuestion it answersMinimum data from each studyMain limitation
Summarizing effect estimatesWhat is the range and distribution of the observed effects?An effect estimate in a common metricThe summary does not weight studies by size or precision and gives no confidence interval for an average effect.
Combining P valuesIs there evidence of an effect in at least one study?An exact P value and the direction of effectThe result says nothing about the size of the effect and is rarely meaningful for health decisions.
Vote counting based on the direction of effectIs there any evidence of an effect?The direction of effectThe count ignores the size and precision of each result, so a trivial benefit and a large one count the same.

Summarizing effect estimates

When the studies in a group report effects in a common metric, such as the standardized mean difference introduced in Section 1, the review can describe their distribution with three statistics. The median is the middle value when the estimates are put in order. The interquartile range (IQR) is the span of the middle half of the estimates, from the 25th to the 75th percentile. The range runs from the smallest to the largest estimate. These statistics describe the spread of observed results and carry no weighting by study size.

Worked example: group-based programs and loneliness

The eight standardized mean differences from Section 1, in order, are −0.55, −0.31, −0.22, −0.20, −0.18, −0.15, −0.02 and 0.03.

Median = (−0.20 + −0.18) ÷ 2 = −0.19, because with eight values the median is the mean of the fourth and fifth.

Interquartile range = −0.24 to −0.12 (computed with the default quantile method in R and Python). Range = −0.55 to 0.03.

By a widely used rule of thumb (Cohen, 1988), a standardized mean difference of about 0.2 is described as small. The summary statement is that the middle half of the trials found loneliness lower in the program group by between 0.12 and 0.24 standard deviations, a small difference.

For the 6 one-to-one befriending or telephone trials, the 5 with estimable results had a median of −0.20 (IQR −0.24 to −0.16; range −0.28 to −0.12). Trial L reported only that there was no significant difference, so the summary must state that it rests on 5 of the 6 trials.

This method needs a standard metric. A difference of 2 points means something different on the 20-item UCLA scale, scored from 20 to 80, than on the three-item scale, scored from 3 to 9, and standard deviation units put the two on a common footing. For binary outcomes, such as the proportion who often feel lonely, the metric would be a ratio such as the odds ratio or risk ratio, which HSCI 230 Lesson 2 teaches.

Combining P values

Methods such as Fisher's method combine the exact P values from several studies into one test of the hypothesis that the program has no effect in any study. A small combined P value suggests a real effect in at least one study but says nothing about its size. The Cedar Valley team did not use these methods, since the planning team asked about the size and consistency of effects.

Vote counting based on the direction of effect

The third method counts how many studies found an effect in the direction of benefit and how many in the direction of harm, whatever the size or significance of each result. It needs the least data, and Section 3 explains how it differs from the flawed practice of counting studies by statistical significance.

2.3 The SWiM Reporting Guideline

SWiM is a reporting guideline: it lists what a review should tell readers about a synthesis that does not pool effects, so that they can judge whether the conclusions follow. It was developed through a consensus process and is designed to be used alongside PRISMA 2020 (Campbell et al., 2020). Its nine items cover the methods (items 1 to 7), the results (item 8) and the discussion (item 9).

Methods section 1Grouping studies for synthesis (1a and 1b) 2Standardized metric and transformations 3Synthesis methods 4Criteria for prioritizing results 5Investigation of heterogeneity 6Certainty of evidence 7Data presentation methods Results section 8 Reporting results: findings and certainty for each comparison and outcome Discussion section 9 Limitations of the synthesis methods and groupings
The nine SWiM items arranged by the part of the review in which each is reported (based on Campbell et al., 2020). Item 1 has two parts: 1a describes the groups and 1b reports changes from the protocol.

The accordion describes each item and shows how the Cedar Valley review reports it.

Item 1: Grouping studies for synthesisv

What to report. Item 1a asks for a description of the groups used in the synthesis, such as groupings of populations, interventions, outcomes or designs, with the rationale for each. Item 1b asks for any changes to these groups after the protocol, with reasons.

Cedar Valley. Studies were grouped by the main component of the program, by outcome domain and by design, because the planning team must choose a type of program. An amendment made before the synthesis began, at the planning team's request, added standardized effect summaries for trials and GRADE ratings for loneliness. The planned category "other" contained no studies.

Item 2: Standardized metric and transformation methodsv

What to report. The standard metric used for each outcome, why it was chosen, and any methods used to transform the results as reported into that metric, citing the guidance followed.

Cedar Valley. Loneliness results were expressed as standardized mean differences because the studies used different scales, with standard deviations derived from standard errors or confidence intervals by methods in the Cochrane Handbook where needed. Where no metric could be calculated, the direction of effect was used.

Item 3: Synthesis methodsv

What to report. The methods used to synthesize the effects for each outcome, with a justification for using them instead of meta-analysis.

Cedar Valley. For groups with three or more estimable results, the team summarized standardized mean differences with the median, interquartile range and range. For all groups, it counted studies by direction of effect and applied a sign test, which Section 3 explains. It did not combine P values. Meta-analysis was not planned, for the reasons given in the protocol.

Item 4: Criteria used to prioritize resultsv

What to report. Any criteria used to select particular studies for the main synthesis or for drawing conclusions, such as design, risk of bias or directness, with a justification.

Cedar Valley. Conclusions about effects rest on the controlled studies. The before-and-after studies and the quantitative strands of the mixed-methods studies are shown in tables and plots but do not support statements about effects, because without a comparison group a change over time cannot be attributed to the program. Risk of bias was used to order studies, never to exclude them.

Item 5: Investigation of heterogeneity in reported effectsv

What to report. The methods used to examine variation in effects across studies when a meta-analysis and its statistical tools for heterogeneity are not used.

Cedar Valley. The team planned to tabulate results by program length, delivery setting and risk of bias, and to describe any pattern as exploratory. Section 2.4 shows the result.

Item 6: Certainty of evidencev

What to report. The methods used to assess the certainty of the synthesis findings.

Cedar Valley. GRADE was applied to loneliness for three comparisons, following the guidance of Murad and colleagues (2017) for rating certainty without a pooled estimate, as Lesson 8 described.

Item 7: Data presentation methodsv

What to report. The tables and graphs used to present effects, and the study characteristics used to order studies in them, with the studies clearly identified.

Cedar Valley. The review uses results tables ordered by design, risk of bias and size, an effect direction plot for community connector programs and a harvest plot of loneliness across the five categories (Section 3).

Item 8: Reporting resultsv

What to report. For each comparison and outcome, a description of the synthesized findings and their certainty, in language consistent with the question the synthesis addresses, stating which studies contribute.

Cedar Valley. Section 2.5 gives the paragraph for group-based programs and loneliness.

Item 9: Limitations of the synthesisv

What to report. The limitations of the synthesis methods and of the groupings, and how they affect the conclusions that can be drawn for the review question.

Cedar Valley. Direction counts ignore the size of effects; most groups hold few controlled studies; one befriending trial had no estimable result; and grouping by main component may hide the contribution of other components in multi-component programs.

SWiM allows any defensible method, provided the review reports it against the nine items. A reader of a review that follows SWiM can find which studies contributed to each statement, what was done with their results, and why.

2.4 Investigating Variation Without Meta-Analysis

Results almost always vary across studies, partly by chance and partly, perhaps, because of real differences between programs, populations or methods. Without the statistical tools of meta-analysis, a review can tabulate or plot results against characteristics that might explain the variation, as SWiM item 5 asks. The characteristics should be few, chosen in the protocol and grounded in a reason to expect a difference, because a search through many characteristics after seeing the results will find patterns by chance.

The Cedar Valley team examined program length for the 8 group-based trials, reasoning that longer programs might allow friendships to form. The table divides the trials at 12 weeks.

Program lengthTrialsStandardized mean differencesMedianRange
12 weeks or lessA, B, D, G, H (5)−0.22, 0.03, −0.55, −0.02, −0.31−0.22−0.55 to 0.03
More than 12 weeksC, E, F (3)−0.18, −0.20, −0.15−0.18−0.20 to −0.15

The medians are similar, and the shorter programs show a wider spread, including both the largest benefit and the two results near zero. With five and three trials, this pattern could easily arise by chance, so the team reported that program length did not explain the variation; tabulations by setting and by risk of bias gave the same conclusion. The unexplained variation is why GRADE rated down for inconsistency in Lesson 8.

2.5 Writing the Synthesis

SWiM item 8 asks that each written statement address a comparison and an outcome, report the findings and their certainty, and name the contributing studies. The tabs contrast a study-by-study draft with a structured synthesis of the same eight trials.

"Trial D found that the group program significantly reduced loneliness. Trial H also found a significant reduction. Trials A, C, E and F found no significant effect. Trial B found no effect, and Trial G found no effect. Overall, the evidence for group programs is mixed."

This draft repeats each study's significance test, never states a size of effect, ignores risk of bias and reaches a conclusion ("mixed") that the data do not support, since seven of the eight estimates point the same way.

"Eight randomized trials (1,236 participants; Trials A to H) compared group-based programs with usual activities or no program. Seven of the eight found lower loneliness in the program group at the end of the program. Standardized mean differences had a median of −0.19 (interquartile range −0.24 to −0.12; range −0.55 to 0.03), a small difference. The estimates varied, and program length, setting and risk of bias did not explain the variation. All trials had some concerns or high risk of bias for this outcome. Group-based programs may reduce loneliness slightly among older adults living in the community (low-certainty evidence)."

This paragraph names the comparison, the studies and the participants, gives the direction and the size of the effects, reports the exploration of variation and the risk of bias, and ends with a GRADE statement in the standard wording from Lesson 8.

Try it: Check a methods paragraph against SWiM

A student review reports its synthesis methods in one paragraph: "Because the studies were heterogeneous, a narrative synthesis was conducted. Studies were grouped by intervention type. Results are presented in tables." List which SWiM items the paragraph addresses, which it omits, and one sentence the student could add for each of two omitted items. Suggested answer: the paragraph partly addresses item 1a (groups, without a rationale) and item 7 (tables, without saying how studies are ordered). It omits items 2 to 6 and gives no basis for items 8 and 9. For item 3, the student could add: "We summarized standardized mean differences with the median and interquartile range and counted studies by direction of effect." For item 4, the student could add: "Conclusions about effects were based on controlled studies; uncontrolled studies are described separately."

Narrative synthesis means no methodClick to explore
A median effect is a pooled effectClick to explore
SWiM tells you which method to useClick to explore
Scoping reviews never synthesize effectsClick to explore

Reflection

Six randomized trials compared peer-support programs with usual care for loneliness among older adults living in the community, with 640 participants in total. Five reported results that the review team converted to standardized mean differences, where negative values mean less loneliness in the program group: Trial 1, −0.40; Trial 2, −0.10; Trial 3, −0.25; Trial 4, 0.05; Trial 5, −0.30. Trial 6 reported only that the difference was "not significant". The review also includes three uncontrolled before-and-after studies of peer support. GRADE rated the evidence from the trials as low certainty. Using these data: (a) calculate the median and range of the five standardized mean differences and count how many trials favour the program; (b) write one sentence for SWiM item 3 (synthesis methods) and one for SWiM item 4 (criteria for prioritizing results); (c) write a results paragraph that meets SWiM item 8; and (d) state one limitation of this synthesis for SWiM item 9.

Model answer

(a) In order, the five values are −0.40, −0.30, −0.25, −0.10 and 0.05, so the median is −0.25 and the range is −0.40 to 0.05. Four of the five trials with estimable results favour the program, and Trial 6 has no estimable direction.

(b) Item 3: "Because the trials used different loneliness scales and a meta-analysis was not planned, we summarized standardized mean differences with the median and range and counted trials by the direction of effect." Item 4: "Conclusions about effect were based on the randomized trials; the three uncontrolled before-and-after studies are described separately because they lack a comparison group."

(c) "Six randomized trials (640 participants) compared peer-support programs with usual care. Four of the five trials with estimable results found lower loneliness in the program group, with a median standardized mean difference of −0.25 (range −0.40 to 0.05), a small difference; Trial 6 reported no usable data. Peer-support programs may reduce loneliness among older adults living in the community (low-certainty evidence)."

(d) The median gives every trial equal weight regardless of size, and the direction count treats the estimate of 0.05, which is close to zero, as a full vote against the program, so the synthesis describes the pattern of results without estimating an average effect with a confidence interval.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: Which question does summarizing effect estimates with the median and interquartile range answer?

Summary statistics describe the distribution of the effects the studies observed. They give each study equal weight and no confidence interval, so they do not provide a weighted average, which is what a meta-analysis estimates. Option a is the question answered by combining P values.

Question 2: Eight standardized mean differences are −0.55, −0.31, −0.22, −0.20, −0.18, −0.15, −0.02 and 0.03. What is their median?

With an even number of values, the median is the mean of the two middle values: (−0.20 + −0.18) ÷ 2 = −0.19. Taking the fourth or fifth value alone gives one of the two middle values, which is the most common error.

Question 3: Which statement describes SWiM item 4?

Item 4 concerns criteria, such as design or risk of bias, used to select studies for the main synthesis or for drawing conclusions. The standard metric is item 2, data presentation is item 7 and certainty of evidence is item 6.

Question 4: What did the Cedar Valley team conclude after tabulating the group-based trial results by program length?

Trials of 12 weeks or less had a median of −0.22 and those of more than 12 weeks a median of −0.18, with overlapping ranges and very few trials in each group, so the team reported that length did not explain the variation. The unexplained variation is why GRADE rated down for inconsistency.
Section 3 of 5

Beyond Vote Counting: Effect Direction Plots and Harvest Plots

⏱ Estimated reading time: 40 minutes
Section 3 of 5

Beyond Vote Counting: Effect Direction Plots and Harvest Plots

This section explains why significance counts mislead and how to count and display results instead.

40 minutes
The problem

Two counts of the same eight trials

2 of 8significant: a significance count calls the evidence mixed
7 of 8favour the program: a direction count shows consistency
Power

Small trials miss small effects

17%power of a 100-person trial for a difference of 0.2
0.3%chance that most of 10 such trials are significant

Majority-rule counts fail more often as underpowered studies accumulate (Hedges and Olkin, 1980).

Direction of effect

Counting by direction, with a sign test

Group-based programs

Seven of 8 trials favour the program (sign test P = 0.07).

One-to-one programs

Five of 5 estimable trials favour the program (P = 0.06), and Trial L is not estimable.

The count ignores the size of each effect, so it is reported with the effect summaries.

Effect direction plot

Community connector programs, five outcome domains

5 of 5controlled studies favour the program for loneliness
2 of 4favour the program for mental health
2 of 4favour the program for service use
Harvest plot

Loneliness across five categories

20controlled studies favour the program
1has an unclear direction
3favour the comparison

The best-designed evidence concerns programs other than community connectors.

Carry forward

From describing to pooling

Counts and plots describe the evidence; a meta-analysis estimates an average effect.

Section 4 asks when pooling is justified.

Learning Objectives for this section

  • Describe vote counting by statistical significance and explain, with reference to statistical power, why it misleads.
  • Explain why a non-significant result is not evidence of no effect, and why the problem grows as the number of small studies grows.
  • Carry out vote counting based on the direction of effect, report the proportion of studies favouring the program, and interpret a sign test.
  • Read and construct an effect direction plot and a harvest plot from tabulated study results.
  • Choose a display suited to the data available in a group of studies.

Introduction

Section 1 ended with a temptation: to see that only two of eight confidence intervals for group-based programs exclude zero and conclude that the programs mostly failed. Lesson 1 showed the same trap with eight simulated trials of one program, of which only three reached significance. This section explains why counting by significance gives wrong answers, presents vote counting based on the direction of effect, and shows two displays designed for synthesis without meta-analysis.

3.1 Vote Counting by Statistical Significance

A result is called statistically significant when its P value falls below a chosen threshold, usually 0.05, or equivalently when its 95 percent confidence interval excludes the value that means no effect (zero for a difference). Vote counting by statistical significance sorts studies into three piles, significant benefit, significant harm and no significant difference, and then reaches a conclusion from the tallies. If most studies are significant and favourable, the program "works"; if most are non-significant, it "does not work"; and if the piles are similar, the evidence is "mixed". The method still appears in published reviews, including one that the Cedar Valley team verified in Lesson 6, and the Cochrane Handbook advises against it (McKenzie and Brennan, 2019).

3.2 Why Counting by Significance Misleads

A non-significant result is not a finding of no effect

Whether a study reaches significance depends on the size of the true effect and on the study's precision, which rises with the number of participants. A small study of an effective program will often have a confidence interval wide enough to include zero. Altman and Bland (1995) summarized the error of reading such a result as proof of no effect in the title of a short paper: "Absence of evidence is not evidence of absence". A non-significant result means that the study could not distinguish the effect from zero. Its confidence interval often includes both no effect and a worthwhile benefit.

The figure plots the Cedar Valley results for group-based programs. It looks like a forest plot, which HSCI 230 Lesson 2 teaches, but it has no pooled estimate at the bottom; SWiM item 7 lists this kind of ordered display as a way to present results without meta-analysis.

Trial (participants) P < 0.05 Standardized mean difference (95% CI) Trial C (210) No Trial F (182) No Trial E (148) No Trial B (120) No Trial A (96) No Trial D (64) Yes Trial G (230) No Trial H (186) Yes Some concerns above the dashed line; high risk below −1.2 −0.9 −0.6 −0.3 0 0.3 0.6 ← Favours the program Favours comparison → Interval excludes zero Interval includes zero
Figure 3.1. Standardized mean differences for loneliness with 95 percent confidence intervals in the 8 illustrative Cedar Valley trials of group-based programs, ordered by risk of bias and size, without a pooled estimate. A count by significance gives two successes and six failures; a count by direction gives seven of eight trials favouring the program.

Six of the eight intervals include zero, and all six of those also include reductions in loneliness of 0.3 standard deviations or more. The non-significant trials are compatible with a small benefit, and seven of the eight point estimates favour the program. A significance count reads this pattern as "two successes, six failures"; a direction count shows seven of eight trials pointing toward benefit, a pattern compatible with a small benefit measured imprecisely.

Small studies have low power

Statistical power is the probability that a study will produce a statistically significant result when the program truly has an effect of a given size. The table shows the power of a two-group trial to detect a standardized mean difference of 0.2, a small effect of the size the Cedar Valley trials suggest, using a two-sided test at the 0.05 level.

Participants in the trial (two equal groups)50100200400
Power to detect a standardized mean difference of 0.211%17%29%52%

Calculated with the normal approximation for a comparison of two means. Most Cedar Valley trials analyzed between 64 and 230 participants.

A trial with 100 participants has about a one-in-six chance of reaching significance even when the program truly reduces loneliness by 0.2 standard deviations, so a significance count across such trials will usually declare an effective program ineffective.

More studies can make the answer worse

One might expect that adding studies would correct the error. Hedges and Olkin (1980) showed the opposite for the majority rule. When each study has power below one half, the proportion of significant results tends toward that power as studies accumulate, so the chance that most studies are significant falls toward zero. The table applies this to trials of 100 participants each, with a true standardized mean difference of 0.2 and power of 17 percent.

Number of trials5102040
Expected number with a significant result0.91.73.46.8
Probability that more than half are significant3.8%0.3%0.01%under 0.01%

With 40 such trials, a vote count is almost certain to conclude that a program with a real effect does not work. As the number of trials grows, a majority-rule count becomes more likely to reach the wrong conclusion.

Size, weight and false conflict

Vote counting by significance also discards information. A trial of 64 people counts the same as a trial of 230, and a large benefit counts the same as a trivial one that reached significance in a very large study. The method also manufactures conflict: when some trials are significant and others are not, reviewers call the evidence "mixed" even when the estimates lie close together, as in Figure 3.1.

3.3 Vote Counting Based on the Direction of Effect

The acceptable alternative counts each study by the direction of its effect estimate alone: favouring the program or favouring the comparison, whatever its size or significance (McKenzie and Brennan, 2019). The review reports the number and proportion of studies favouring the program, ideally with a confidence interval for that proportion, and can apply a sign test. The sign test asks how likely a split at least as uneven as the observed one would be if the program had no effect, so that each study were as likely to point one way as the other, like a fair coin.

Worked example: direction counts in the Cedar Valley trials

Group-based programs. Seven of 8 trials favour the program (88 percent; 95 percent confidence interval for the proportion 53 to 98 percent, Wilson method). Two-sided sign test: P = 0.07.

One-to-one befriending or telephone programs. Five of the 5 trials with an estimable direction favour the program, and Trial L, which reported only "no significant difference", has no estimable direction. Two-sided sign test on 5 trials: P = 0.06. No trial in this group reached statistical significance, so a significance count would have recorded zero successes in six.

These P values need care. With only five studies, the sign test cannot produce a two-sided P value below 0.06, because even five out of five has a probability of 2 in 32 under the hypothesis of no effect. Reading such a P value as proof of no effect would repeat the error of Section 3.2, so the team reports the counts, the proportion with its interval and the P value together, and lets the effect summaries from Section 2 describe size.

Direction counting has its own limits, which SWiM item 9 asks reviews to state. It ignores the size of effects, so an estimate of 0.03 counts as a full vote against the program although it is essentially zero. It gives a small study the same vote as a large one. It requires every result to be aligned in direction before counting (Section 1, Rule 6). And it counts studies, so a study that reports five loneliness measures still casts one vote in each synthesis, chosen by the decision rules.

Try it: Recount the befriending trials

The one-to-one trials reported these standardized mean differences (95 percent confidence intervals): Trial I −0.28 (−0.70 to 0.14); Trial J −0.12 (−0.53 to 0.29); Trial K −0.24 (−0.64 to 0.16); Trial L not estimable; Trial M −0.16 (−0.52 to 0.20); Trial N −0.20 (−0.61 to 0.21). (1) What would a significance count conclude? (2) What does a direction count show? (3) Write one sentence for the review. Suggested answers: (1) none of the six trials is significant, so a significance count would conclude that befriending does not reduce loneliness. (2) All five trials with an estimable result favour the program, with a median standardized mean difference of −0.20, and one trial has no estimable direction. (3) "All five trials with estimable results found lower loneliness in the befriending group (median standardized mean difference −0.20), but each was small and imprecise, and one further trial reported no usable data."

3.4 Effect Direction Plots

Direction counts become harder to follow when studies report several outcomes measured in different ways. The effect direction plot, developed by Thomson and Thomas (2013) for reviews of complex public health interventions, displays the direction of effect for every study and every outcome domain in one grid. Each row is a study and each column an outcome domain. An upward arrow shows an effect favouring the intervention, a downward arrow an effect favouring the comparison, and a sideways arrow a domain in which the study's several outcomes conflict. The size of the arrow shows the study's sample size, and rows are ordered or shaded by design and risk of bias. Hilton Boon and Thomson (2021) revised the plot to follow the 2019 Cochrane Handbook guidance, so that arrows record the direction of effect regardless of statistical significance and each column can be summarized with a sign test.

When a study reports several outcomes within one domain, such as depressive symptoms and anxiety within mental health, the team needs a rule for the domain's arrow. The Cedar Valley team gave a domain a direction when at least 70 percent of its outcomes in the study pointed the same way and marked it as conflicting otherwise, a threshold of the kind Thomson and Thomas (2013) described. It also defined the direction of each domain in advance: for health service use, fewer unplanned visits to primary care or emergency departments counted as benefit.

StudynRisk of bias or design Loneliness Socialisolation Well-being Mentalhealth Serviceuse Controlled studies (non-randomized; ROBINS-I risk of bias) Uncontrolled studies (described; not used for statements about effect) P 420 Moderate Q 310 Moderate R 260 Serious S 290 Serious T 200 Serious BA1 140 Before-and-after BA2 85 Before-and-after BA3 60 Before-and-after MM1 120 Mixed methods MM2 48 Mixed methods Controlled: favour program All studies: favour program 5 of 5 9 of 10 3 of 3 5 of 6 4 of 4 7 of 8 2 of 4 4 of 7 2 of 4 3 of 5 Benefit Harm Conflicting within domain Not measured Arrow size: large 300 or more participants; medium 100 to 299; small fewer than 100.
Figure 3.2. Effect direction plot for the 10 illustrative Cedar Valley studies of community connector programs: 5 non-randomized controlled studies (P to T), 3 before-and-after studies (BA1 to BA3) and the quantitative strands of 2 mixed-methods studies (MM1 and MM2). Upward arrows favour the program regardless of statistical significance.

For loneliness, all five controlled studies favour the program, as do four of the five uncontrolled studies. On a two-sided sign test, five of five gives P = 0.06, the smallest value possible with five studies. Social isolation and well-being point the same way in every controlled study that measured them. Mental health and health service use are mixed: Studies Q and T recorded more unplanned visits among program participants, which the team interpreted with caution, since connectors may link people to care they had been going without. Because the controlled studies are non-randomized and three of five are at serious risk of confounding, the plot supports a statement about consistency of direction, and the GRADE rating from Lesson 8 (very low certainty) governs how strongly the review can state it.

Following SWiM item 4, the uncontrolled studies appear below a labelled divider, visible to readers, and the review's statements about effect rest on the controlled studies above it.

3.5 Harvest Plots

The harvest plot was developed by Ogilvie and colleagues (2008) to synthesize evidence about the differential effects of interventions, such as whether they narrow or widen health inequalities, and is now used for many kinds of synthesis without meta-analysis. It is a matrix. Rows usually represent outcomes, hypotheses or categories of intervention, and columns represent categories of finding. Each study is drawn as a bar in the cell that matches its finding, and the height and shading of the bar encode study characteristics chosen by the review team, such as design and risk of bias, with a label above the bar identifying the study. The plot combines the information of a vote count with a picture of the strength of the evidence behind each vote.

The Cedar Valley team drew one harvest plot for loneliness, covering the 24 controlled studies across the five intervention categories. It placed studies by direction of effect, consistent with Section 3.3, so that the plot shows a direction-based vote count with design and risk of bias added.

Category Favours the program Uncleardirection Favours thecomparison Group activity (8) C F E A D G H B One-to-one befriending (6) J M I K N L Community connector (5) P Q R S T Intergenerational (2) U V Technology-based (3) X Y W Randomized trial Non-randomized controlled study Some concerns (RoB 2) or moderate risk (ROBINS-I) High risk (RoB 2) or serious risk (ROBINS-I)
Figure 3.3. Harvest plot of loneliness results from the 24 illustrative Cedar Valley controlled studies. Bar height shows design (tall: randomized trial; short: non-randomized controlled study), colour shows risk of bias, and the letter identifies the study in the results tables.

Twenty of the 24 controlled studies favour the program, one has an unclear direction and three favour the comparison. The tall bars, the randomized trials, are concentrated in the group and one-to-one rows. The community connector row holds only short bars, and three of its five are red. The intergenerational and technology rows hold so few studies, mostly at serious risk of bias, that no pattern can be read from them. A planning team can see in one figure that the most consistent and best-designed evidence concerns programs other than the one it plans to launch, which is the same message the cross-tabulation in Section 1 and the GRADE ratings in Lesson 8 gave.

3.6 Choosing a Display

The three displays suit different data, and a review often uses more than one. The table compares them. Other displays exist, such as the albatross plot, which plots each study's P value against its sample size (Harrison et al., 2017), and the choice among them should follow the question and the data, as SWiM item 7 asks reviews to explain.

DisplayWhat it showsSuited toCedar Valley use
Ordered estimate plot without a pooled resultEach study's effect estimate and confidence interval on a common metricGroups in which the studies report a common or convertible metricGroup-based trials (Figure 3.1)
Effect direction plotDirection of effect for each study in several outcome domains, with sample sizeStudies that measure many outcomes in different waysCommunity connector programs (Figure 3.2)
Harvest plotWhere studies fall by direction across several categories, with design and risk of biasComparing categories, outcomes or hypotheses at a glanceLoneliness across five categories (Figure 3.3)
Counting outcomes instead of studiesClick to explore
Reading a sideways arrow as no effectClick to explore
Choosing arrows by significanceClick to explore
Hiding studies with no usable directionClick to explore

Each of these methods describes a body of evidence without combining its results into a single estimate. Section 4 asks when a review should go further and pool the results statistically, and what HSCI 230 Lesson 2 adds when it does.

Reflection

A published review of a home-visiting program for isolated older adults included 12 randomized trials, each with between 40 and 120 participants. Three trials found a statistically significant reduction in loneliness and nine found no significant difference. The authors concluded that "most trials found no effect, so the program is ineffective". The review's supplementary table shows that 10 of the 12 point estimates favoured the program and 2 favoured the comparison, and that most confidence intervals included both zero and a reduction of 0.3 standard deviations or more. For context, a two-group trial of 80 participants has about 15 percent power to detect a standardized mean difference of 0.2, and a two-sided sign test for 10 of 12 studies favouring one direction gives P = 0.04. Explain why the authors' conclusion is not justified, state what a vote count based on the direction of effect shows, and describe a figure that would present these trials more informatively, including what its axes or columns would show.

Model answer

The authors used vote counting by statistical significance, which treats each non-significant trial as evidence of no effect. A non-significant result only means that a trial could not distinguish the effect from zero. These trials were small: a trial of 80 participants has about 15 percent power to detect a small effect, so even if the program works, most such trials will be non-significant. Hedges and Olkin (1980) showed that when power is this low, a majority-rule count becomes less likely to detect a real effect as more trials are added. The supplementary table supports this reading, since most confidence intervals include a worthwhile reduction as well as zero.

A count based on direction shows that 10 of the 12 trials (83 percent) favour the program, and the sign test gives P = 0.04, so a split this uneven would be unlikely if the program had no effect. The count says nothing about size, so the review should also summarize the standardized mean differences with a median and interquartile range.

An ordered estimate plot without a pooled result would show each trial as a row, with its standardized mean difference and confidence interval on a horizontal axis, a vertical line at zero, and rows ordered by risk of bias and sample size. Readers would see that the estimates cluster on the side favouring the program. A harvest plot by outcome domain would be an alternative strong answer.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: A trial of 100 participants tests a program whose true standardized mean difference is 0.2. About how likely is the trial to reach statistical significance at the 0.05 level?

Power for a small effect in a trial of 100 participants is about 17 percent, so most such trials of an effective program will be non-significant. A trial of 400 participants has about 52 percent power, and 5 percent is the false-positive rate when there is no effect.

Question 2: What did Hedges and Olkin (1980) show about majority-rule vote counting when individual studies have low power?

When power is below one half, the proportion of significant studies tends toward that power as studies accumulate, so the chance that most are significant falls toward zero. Adding studies therefore makes a majority-rule count more likely to miss a real effect.

Question 3: In an effect direction plot, what does a sideways arrow mean?

A sideways arrow marks a domain in which a study's outcomes conflict, for example depressive symptoms improving while anxiety worsens. Arrows follow the direction of effect regardless of significance, so option a would reintroduce counting by significance; domains that were not measured are left blank.

Question 4: In the Cedar Valley harvest plot, what do bar height and bar colour show?

Tall bars are randomized trials and short bars non-randomized controlled studies; red bars are at high or serious risk of bias. The column in which a bar sits shows the direction of effect, and the letter above it identifies the study.
Section 4 of 5

When Pooling Is Justified

⏱ Estimated reading time: 35 minutes
Section 4 of 5

When Pooling Is Justified

This section covers what pooling adds, five questions before pooling and the Cedar Valley decisions.

35 minutes
Adds and cannot fix

What pooling does

It adds

Pooling adds precision, an effect size with a confidence interval, and tools for variation.

It cannot fix

Pooling cannot fix shared bias, an average of unlike programs, or missing studies.

Five questions

Before pooling

  • The studies address the same question.
  • Compatible data are available for each study.
  • The designs are compatible and kept apart where they differ.
  • Risk of bias is acceptable or handled.
  • Pooling was planned, with enough studies to estimate variation.
Misconceptions

Beliefs that lead to poor decisions

Diversity

Some diversity is expected, and the test is whether the average is meaningful.

Variation

Large variation calls for investigation and a report of the spread of effects.

Cedar Valley

Pool or not, comparison by comparison

Group-based

Pooling would be justified in a full review, but the protocol planned none.

One-to-one

Five usable trials split by comparator are too few for a useful pooled estimate.

Connector

Seriously confounded and differently adjusted studies should not be pooled.

Statements and next steps

From synthesis to the evidence brief

Group-based

These programs may reduce loneliness slightly (low certainty).

One-to-one

These programs may reduce loneliness (low certainty).

Connector

The evidence is very uncertain (very low certainty).

Complete the reflection and the final assessment below.

Learning Objectives for this section

  • Explain what a meta-analysis adds to a synthesis and what problems pooling cannot correct.
  • Apply five questions to decide whether pooling the results of a group of studies is justified.
  • Correct common misconceptions about heterogeneity, study similarity and the purpose of pooling.
  • Justify a decision to pool or not to pool for each comparison in a review, and report the decision transparently.
  • Identify what HSCI 230 Lesson 2 teaches about meta-analytic methods and how the tables built in this lesson prepare the data for them.

Introduction

The methods in Sections 2 and 3 describe a body of evidence without combining its results into a single number. A meta-analysis goes further: it combines the effect estimates from several studies into one weighted average, with a confidence interval, using statistical methods. Gene Glass (1976) coined the term to describe "the analysis of analyses", and the method is now standard in reviews of the effects of health interventions. This section does not teach how to compute a meta-analysis. HSCI 230 Lesson 2 does that, covering fixed-effect and random-effects models, forest plots, statistical heterogeneity, publication bias and influential studies. This section teaches the decision that comes first: whether pooling is justified for a given group of studies, and how to report the decision either way.

4.1 What Pooling Adds and What It Cannot Fix

Pooling adds three things. First, it combines the information in several imprecise studies into one more precise estimate, which addresses the problem of small studies that Section 3 described. The eight Cedar Valley trials of group-based programs, none with more than 230 participants, would together give an estimate based on 1,236. Second, it estimates the size of the effect with a confidence interval, which a direction count cannot do. Third, it provides statistical tools for describing how much results vary between studies and for exploring why, as well as for examining signs of publication bias.

Pooling cannot fix three other problems. It cannot remove bias from the included studies: if every trial overstates the benefit because participants who know their group report less loneliness, the pooled estimate will overstate it too, with a narrower confidence interval that makes the error look more certain. It cannot give meaning to an average of things that should not be averaged: a pooled effect of "community programs" that mixes group activities, befriending calls and connector referrals would describe no program anyone could implement. And it cannot recover studies that were never published or found, which is why the searching methods in Lessons 4 and 5 matter. These limits explain why the decision to pool depends on judgement about the studies as well as on whether the numbers are available.

4.2 Five Questions Before Pooling

The Cochrane Handbook treats the decision as a matter of whether a meta-analysis would be meaningful and appropriate for a given synthesis, considering the diversity of the studies, the data available and the risk of bias (Deeks, Higgins and Altman, 2019; McKenzie and Brennan, 2019). Those considerations can be organized as five questions, which the Cedar Valley team asked of each comparison.

1. Do the studies address the same question?v

The studies should share the population, intervention, comparator and outcome defined for the synthesis (Section 1.2), closely enough that an average effect would answer a question someone is asking. Reviewers call differences in these features clinical diversity; in public health the same idea applies to programs, settings and populations. Some diversity is expected and acceptable, and the question is whether the average would mean something. A pooled effect of group programs for older adults living in the community answers a planning question; a pooled effect of every community intervention does not.

2. Are compatible data available?v

Each study must provide an effect estimate and a measure of its precision (a standard error or confidence interval), or data from which these can be calculated, in a metric that can be converted to a common one. Studies that report only "no significant difference" or only a P value without direction cannot contribute. If many studies lack usable data, a meta-analysis of the rest may be unrepresentative, and a direction count or effect direction plot, which can include more studies, may give a fairer picture.

3. Are the designs compatible?v

Randomized and non-randomized studies are subject to different biases, and Cochrane guidance advises against combining them in one meta-analysis; when both are included, they are analyzed separately. Within a design, problems of the unit of analysis must be handled, such as cluster trials analyzed as if individuals had been randomized and trials with several arms sharing one control group (Section 1.4). Non-randomized studies should contribute estimates adjusted for important confounders.

4. Is the risk of bias acceptable or handled?v

If most studies are at high risk of bias in the same direction, a pooled estimate will be precise and probably wrong. Reviews can restrict the main analysis to studies at lower risk or run a sensitivity analysis that excludes studies at high risk, both of which are planned in the protocol (Lesson 8). A review should not pool studies at critical risk of bias by ROBINS-I, which the tool's developers advise excluding from synthesis.

5. Was pooling planned, and are there enough studies?v

The protocol should state whether and how effects will be pooled, so that the choice is not made after seeing which option gives the more favourable answer. Two studies are enough to compute a pooled estimate, but with very few studies the amount of variation between them cannot be estimated reliably, which affects random-effects results; HSCI 230 Lesson 2 explains why. With few studies, reviewers should report the pooled result with caution or describe the studies individually.

1. Do the studies address the same question? 2. Are compatible data available? 3. Are the designs compatible? 4. Is risk of bias acceptable or handled? 5. Was pooling planned, with enough studies to estimate variation? YesYesYesYesYes Pool, using the methods of HSCI 230 Lesson 2 NoNoNoNoNo Synthesize without meta-analysis Sections 2 and 3: summaries, direction counts and plots, reported with SWiM
Five questions to ask of each group of studies before pooling. A "no" at any step leads to synthesis without meta-analysis, and the review reports which question was decisive. The questions summarize considerations in the Cochrane Handbook (Deeks, Higgins and Altman, 2019; McKenzie and Brennan, 2019).

4.3 Common Misconceptions About Pooling

Students and some review authors hold beliefs about pooling that lead to poor decisions in both directions. The cards describe five of them.

Studies must be identical to poolClick to explore
High heterogeneity forbids poolingClick to explore
Pooling rescues non-significant studiesClick to explore
Pooling corrects for biasClick to explore
Without pooling there is no conclusionClick to explore

4.4 Worked Example: Should the Cedar Valley Team Pool?

The team applied the five questions to each comparison for loneliness. The table records its reasoning.

Comparison (controlled studies)Questions 1 to 5Would pooling be justified?
Group-based programs versus usual activities or no program (8 randomized trials; 1,236 participants)Similar programs, populations and comparators; standardized mean differences available for all 8, with clustering handled in Trials C and G; all randomized; no trial at low risk of bias, with Trials G and H at high risk, so a sensitivity analysis excluding them would be needed; not planned in the protocol.Yes in a full systematic review, with a random-effects model and a sensitivity analysis by risk of bias. Not done in this review because the protocol planned none.
One-to-one befriending or telephone programs versus usual care or a waiting list (6 randomized trials; 742 participants)Similar programs; two kinds of comparator, which would need separate analysis because waiting-list comparisons may inflate self-reported benefit; Trial L has no usable data, leaving 5; three trials at high risk of bias; not planned.Possible in principle, but with 5 usable trials split by comparator, each analysis would rest on very few studies. Synthesis without meta-analysis is more appropriate here.
Community connector programs versus usual care (5 non-randomized controlled studies; 1,480 participants)Programs differ in referral pathways, connector caseloads and duration; adjusted estimates available in 4 of 5, using different sets of confounders; three studies at serious risk of confounding; designs cannot be combined with trials.No. A pooled estimate would combine differently adjusted, seriously confounded results and would suggest more precision than the evidence holds.
Intergenerational (2) and technology-based (3) programsDiverse programs and outcomes; very few studies; most at serious risk of bias.No.
All 24 controlled studies as one groupDifferent programs, comparators and designs; the average would not inform the choice among programs.No.

The first row needed the most thought. The group-based trials meet the conditions that the Cochrane Handbook sets out, and a full systematic review of these programs would normally pool them. The team considered amending its protocol to add a meta-analysis and decided against it for two reasons. The decision facing the planning team concerns community connector programs, for which pooling is not justified, and a precise estimate for group programs would not change that decision. And an amendment made after the results table had been seen would be open to the criticism that the method was chosen by the result. The review reports that the data for group-based programs would support a meta-analysis and recommends one in any full review. A team with different aims might reasonably decide the other way, provided it reports the change and its reason, as SWiM item 1b and PRISMA 2020 require.

4.5 From Synthesis to Statements for Decision-Makers

The synthesis ends in statements that combine the findings from Sections 2 and 3 with the GRADE ratings from Lesson 8, written in the standard wording of Santesso and colleagues (2020). The table gives the Cedar Valley statements for loneliness that will go into the evidence brief in Lesson 12.

ComparisonSynthesis findingsStatement and certainty
Group-based programs7 of 8 trials favour the program; median standardized mean difference −0.19 (interquartile range −0.24 to −0.12)Group-based programs may reduce loneliness slightly among older adults living in the community (low certainty).
One-to-one befriending or telephone programs5 of 5 trials with estimable results favour the program; median −0.20; one trial with no usable dataOne-to-one befriending or telephone programs may reduce loneliness (low certainty).
Community connector programs5 of 5 controlled studies favour the program; mixed results for mental health and service useThe evidence is very uncertain about the effect of community connector programs on loneliness (very low certainty).
Intergenerational and technology-based programs3 of 5 controlled studies favour the program; most at serious risk of biasToo few controlled studies to support a statement; certainty not rated.

These statements show how far a careful synthesis without meta-analysis can go. The planning team learns which kinds of program have the most consistent evidence, roughly how large their effects appear to be, how confident it can be, and that the evidence for the model it intends to launch is very uncertain. Lesson 8 drew the practical implication: the Cedar Valley program should be launched with an evaluation built in, so that the program adds to the evidence about community connector programs.

4.6 Connecting to Meta-Analysis

Students who take HSCI 230 Lesson 2, before or after this course, will find that the work of this lesson is the preparation a meta-analysis needs. The groups for synthesis define which studies enter each meta-analysis. The decision rules choose one result per study. The results table, with standardized mean differences, confidence intervals and risk-of-bias judgements, is the data set the meta-analysis uses, and the ordered estimate plot in Figure 3.1 becomes a forest plot once a pooled estimate is added. The exploration of variation in Section 2.4 becomes subgroup analysis, and the restriction to studies at lower risk becomes a sensitivity analysis. HSCI 230 Lesson 2 adds the statistical model, the weights and the heterogeneity statistics, and its section on when pooling is inappropriate returns to the questions asked here.

Reflection

A review team is deciding whether to pool results for three groups of studies of programs for older adults. Group A: 7 randomized trials of telephone befriending compared with usual care, all measuring loneliness; standardized mean differences with confidence intervals are available for all 7; 2 trials are at high risk of bias and 5 have some concerns; the protocol planned a random-effects meta-analysis if at least five trials were comparable. Group B: 4 non-randomized studies of community workshop programs for older men; 2 measure loneliness and all 4 measure well-being with different scales; all report only unadjusted estimates, and 3 are at serious risk of confounding. Group C: a team member proposes pooling all 11 studies from Groups A and B into one estimate of the effect of "community programs" on loneliness. The five questions for pooling are: (1) Do the studies address the same question? (2) Are compatible data available? (3) Are the designs compatible? (4) Is risk of bias acceptable or handled? (5) Was pooling planned, with enough studies to estimate variation? Apply the questions to each group, decide whether pooling is justified, and state what the review should do instead where it is not.

Model answer

Group A. The trials share a population, intervention, comparator and outcome, so an average effect of telephone befriending on loneliness would be meaningful (question 1). Effect estimates with confidence intervals are available for all seven (question 2), and all are randomized (question 3). Risk of bias can be handled with a sensitivity analysis that excludes the two trials at high risk (question 4), and the protocol planned a meta-analysis with a threshold this group meets (question 5). Pooling is justified, using the methods in HSCI 230 Lesson 2, with heterogeneity reported and the sensitivity analysis shown.

Group B. Only two studies measure loneliness, the estimates are unadjusted, three studies are at serious risk of confounding, and the well-being scales differ. Questions 2 and 4 fail, and question 5 fails for loneliness. Pooling is not justified. The review should tabulate the four studies, present an effect direction plot for loneliness and well-being, and report the synthesis with SWiM.

Group C. Pooling all eleven fails questions 1 and 3: it mixes different programs and combines randomized with non-randomized designs, which Cochrane guidance advises against. The average would describe no program anyone could deliver. The review should keep the groups separate and, if readers need an overview, use a harvest plot across the two categories.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: Which problem does pooling leave uncorrected?

Pooling combines imprecise estimates into a more precise one, gives an estimate with a confidence interval and offers tools for describing variation. It cannot remove a bias that the studies share; pooling biased studies produces a precise estimate of the wrong answer.

Question 2: Why did the Cedar Valley team judge pooling unjustified for the community connector studies?

Three of the five non-randomized studies were at serious risk of confounding, their adjusted estimates used different sets of confounders, and the programs differed in referral pathways and intensity. Two studies are enough to compute a pooled estimate, and SWiM is a reporting guideline that does not forbid any method.

Question 3: What does Cochrane guidance advise about randomized trials and non-randomized studies of the same intervention?

Randomized and non-randomized studies are subject to different biases, so Cochrane guidance advises against combining them in one meta-analysis and recommends separate analyses. Reviews may include both designs, as the Cedar Valley review does.

Question 4: A review's results show large variation between trials of similar programs. What is the most appropriate response?

Variation is a reason to investigate its sources, using characteristics planned in the protocol, and to report the spread of effects with any summary. A random-effects model allows for variation without removing it, and describing studies one at a time abandons synthesis.
Section 5 of 5

Final Assessment

⏱ Estimated time: 25 minutes

Bringing It All Together

This lesson followed the Cedar Valley evidence team from extracted data to statements about effect. The team first made the evidence visible: a characteristics table described every study, a cross-tabulation showed that all 14 randomized trials concerned group or one-to-one programs while the 12 community connector studies were all non-randomized or uncontrolled, and results tables set out one result per study, chosen by rules written in advance and aligned in direction. Groups for synthesis, defined by population, intervention, comparator, outcome and design, kept unlike comparisons apart.

Because the review did not pool results, it used the methods that the Cochrane Handbook accepts for synthesis without meta-analysis and reported them against the nine SWiM items. Summaries of standardized mean differences described the size and spread of effects, and counts by the direction of effect, with sign tests, replaced the misleading practice of counting studies by statistical significance, which in the Cedar Valley trials would have reported two successes in eight and none in six. Effect direction plots and harvest plots displayed these counts alongside design, sample size and risk of bias.

Finally, five questions about the question, the data, the designs, the risk of bias and the plan determined whether pooling was justified. The group-based trials could support a meta-analysis of the kind HSCI 230 Lesson 2 teaches, while the connector studies could not. Combined with the GRADE ratings from Lesson 8, the synthesis produced clear statements for decision-makers: group and one-to-one programs may reduce loneliness, and the evidence for community connector programs is very uncertain.

Key Takeaways from this lesson

  • Synthesis brings the results of several studies together to answer a review question, and a quantitative synthesis begins with structured tables that make the studies comparable.
  • A characteristics table describes every study, a cross-tabulation shows where evidence is concentrated and where gaps lie, and a results table sets out one result per study for each comparison and outcome.
  • Groups for synthesis are defined by population, intervention, comparator, outcome and design, planned in the protocol, and any later change is reported with its reason.
  • Decision rules written before extraction select one result per study when studies report several scales, time points or analyses, and every effect is aligned so that the same sign means the same thing.
  • When effects are not pooled, the Cochrane Handbook accepts summarizing effect estimates, combining P values and vote counting based on the direction of effect, each of which answers a different question.
  • The SWiM guideline asks reviews to report nine items covering grouping, metrics, synthesis methods, prioritization, heterogeneity, certainty, data presentation, results and limitations.
  • Vote counting by statistical significance misleads because small studies have low power, non-significant results are not evidence of no effect, and the error grows as underpowered studies accumulate.
  • Vote counting based on the direction of effect, reported with the proportion favouring the program and a sign test, is acceptable but ignores the size and precision of each result.
  • Effect direction plots display direction across several outcome domains for each study, and harvest plots display where studies fall across categories with design and risk of bias encoded in each bar.
  • Pooling is justified when the studies address the same question, provide compatible data, share a design, have risk of bias that is acceptable or handled, and were planned for pooling in sufficient number; HSCI 230 Lesson 2 teaches the computation.

Core Concepts Reviewed

Section 1: synthesis, characteristics tables, cross-tabulations, results tables, groups for synthesis, lumping and splitting, decision rules for selecting results, alignment of direction and the standardized mean difference.

Section 2: narrative synthesis, reasons for synthesis without meta-analysis, summarizing effect estimates with the median and interquartile range, combining P values, the nine SWiM items and the investigation of variation without meta-analysis.

Section 3: statistical significance, statistical power, the failure of vote counting by significance shown by Hedges and Olkin, vote counting based on direction with a sign test, effect direction plots and harvest plots.

Section 4: what meta-analysis adds and cannot fix, five questions before pooling, the separation of randomized and non-randomized designs, misconceptions about heterogeneity, and the link to HSCI 230 Lesson 2.

The final reflection asks you to explain the lesson's main ideas to a decision-maker who has read a misleading headline about the evidence.

Reflection

The director of a health authority planning a community connector (social prescribing) program for older adults has read a news story headlined "Social prescribing doesn't work: most studies found no significant effect". The director asks you, the review team's intern, for a reply of about 200 words. Your review found the following for loneliness. Group-based programs: 8 randomized trials, 2 statistically significant, 7 of 8 point estimates favouring the program, median standardized mean difference −0.19 (interquartile range −0.24 to −0.12), low certainty. One-to-one befriending or telephone programs: 6 randomized trials, none statistically significant, 5 of the 5 trials with estimable results favouring the program, low certainty. Community connector programs: 5 non-randomized controlled studies, all favouring the program, three at serious risk of confounding, very low certainty. The review did not pool results, because its protocol planned no meta-analysis and pooling was not justified for the connector studies. Write the reply. It should explain what is wrong with the reasoning in the headline, state what the review found for each kind of program and how certain those findings are, explain why the review did not produce a single pooled number, and say what the findings imply for the planned program.

Model answer

The headline counts studies by statistical significance. Most studies of these programs are small, and a small study will often miss a real but modest effect, so "no significant effect" in a single study usually means that the study could not tell whether the program worked. Our review counted studies by the direction of their results and summarized the size of their effects instead.

Seven of eight trials of group-based programs found less loneliness in the program group, and the typical difference was small (a median of about 0.2 standard deviations); these programs may reduce loneliness slightly, with low certainty. All five befriending or telephone trials with usable results pointed toward benefit, also with low certainty, although none was significant on its own. All five controlled studies of community connector programs also pointed toward benefit, but they were non-randomized and mostly at serious risk of confounding, so the evidence is very uncertain.

We did not combine the results into one number because the connector studies differ too much and are too open to bias for an average to be trustworthy, and our protocol planned no meta-analysis.

The evidence leaves the effect of connector programs uncertain. The planned program should therefore be launched with an evaluation built in, for example by introducing it across clinics in stages, so that Cedar Valley adds to the evidence.

Minimum 30 characters required.

✓ Reflection saved

Final Knowledge Assessment

Final Assessment, this lesson: Synthesizing Quantitative Findings (15 Questions)

Question 1: How many Cedar Valley studies does this lesson's quantitative synthesis cover, and where are the qualitative findings synthesized?

The quantitative synthesis covers the 33 quantitative studies and the quantitative strands of the 3 mixed-methods studies, 36 in all. The 6 qualitative studies and the qualitative strands go to the qualitative evidence synthesis in Lesson 10. The 24 controlled studies are the subset used for statements about effect.

Question 2: A review team places telephone befriending and home visiting in separate groups for synthesis, although each group then holds only two studies. What kind of grouping choice is this?

Splitting places studies in narrow groups, giving specific answers at the cost of few studies per synthesis. Lumping would combine the two kinds of program into one larger group. The right level depends on the decision the review supports.

Question 3: Which SWiM items are reported in the methods section of a review?

Items 1 to 7 cover grouping, metrics, synthesis methods, prioritization, heterogeneity, certainty and data presentation, all reported in the methods. Item 8 (reporting results) belongs in the results and item 9 (limitations of the synthesis) in the discussion.

Question 4: A review states that "three of ten trials found a significant benefit, so the evidence is inconclusive". What is the main flaw in this reasoning?

Counting by significance treats non-significant results as evidence of no effect, although small trials often lack power to detect a real effect and their estimates may still favour the program. The review should count by direction and summarize effect sizes.

Question 5: A sign test on 5 studies that all favour the program gives a two-sided P value of 0.06. What is the best interpretation?

With five studies, the smallest possible two-sided P value is 2 in 32, or about 0.06, so five of five is the most extreme split possible. Reading P above 0.05 as no effect would repeat the error of significance counting, and a direction count says nothing about size.

Question 6: Which display best shows the direction of effect in several outcome domains for each of ten studies that used different measures?

An effect direction plot has a row for each study and a column for each outcome domain, with arrows for direction and arrow size for sample size. An ordered estimate plot needs a common metric, which studies with different measures across many outcomes often lack.

Question 7: In the Cedar Valley effect direction plot, Studies Q and T recorded more unplanned health service visits among program participants. How did the team handle these results?

The protocol defined fewer unplanned visits as benefit, so more visits were recorded as downward arrows. The text then noted that connectors may link people to care they had been going without. Redefining direction after seeing results would be a post hoc change.

Question 8: Why did the Cedar Valley team place the uncontrolled studies below a labelled divider in its effect direction plot?

Under SWiM item 4, the team prioritized controlled studies for statements about effect, since a change over time without a comparison group cannot be attributed to the program. The uncontrolled studies remain visible so that readers see all the evidence.

Question 9: Trial L reported only "no significant difference" with no numbers. How should a synthesis treat it?

A study without an estimable direction keeps its row as not estimable and is reported alongside the direction counts, as in "five of the five trials with estimable results". Leaving it out makes the evidence look more complete and consistent than it is.

Question 10: What does pooling in a meta-analysis add that a direction count cannot provide?

A meta-analysis estimates the size of the average effect with a confidence interval, which a direction count cannot do. It cannot remove shared bias, cannot use studies without data, and its result may or may not be significant.

Question 11: Why did the Cedar Valley team decide against pooling the group-based trials, even though the data would support a meta-analysis?

The protocol planned no meta-analysis, an amendment after seeing the results table could appear to choose the method by the result, and a pooled estimate for group programs would not change the decision about connector programs. Different scales can be handled with standardized mean differences, and high-risk trials can be handled with a sensitivity analysis.

Question 12: Which GRADE certainty level matches the statement "Group-based programs may reduce loneliness slightly among older adults living in the community"?

In the standard wording of Santesso and colleagues (2020), "may" signals low certainty. High certainty uses a plain verb ("reduces"), moderate uses "probably", and very low certainty is stated as "the evidence is very uncertain".

Question 13: When a review later carries out a meta-analysis, what role do its groups for synthesis, decision rules and results tables play?

The groups decide which studies enter each meta-analysis, the decision rules choose one result per study, and the results table is the data set; HSCI 230 Lesson 2 adds the model, weights and heterogeneity statistics. None of these steps removes the need for appraisal or guarantees an unbiased result.

Question 14: Which pattern in the Cedar Valley harvest plot matters most for the planning team?

The tall bars, the randomized trials, sit in the group and one-to-one rows, while the connector row holds only non-randomized studies, three of them at serious risk of bias. The model Cedar Valley plans to launch has the least secure evidence, which supports building an evaluation into the launch.

Question 15: A student argues that because the studies in a group differ somewhat in populations and program length, pooling is never appropriate. What is the best reply?

Meta-analysis expects some diversity, and the random-effects model assumes that true effects vary. The question is whether an average effect would answer a meaningful question, alongside the other questions about data, design, risk of bias and planning.
✦ Complete the final reflection above before submitting

Congratulations!

You have successfully completed this lesson: Synthesizing Quantitative Findings.

You can now tabulate and group the studies in a review, choose one result per study by rules set in advance, summarize effects without pooling them, count studies by direction instead of significance, draw effect direction plots and harvest plots, report the synthesis against the SWiM items, and decide whether a meta-analysis would be justified.

Lesson 10 turns to scoping reviews, rapid reviews and qualitative evidence synthesis. It returns to the JBI charting method and PRISMA-ScR, examines the shortcuts that rapid reviews take, and introduces thematic synthesis, framework synthesis and meta-ethnography, with GRADE-CERQual, for the qualitative studies that this lesson set aside.

Reference

Glossary: Key Terms, People & Frameworks

📚 Reference page, available throughout the lesson

These terms, tools and people appear in this lesson on synthesizing quantitative findings.

Core Concepts
Synthesis The process of bringing the results of several studies together to answer a review question. A quantitative synthesis combines numerical results, and a qualitative synthesis combines qualitative findings.
Characteristics table A table that describes each included study in one row, including its design, setting, participants, program, comparator, outcomes and time points.
Cross-tabulation A table that counts studies in the cells formed by two characteristics, such as intervention category by design, showing where evidence is concentrated and where gaps lie.
Results table A table that lists one result per study for a single comparison and outcome, in a consistent form, with the information needed to weigh each result.
Groups for synthesis Sets of studies defined by a specific population, intervention, comparator, outcome and often design, each synthesized separately; the Cochrane Handbook calls this the PICO for each synthesis.
Lumping and splitting The trade-off between placing broadly similar studies in one large group (lumping) and placing them in narrow groups that give specific answers with fewer studies each (splitting).
Decision rules Rules written before extraction that select one result per study when a study reports several scales, time points or analyses, to prevent selective choice.
Standardized mean difference The difference in mean scores between two groups divided by a standard deviation, which expresses results from different scales in standard deviation units.
Interquartile range The span of the middle half of a set of values, from the 25th to the 75th percentile, used with the median to describe the spread of effect estimates.
Narrative synthesis A structured approach to synthesizing findings mainly in words and tables, described by Popay and colleagues (2006); the term is also used loosely for any synthesis that is not a meta-analysis.
Synthesis without meta-analysis Synthesis of quantitative effects that does not pool them statistically, using methods such as summarizing effect estimates or vote counting based on direction.
Statistical significance The property of a result whose P value falls below a chosen threshold, usually 0.05, or whose 95 percent confidence interval excludes the value meaning no effect.
Statistical power The probability that a study will produce a statistically significant result when an effect of a given size truly exists; small studies of small effects have low power.
Vote counting by statistical significance Tallying studies as significant benefit, significant harm or no significant difference and concluding from the tallies, a practice that misleads and that Cochrane advises against.
Vote counting based on direction of effect Counting how many studies favour the intervention and how many favour the comparison, regardless of size or significance, and reporting the proportion with a sign test.
Sign test A test of whether the split of studies between two directions is more uneven than would be expected if each study were equally likely to point either way.
Meta-analysis A statistical method that combines effect estimates from several studies into a weighted average with a confidence interval; HSCI 230 Lesson 2 teaches its methods.
Clinical diversity Differences between studies in their participants, interventions, comparators or outcomes, which affect whether an average effect would be meaningful.
Heterogeneity Variation in results between studies beyond what would be expected from chance alone; without meta-analysis, reviews explore it by tabulating results against planned characteristics.
Unit-of-analysis error An error that arises when the unit analyzed differs from the unit randomized or counted, as in a cluster trial analyzed as if individuals had been randomized, or a control group used twice.
Frameworks & Tools
SWiM (Synthesis Without Meta-analysis) A reporting guideline of nine items for reviews that synthesize quantitative effects without meta-analysis, used alongside PRISMA 2020 (Campbell et al., 2020).
Effect direction plot A grid with a row for each study and a column for each outcome domain, in which arrows show the direction of effect and arrow size shows sample size (Thomson and Thomas, 2013; Hilton Boon and Thomson, 2021).
Harvest plot A matrix in which each study is drawn as a bar in the cell matching its finding, with bar height and shading showing study characteristics such as design and risk of bias (Ogilvie et al., 2008).
Ordered estimate plot A plot of each study's effect estimate and confidence interval on a common metric, ordered by a characteristic such as risk of bias, without a pooled estimate.
Albatross plot A plot of each study's P value against its sample size, used to display results from studies that report their findings in different ways (Harrison et al., 2017).
Combining P values Methods such as Fisher's method that combine exact P values from several studies to test whether there is an effect in at least one study, without estimating its size.
Cochrane Handbook, Chapter 12 The chapter of the Cochrane Handbook for Systematic Reviews of Interventions on synthesizing and presenting findings using methods other than meta-analysis (McKenzie and Brennan, 2019).
Five questions before pooling This lesson's summary of whether pooling is justified: the same question, compatible data, compatible designs, acceptable or handled risk of bias, and a plan with enough studies.
Key People
Mhairi Campbell A researcher at the University of Glasgow who led the 2019 assessment of how reviews report synthesis without meta-analysis and the development of the SWiM guideline (2020).
Joanne E. McKenzie A biostatistician at Monash University and lead author, with Sue E. Brennan, of the Cochrane Handbook chapter on synthesis without meta-analysis, and of its chapter on defining groups for synthesis.
Jennie Popay A professor of sociology and public health at Lancaster University who led the 2006 guidance on the conduct of narrative synthesis in systematic reviews.
Hilary Thomson A public health researcher at the University of Glasgow who developed the effect direction plot with Sian Thomas (2013) and revised it with Michele Hilton Boon (2021).
David Ogilvie A public health researcher at the University of Cambridge and lead author of the 2008 paper that introduced the harvest plot.
Larry V. Hedges A statistician whose 1980 paper with Ingram Olkin showed that vote counting by significance can become less likely to detect a real effect as studies accumulate.
Ingram Olkin A statistician at Stanford University who co-authored, with Larry Hedges, the 1980 critique of vote counting and the 1985 book Statistical Methods for Meta-Analysis.
Gene V. Glass An educational researcher who introduced the term meta-analysis in 1976, describing it as the analysis of analyses.
Douglas G. Altman A British medical statistician who, with J. Martin Bland, wrote the 1995 note "Absence of evidence is not evidence of absence".
No matching entries. Try a different search term.