Scoping Reviews, Rapid Reviews and Qualitative Evidence Synthesis
Finding & Synthesizing Health Evidence
Learning objectives for this lesson:
- Explain when a scoping review is the appropriate design, and describe the Arksey and O'Malley (2005) framework, the refinements of Levac, Colquhoun and O'Brien (2010) and the JBI method.
- Complete a charting table and a descriptive numerical summary, and report a scoping review with the PRISMA-ScR checklist.
- Summarize the Cochrane Rapid Reviews Methods Group guidance (Garritty et al., 2021) and distinguish shortcuts that narrow what a review can find from shortcuts that reduce independent checking.
- Calculate the workload that a rapid review shortcut saves, weigh it against the risk the shortcut carries, and report rapid review methods transparently.
- Describe thematic synthesis, framework synthesis and meta-ethnography, and choose among them for a given question, audience, timeline and body of studies.
- Assess confidence in qualitative review findings with the four components of GRADE-CERQual and present the results in a summary of qualitative findings table.
- Build an evidence and gap map from a charting table, reconcile its counts, and interpret its clusters and gaps with appropriate caution.
- Write a rapid review shortcuts table that compares a protocol with what was done and states the main risk of each shortcut.
This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University, drawing on the Cochrane Handbook for Systematic Reviews of Interventions and the JBI Manual for Evidence Synthesis.
Scoping Reviews: The JBI Method, Charting and PRISMA-ScR
Learning Objectives for this section
- Describe the purposes for which a scoping review is the appropriate design, and distinguish a scoping review from a systematic review of effects.
- Describe the five stages and optional consultation stage of the Arksey and O'Malley (2005) framework and the refinements that Levac, Colquhoun and O'Brien (2010) recommended.
- Outline the JBI method for scoping reviews, including the PCC question, the three-step search and the descriptive approach to analysis.
- Complete a charting table and produce a descriptive numerical summary of the included sources.
- Report a scoping review with the PRISMA-ScR checklist and explain how its items differ from those of PRISMA 2020.
Introduction
Lessons 7 to 9 took the Cedar Valley evidence team from search results to included studies, completed forms, appraisal judgements and a synthesis of quantitative results. This lesson turns to three kinds of review that a health sciences student will meet often: the scoping review, the rapid review and the qualitative evidence synthesis. It ends with the evidence and gap map, a picture of where evidence exists and where it is missing. Lesson 1 introduced these review types, and this lesson shows how each is carried out.
Before launching a community connector (social prescribing) program for older adults, the planning team of the fictional Cedar Valley Health Authority in British Columbia asked a small evidence team (an evidence officer, a university librarian and a student intern) for a rapid scoping review and environmental scan within twelve weeks. The review asks which community-based interventions have been evaluated for reducing loneliness or social isolation among adults aged 65 and older, and with what outcomes. It includes 42 studies in 44 reports (38 from the databases and 4 from citation chasing): 14 randomized trials, 10 non-randomized controlled studies, 9 uncontrolled before-and-after studies, 6 qualitative studies and 3 mixed-methods studies. Numbers beyond these shared figures are illustrative.
1.1 What a Scoping Review Is For
A scoping review maps the extent, range and nature of the evidence on a topic: what kinds of evidence exist, how the topic has been studied and where the gaps lie. A systematic review of effects asks a narrower question, usually whether an intervention works and by how much. Munn and colleagues (2018) described six purposes for which a scoping review suits: to identify the types of evidence available in a field, to clarify key concepts or definitions, to examine how research on a topic has been conducted, to identify key characteristics or factors related to a concept, to serve as a precursor to a systematic review, and to identify and analyze gaps in knowledge.
The Cedar Valley question fits several of these purposes, because the planning team wants to know which types of program have been evaluated, with which outcome measures and in which groups of older adults. It also wants to know whether the programs work, and Lesson 8 showed how the team gave a provisional answer for three intervention categories while labelling that work as going beyond a typical scoping review.
| Feature | Scoping review | Systematic review of effects |
|---|---|---|
| Typical question | What has been studied, how, in whom, and where are the gaps? | Does the intervention work, for whom, and by how much? |
| Question framework | PCC (Population, Concept and Context) | PICO (Population, Intervention, Comparison and Outcome) |
| Appraisal of sources | Optional and usually omitted | Required, with tools matched to design |
| Analysis | Descriptive counts and basic content analysis | Synthesis of effects, with or without meta-analysis |
| Reporting guideline | PRISMA-ScR (Tricco et al., 2018) | PRISMA 2020 (Page et al., 2021) |
Both designs start from a protocol, search systematically and select sources with explicit criteria applied by more than one person, so the label of scoping review carries the same expectation of a documented and reproducible search.
1.2 The Arksey and O'Malley Framework and the Levac Refinements
Hilary Arksey and Lisa O'Malley (2005) published the first methodological framework for what they called scoping studies, which they saw as a way to examine the extent, range and nature of research activity, to decide whether a full systematic review would be worthwhile, to summarize and disseminate findings, and to identify gaps in the literature. Their framework has the five stages shown in Figure 1.1 and an optional sixth, a consultation exercise, in which people with practical knowledge of the topic suggest sources, comment on the findings and add their perspectives. They did not assess the quality of included studies, which they saw as consistent with mapping a field.
Danielle Levac, Heather Colquhoun and Kelly O'Brien (2010) drew on three scoping studies to recommend refinements at every stage (Figure 1.1). They asked teams to link the purpose of the review to its question, to balance breadth against feasibility, to have two reviewers select studies independently, and to develop the charting form together, with two people charting the first five to ten studies. They divided the fifth stage into analysis, reporting and consideration of implications for research, practice and policy, and argued that consultation should be an essential stage with a clear purpose.
The Cedar Valley team consulted twice. In week 2 it met the planning team, two community connectors, two members of the older adults' advisory group and a member of the Indigenous health team to settle the scope (Lesson 2). In week 10 it showed the same group its descriptive summary and a draft evidence and gap map. The connectors noted that the charting form had no field for whether a program offered transport, which the team then added and charted from the program descriptions.
1.3 The JBI Method
JBI (an evidence synthesis organization based at the University of Adelaide, formerly the Joanna Briggs Institute) built on these frameworks to produce detailed guidance for scoping reviews (Peters et al., 2015; Peters et al., 2020), with later guidance on protocols (Peters et al., 2022) and on extracting, analyzing and presenting results (Pollock et al., 2023). The team registers a protocol before searching and frames its question with PCC (Population, Concept and Context) and the types of evidence source to be included (Lesson 2). The search follows three steps: a limited search of at least two databases, an analysis of the words and index terms of relevant records, and a full search with all identified terms in every database, followed by searching the reference lists of included sources. At least two reviewers select sources, appraisal is not generally required, and the analysis is mainly descriptive. The table sets out the JBI steps with the Cedar Valley decision for each.
| JBI step | Cedar Valley decision |
|---|---|
| Define and align the objectives and question | A PCC question with five sub-questions (Lesson 2) |
| Develop and align the inclusion criteria | Adults aged 65 and older in the community; high-income countries; reports from 2010 onward |
| Describe the planned approach in a protocol | Protocol registered on OSF Registries before searching |
| Search for the evidence | Five databases, grey literature, websites and citation chasing (Lessons 4 and 5) |
| Select the evidence | Dual screening at both stages (Lesson 7) |
| Extract (chart) the evidence | A piloted form built on the sub-questions, TIDieR and PROGRESS-Plus (Lesson 8) |
| Analyze the evidence | Counts by design, category, country and outcome; categories for referral pathways |
| Present the results | A charting table, summary tables and an evidence and gap map (Section 4) |
| Summarize, conclude and note implications | Implications for program design and evaluation, without graded recommendations |
Critical appraisal in a scoping review
JBI guidance does not generally require scoping reviews to appraise their sources, because the purpose is to describe what evidence exists. A team may still appraise when its users need to know how strong the designs are, as the Cedar Valley team did (Lesson 8), provided it states the reason in the protocol and does not present the scoping review as a test of effectiveness.
1.4 Charting the Data
Data charting is the JBI term for extracting information in a scoping review. It is usually more descriptive than extraction for a review of effects, because the questions concern what was done, with whom and how it was measured. It is also more iterative, because a team often learns only while charting which categories its sources need, so JBI guidance asks teams to pilot the form, revise it as charting proceeds and report the changes. Every field should answer a sub-question, and every sub-question should have at least one field. The completed forms become the charting table, with one row per source and one column per field. The table shows six rows of the Cedar Valley charting table, with columns shortened for display.
| Study and country | Design and participants | Intervention category and delivery | Loneliness or isolation measure | Other outcomes | Equity characteristics reported | Implementation information |
|---|---|---|---|---|---|---|
| CV-004, United Kingdom | Randomized trial; 156 randomized, 139 analyzed | One-to-one telephone: trained volunteers call weekly for 12 weeks; general practice referral | Three-item loneliness scale (Hughes et al., 2004) | Depressive symptoms; general practice visits | Gender, living alone, area deprivation | Proportion of planned calls completed |
| CV-017, Canada | Cluster-randomized trial in 12 seniors' centres; 240 randomized, 210 analyzed | Group activity: weekly volunteer-led walks for 16 weeks | Six-item De Jong Gierveld Loneliness Scale | Well-being; physical activity | Gender, living alone | Session attendance |
| CV-023, United Kingdom | Non-randomized controlled study; 420 participants | Community connector: primary care referral, three to six meetings, links to community groups | Revised UCLA Loneliness Scale | Well-being; general practice visits | Gender, ethnicity, area deprivation | Uptake after referral; connector cost |
| CV-028, Netherlands | Non-randomized controlled study; 96 participants | Intergenerational: weekly visits by university students for six months | Eleven-item De Jong Gierveld scale; Lubben Social Network Scale | Life satisfaction | Gender only | Not reported |
| CV-031, Australia | Before-and-after study; 58 participants | Technology-based: tablet and video-calling lessons in eight weekly sessions | Revised UCLA Loneliness Scale; Lubben Social Network Scale | Well-being | Gender, education, language at home | Retention; participants' ratings of the training |
| CV-036, Canada (British Columbia) | Qualitative: interviews with 18 participants and 6 connectors | Community connector: community agency program with primary care referral | Participants' accounts of loneliness and new contacts | Not applicable | Gender, rural residence | Views on referral, transport and program endings |
The table uses closed categories where the team will count, keeps the authors' terms where readers need detail, and records missing information as "not reported", which is itself a finding for the sub-question on implementation. The identifiers link each row to its full form, appraisal and reports.
An evaluation report describes a weekly "tech café" in a small British Columbia town, where library volunteers help older adults use smartphones to call family. It reports a before-and-after survey of 31 participants using the three-item loneliness scale, attendance figures, and the participants' ages and genders, but no other characteristics. Write its row in the seven columns of the Cedar Valley charting table, using "not reported" where appropriate.
1.5 Summarizing the Charted Data
JBI guidance describes the analysis in most scoping reviews as descriptive (Pollock et al., 2023). A descriptive numerical summary counts sources by the characteristics charted, and basic qualitative content analysis sorts charted text into categories, such as the referral pathways of connector programs. Content analysis of this kind summarizes what sources say without the interpretation that a qualitative evidence synthesis aims for (Section 3). The table cross-tabulates the 42 Cedar Valley studies by category and design.
| Intervention category | Randomized trials | Non-randomized controlled | Before-and-after | Qualitative | Mixed methods | Total |
|---|---|---|---|---|---|---|
| Group activity | 8 | 0 | 3 | 2 | 1 | 14 |
| One-to-one befriending or telephone | 6 | 0 | 1 | 1 | 0 | 8 |
| Community connector or social prescribing | 0 | 5 | 3 | 2 | 2 | 12 |
| Intergenerational | 0 | 2 | 1 | 1 | 0 | 4 |
| Technology-based | 0 | 3 | 1 | 0 | 0 | 4 |
| Total | 14 | 10 | 9 | 6 | 3 | 42 |
The other counts answer the remaining sub-questions. The studies came from the United Kingdom (11), the United States (9), Canada (6, two from British Columbia), Australia (5), the Netherlands (4), other European countries (5) and Japan (2). Of the 36 studies with quantitative outcomes, 33 measured loneliness, most often with the Revised UCLA Loneliness Scale (16) or a De Jong Gierveld scale (10), and less often with the three-item short scale derived from the UCLA scale (5) or a single question (2); the other 3 measured social isolation only. Gender was reported in 39 studies, living alone in 27, income or education in 18, ethnicity in 11, rural residence in 7, disability in 4 and language in 3. Nine of the 12 connector studies described referral from primary care, 2 described self-referral or community referral, and 1 did not say. The counts show that connector programs have been evaluated only with non-randomized and uncontrolled designs, and that few studies describe participants' income, ethnicity, language or disability.
1.6 Reporting with PRISMA-ScR
PRISMA-ScR (the Preferred Reporting Items for Systematic reviews and Meta-Analyses extension for Scoping Reviews) was published by Andrea Tricco and colleagues (2018) after a consensus process with an expert panel. It has 20 essential items and 2 optional items, the optional ones being critical appraisal of the included sources in the methods (item 12) and results (item 16). It also changes the vocabulary of PRISMA, speaking of sources of evidence because a scoping review may include documents that are not research studies, of the data charting process in place of data extraction, and of critical appraisal in place of risk of bias. Lesson 7 adapted the flow diagram, and Lesson 12 returns to the full report.
The title identifies the report as a scoping review, the abstract is structured, the rationale explains why a map is needed, and the objectives state the question with its PCC elements. The Cedar Valley title names both the design and its rapid methods, and the objectives list all five sub-questions.
The methods report the protocol and registration, eligibility criteria, information sources, the full search for at least one database, the selection process, the data charting process and data items, any critical appraisal, and the methods for summarizing the data. The Cedar Valley report states that the charting form was piloted and that a transport field was added after consultation.
The results give the selection numbers with a flow diagram, the characteristics of the sources, any appraisal results, the relevant data from each source and the summary that answers each sub-question. The Cedar Valley report puts the charting table in an appendix and the map in the main text.
The discussion summarizes the evidence, states the limitations of the scoping process and gives conclusions with implications, and the last item reports funding. The Cedar Valley report states that the planning team agreed the scope but took no part in selection or charting.
What Carries Forward
The Cedar Valley team now has a charting table and a descriptive summary that answer its sub-questions. It produced them in twelve weeks by streamlining parts of the method, and the next section examines the rapid review shortcuts that made this possible and what each one costs.
Reflection
A student is planning a rapid scoping review for a health authority on peer support programs for new parents in rural British Columbia. The PCC question is as follows. Population: parents of infants under one year who live in rural communities. Concept: peer support programs delivered in person, by telephone or online, and the outcomes their evaluations report. Context: high-income countries. There are three sub-questions: (1) What types of peer support program have been evaluated, and how are they delivered? (2) Which outcomes and measures have been used? (3) To what extent do evaluations describe fathers, Indigenous parents and parents with low incomes? Design a charting form of eight to ten fields that answers these sub-questions. For each field, state whether it is a closed field (give its allowed values) or a free-text field, and name the sub-question it serves. Then explain which two PRISMA-ScR items are optional and whether you would complete them for this review, with a reason.
(1) Source identifier and linked reports (free text; all sub-questions). (2) Country and setting (closed: Canada, United States, United Kingdom, Australia, other high-income country; rural only or mixed rural and urban; sub-question 1). (3) Design (closed: randomized trial, non-randomized controlled, before-and-after, qualitative, mixed methods, program report; sub-question 2). (4) Program type (closed: lay peer support, professionally led peer group, online peer forum, other; sub-question 1). (5) Delivery (closed: in person, telephone, online, combined; plus free-text description of who delivers it and how often; sub-question 1). (6) Outcomes measured (free text listing each outcome with its instrument, for example the Edinburgh Postnatal Depression Scale; sub-question 2). (7) Outcome domain (closed: parental mental health, breastfeeding, parenting confidence, social support, service use, experience; sub-question 2). (8) Fathers included (closed: yes, no or not reported; sub-question 3). (9) Indigenous parents described (closed: yes, no or not reported; sub-question 3). (10) Income or socioeconomic status reported (closed: yes, no or not reported; sub-question 3).
The optional PRISMA-ScR items are item 12 (critical appraisal of individual sources, methods) and item 16 (critical appraisal within sources, results). I would not complete them, because the health authority asked what has been evaluated and for whom, and appraisal is not generally required for that purpose. I would state this decision in the protocol. If the authority later asked whether peer support improves maternal mental health, a review of effects with appraisal would be needed.
Minimum 20 characters required.
Question 1: A planning team asks which kinds of community programs for loneliness have been evaluated, in which settings, and with which outcome measures. Which design fits this question best?
Question 2: What did Levac, Colquhoun and O'Brien (2010) recommend about the consultation stage of the Arksey and O'Malley framework?
Question 3: Which statement about PRISMA-ScR is accurate?
Question 4: While charting, two reviewers record different intervention categories for a program that combines weekly home visits with a monthly group lunch. What does the JBI approach to charting suggest?
Rapid Reviews: The Cochrane Guidance and the Trade-offs of Shortcuts
Learning Objectives for this section
- Define a rapid review and explain why decision-makers commission rapid reviews.
- Summarize the Cochrane Rapid Reviews Methods Group guidance (Garritty et al., 2021) for each stage of a review, and note the main changes in its 2024 update.
- Distinguish shortcuts that narrow what a review can find from shortcuts that reduce independent checking, and describe the risk each carries.
- Calculate the workload saved by a screening shortcut and judge whether the saving justifies the risk.
- Report the shortcuts in a rapid review so that readers can judge their likely effect on the findings.
Introduction
A full systematic review commonly takes a year or more. Decision-makers who must choose a program, a policy or a clinical approach within weeks cannot wait that long, and they often commission a rapid review instead. Lesson 1 chose a rapid scoping review for the fictional Cedar Valley project, and Lessons 2, 7 and 8 mentioned shortcuts the team considered. This section examines what the main guidance recommends, the trade-off each shortcut makes, and how a team chooses and reports its shortcuts.
2.1 What Makes a Review Rapid
Hamel and colleagues (2021) analyzed published definitions and proposed that a rapid review is a form of knowledge synthesis that speeds up the process of a systematic review by streamlining or omitting some of its methods, so that evidence can be produced with fewer resources. Two features follow from this definition. A rapid review keeps the logic of a systematic review, with a protocol, an explicit search, stated eligibility criteria and a structured synthesis. And it departs from the standard methods at identified points, which the team chooses in advance and reports.
Rapid reviews vary widely, and a scoping review of rapid review methods found many different shortcuts, often poorly described (Tricco et al., 2015). A practical guide from the World Health Organization's Alliance for Health Policy and Systems Research emphasizes working closely with knowledge users, the managers, clinicians, policy-makers and community members who will act on the findings (Tricco, Langlois and Straus, 2017). For Cedar Valley, they are the planning team and, through the advisory group, the older adults the program will serve.
2.2 The Cochrane Rapid Reviews Methods Group Guidance
Cochrane established a Rapid Reviews Methods Group to develop methods for rapid reviews. The group published interim guidance in 2020, and the journal version appeared the following year (Garritty et al., 2021). The guidance makes 26 recommendations, drawn from methodological studies of the effect of shortcuts and from the experience of rapid review producers, and it is intended mainly for rapid reviews of the effects of interventions. The table summarizes its recommendations by stage, in paraphrase.
| Stage | What the 2021 guidance recommends (paraphrased) |
|---|---|
| Question and eligibility | Involve knowledge users in setting and refining the question, eligibility criteria and outcomes. Limit the number of interventions, comparators and outcomes to those that matter most for the decision. Use date limits only with a clear justification, and consider limiting the review to English-language publications unless other languages are justified. Consider giving priority to stronger designs, for example by looking first for existing systematic reviews. |
| Searching | Involve an information specialist, and consider peer review of the search strategy. Search CENTRAL (the Cochrane Central Register of Controlled Trials), MEDLINE and Embase where available, with no more than one or two specialized databases. Limit grey literature and supplementary searching. |
| Title and abstract screening | Pilot a screening form with the whole team on the same 30 to 50 abstracts. Have two reviewers screen at least 20 percent of abstracts independently, with conflicts resolved. One reviewer then screens the remaining abstracts, and a second reviewer screens all the abstracts excluded. |
| Full-text screening | Pilot a full-text form on 5 to 10 articles. One reviewer screens the full texts, and a second reviewer screens all those excluded. |
| Data extraction | Use a piloted form and limit extraction to a minimal set of required items. One reviewer extracts, and a second checks for accuracy and completeness. Consider using data from existing systematic reviews. |
| Risk of bias | Use a valid tool for each design. Limit appraisal to the most important outcomes. One reviewer appraises, and a second verifies every judgement. |
| Synthesis and certainty | Synthesize the evidence narratively, and use meta-analysis only where the studies are similar enough to pool. Where certainty is rated with GRADE, one reviewer rates and a second verifies the judgements and their reasons. |
The group updated its guidance for rapid reviews of effectiveness in 2024, reducing the recommendations to 24 (Garritty et al., 2024). The update keeps the structure of the 2021 guidance. Among its changes, it suggests that a team may move to single screening after a dual-screened sample of about 20 percent of records only when the two reviewers agree closely, it recommends narrative synthesis reported with the SWiM guideline (Lesson 9), and it asks every rapid review to describe its restricted methods and discuss their possible effect on the findings. Teams should cite the version they followed.
2.3 Two Kinds of Shortcut
The recommendations contain two kinds of shortcut, and the distinction helps a team judge their risks. Some shortcuts narrow what the review can find: a narrower question, fewer databases, date and language limits, and less grey literature. Their risk is that relevant studies never enter the review, and the risk is larger when the missing studies differ systematically from those found, for example when negative results are more often published in reports or in other languages. Other shortcuts reduce independent checking: single screening, single extraction and single appraisal, each with some verification. Their risk is that the team makes errors in handling studies it has found, such as excluding an eligible study or copying a number incorrectly. Figure 2.1 sets out where each kind of shortcut falls.
2.4 The Trade-off Each Shortcut Makes
Methodological studies have tested several shortcuts by applying them to completed systematic reviews and asking whether the conclusions would have changed, or by comparing reviewers in controlled experiments. The accordion summarizes what is known, in general terms.
Nussbaumer-Streit and colleagues (2018) examined a sample of Cochrane reviews to see which of their included studies abbreviated searches, such as a few major databases combined with checking reference lists, would have found, and concluded that most reviews would have reached the same conclusions, although not all. The risk is greater for topics whose literature sits outside the major biomedical databases. Loneliness research appears in nursing, psychology, gerontology and social science journals, which is why the Cedar Valley team kept CINAHL, PsycINFO and Web of Science.
Nussbaumer-Streit and colleagues (2020) found that excluding publications in languages other than English rarely changed the conclusions of the reviews they examined. The finding may not hold for topics studied mainly in other languages, and a review of programs designed for one country loses little by excluding them. Date limits save screening time and are defensible when the intervention or its context has changed, or when earlier reviews cover the older studies; they are hard to defend when older trials remain the best evidence.
Screening errors are common. A methodological review found that single screening missed more eligible studies than dual screening (Waffenschmidt et al., 2019), and in a randomized trial of abstract screening, single reviewers missed 13 percent of relevant studies (Gartlehner et al., 2020). A second reviewer who screens every excluded record protects against the error that matters most, a missed study, at little saving in time, as the worked example shows.
Buscemi and colleagues (2006) found that single extraction with verification by a second person produced more errors than independent double extraction, although it took less time. The errors matter most in numerical results, which is why the Cedar Valley team extracted the results of its comparative studies twice (Lesson 8).
Appraising only the most important outcomes, with one reviewer and a verifier, reduces time with a modest risk of inconsistent judgements. Narrative synthesis is the usual approach in rapid reviews and needs the discipline described in Lesson 9 to avoid vote counting. Marshall and colleagues (2019) simulated several rapid review methods on existing meta-analyses and found that rapid methods sometimes produced results that differed from those of the full reviews, which is why a rapid review should report its shortcuts as limitations.
Decision-makers differ in how much uncertainty they will accept for speed. In an international survey, guideline developers and policy-makers reported that they would accept some loss of certainty in exchange for a faster answer, within limits (Wagner et al., 2017). A team should therefore ask its knowledge users which errors would matter most to them. A planning team choosing between program models might accept missing an older trial of a model it has already ruled out, but would be poorly served by missing a recent evaluation of the model it intends to launch.
2.5 Worked Example: The Cedar Valley Shortcuts
The Cedar Valley protocol (Lesson 2) stated its shortcuts in advance, and the team amended two of them as the work proceeded. The table compares the protocol with the 2021 guidance and with what the team did.
| Step | 2021 Cochrane suggestion | Cedar Valley protocol | What the team did | Main risk and response |
|---|---|---|---|---|
| Question | Involve knowledge users; limit interventions and outcomes | PCC question set with the planning team; community settings in high-income countries; loneliness and isolation as main outcomes | As planned, with a second consultation in week 10 | Evidence from long-term care and lower-income settings lies outside the map; stated as a limitation |
| Dates and languages | Limits only with justification; English unless justified | Reports from 2010 onward; English and French, with other languages listed | As planned | Older and non-English evaluations are missing; existing reviews cited for earlier work |
| Databases and grey literature | Core databases and one or two others; limited grey literature | Five databases, targeted websites and citation chasing | As planned; 26 documents and 4 studies added | Little risk from this step; the time was recovered elsewhere |
| Title and abstract screening | Dual screening of at least 20 percent, then one reviewer with all exclusions checked | Dual screening of the first 374 records, then single screening of 1,496 with exclusions checked | Amended on 20 February 2026 (end of week 5) to dual screening of all 1,870 records (Lesson 7) | No residual risk from this shortcut; amendment dated and reported |
| Full-text screening | One reviewer, with exclusions checked | Two reviewers independently | As planned for 142 full texts | None from a shortcut |
| Charting | One extracts a minimal set; a second checks | One charts; a second verifies | Amended in week 7: results fields of the 24 comparative studies charted twice, and authors contacted when key results were missing; the rest charted once and verified (Lesson 8) | Errors in descriptive fields; reduced by verification |
| Appraisal | One appraises key outcomes; a second verifies | Optional; design-matched tools for description | Two reviewers appraised the 14 trials independently (Lesson 8) | Appraisal describes the evidence and was not used to exclude studies |
| Synthesis and certainty | Narrative synthesis; GRADE with verification | Descriptive mapping; no certainty ratings | Provisional GRADE ratings for loneliness in three comparisons (Lesson 8) and a thematic synthesis with GRADE-CERQual (Section 3) | Additions labelled as going beyond a scoping review, with reasons |
The screening amendment illustrates how a team can test a shortcut against its own numbers. The calculation below counts screening decisions, treating each reading of a record by one reviewer as one decision. In the first 374 records, which both reviewers screened as the protocol planned, 345 (92 percent) were excluded, and the calculation assumes that the same proportion would be excluded among the remaining records. The final count bore this out: 1,725 of the 1,870 records (92 percent) were excluded at title and abstract.
How much does the screening shortcut save?
Dual screening of every record: 2 × 1,870 = 3,740 decisions.
Protocol plan: 20 percent of 1,870 = 374 records screened by two reviewers (748 decisions); the remaining 1,496 screened by one reviewer (1,496 decisions); a second reviewer checks the excluded records among those, about 1,496 × 345 ÷ 374 = 1,380 decisions. Total: 748 + 1,496 + 1,380 = 3,624 decisions, a saving of 116 decisions, or about 3 percent.
Single screening after the dual sample, with no check of exclusions: 748 + 1,496 = 2,244 decisions, 40 percent fewer than dual screening, at the cost of the missed studies that single screening produces.
The comparison shows why the team amended its plan. When nine records in ten are excluded, checking every exclusion costs almost as much as screening every record twice, so the protocol's shortcut saved little. Larger savings would have required single screening without a check, which carries the risk documented by Gartlehner and colleagues (2020). With 1,870 records, the team judged that the extra 116 decisions were affordable, and on 20 February 2026, at the end of week 5, it amended its protocol to screen every remaining record in duplicate. A second amendment in week 7 had the results fields of the 24 comparative studies charted twice, for the reason Lesson 8 gives. At the full-text stage the arithmetic differs, because a larger share of reports is included. Of the 142 full-text reports, 40 (describing 38 studies) were included, so checking only the 102 exclusions would have saved 40 of 142 second readings, or about 28 percent.
A health authority asks two analysts for a rapid review of the effects of school-based programs to prevent vaping among students aged 12 to 18, due in four weeks. A preliminary search suggests about 3,000 records across MEDLINE, Embase, PsycINFO, ERIC and CINAHL, and a systematic review published three years ago. List three shortcuts you would accept and one you would refuse, and give a reason for each that refers to the kind of risk it carries (narrowing what can be found or reducing independent checking). Then write one sentence for the limitations section that reports your shortcuts.
2.6 Reporting a Rapid Review
A rapid review is reported with PRISMA 2020 for a review of effects or PRISMA-ScR for a rapid scoping review. The title or abstract should identify the review as rapid, the methods should state each shortcut where it occurs (for example, "one reviewer screened 80 percent of titles and abstracts after dual screening of a 20 percent sample"), and the flow diagram keeps the standard boxes (Lesson 7). The limitations should explain the likely effect of each shortcut, and the protocol and amendments should be available, so that a decision-maker can weigh speed against certainty.
What Carries Forward
The Cedar Valley review is now complete in its quantitative and descriptive parts. Nine of its sources, the 6 qualitative studies and the qualitative strands of the 3 mixed-methods studies, describe what older adults and providers experienced. The next section shows how a team synthesizes such evidence and judges how much confidence each finding deserves.
Reflection
A provincial health ministry asks two analysts for a rapid review, due in five weeks, of whether school-based programs reduce vaping among students aged 12 to 18. A preliminary search of five databases suggests about 2,600 records, of which about 95 percent are likely to be excluded at title and abstract, and a systematic review of the topic was published three years ago. The team is considering five shortcuts: (a) searching only from the existing review's search date onward and using the review as the source of earlier studies; (b) restricting to English-language reports; (c) dual screening a 20 percent sample of titles and abstracts, then single screening the rest with a second reviewer checking every exclusion; (d) single data extraction with verification by the second analyst; and (e) skipping risk-of-bias assessment entirely. For each shortcut, say whether it narrows what the review can find or reduces independent checking, and whether you would accept it, with a reason. Then calculate the number of screening decisions under shortcut (c) and under full dual screening, assuming that 95 percent of the singly screened records are excluded, and comment on the saving.
(a) Narrows what can be found. I would accept it if the existing review is trustworthy and its search was comprehensive, because it covers the earlier studies; I would check its methods first. (b) Narrows what can be found. I would accept it, since methodological studies suggest language restrictions rarely change conclusions and the ministry needs evidence transferable to Canadian schools, and I would list non-English reports found. (c) Reduces independent checking. I would accept it only with a pilot and good agreement in the 20 percent sample. (d) Reduces independent checking. I would accept it, with both analysts extracting the outcome data independently because errors there change the answer. (e) This removes a step entirely. I would refuse it, because a question about effects needs an assessment of bias; I would instead have one analyst appraise the main outcome and the other verify.
Under (c), 20 percent of 2,600 = 520 records are screened twice (1,040 decisions), 2,080 are screened once (2,080 decisions), and about 0.95 × 2,080 = 1,976 exclusions are checked, for 1,040 + 2,080 + 1,976 = 5,096 decisions. Full dual screening needs 2 × 2,600 = 5,200, so the shortcut saves only 104 decisions, about 2 percent. With so many exclusions, I would dual screen everything and save time elsewhere.
Minimum 20 characters required.
Question 1: According to the 2021 Cochrane Rapid Reviews Methods Group guidance, how should title and abstract screening be done in a rapid review?
Question 2: Which of the following shortcuts narrows what a review can find, as distinct from reducing independent checking?
Question 3: The Cedar Valley protocol planned to dual screen 20 percent of the 1,870 titles and abstracts. How many records would two reviewers screen, and how many would one reviewer screen alone?
Question 4: Why would checking every excluded record have saved the Cedar Valley team so little time at the title and abstract stage?
Qualitative Evidence Synthesis and GRADE-CERQual
Learning Objectives for this section
- Explain what qualitative evidence synthesis adds to a review of interventions and the kinds of question it answers.
- Describe the stages of thematic synthesis (Thomas and Harden, 2008) and apply them to a small set of qualitative findings.
- Describe framework synthesis and best fit framework synthesis, and explain when an existing framework helps or constrains a synthesis.
- Describe the seven phases of meta-ethnography (Noblit and Hare, 1988) and its three forms of synthesis.
- Assess confidence in a review finding with the four components of GRADE-CERQual and present the result in a summary of qualitative findings table.
Introduction
Decision-makers need to know how people experience a program, why some take part and others do not, and what makes it feasible to deliver. A qualitative evidence synthesis brings together the findings of several qualitative studies, which analyze data such as interviews and focus groups, to give an account that no single study could give. This section introduces thematic synthesis, framework synthesis and meta-ethnography, and the GRADE-CERQual approach for judging confidence in the findings. HSCI 841 Lesson 2 builds on this section for graduate students, adding critical interpretive synthesis, qualitative meta-summary, the choice of method with the RETREAT criteria and a worked GRADE-CERQual assessment.
3.1 What Qualitative Evidence Synthesis Adds
Cochrane guidance describes several roles for qualitative evidence in reviews of interventions (Noyes et al., 2019). A synthesis can explore how people experience an intervention, how acceptable and feasible it is, how it is implemented, and why it may work for some groups and not for others, and it can identify outcomes that matter to participants but were not measured in trials.
A scoping review usually stops at the basic content analysis described in Section 1. In week 8, the Cedar Valley planning team asked what older adults and the people who deliver programs say helps or hinders connection. The team amended its protocol, with the date and reason recorded, to add a brief thematic synthesis of its 9 qualitative sources (the 6 qualitative studies and the qualitative strands of the 3 mixed-methods studies), assessed with GRADE-CERQual and labelled as going beyond a standard scoping review. The sources include about 150 older adults and about 40 volunteers, connectors and program staff.
3.2 Three Methods Compared
Methods of qualitative evidence synthesis differ in how much they interpret, how far they rely on an existing framework, and how much time and expertise they need. The tabs compare the three methods named in this lesson.
Origin. James Thomas and Angela Harden (2008) described thematic synthesis from their work at the EPPI-Centre in London on reviews of children's views about healthy eating. The method adapts thematic analysis of primary data to the findings of published studies.
Steps. The reviewers code the text of each study's findings line by line, organize the codes into descriptive themes that stay close to the studies, and then develop analytical themes that go beyond the studies to answer the review question.
Suits. Reviews that must inform intervention design within a fixed time.
Origin. Framework synthesis adapts framework analysis, a method for applied policy research (Ritchie and Spencer, 1994), and was developed for reviews at the EPPI-Centre (Brunton, Oliver and Thomas, 2020). Carroll, Booth and Cooper (2011) described a variant called best fit framework synthesis.
Steps. The reviewers index each study's findings against a framework chosen before coding, chart them in a matrix, and interpret the patterns. In the best fit variant, findings that do not fit are analyzed thematically to create new themes.
Suits. Questions for which a relevant model or theory already exists.
Origin. George Noblit and R. Dwight Hare (1988) developed meta-ethnography in education research to synthesize ethnographic studies.
Steps. The reviewers work through seven phases, the central one being the translation of studies into one another, in which the concepts and metaphors of each study are compared with those of the others to build a new interpretation.
Suits. Reviews that aim to develop new concepts or theory from rich studies, with an experienced team and more time.
Other methods exist, including JBI's meta-aggregation, which groups findings into categories and synthesized findings that can guide action without reinterpreting the primary studies (Lockwood, Munn and Porritt, 2015). Booth and colleagues (2018) proposed seven criteria for choosing a method, known as RETREAT: the review question, epistemology, time, resources, expertise, audience and purpose, and type of data. The Cedar Valley team chose thematic synthesis because its audience needed practical implications within four weeks, its sources were mostly short evaluations with modest depth, and its members had general qualitative training.
3.3 Thematic Synthesis in Practice
Thomas and Harden (2008) treated the text of each study's results or findings section, including participants' quotations and the authors' interpretations, as the data for synthesis. The Cedar Valley team followed their three stages.
Stage 1: line-by-line coding. The intern and the evidence officer coded the same two sources, agreed on a shared approach, then divided the remaining seven and checked each other's coding. Each code describes the meaning of a passage, such as "same volunteer every week", and the nine sources produced 71 codes.
Stage 2: descriptive themes. The reviewers grouped related codes and named each group, producing six descriptive themes that stay close to what the studies reported: a consistent and trusted person, programs that end, practical barriers to getting there, being accompanied at the start, roles that let people contribute, and interests before labels.
Stage 3: analytical themes. The reviewers then asked what the descriptive themes meant for the planning team's question about program design. This step goes beyond the content of the studies, and Thomas and Harden described it as the hardest stage to describe and the one most dependent on the reviewers' judgement and insight. It produced three analytical themes, each with an implication for the Cedar Valley program, as Figure 3.1 shows.
The analytical themes suggest design choices: keeping the same connector with each person and planning how contact ends, budgeting for transport and accompaniment to a first activity, and framing referrals around interests and roles in which older adults help others. These are hypotheses for the program to test, because they rest on experiences in other programs.
3.4 Framework Synthesis and the Best Fit Approach
Framework synthesis starts from a framework chosen or built before coding. Its stages follow framework analysis (Ritchie and Spencer, 1994): familiarization, identifying a framework, indexing the findings against it, charting them in a matrix, and mapping and interpretation. In best fit framework synthesis, the team codes findings against the published model closest to its question and analyzes the findings that do not fit to create new themes, producing a revised framework (Carroll, Booth and Cooper, 2011; Carroll et al., 2013).
To compare the methods, the intern coded the same 71 codes against a published classification of loneliness interventions. Masi and colleagues (2011), in a meta-analysis of interventions to reduce loneliness, grouped them by four strategies: improving social skills, enhancing social support, increasing opportunities for social contact, and addressing maladaptive social cognition (the habitual negative thoughts about social relationships that can sustain loneliness). The table shows how the Cedar Valley findings fitted.
| Framework category | Cedar Valley findings indexed to it | Fit |
|---|---|---|
| Improving social skills | Confidence to join conversations after learning new skills in a group, from two sources | Limited |
| Enhancing social support | Relationships with volunteers and connectors, continuity, the loss felt when contact ended | Good |
| Increasing opportunities for social contact | Group activities, links to community groups made by connectors, shared interests | Good |
| Addressing maladaptive social cognition | One passage on worry about being a burden to others | Little data |
| New themes from findings outside the framework | Practical barriers to getting there; being accompanied at the start; contributor roles; reluctance to be called lonely | Added |
The framework organized the findings quickly and linked them to established research, yet about a third of the codes did not fit. They concerned access and identity, matters of delivery that a classification of intervention strategies was not designed to capture, and a team that indexed only against the framework would have lost the findings most useful to the planning team. The best fit approach guards against this by requiring the team to analyze what does not fit.
3.5 Meta-ethnography
Meta-ethnography is the most interpretive of the three methods, and its central task is to translate the concepts of one study into the terms of another (Noblit and Hare, 1988). Britten and colleagues (2002), in a worked example on how people take medicines, distinguished first-order constructs (participants' own understandings), second-order constructs (the study authors' interpretations) and third-order constructs (the reviewers' new interpretations). The accordion lists Noblit and Hare's seven phases.
The team identifies an area of interest that qualitative research can inform and decides which studies to include. Meta-ethnographies often use purposive searches and select studies for the richness of their concepts, which differs from the exhaustive search of a scoping review.
The team reads each study repeatedly, records its key concepts and metaphors, lists them side by side, and judges whether the studies describe the same things in different terms, contradict each other, or describe different parts of a larger picture.
The team compares each study's concepts with those of the others, usually in chronological order or starting from an index study, and records how each concept is expressed in each study.
The team brings the translations together into a new interpretation. Noblit and Hare described three forms. A reciprocal translation applies when studies describe similar things, and expresses them in shared concepts. A refutational synthesis applies when studies contradict each other, and explores and explains the contradiction. A lines-of-argument synthesis applies when studies describe different aspects of a phenomenon, and assembles them into a whole, much as a researcher builds an interpretation from parts of a single ethnography.
The team communicates the synthesis in a form its audience can use, such as a conceptual model or a set of propositions. The eMERGe reporting guidance (France et al., 2019) sets out what a meta-ethnography report should contain.
The Cedar Valley findings show how the three forms would apply. Two studies described "a familiar face" and "someone who remembers my week", which a reciprocal translation would express as continuity of a known person. One study described participants who welcomed a group openly aimed at loneliness while three described reluctance to be labelled lonely, a contradiction that a refutational synthesis would try to explain. Studies of referral, first attendance, staying involved and endings each covered part of a person's path through a program, which a lines-of-argument synthesis would assemble into a model. The team judged its sources too thin and its time too short for a full meta-ethnography.
3.6 How Much Confidence? GRADE-CERQual
A finding from a qualitative evidence synthesis also needs a statement of how much confidence a reader can place in it. GRADE-CERQual (Confidence in the Evidence from Reviews of Qualitative research) was developed by Simon Lewin and colleagues for this purpose (Lewin et al., 2015; Lewin et al., 2018). It assesses each review finding separately, as GRADE assesses each outcome (Lesson 8), and it asks whether the finding is a reasonable representation of the phenomenon of interest. The flip cards describe its four components.
For each component, the reviewers judge whether there are no or very minor, minor, moderate or serious concerns. They then judge overall confidence, starting from the assumption that the finding is a reasonable representation of the phenomenon and moving down as concerns accumulate. The four levels are high (it is highly likely that the finding is a reasonable representation), moderate (likely), low (possible) and very low (it is not clear whether it is). The developers have also discussed dissemination bias as a possible further component. The results are presented in a summary of qualitative findings table, of which the Cedar Valley version follows.
| Summarized review finding | Contributing sources | Confidence | Explanation |
|---|---|---|---|
| Older adults valued meeting or speaking with the same volunteer or connector over time, and several described the end of a program as a loss. | 7 (5 qualitative, 2 mixed methods) | Moderate | Minor concerns about methodological limitations (two studies gave little detail on recruitment) and about adequacy (several reported the point briefly); no or very minor concerns about coherence and relevance. |
| Transport, cost, and hearing or mobility difficulties limited attendance at group programs, and being accompanied to a first session helped people join. | 6 (4 qualitative, 2 mixed methods) | Moderate | Moderate concerns about methodological limitations (three sources had MMAT concerns about sampling or analysis); no or very minor concerns about coherence, adequacy and relevance, with two sources from small towns similar to Cedar Valley's communities. |
| Some participants did not want to be described as lonely and preferred programs framed around shared interests or roles in which they could help others. | 4 (3 qualitative, 1 mixed methods) | Low | Moderate concerns about coherence (one source described participants who welcomed a program openly aimed at loneliness) and about adequacy (thin data); minor concerns about methodological limitations and relevance. |
| Connectors described heavy caseloads and few community activities in small towns as reasons why some referrals did not lead to new social contacts. | 2 (1 qualitative, 1 mixed methods) | Low | Serious concerns about adequacy (two sources, with brief data in one); minor concerns about methodological limitations and relevance (one source from British Columbia and one from Australia); no or very minor concerns about coherence. |
Confidence differs between findings from the same synthesis, so each finding needs its own rating. The findings also complement Lesson 8, which found with low certainty that group programs may reduce loneliness slightly and one-to-one programs may reduce it; the qualitative evidence suggests, with moderate confidence, that continuity and practical access shape whether people take part and stay. Syntheses can be reported with ENTREQ, Enhancing Transparency in Reporting the Synthesis of Qualitative Research (Tong et al., 2012).
A finding states that "older adults using video-calling programs valued a patient volunteer when learning the technology". Of its three studies, two have no important limitations and one recruited only participants nominated by staff. All three report the point consistently with detailed quotations. Two were in community settings in Canada and one in a residential care home in Japan, while the question concerns community-dwelling older adults in British Columbia. Judge the four CERQual components, give an overall confidence level, and write a one-sentence explanation.
What Carries Forward
The Cedar Valley team now has a descriptive summary, a record of its rapid methods and four rated qualitative findings. Section 4 turns the charting table into an evidence and gap map.
Reflection
A team is synthesizing seven qualitative studies of older adults' experiences of telephone befriending services for a health authority that must decide within ten weeks whether to fund such a service. Five studies are short evaluation reports with brief findings, and two are in-depth interview studies; the team members have general qualitative training. (1) Choose thematic synthesis, framework synthesis or meta-ethnography, and justify the choice with reference to the question, the time, the team's expertise, the audience and the data. (2) Assess this finding with GRADE-CERQual: "Participants valued calls from the same volunteer and felt let down when volunteers changed." It draws on four studies. Two of them had concerns about recruitment, because staff selected the participants. All four report the finding consistently, but three report it in one or two sentences with a single quotation. Three were conducted with community-dwelling older adults in Canada, and one with people recently discharged from hospital in England, while the review question concerns community-dwelling older adults in Canada. For each of the four components, choose no or very minor, minor, moderate or serious concerns; give an overall confidence level (high, moderate, low or very low); and write the explanation for a summary of qualitative findings table.
(1) I would choose thematic synthesis. The health authority needs practical implications for a funding decision within ten weeks, and the analytical themes of a thematic synthesis can be directed at that question. The team has general qualitative skills, while meta-ethnography needs more experience and time and suits the development of new theory. Most of the studies are short reports with brief findings, which give too little conceptual depth for translation between studies. Framework synthesis would be a strong alternative if a suitable model of befriending existed, using the best fit approach so that findings outside the model are kept.
(2) Methodological limitations: moderate concerns, because two of the four studies let staff select participants, which may favour satisfied users. Coherence: no or very minor concerns, because all four report the finding consistently. Adequacy of data: moderate concerns, because three studies report the point briefly with one quotation each. Relevance: minor concerns, because one study concerns people discharged from hospital in England. Overall: low confidence. Explanation: "Moderate concerns about methodological limitations (two studies with staff-selected participants) and adequacy (brief data in three of four studies); minor concerns about relevance (one study of people discharged from hospital in England); no or very minor concerns about coherence." A team could defend moderate confidence if it judged the consistency across settings to outweigh the thin data, provided it explained the judgement.
Minimum 20 characters required.
Question 1: What are the three stages of thematic synthesis described by Thomas and Harden (2008)?
Question 2: A team codes qualitative findings against a published model of loneliness interventions and creates new themes for the findings that do not fit the model. Which method is the team using?
Question 3: In meta-ethnography, what is a refutational synthesis?
Question 4: Which set lists the four components of GRADE-CERQual?
Evidence and Gap Maps
Learning Objectives for this section
- Define an evidence and gap map and distinguish it from a systematic review and from the tables of a scoping review.
- Describe the parts of a map, including its framework, cells, bubbles and filters, and distinguish absolute gaps from synthesis gaps.
- Outline the steps for building a map, from agreeing its framework with interest holders to coding and drawing it.
- Build a simple evidence and gap map from a charting table and check that its counts reconcile with the table.
- Interpret clusters and gaps in a map cautiously and explain what a map can and cannot tell a decision-maker.
Introduction
A charting table holds everything a scoping review found, but a planning team cannot easily see the pattern in 42 rows and 30 columns. An evidence and gap map presents the same information as a picture: a grid in which each cell shows how much evidence exists for one combination of intervention and outcome. Decision-makers use maps to see where evidence is concentrated and where it is missing, and research funders use them to set priorities. This section explains what a map is, how one is built, and how the Cedar Valley team built and read its own.
4.1 What an Evidence and Gap Map Is
The International Initiative for Impact Evaluation (3ie) developed evidence and gap maps to show the evidence on development programs (Snilstveit et al., 2016). A map is usually a matrix with interventions as rows and outcomes as columns, built from a systematic search and coding of studies, and it shows in each cell the number of studies, often with an indication of their design or of the confidence that can be placed in them. The Campbell Collaboration has since published guidance for producing evidence and gap maps (White et al., 2020), and maps are now common in health, social welfare and education. Related products go by other names. Miake-Lye and colleagues (2016) reviewed published evidence maps in health and found wide variation in methods, with a systematic search and a visual display, often a bubble plot, as common features. In environmental research, systematic maps follow guidance from the Collaboration for Environmental Evidence (James, Randall and Haddaway, 2016).
A map differs from a systematic review in what it reports. A systematic review synthesizes findings to say what the evidence shows. A map describes where evidence exists and leaves the findings to reviews, so a large bubble means that many studies exist, whatever they found. A map is close in purpose to a scoping review, and many scoping reviews, including the Cedar Valley review, present a map as one of their main figures.
4.2 The Parts of a Map
Many Campbell maps include both systematic reviews and primary studies, and rate the confidence that can be placed in each review, for example with the AMSTAR 2 tool, showing the rating by colour. A map of primary studies alone, like the Cedar Valley map, can show absolute gaps but cannot show synthesis gaps, so the team listed the relevant existing reviews beside its map instead.
4.3 Building a Map
The steps for building a map follow those of a scoping review, with particular attention to the framework and to coding. The accordion describes each step.
The team drafts the rows and columns from the review question and its sub-questions and tests them with the people who will use the map. Campbell guidance recommends involving users in this step (White et al., 2020). The Cedar Valley rows are the intervention categories already used in the charting form (Lesson 8), and the columns are the outcome groups set in the protocol (Lesson 2). The team also looked at how a published Campbell map of digital interventions for social isolation and loneliness in older adults had structured its framework (Welch et al., 2023).
A map rests on a systematic search and explicit criteria, as any scoping review does. The Cedar Valley map uses the 42 included studies. The 26 grey-literature documents report their evaluation methods too briefly to classify by design, so they are summarized separately, feed the environmental scan (Lesson 11) and are left out of the map.
Each study is coded for its row, every column in which it reports an outcome, and any filter dimensions. Coding rules are needed for studies that fit several rows; the Cedar Valley team used the primary-category rule from Lesson 8, so each study has one row. Two people should check the coding, since a miscoded study moves a bubble.
Maps that include systematic reviews often appraise them. The Cedar Valley map shows design by bubble position and colour, and it does not show risk of bias, which is reported separately (Lesson 8).
The team counts studies in each cell, draws the map with a legend, and checks the counts against the charting table before anyone interprets the map.
The team reads the map with its users, asking which gaps matter for the decision and which reflect outcomes that nobody would expect a program to change.
4.4 Worked Example: The Cedar Valley Evidence and Gap Map
The Cedar Valley map has five rows for the intervention categories and six columns for the outcome domains. The protocol's outcome groups were merged slightly for display: mental health and well-being form one column, and implementation outcomes (uptake, cost and acceptability) and participants' experiences form another. In each cell, the left bubble counts randomized trials, the centre bubble counts non-randomized controlled studies, and the right bubble counts other designs, meaning uncontrolled before-and-after, qualitative and mixed-methods studies. Figure 4.1 shows the map, and the table that follows gives the same counts in text form.
| Intervention category (studies) | Loneliness | Social isolation | Mental health and well-being | Physical health and function | Health service use | Uptake and experience |
|---|---|---|---|---|---|---|
| Group activity (14) | 13 (8 R, 5 O) | 6 (3 R, 3 O) | 9 (6 R, 3 O) | 4 (3 R, 1 O) | 0 | 6 (2 R, 4 O) |
| One-to-one befriending or telephone (8) | 7 (6 R, 1 O) | 3 (2 R, 1 O) | 6 (5 R, 1 O) | 1 (1 R) | 2 (2 R) | 3 (1 R, 2 O) |
| Community connector or social prescribing (12) | 12 (5 N, 7 O) | 6 (3 N, 3 O) | 10 (5 N, 5 O) | 1 (1 N) | 5 (4 N, 1 O) | 7 (2 N, 5 O) |
| Intergenerational (4) | 3 (2 N, 1 O) | 2 (1 N, 1 O) | 2 (1 N, 1 O) | 0 | 0 | 1 (1 O) |
| Technology-based (4) | 4 (3 N, 1 O) | 3 (2 N, 1 O) | 2 (1 N, 1 O) | 0 | 0 | 2 (1 N, 1 O) |
| Studies in column | 39 | 20 | 29 | 6 | 7 | 19 |
R, randomized trials; N, non-randomized controlled studies; O, other designs (before-and-after, qualitative and mixed methods). Each study has one row, and it appears in every column for which it reports an outcome.
Before interpreting the map, the team checked that its counts reconciled with the charting table. Three checks were made. First, the row totals in the label column must match the cross-tabulation in Section 1 (14, 8, 12, 4 and 4 studies, totalling 42). Second, the eligibility criteria required every study to report loneliness or social isolation, so every study must appear in at least one of the first two columns. The loneliness column holds 39 studies, and the 3 studies outside it (one before-and-after study each in the group, one-to-one and intergenerational rows) all appear in the social isolation column, which accounts for all 42. Third, the counts within a row can exceed the row total, because studies report several outcomes. The connector row has 12 studies and 41 entries across its six cells, three or four outcome domains per study, which matches the charting table.
A late-identified report describes a non-randomized controlled study of a community connector program that measured loneliness with the De Jong Gierveld scale, depressive symptoms and emergency department visits, and reported the cost per participant. List the cells whose counts would change, state which bubble in each cell would grow, and give the new row total. Then check that your answer keeps the reconciliation rules above.
4.5 Reading the Map Carefully
The map shows several patterns the planning team can use. Evidence clusters in the loneliness column for group activities (13 studies, 8 of them randomized) and one-to-one programs (7 studies, 6 randomized), which are the bodies of evidence rated in Lesson 8. Community connector programs, the model the health authority intends to launch, have 12 studies and no randomized trial; their evidence for loneliness comes from 5 non-randomized controlled studies and 7 studies of other designs. Connector studies measured health service use more often (5 of 12) than studies of one-to-one programs (2 of 8), and no study of group activities measured it, which reflects the interest of primary care in whether referral reduces visits. Physical health and function and health service use are absolute gaps for intergenerational and technology-based programs, health service use is also an absolute gap for group activities, and the intergenerational row as a whole is thin. Most evidence on uptake and experience comes from other designs: 13 of the 19 entries in that column are before-and-after, qualitative or mixed-methods studies, which suits questions about experience and acceptability.
What a map cannot tell you
Bubble size counts studies. It says nothing about the size or direction of effects, the quality of the studies or the number of participants, so a large bubble may hold many small, weak studies. An empty cell shows that no included study measured the outcome; it is evidence of a gap in research and gives no information about whether the program affects that outcome. The framework decides which gaps can appear, so a map without equity columns cannot show that few studies reported participants' income, ethnicity, language or disability, which the Section 1 counts revealed. And a gap is a research priority only if the outcome matters to the decision; few people would expect a video-calling program to change physical function.
The team drew three implications for the planning team, which the evidence brief in Lesson 12 develops. The connector program should be launched with an evaluation built in, a recommendation that the very low certainty rating in Lesson 8 and the absence of randomized trials in this map both support. The evaluation should measure loneliness with a widely used scale, such as the Revised UCLA Loneliness Scale or the De Jong Gierveld scale, so that its results can be compared with the 26 included studies that used one of those two scales. And it should record the equity characteristics of participants, uptake and experience, together with health service use, since these are the outcomes primary care partners will ask about and those on which the existing evidence is weakest. The map also helped the planning team decide what it did not need to commission: further evidence on group programs and loneliness already exists and is summarized in existing reviews.
Reflection
A draft evidence and gap map for a scoping review of interventions to increase physical activity among adults with arthritis has three rows and four columns. The rows are exercise classes (12 studies), walking programs (7 studies) and app-based programs (5 studies). The columns are physical activity, pain, quality of life and cost. The cell counts are: exercise classes 12, 10, 8 and 1; walking programs 7, 3, 2 and 0; app-based programs 5, 0, 1 and 0. All 12 exercise-class studies and 4 of the 7 walking studies are randomized trials, and all 5 app-based studies are uncontrolled before-and-after studies. (1) Identify two clusters and two gaps in the map. (2) Explain to a program manager what the map can and cannot tell her about whether app-based programs reduce pain. (3) Name one check you would make on the counts before sharing the map, and one question you would ask the map's users about its gaps.
(1) The largest cluster is exercise classes and physical activity (12 randomized trials), followed by exercise classes and pain (10 studies). Walking programs and physical activity (7 studies, 4 randomized) form a smaller cluster. The clearest gaps are app-based programs and pain (no studies) and cost for walking and app-based programs (no studies), and cost is nearly a gap for exercise classes as well (1 study).
(2) The map tells the manager that no included study of an app-based program measured pain, so the review can say nothing about whether these programs reduce it. The empty cell is a gap in research and gives no evidence that the apps fail to reduce pain. Even in cells with studies, the map counts studies and does not show their findings, and the app-based studies are all uncontrolled before-and-after designs, which cannot separate program effects from natural change. For evidence on effects she would need a systematic review of effects or a new controlled evaluation.
(3) Every study had to report an eligible outcome, so I would check that each of the 24 studies appears in at least one column and that the row totals match the charting table; here every study appears in the physical activity column, so the rows reconcile. I would ask the users whether cost and pain matter for their decision about app-based programs, since a gap is a research priority only if it bears on a decision.
Minimum 20 characters required.
Question 1: In a typical evidence and gap map, what does the size of a bubble show?
Question 2: The Cedar Valley map has no studies of technology-based programs that measured physical health and function. What can the team conclude?
Question 3: What is a synthesis gap in an evidence and gap map?
Question 4: The Cedar Valley connector row lists 12 studies, but its six cells hold 41 entries in total. Why?
Final Assessment
Bringing It All Together
This lesson followed the fictional Cedar Valley evidence team through three kinds of review and the map that brings them together. As a scoping review, the project mapped the extent, range and nature of the evidence, following the stages of Arksey and O'Malley (2005), the refinements of Levac, Colquhoun and O'Brien (2010) and the JBI method. Its charting table and descriptive summary answered the sub-questions, and PRISMA-ScR set out what the report must contain, with critical appraisal as an optional addition.
As a rapid review, the project streamlined parts of the method under the Cochrane Rapid Reviews Methods Group guidance (Garritty et al., 2021). Some shortcuts narrow what a review can find and others reduce independent checking, and the team's own arithmetic showed that checking every excluded record saved little when nine records in ten were excluded, which led it to dual screen everything. As a qualitative evidence synthesis, a brief thematic synthesis of 9 sources produced analytical themes about continuity, access and contribution, and GRADE-CERQual showed how much confidence each finding deserves. The evidence and gap map then displayed all 42 studies by intervention and outcome, showing clusters of randomized evidence for group and one-to-one programs, no randomized trials of community connector programs, and absolute gaps for some outcomes.
Key Takeaways from this lesson
- A scoping review maps the extent, range and nature of the evidence on a topic and is suited to questions about what has been studied, how and where the gaps lie.
- Arksey and O'Malley (2005) proposed five stages and an optional consultation exercise, and Levac, Colquhoun and O'Brien (2010) refined every stage and argued that consultation should be essential.
- The JBI method adds a registered protocol, a PCC question, a three-step search, selection by at least two reviewers, iterative charting and a mainly descriptive analysis.
- PRISMA-ScR has 20 essential and 2 optional items, and it speaks of sources of evidence, data charting and critical appraisal to fit the purposes of a scoping review.
- The Cochrane Rapid Reviews Methods Group guidance (Garritty et al., 2021) recommends specific shortcuts at each stage, and its 2024 update asks reviews to report their restricted methods and discuss their likely effect.
- Shortcuts that narrow what a review can find risk missing studies, while shortcuts that reduce independent checking risk errors in handling the studies found, and a team should calculate what each shortcut saves before adopting it.
- Thematic synthesis moves from line-by-line codes to descriptive themes and then to analytical themes that answer the review question.
- Framework synthesis indexes findings against a prior framework, and the best fit approach creates new themes from findings that do not fit, while meta-ethnography translates studies into one another to build new interpretations.
- GRADE-CERQual rates confidence in each qualitative finding from its methodological limitations, coherence, adequacy of data and relevance.
- An evidence and gap map shows where evidence exists and where it is missing, and an empty cell is a gap in research that says nothing about whether an intervention works.
Core Concepts Reviewed
Section 1: scoping reviews, the Arksey and O'Malley framework, the Levac refinements, the JBI method, PCC, data charting, the charting table, descriptive numerical summary, basic qualitative content analysis and PRISMA-ScR.
Section 2: rapid reviews, knowledge users, the Cochrane Rapid Reviews Methods Group guidance and its 2024 update, shortcuts that narrow what can be found and shortcuts that reduce independent checking, and the reporting of shortcuts.
Section 3: qualitative evidence synthesis, thematic synthesis with descriptive and analytical themes, framework and best fit framework synthesis, meta-ethnography and its three forms of synthesis, meta-aggregation, GRADE-CERQual and the summary of qualitative findings table.
Section 4: evidence and gap maps, their framework, cells, bubbles and filters, absolute and synthesis gaps, building and reconciling a map, and cautious interpretation.
The final reflection asks you to explain to a decision-maker how much weight to place on a rapid scoping review, drawing on all four sections.
Reflection
The director of the fictional Cedar Valley Health Authority's planning team asks for a short explanation of how much weight to place on the evidence review. The facts are as follows. The review is a rapid scoping review of 42 studies (14 randomized trials, 10 non-randomized controlled studies, 9 before-and-after studies, 6 qualitative studies and 3 mixed-methods studies). It included reports from 2010 onward in English and French, and it dual screened all 1,870 titles and abstracts after amending a plan for partial single screening. Community connector programs, the model the authority plans to launch, have 12 studies, none of them randomized. GRADE rated the certainty of evidence on loneliness as low for group and one-to-one programs and very low for connector programs. A thematic synthesis of 9 qualitative sources found, with moderate confidence, that continuity with the same volunteer or connector and practical access (transport and accompaniment to a first activity) shaped whether people took part. The evidence and gap map shows no studies of physical health or health service use for intergenerational or technology-based programs. Write 200 to 300 words for the director that (a) states what kind of review this is and what it can and cannot answer, (b) names the two shortcuts most likely to matter and why, (c) explains what the qualitative findings and the map add to the GRADE ratings, and (d) recommends what the program's own evaluation should measure.
This is a rapid scoping review. It maps which community programs for loneliness among older adults have been evaluated, how and with what outcomes, and it can describe the range and strength of that evidence. It was not designed to estimate how much each program reduces loneliness, so its provisional GRADE ratings should be read with caution.
Two shortcuts matter most. The review included reports only from 2010 onward, so older evaluations appear only through the existing reviews it cites, and it included only English and French reports, which may omit some European and Asian evaluations, although language limits rarely change conclusions. Screening was not shortened.
The GRADE ratings show that the evidence on loneliness is uncertain for every program type and very uncertain for connector programs, which have no randomized trials. The qualitative findings suggest what makes a program work in practice: continuity with one connector, transport and accompaniment to a first activity. The map shows where evidence is missing, including randomized evaluations of connector programs.
I recommend launching the program with an evaluation built in, for example by introducing it in stages across clinics. It should measure loneliness with the Revised UCLA Loneliness Scale or the De Jong Gierveld scale, record participants' income, ethnicity, language, disability and rural residence, track uptake, continuity with connectors and health service use, and interview participants about access and endings.
Minimum 30 characters required.
Final Knowledge Assessment
Question 1: A health authority needs, within ten weeks, a description of which peer support programs for new parents have been evaluated in Canada and which outcomes were measured. Which design fits best?
Question 2: Which feature does the JBI method for scoping reviews share with a systematic review of effects?
Question 3: Which pair correctly matches authors with their contribution to scoping review methods?
Question 4: Why are the two critical appraisal items in PRISMA-ScR optional?
Question 5: A rapid review team drops two specialized databases and restricts reports to English. What risk do these shortcuts carry?
Question 6: In the randomized trial of abstract screening by Gartlehner and colleagues (2020), what share of relevant studies did single reviewers miss?
Question 7: Which statement about the 2024 update of the Cochrane rapid review guidance is accurate?
Question 8: A team has nine qualitative sources, four weeks, general qualitative training and an audience that needs practical implications for program design. Which method fits best?
Question 9: In GRADE-CERQual, what does the component called adequacy of data assess?
Question 10: A qualitative review finding is rated low confidence with GRADE-CERQual. What does that mean?
Question 11: How do GRADE (Lesson 8) and GRADE-CERQual differ?
Question 12: Which statement about an evidence and gap map is accurate?
Question 13: Every Cedar Valley study had to report loneliness or social isolation, and the loneliness column of the map holds 39 of the 42 studies. How many studies must appear in the social isolation column without appearing in the loneliness column?
Question 14: Which implication did the Cedar Valley team draw from its map and ratings for the planning team?
Question 15: A student's rapid scoping review says only that "records were screened by one reviewer", and its conclusion states that intergenerational programs "do not improve physical health" because the map has an empty cell. Which revision is most needed?
Glossary: Key Terms, People & Frameworks
📚 Reference page, available throughout the lesson
These terms, tools and people appear in this lesson on scoping reviews, rapid reviews, qualitative evidence synthesis and evidence and gap maps.