Managing Records and Screening Studies
Finding & Synthesizing Health Evidence
Learning objectives for this lesson:
- Explain why a review team gathers its search results in a reference manager and keeps a written log of counts from export to screening.
- Distinguish records, reports and studies as PRISMA 2020 defines them, and explain why duplicate records are merged while multiple reports of one study are linked.
- Import database exports into Zotero, check that every batch arrived, and remove duplicates with automated and manual passes.
- Compare Covidence and Rayyan and set up a screening project with a screening guide, ordered exclusion reasons and blinded dual review.
- Apply eligibility criteria at the title and abstract stage and the full-text stage, recording one ordered reason for each full-text exclusion.
- Calculate and interpret percent agreement and Cohen's kappa for a pilot of dual screening, and explain why kappa can be low when agreement is high.
- Draw a complete PRISMA 2020 flow diagram for a review that searched databases and other sources, and check that every number reconciles.
This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University, drawing on the Cochrane Handbook for Systematic Reviews of Interventions and the JBI Manual for Evidence Synthesis.
Reference Managers and De-duplication
Learning Objectives for this section
- Explain why a review team gathers its search results in a reference manager and keeps a written record of counts from export to screening.
- Distinguish records, reports and studies as PRISMA 2020 defines them, and explain why de-duplication works on records while the linking of reports to studies waits until full-text screening.
- Export records from bibliographic databases in RIS format, import them into Zotero by source, and check that every batch arrived.
- Identify and merge duplicate records with automated and manual checks, and recognize pairs of records that look alike but describe different reports.
- Keep an audit trail of de-duplication that supplies the first numbers of the PRISMA 2020 flow diagram.
Introduction
By the end of Lessons 4 and 5, a review team holds the results of several database searches, documents from websites and organizations, and references found by citation chasing, in different file formats and with overlapping content. Before anyone can screen them, the team has to gather them in one place, count them, remove the copies, and move a clean set into a screening platform without losing a record. This work is clerical, and it is also part of the method: a review that cannot say how many records it found and how many duplicates it removed cannot complete its flow diagram, and a review that loses records during import has quietly narrowed its own search.
This section explains the units that the work counts, describes what a reference manager does, and walks through import and de-duplication with the running case. Sections 2 to 4 cover screening platforms, screening itself and the PRISMA 2020 flow diagram.
The fictional Cedar Valley Health Authority in British Columbia serves about 210,000 residents, about 46,000 of them aged 65 and older, through 24 primary care clinics. Before launching a community connector (social prescribing) program for older adults, its planning team asked a small evidence team for a rapid scoping review and environmental scan within twelve weeks. The team is a health authority evidence officer, a university librarian who works with the team, and a student intern. The review asks which community-based interventions have been evaluated for reducing loneliness or social isolation among adults aged 65 and older, and with what outcomes.
The librarian has run the final searches in MEDLINE, Embase, CINAHL, PsycINFO and Web of Science, which together retrieved 2,480 records. The intern has been asked to bring these records together, remove duplicates, and prepare them for screening. All numbers about Cedar Valley in this lesson are illustrative.
1.1 Records, Reports and Studies
Three words recur throughout this lesson, and PRISMA 2020 gives each a specific meaning (Page et al., 2021). A record is the title, abstract or both of a report as it is indexed in a database or on a website. When MEDLINE returns a citation with its abstract, that citation is a record. A report is a document that supplies information about a study, such as a journal article, a preprint, a conference abstract, a thesis, a trial registry entry or a government report. A study is the investigation itself, for example a trial of a befriending program with a defined group of participants, an intervention and outcomes. One study can produce several reports, and one report can be indexed as a record in several databases.
These distinctions matter because different stages of the review count different units. De-duplication works on records: its job is to make sure that each report enters screening once, however many databases indexed it. Linking reports to studies happens later, usually at full-text screening, when the reviewer can read enough to see that a conference abstract from 2020 and a journal article from 2022 describe the same trial. If the team tried to merge reports of the same study during de-duplication, it would have to judge study identity from titles and abstracts alone, and it would risk discarding a report that holds information the full article leaves out.
1.2 What a Reference Manager Does
Background: Zotero basics
A reference manager is software that stores the details of each source as structured fields (authors, title, journal, year, abstract, DOI and so on), inserts citations as you write and builds the reference list in a chosen style. Zotero, the example in this lesson, is free, open-source software maintained by the Corporation for Digital Scholarship. Its basic parts are the library, which holds every item; collections, which are folders inside the library, with any item able to sit in several of them; the Zotero Connector, a browser extension that saves the article or page you are viewing, alongside an Add Item by Identifier button that retrieves an item's details from its DOI or ISBN; and the word-processor plug-in for Microsoft Word and LibreOffice (with a Zotero menu in Google Docs), which inserts citations and refreshes the bibliography. Imported details are often imperfect, so each entry should be checked. HSCI 207 Lesson 12, Section 4.4 (Writing Up Research), walks through installation and these steps in detail and is optional reading for anyone who has not used Zotero.
In a review, the same software does a different job: it serves as the team's central store of records between the search and the screening platform, holding each export in its own collection, finding duplicates and exporting a clean file for screening. Zotero's online group libraries let the whole team work from one shared library. EndNote (a commercial product from Clarivate) and Mendeley (from Elsevier) offer similar functions, and many librarians who support reviews work in EndNote. The principles in this section apply to any of them.
Teams organize this stage in one of three ways, and the choice depends on the size of the search, the tools the institution supports and the preferences of the librarian.
The team imports every export into a reference manager such as Zotero, removes duplicates there, and exports one de-duplicated file to the screening platform. This route gives full control over de-duplication and a permanent library for later updates, at the cost of an extra transfer whose counts must be checked.
Many health sciences librarians de-duplicate in EndNote with a published sequence of field comparisons and manual checks, designed to remove most duplicates while keeping false merges rare (Bramer et al., 2016). Teams working with such a librarian often receive a de-duplicated file.
Covidence and Rayyan both look for duplicates on import, so a small team may import the raw exports directly. This route is quick, and it leaves the team dependent on the platform's matching rules.
1.3 Exporting and Importing Records
Each database interface exports records in its own way, and the team's first task is to export every record from each final search, with abstracts, in a format the reference manager can read. The RIS format is a plain-text format in which each field sits on its own line with a two-letter tag, such as TY for the type of reference, AU for an author, TI for the title, PY for the year and ER for the end of the record. A fictional record exported in RIS looks like this.
TY - JOUR AU - Lindqvist, Maren AU - Osei, Kwame TI - A telephone befriending service for community-dwelling older adults: a pilot evaluation JO - Journal of Community Ageing PY - 2021 VL - 14 SP - 112 EP - 124 DO - 10.0000/fictional.2021.0112 AB - Loneliness was measured with a three-item scale at baseline and twelve weeks... ER -
Several practical rules prevent lost records. Some interfaces cap the number of records per export, so a large result set may need several batches whose sizes should add up to the search total. Each file should be named with the database and the date of the search (for example, CINAHL_2026-02-10.ris), kept unchanged as the original export, and imported into its own collection, and the number that arrives should be compared with the number exported. A shortfall usually means that a batch was skipped or that the import stopped at a malformed record.
| Database | Records exported | Records in Zotero collection | Check |
|---|---|---|---|
| MEDLINE | 742 | 742 | Matches |
| Embase | 816 | 816 | Matches (exported in two batches) |
| CINAHL | 388 | 388 | Matches |
| PsycINFO | 296 | 296 | Matches |
| Web of Science | 238 | 238 | Matches |
| Total | 2,480 | 2,480 | All records accounted for |
The Cedar Valley import log above shows the check the intern completed before de-duplication began. The first import of the Embase file brought in only 500 records, because the librarian had exported the results in two batches and the second file had been saved in a different folder. The log made the shortfall obvious, and the second batch was imported before anyone started screening.
1.4 Finding and Removing Duplicates
A duplicate record is a second or later copy of the same report. Duplicates arise mainly because databases overlap: a gerontology article in a well-known journal is likely to be indexed in MEDLINE, Embase, CINAHL and PsycINFO, and to appear in Web of Science as well. Duplicates also arise within a single database when a report is indexed twice, for example once as an online-first version and again in its print issue.
De-duplication can go wrong in two directions. A false-positive merge joins two records that describe different reports, and the second report disappears from the review before anyone screens it. This error reduces recall, which Lesson 3 identified as the property that reviews protect most carefully. A false-negative leaves two copies of the same report in the set, so screeners see it twice. That error costs time and can produce inconsistent decisions on the same report, but it does not remove evidence, and it is usually caught later. For this reason, teams set automated matching to be cautious and send uncertain pairs to a person.
Matching on fields
Software decides whether two records match by comparing fields. The DOI (digital object identifier) is the strongest single identifier, but older records and conference abstracts often lack it. Titles differ across databases in punctuation, capitalization, spelling and the handling of subtitles, and author names appear as full names in one database and as initials in another. Good matching therefore normalizes fields (for example, by removing punctuation and converting text to lower case) and combines several fields, such as title, first author, year and journal.
The Cedar Valley de-duplication
The intern de-duplicated in three passes and recorded the number removed at each pass.
| Pass | What was done | Duplicates removed | Records remaining |
|---|---|---|---|
| Start | All five database exports imported into Zotero. | None | 2,480 |
| 1. Automated | Zotero's Duplicate Items view; each suggested group was checked and merged. | 571 | 1,909 |
| 2. Manual | Records sorted by title and then by first author, and adjacent near-matches inspected. | 27 | 1,882 |
| 3. Platform check | De-duplicated file imported into Covidence, which flagged further matches for confirmation. | 12 | 1,870 |
| Total | 610 | 1,870 |
The 27 duplicates found in the manual pass were mostly pairs in which one record lacked a DOI and the titles differed in punctuation or spelling (for example, "behaviour" in one database and "behavior" in another). The 12 found in Covidence were pairs whose author lists had been formatted differently. During the manual pass the intern also found nine pairs that looked alike but were different reports, including a conference abstract and its later journal article, and kept both records in each pair with a note for the full-text stage.
These are two reports of one study. Keep both records, link them to the same study at full text, and check the abstract for any results the article omits.
These are also two reports of one study. Keep both, note the relationship, and treat the published version as the primary report.
A protocol and the trial's results share authors and similar titles, and they are separate reports. The protocol is usually excluded at full text or kept as a supporting report, depending on the review's rules.
A notice titled "Correction to" followed by the original title is a separate document. Most teams keep it with the article it corrects.
Series titles such as "part 1" and "part 2" can fool automated matching. Checking the year, pages and abstract usually separates them.
The intern's manual pass brought up the six fictional records below. Decide which pairs are duplicate records of one report, which are different reports of one study, and which are unrelated.
| ID | Record |
|---|---|
| R1 | Okafor A, Byrne L. Intergenerational gardening and loneliness in older adults: a mixed-methods evaluation. Ageing Soc Res. 2023;9:41–55. DOI present. |
| R2 | Okafor, Adaeze; Byrne, Liam. Intergenerational gardening and loneliness in older adults - a mixed methods evaluation. 2023. No DOI. |
| R3 | Okafor A. Intergenerational gardening and loneliness: early findings. Conference abstract, 2021. |
| R4 | Chen M, Patel S. Loneliness in later life: part 1, measurement. 2019. |
| R5 | Chen M, Patel S. Loneliness in later life: part 2, interventions. 2019. |
| R6 | Correction to: Intergenerational gardening and loneliness in older adults. 2024. |
R1 and R2 are duplicate records of the same report: the authors, title and year match once punctuation and name formats are normalized, and R2 should be merged into R1, which carries the DOI. R3 is a different report (a conference abstract) of what is probably the same study, so it stays in the set and is linked at full text. R4 and R5 are different reports in a series and are unrelated as duplicates, even though the start of the title matches. R6 is a correction notice for R1; it is a separate document that most teams keep with R1 so that the full-text reviewer reads both.
1.5 Keeping an Audit Trail
An audit trail is a written record that would let another person repeat the steps and arrive at the same numbers. For record management it has three parts: the original export files, kept unchanged; a log of counts for each source and each de-duplication pass; and a copy of the library exported before merging, so that any merge can be checked later.
Counts to record before screening begins
The number of records retrieved from each database and register, and their total, become the first box of the PRISMA 2020 flow diagram. The number of duplicates removed becomes the "records removed before screening" box, together with any records removed by automation tools or for other reasons. The number of records sent to the screening platform becomes the "records screened" box. For Cedar Valley these are 2,480 records identified, 610 duplicates removed and 1,870 records screened, and the arithmetic 2,480 − 610 = 1,870 must hold.
The records found through grey-literature searching and citation chasing are kept in their own collections and counted separately, because PRISMA 2020 reports them in a second column of the flow diagram. Records from citation chasing should be checked against the database records before they are counted, so that a study the databases already found is not counted twice. Section 4 returns to these numbers.
With 1,870 unique records counted and logged, the Cedar Valley team is ready to move them into a screening platform. Section 2 describes the two platforms most often used for that step.
Reflection
You are de-duplicating 1,240 records exported for a review of walking programs for older adults: 520 from MEDLINE, 410 from CINAHL and 310 from PsycINFO. Under PRISMA 2020, a record is the indexed title or abstract of a report, a report is a document that supplies information about a study (such as an article, conference abstract or protocol), and a study is the investigation itself. Zotero's Duplicate Items view suggests four groups. Group 1: two records with the same authors, year, journal and pages, one with a DOI and one without, whose titles differ only by a hyphen. Group 2: a 2021 conference abstract and a 2023 journal article by the same first author with similar titles. Group 3: a 2022 trial protocol and a 2024 results paper with the same trial name in their titles. Group 4: two records titled "Walking for health: part 1" and "Walking for health: part 2" by the same authors in the same year. For each group, say whether you would merge the records and explain why, using the terms record, report and study. Then state the counts you would enter in your log if the automated pass removed 268 duplicates and the manual pass removed 14 more.
Group 1 should be merged. The two items are duplicate records of one report: the authors, year, journal and pages match, and a hyphen is the kind of difference that databases introduce when they index the same title. I would keep the record that carries the DOI. Group 2 should not be merged. The conference abstract and the article are two different reports that probably describe one study, so both stay in the set with a note, and I would link them to the same study at full text, checking the abstract for results the article omits. Group 3 should also be kept apart. The protocol and the results paper are separate reports of one trial; the protocol will probably be excluded at full text for having no outcome data, or kept as a supporting report. Group 4 contains two different reports in a series, so the records are unrelated as duplicates despite the shared title.
For the log, I would record 520, 410 and 310 records by database, totalling 1,240 records identified; 268 duplicates removed in the automated pass and 14 in the manual pass, totalling 282 duplicates; and 1,240 − 282 = 958 records sent to screening. I would also keep the original export files and a copy of the library from before merging, so that any merge can be checked later.
Minimum 20 characters required.
Question 1: According to PRISMA 2020, how does a record differ from a report?
Question 2: During de-duplication, the intern finds a 2020 conference abstract and a 2022 journal article that appear to describe the same befriending trial. What should she do?
Question 3: Why do review teams set automated de-duplication to be cautious and send uncertain pairs to a person?
Question 4: The intern exported 816 Embase records in two batches, but only 500 arrived in the Zotero collection. What is the most likely explanation, and what should she do?
Screening Platforms: Covidence and Rayyan
Learning Objectives for this section
- Describe the functions that screening platforms provide and explain why most teams prefer them to spreadsheets for dual screening.
- Compare Covidence and Rayyan by cost model, workflow, blinding, conflict handling and recording of exclusion reasons.
- Set up a screening project with a screening guide, an ordered list of exclusion reasons, keyword highlights and blinded dual review.
- Write screening questions for the Cedar Valley review that turn the eligibility criteria into decisions that two reviewers can apply consistently.
- Identify the privacy, licensing and record-keeping matters that a team should settle before screening begins.
Introduction
Screening is the process of deciding, record by record and then report by report, which items meet the review's eligibility criteria. With 1,870 records to screen, the fictional Cedar Valley team needs a place where two people can make independent decisions, where disagreements are found automatically, where reasons for exclusion are recorded, and where the numbers at each stage can be counted at the end. A spreadsheet can do some of this for a small review, but it makes blinding difficult, invites errors when people sort or overwrite rows, and leaves the team to count everything by hand. Dedicated screening platforms were built to solve these problems.
This section describes what screening platforms do, compares the two most widely used in health sciences, Covidence and Rayyan, and walks through setting up the Cedar Valley screening project. The decisions themselves, and how to measure whether two screeners agree, are the subject of Section 3.
2.1 What Screening Platforms Do
Most screening platforms share a core set of functions. They import records in RIS and similar formats and check for duplicates on the way in. They present one record at a time, with the title and abstract laid out for reading and with chosen keywords highlighted. They let each reviewer record a decision (usually include, exclude or maybe) without seeing the other reviewer's decision, a feature called blinding. They compare the two sets of decisions and list the conflicts, meaning the records on which the reviewers disagree, for resolution. They store reasons for exclusion and notes, and they export decisions so that the team can keep its own copy and report its numbers.
Keyword highlighting deserves a caution. Highlighting the word "residential" in red helps a reviewer notice that a study may have been set in long-term care, but in the record above the word appears in a sentence that contrasts the program with residential care. Highlights direct attention; they do not make decisions, and a reviewer who screens by colour alone will make errors.
2.2 Covidence
Covidence is a web-based platform produced by Veritas Health Innovation, a not-for-profit organization in Melbourne, Australia, and it is widely used by Cochrane review teams. It is a subscription service, and many universities and health organizations hold licences for their staff and students. Covidence organizes a review as a fixed sequence of stages: import, title and abstract screening, full-text review, and then data extraction and quality assessment, which Lesson 8 covers.
At the title and abstract stage, each reviewer votes yes, no or maybe on each record. By default Covidence asks for two votes on every record, keeps each reviewer's votes hidden from the other, and moves a record forward or out when the votes agree. Records with disagreeing votes go to a list of conflicts for the team to resolve. At the full-text stage, reviewers upload or attach the full text, and an exclusion must be accompanied by a reason chosen from a list that the team writes in advance. Covidence keeps a running count of records and reports at each stage and can produce a PRISMA flow diagram from those counts, which the team should still check against its own log. It also lets the team mark several reports as belonging to one study.
2.3 Rayyan
Rayyan was developed at the Qatar Computing Research Institute and described by Ouzzani and colleagues as a web and mobile application for systematic reviews (Ouzzani et al., 2016). It offers a free plan with limits on features and paid plans for larger teams. Rayyan presents all imported records in a single screening space, where each reviewer marks records as include, exclude or maybe and can add exclusion reasons, labels and notes. In blind mode, which collaborative reviews use during screening, no reviewer can see another's decisions; when the team turns blind mode off, the platform shows where decisions conflict, and the team resolves them.
Rayyan is more flexible than Covidence and less structured. It does not force a sequence of stages, so the team decides how to run full-text screening, for example as a second round in a separate review. Its filters, keyword highlights and labels make it easy to sort records, and it detects possible duplicates on import. The flexibility suits small teams and student projects, and it places more responsibility on the team to keep its stages and counts organized.
| Feature | Covidence | Rayyan |
|---|---|---|
| Cost model | Paid subscription, often through an institutional licence. | Free plan with limits, and paid plans with more features. |
| Workflow | Fixed stages from import to extraction. | One flexible screening space; the team defines its rounds. |
| Blinding | Votes are hidden from the other reviewer by default. | Blind mode hides decisions until the team turns it off. |
| Conflicts | Listed automatically for resolution. | Shown when blind mode is off, and filtered for resolution. |
| Exclusion reasons | Required at full text, from the team's list. | Added by reviewers; requiring them is a team rule. |
| Reports and studies | Several reports can be merged into one study. | Relationships are recorded with labels and notes. |
| Flow diagram counts | Counts kept by stage and drawn as a PRISMA diagram. | Decisions can be exported; the team compiles or checks the counts. |
| Machine learning | Both platforms offer features that rank records or predict relevance from earlier decisions. Lesson 6 explains how such features work and how far to rely on them. | |
Both platforms change often, and the features above should be checked against the current version before a team commits to one. For the Cedar Valley review, the librarian's university holds a Covidence licence, and the team chose Covidence because its fixed stages and required exclusion reasons would make the PRISMA counts straightforward for a team that included a first-time screener. A student team without a licence could run the same process in Rayyan's free plan by setting its own rules for reasons and rounds.
2.4 Setting Up a Screening Project
A screening project should be set up before anyone screens a record, because changing the rules part of the way through makes earlier decisions inconsistent with later ones. The steps below apply to any platform.
Create one review, invite each screener, and decide who will resolve conflicts. In Cedar Valley, the evidence officer and the intern screen, and the librarian acts as third reviewer for conflicts that discussion does not settle.
Set two independent reviewers per record at both stages unless the protocol states otherwise. Section 3 discusses when a rapid review might reduce this.
Import the de-duplicated file and confirm that the platform shows the number in the log. Record any further duplicates the platform finds. For Cedar Valley, 1,882 records were imported, Covidence flagged 12 duplicates, and 1,870 records entered screening.
Enter the eligibility criteria and the screening questions where screeners will see them. Covidence and Rayyan both allow the criteria to be displayed or attached to the project.
Write a short list of exclusion reasons that match the criteria, and put them in a fixed order. A report that fails several criteria receives the first reason in the list that applies, which keeps the counts in the flow diagram consistent.
Add inclusion terms such as lonel*, isolat*, befriend* and social prescri*, and exclusion terms such as nursing home, inpatient and adolescent. Check that the platform handles truncation in the way the team expects.
Choose a random sample of records for both screeners to screen first, and agree to meet afterwards to compare decisions and revise the guide. Section 3 describes the pilot and how to measure agreement.
The Cedar Valley screening guide
A screening guide turns eligibility criteria into questions that a screener can answer from a title, an abstract or a full text. Each question should have a clear rule for the common borderline cases, because borderline cases cause most disagreements. The Cedar Valley guide below follows the review's Population, Concept and Context (PCC) question from Lesson 2: adults aged 65 and older who live in the community, community-based interventions evaluated for their effects on loneliness or social isolation, and studies conducted in high-income countries.
| Question | Rule for borderline cases | Exclusion reason if the answer is no |
|---|---|---|
| 1. Are the participants adults aged 65 and older who live in the community? | Include if at least 80 percent of participants are 65 or older, if results are reported separately for this group, or, when the age distribution is not reported, if the mean or median age is 65 or older. Residents of long-term care and hospital inpatients are excluded. | 1. Population outside criteria |
| 2. Is the intervention delivered in the community? | Homes, community centres, libraries, faith settings, telephone and online delivery all count. Programs delivered only in hospitals or long-term care do not. | 2. Intervention or setting not community-based |
| 3. Was loneliness or social isolation measured or explored as an outcome? | Any validated scale, single question or qualitative exploration counts. A study of depression alone does not count. | 3. No loneliness or social isolation outcome |
| 4. Was the study conducted in a high-income country? | Use the World Bank income classification for the year of the study; multi-country studies count if results are reported for high-income countries. | 4. Not conducted in a high-income country |
| 5. Is the report an eligible type of evidence source? | Primary studies of any design and program evaluation reports are eligible. Commentaries, editorials and protocols without results are excluded. Reviews are excluded, and their reference lists are checked in citation chasing. | 5. Ineligible publication type |
Why the order of reasons matters
A commentary about a hospital-based program for adults aged 50 and over fails questions 1, 2 and 5. Under the ordered list it is recorded once, as "population outside criteria", because that is the first reason that applies. Without a fixed order, one screener might record the population and another the publication type, and the counts in the flow diagram would depend on who screened which report. The order should be stated in the review's methods.
2.5 Privacy, Licences and Records
Bibliographic records do not contain information about study participants, so screening carries little privacy risk. Three related matters still need a decision before screening begins. The first is the institution's policy on cloud-based tools, which some health authorities restrict, and which may determine whether staff can use a given platform. The second is the licence under which full texts are obtained: PDFs supplied through a university library are licensed to that library's users, and the team should share them only among members entitled to access them. The third is the team's own copy of its decisions. Platforms can change their plans, and accounts can lapse, so the team should export the screening decisions and exclusion reasons at the end of each stage and store them with the audit trail described in Section 1.
A review of walking programs for older adults has these eligibility criteria: participants are adults aged 60 and older living in the community; the program is a structured walking program; physical activity or mobility is measured as an outcome; and the report is a primary study with results. Write three screening questions in the style of the Cedar Valley guide, each with a rule for one borderline case that screeners are likely to meet often, and an exclusion reason. Put the reasons in the order you would apply them, and write one sentence explaining the order.
With the platform set up and the guide written, the Cedar Valley screeners are ready to begin. Section 3 follows them through a pilot, through title and abstract screening and full-text screening, and through the calculation of their agreement.
Reflection
A three-person student team will screen about 600 records for a scoping review of school-based mental health programs. The team wants blinded dual screening, a recorded reason for every full-text exclusion, and accurate counts for a PRISMA 2020 flow diagram. Its university holds no Covidence licence, and the team has no budget. The two options are as follows. Covidence is a paid subscription service, usually accessed through an institutional licence; it moves records through fixed stages (title and abstract, full text, extraction), hides each reviewer's votes from the other by default, lists conflicts automatically, requires an exclusion reason from the team's list at full text, and produces PRISMA counts. Rayyan offers a free plan with limits and paid plans; it has one flexible screening space, a blind mode that hides decisions until the team turns it off, and exclusion reasons and labels that reviewers add, and the team compiles or checks its own counts from exported decisions. The review's criteria are: (1) participants are school-aged children or adolescents; (2) the program is delivered in schools; (3) a mental health outcome is measured; (4) the report is a primary study or a program evaluation report. Choose a platform and justify the choice, write an ordered list of exclusion reasons and explain why the order matters, and describe two team rules that would make the chosen platform's process consistent.
The team should use Rayyan's free plan. Without a licence or budget, Covidence is not available, and Rayyan's blind mode supports independent dual screening of 600 records. Rayyan's flexibility means the team must supply some of the structure that Covidence would impose.
The ordered exclusion reasons would be: 1, population outside criteria (not school-aged); 2, setting outside criteria (program not delivered in schools); 3, no mental health outcome; 4, ineligible publication type. The order follows the screening questions, and it matters because a report that fails several criteria must receive one reason only, the first that applies. Without a fixed order, two screeners could record different reasons for the same kind of report and the counts in the flow diagram would depend on who screened it. The order would be stated in the methods.
The first rule is that every full-text exclusion must carry exactly one reason label from the list, which the team will check before closing the stage. The second rule is that full-text screening runs as a separate round, with blind mode on, and that decisions and reasons are exported at the end of each stage and stored with the team's log, so that the flow diagram is built from exported counts rather than from memory. A further rule worth adding is that maybe counts as include at the title and abstract stage.
Minimum 20 characters required.
Question 1: In Rayyan, what does blind mode do during collaborative screening?
Question 2: Which feature distinguishes Covidence's workflow from Rayyan's, as this section describes them?
Question 3: A full text is a commentary about a hospital-based program for adults aged 50 and older. The Cedar Valley exclusion reasons are ordered 1 population, 2 intervention or setting, 3 outcome, 4 country, 5 publication type. Which reason is recorded?
Question 4: Why should keyword highlighting be treated with caution during screening?
Screening Studies: Stages, Dual Screening and Agreement
Learning Objectives for this section
- Describe the purposes of title and abstract screening and full-text screening, and apply the decision rules appropriate to each stage.
- Run a pilot screening round and use its results to revise a screening guide.
- Explain why reviews use two independent screeners, and describe how conflicts are resolved and recorded.
- Calculate percent agreement, chance-expected agreement and Cohen's kappa from a two-by-two table of screening decisions.
- Interpret kappa with care, including the reason kappa can be low when percent agreement is high.
Introduction
Screening moves through two stages. At title and abstract screening, reviewers read the short record for each item and remove those that clearly fail the eligibility criteria. At full-text screening, reviewers obtain and read the full report for each item that survived the first stage, and decide whether it meets every criterion. The two stages exist because reading 1,870 full texts would take months, while most records can be excluded safely in a few seconds from the title and abstract. The cost of the two-stage design is that the first stage must be inclusive: a relevant study excluded on its abstract never reaches a reader who could see that it belonged.
This section follows the fictional Cedar Valley screeners through both stages, the pilot that came before them, and the agreement statistics they used to check that two people were applying the same rules.
3.1 Title and Abstract Screening
The working rule at the title and abstract stage is that a record moves forward unless it clearly fails at least one criterion. When the abstract is silent about a criterion, the record moves forward. When a record has no abstract, the screener decides from the title, and moves the record forward if the title leaves the question open. Most teams use three options, include, exclude and maybe, and treat maybe as include for the purpose of moving records to the next stage. The value of the maybe option is that it tells the full-text reviewer which records were uncertain and why.
Reasons for exclusion are usually not recorded at this stage. PRISMA 2020 asks only for the number of records excluded at title and abstract, and recording reasons for more than a thousand quick exclusions would slow the work without improving the report. Reasons become mandatory at full text.
Title: Effects of a peer-led walking group on loneliness and social participation among adults aged 70 and older in a mid-sized Canadian city (fictional).
Abstract summary: A controlled before-and-after evaluation of a twelve-week walking group run by a seniors' centre, with loneliness measured on a three-item scale.
Decision: Include. The participants are older adults in the community, the program runs in the community, loneliness is an outcome, and Canada is a high-income country. Every question in the guide is answered yes.
Title: Loneliness among hospital inpatients aged 65 and older: a cross-sectional survey (fictional).
Abstract summary: A survey of loneliness among older adults admitted to medical wards, with no intervention.
Decision: Exclude. Hospital inpatients fall outside the population, and the study evaluates no intervention. The abstract settles the question, so there is no need to read the full text.
Title: A community arts program for older people: outcomes from the first year (fictional).
Abstract: None available.
Decision: Maybe, which moves the record forward. The title suggests a community-based program for older adults with reported outcomes, but it does not say how old the participants were or whether loneliness was measured. Only the full text can answer those questions.
3.2 Full-Text Screening
At full text the team first has to obtain the reports. Most come through the university library; others come through interlibrary loan, from the authors, or from organization websites. Reports that cannot be obtained after reasonable effort are counted as reports not retrieved and listed in the review, which is more honest than leaving them out of the count. The Cedar Valley team sought 145 reports and could not retrieve three: a doctoral dissertation indexed in PsycINFO that was not available from the awarding university, and two articles from journals the library could not supply within the review's timeline after requests to the authors went unanswered. That left 142 reports to assess.
Full-text screening applies every criterion strictly. Each excluded report receives exactly one reason, the first reason in the ordered list from Section 2 that applies. This is also the stage at which reports are linked to studies: when two or more reports describe the same study, the reviewer links them so that the study is counted once. The Cedar Valley team included 40 reports, which described 38 studies, because two studies had each been reported in both a conference abstract and a journal article.
| Exclusion reason (in the order applied) | Reports excluded | Example from the Cedar Valley full texts |
|---|---|---|
| 1. Population outside criteria | 32 | Participants aged 50 and older, with no separate results for those aged 65 and older. |
| 2. Intervention or setting not community-based | 21 | A visiting program delivered only to residents of long-term care homes. |
| 3. No loneliness or social isolation outcome | 27 | A group exercise program that measured only physical function and falls. |
| 4. Not conducted in a high-income country | 9 | A community program evaluated in a middle-income country. |
| 5. Ineligible publication type | 13 | A protocol for a trial that has not yet reported results. |
| Total excluded | 102 | 142 reports assessed − 102 excluded = 40 reports included. |
3.3 Piloting the Screening Guide
A pilot (sometimes called a calibration exercise) is a small round of screening that both screeners complete before the main screening begins. Each screener works through the same random sample of records independently, and the team then compares the decisions, discusses every disagreement, and revises the guide. The pilot catches criteria that are ambiguous, rules that the guide leaves unstated, and differences in how screeners read the same words. Teams repeat the pilot with a fresh sample if agreement is poor or if the guide changed substantially.
The evidence officer and the intern each screened the same random sample of 200 of the 1,870 records, with maybe counted as include. They agreed on 186 records and disagreed on 14. When they discussed the disagreements, they found that nine involved samples of adults aged 60 and older with no separate results for those aged 65 and older, three involved telephone befriending programs that one screener had not counted as community-based, and two were simple misreadings. The first version of the guide had shortened the protocol's age rule and left out its statement on telephone delivery, so the team restored the 80 percent rule with its mean or median age fallback and the statement that telephone and online delivery count as community-based, which are the rules shown in the Section 2 guide. The two screeners then screened the remaining 1,670 records independently with the revised guide.
3.4 Dual Independent Screening and Resolving Conflicts
In dual independent screening, two reviewers screen every record without seeing each other's decisions, and the team resolves each disagreement afterwards. The method rests on a simple observation: careful people make errors when they read hundreds of abstracts, and two people rarely make the same error on the same record. Methodological studies support the practice. A methodological systematic review comparing single with double screening found that single screening missed more eligible studies (Waffenschmidt et al., 2019), and a randomized trial of abstract screening reached the same conclusion for single-reviewer screening (Gartlehner et al., 2020). The Cochrane Handbook recommends that at least two people, working independently, decide whether each study meets the eligibility criteria (Higgins et al., 2023).
Conflicts are resolved in one of three ways. The two screeners can discuss the record and agree; a third reviewer can decide; or the team can use discussion first and refer anything unresolved to the third reviewer, which is the Cedar Valley approach. Whatever the method, the protocol should state it, and the platform should record the final decision. A record of conflicts is also useful in its own right, because a run of conflicts on one criterion signals that the guide needs another rule.
Rapid reviews sometimes reduce dual screening, for example by having one person screen most records while a second person screens a sample or checks all exclusions. These shortcuts save time at a known cost in missed studies, and Lesson 10 examines the guidance on when they are acceptable. The Cedar Valley protocol planned a shortcut of this kind: two reviewers would screen the first 20 percent of titles and abstracts (374 records, starting with the pilot sample), and one reviewer would then screen the rest while a second checked every exclusion. In the first 374 records, 345 (92 percent) were excluded, so checking every exclusion would have saved only about 3 percent of the work. The team amended its protocol on 20 February 2026, at the end of week 5, and screened every remaining record in duplicate, because the number of records was manageable and the planning team would rely on the review to choose a program model. Full texts were screened by two reviewers independently, as the protocol planned.
3.5 Measuring Agreement Between Screeners
Teams often summarize how well two screeners agreed, either for the pilot, to decide whether the guide is ready, or for each stage, to report in the methods. The two most common statistics are percent agreement and Cohen's kappa. Both start from a two-by-two table that cross-classifies the two screeners' decisions.
| Cedar Valley pilot (200 records) | Officer: include | Officer: exclude | Intern total |
|---|---|---|---|
| Intern: include | 22 (both include) | 8 | 30 |
| Intern: exclude | 6 | 164 (both exclude) | 170 |
| Officer total | 28 | 172 | 200 |
Background: percent agreement
Percent agreement (also called observed agreement, po) is the proportion of items on which two people who classify the same items independently make the same decision. In the Cedar Valley pilot the screeners agreed on 22 includes and 164 excludes, so observed agreement is (22 + 164) ÷ 200 = 186 ÷ 200 = 0.93, or 93 percent. The statistic is easy to compute and explain, and it has two limits. First, some agreement arises by chance: when most records are obvious exclusions, two screeners will agree on most of them whatever their skill, so high percent agreement can hide poor agreement on the few records that matter. Second, agreement measures consistency and does not measure accuracy: two screeners who misread a criterion in the same way agree perfectly and are both wrong. HSCI 207 Lesson 8, Section 4.5 (Collecting Quantitative Data), teaches percent agreement on 30 pilot charts abstracted by two people, with a worked example of agreement that arises by chance, and is optional reading. The rest of this section deals with the first limit through Cohen's kappa; the screening guide, the pilot and the discussion of conflicts deal with the second.
Agreement expected by chance
Cohen's kappa corrects for this by asking how much agreement two screeners would reach if each made decisions at their own overall rate but independently of the content of the records. The intern included 30 of 200 records (0.15) and the officer included 28 of 200 (0.14). If their decisions were unrelated, the expected proportion of records that both include is 0.15 × 0.14 = 0.021, and the expected proportion that both exclude is 0.85 × 0.86 = 0.731. The chance-expected agreement is the sum, 0.021 + 0.731 = 0.752. In counts, this is 4.2 expected joint includes and 146.2 expected joint excludes, or 150.4 of 200 records.
Cohen's kappa
Cohen's kappa (Cohen, 1960) is the agreement achieved beyond chance, divided by the largest possible agreement beyond chance.
Formulas for agreement between two screeners
Observed agreement: po = (both include + both exclude) ÷ total records
Chance-expected agreement: pe = (proportion included by screener 1 × proportion included by screener 2) + (proportion excluded by screener 1 × proportion excluded by screener 2)
Cohen's kappa: κ = (po − pe) ÷ (1 − pe)
Cedar Valley pilot: κ = (0.93 − 0.752) ÷ (1 − 0.752) = 0.178 ÷ 0.248 = 0.72
Interpreting kappa
Kappa equals 1 when agreement is perfect, 0 when agreement is no better than chance, and falls below 0, to a minimum of −1, when screeners agree less often than chance would predict. The most often cited labels for values in between come from Landis and Koch (1977), who proposed them as convenient benchmarks for discussion. McHugh (2012) argued that these labels are too generous for health research, where a kappa in the 0.60s still leaves a substantial share of decisions in doubt. Labels are therefore a starting point for judgement.
| Kappa | Label proposed by Landis and Koch (1977) |
|---|---|
| Below 0 | Poor |
| 0.00 to 0.20 | Slight |
| 0.21 to 0.40 | Fair |
| 0.41 to 0.60 | Moderate |
| 0.61 to 0.80 | Substantial |
| 0.81 to 1.00 | Almost perfect |
The Cedar Valley pilot's kappa of 0.72 falls in the band Landis and Koch called substantial. The more useful reading comes from the table itself. Of the 36 records that at least one screener wanted to include (22 + 8 + 6), the screeners agreed on only 22, which is the agreement that matters most at this stage. That observation, together with the pattern in the disagreements, is what led the team to revise its guide before continuing.
High agreement, low kappa
Kappa depends on the proportion of records that are included. When very few records are relevant, chance-expected agreement is already close to 1, so even a small number of disagreements produces a low kappa. Feinstein and Cicchetti (1990) described this as a paradox of high agreement and low kappa. It is common in screening, where included records are usually a small minority: in Cedar Valley, the 40 included reports came from 1,870 screened records, about 2 percent. For this reason, a review should report the two-by-two table or the number of conflicts alongside kappa, and should judge agreement mainly by how the screeners did on the records that either of them included.
Two screeners on another team each screen the same 200 records. Both include 2; screener 1 includes 5 that screener 2 excludes; screener 2 includes 4 that screener 1 excludes; both exclude 189. Calculate percent agreement, chance-expected agreement and kappa, and say what the results suggest.
Observed agreement is (2 + 189) ÷ 200 = 0.955. Screener 1 included 7 of 200 (0.035) and screener 2 included 6 of 200 (0.03). Chance-expected agreement is (0.035 × 0.03) + (0.965 × 0.97) = 0.00105 + 0.93605 = 0.9371. Kappa is (0.955 − 0.9371) ÷ (1 − 0.9371) = 0.0179 ÷ 0.0629 = 0.28. Percent agreement looks excellent, but of the 11 records that either screener wanted to include, the screeners agreed on only 2. The low kappa reflects a real problem: the screeners are applying the criteria differently to exactly the records that matter, and the team should discuss the 9 conflicts and revise its guide before going further.
A clear methods statement names the stage, the number of records, the statistic and the way disagreements were resolved, for example: "Two reviewers independently screened a pilot sample of 200 records (agreement 93 percent, kappa 0.72) and, after revising the screening guide, screened all remaining records independently; disagreements were resolved by discussion or by a third reviewer."
When three or more screeners each rate the same records, Fleiss' kappa extends the same idea to several raters. When different pairs of screeners share the work, teams usually report agreement for each pair or for the pilot only.
Some authors also report the prevalence-adjusted bias-adjusted kappa (PABAK), which equals 2 × po − 1 for two categories (Byrt et al., 1993). For the Cedar Valley pilot, PABAK is 2 × 0.93 − 1 = 0.86. It removes the effect of a low proportion of includes, which is useful for comparison, and for the same reason it can make agreement on the rare includes look better than it was.
Two screeners who make the same mistake agree perfectly. Kappa measures consistency between people, and the screening guide, the pilot and the discussion of conflicts are what bring the decisions in line with the criteria.
With 38 studies included from the database search, and the reports and exclusions counted at each stage, the Cedar Valley team has the numbers it needs for the left-hand column of its flow diagram. Section 4 adds the grey literature and citation chasing and draws the complete PRISMA 2020 flow diagram.
Reflection
Two reviewers independently assessed 142 full-text reports. Both included 36 reports; reviewer A included 4 that reviewer B excluded; reviewer B included 3 that reviewer A excluded; both excluded 99. Observed agreement is po = (both include + both exclude) ÷ total. Chance-expected agreement is pe = (proportion included by A × proportion included by B) + (proportion excluded by A × proportion excluded by B). Cohen's kappa is (po − pe) ÷ (1 − pe). Landis and Koch (1977) labelled kappa values of 0.61 to 0.80 substantial and 0.81 to 1.00 almost perfect. Calculate po, pe and kappa, showing your working. Interpret the result, commenting on agreement among the reports that at least one reviewer wanted to include. Explain what the team should do with the disagreements, state how many reports it would include if discussion led it to include 4 of the disputed reports, and write one sentence reporting this agreement in the methods.
Observed agreement is (36 + 99) ÷ 142 = 135 ÷ 142 = 0.9507. Reviewer A included 40 of 142 reports (0.2817) and reviewer B included 39 (0.2746). Chance-expected agreement is (0.2817 × 0.2746) + (0.7183 × 0.7254) = 0.0774 + 0.5210 = 0.5984. Kappa is (0.9507 − 0.5984) ÷ (1 − 0.5984) = 0.3523 ÷ 0.4016 = 0.88.
A kappa of 0.88 falls in the band that Landis and Koch called almost perfect. Agreement is also good where it matters most: of the 43 reports that at least one reviewer wanted to include (36 + 4 + 3), the reviewers agreed on 36. Because more than a quarter of full texts were included, chance-expected agreement is much lower than at the title and abstract stage, and kappa gives a fair summary here.
The 7 disagreements should be resolved by discussion, with a third reviewer deciding any that remain, and each final decision and its reason should be recorded in the platform. If discussion led the team to include 4 of the 7, it would include 36 + 4 = 40 reports. A methods sentence could read: "Two reviewers independently assessed all 142 full-text reports (agreement 95 percent, kappa 0.88); the 7 disagreements were resolved by discussion or by a third reviewer."
Minimum 20 characters required.
Question 1: Which rule best describes title and abstract screening?
Question 2: In a pilot of 200 records, two screeners both include 22 records and both exclude 164. What is the observed agreement?
Question 3: Why can kappa be low when percent agreement is high in screening?
Question 4: The Cedar Valley team sought 145 full-text reports and could not obtain 3 of them. How should those 3 be handled?
Reporting Study Selection with the PRISMA 2020 Flow Diagram
Learning Objectives for this section
- Explain what the PRISMA 2020 flow diagram reports and which PRISMA 2020 items it supports.
- Describe how the 2020 diagram differs from the 2009 version, including the distinction between records, reports and studies and the separate column for other methods.
- Choose the PRISMA 2020 template that fits a new or updated review.
- Draw a complete flow diagram for the Cedar Valley review and check that every number reconciles with the team's log.
- Adapt the diagram for a scoping review and recognize the errors that most often appear in published flow diagrams.
Introduction
The PRISMA 2020 statement (Preferred Reporting Items for Systematic reviews and Meta-Analyses) is a 27-item checklist of what a systematic review report should contain, published with an explanation and elaboration paper and a set of flow diagram templates (Page et al., 2021). Item 16a asks authors to describe the results of the search and selection process, from the number of records identified to the number of studies included, ideally with a flow diagram. Item 16b asks authors to cite studies that might appear to meet the inclusion criteria but were excluded, and to explain why. The flow diagram answers item 16a in a single figure: it shows how many records the searches found, how many were removed at each stage and why, and how many studies and reports the review finally included.
The diagram is the public face of the record management and screening work in Sections 1 to 3. Every number in it should come straight from the team's log, and a reader should be able to add and subtract their way from the top of the diagram to the bottom. This section explains the structure of the 2020 diagram, draws the complete diagram for the fictional Cedar Valley review, and shows how to check it.
4.1 What the Flow Diagram Shows
A reader looks at a flow diagram to answer practical questions. How large was the search? How many records were duplicates? What share of records survived title and abstract screening? Were any reports impossible to obtain? Why were reports excluded at full text, and was any one reason dominant? How many studies did the review include, and how many reports described them? Answers to these questions help the reader judge whether the search was broad enough, whether the criteria were applied as described, and how the review's results might change if, for example, the unretrieved reports had been found.
The diagram does not describe the search strategies themselves. Those belong in the methods and appendices, reported according to PRISMA-S (Rethlefsen et al., 2021), which Lesson 4 introduced. The diagram also does not replace the list of excluded studies that item 16b asks for; it summarizes the reasons as counts, while the list names the specific studies that a reader might have expected to see included.
4.2 From PRISMA 2009 to PRISMA 2020
The original PRISMA statement (Moher et al., 2009) included a four-phase flow diagram labelled identification, screening, eligibility and included. The 2020 update kept the top-to-bottom design and changed its content in several ways that reflect how reviews are now conducted.
4.3 Choosing a Template
PRISMA 2020 offers templates for two kinds of review and two kinds of search. A new review starts from nothing; an updated review builds on a previous version and reports both the earlier and the new studies. A review that searched only databases and registers needs one column; a review that also searched websites, contacted organizations or chased citations needs the version with a second column for other methods.
| Template | Use it when |
|---|---|
| New review, databases and registers only | The review is new and every record came from bibliographic databases or study registers. |
| New review, databases, registers and other sources | The review is new and the team also searched websites, contacted organizations, chased citations or used other methods. The Cedar Valley review uses this template. |
| Updated review, databases and registers only | The review updates an earlier version and the new search used only databases and registers. |
| Updated review, databases, registers and other sources | The review updates an earlier version and the new search also used other methods. |
4.4 The Cedar Valley Flow Diagram
The Cedar Valley review searched five databases (Sections 1 to 3), searched grey literature, trial registries and websites and contacted organizations (Lesson 5), and chased citations from the included studies and from relevant reviews (Lesson 5). It therefore uses the template for new reviews with other sources. The left-hand column summarizes Sections 1 to 3 of this lesson. The right-hand column summarizes the other methods. The grey-literature searches produced 771 records and items, of which 214 were duplicates; the remaining 557 were screened, and 61 were sought as full reports. Organizations contacted during the environmental scan sent 7 documents, all of which were sought. Citation chasing retrieved 2,448 records, of which 1,288 were duplicates or had already been screened; the remaining 1,160 were screened, and 23 were sought. In all, 91 reports were sought. Two could not be obtained (a program report that an organization had withdrawn from its website and an internal evaluation that its authors could not share), 89 were assessed, and 30 were included: 26 grey-literature documents, such as program evaluation reports and agency reports, and 4 reports of 4 further studies found by citation chasing. PRISMA 2020 places study registers beside databases in the left-hand column. The Cedar Valley team searched the trial registries in week 5 with its other grey-literature sources and screened their records in that set, so it counted them in the right-hand column and explained the adaptation in its methods; none of the registry records led to an included study.
The final box reports three numbers. The review includes 42 studies (38 from the databases and 4 from citation chasing), described in 44 reports (40 and 4). It also includes 26 grey-literature documents. The team decided at the protocol stage to chart grey-literature documents separately from research studies because their methods are often reported briefly, and it adapted the final box to show them. PRISMA 2020 presents its flow diagram as a template that authors can adapt, and an adaptation of this kind should be explained in the methods.
4.5 Checking That the Numbers Reconcile
Before a flow diagram goes into a report, someone other than the person who drew it should check every number against the log. The check is arithmetic: each box should equal the box above it minus the exclusions beside it, the reasons for exclusion should add up to their total, and the included counts should add up across columns.
Reconciliation checks for the Cedar Valley diagram
Database sources: 742 + 816 + 388 + 296 + 238 = 2,480 records identified.
Before screening: 2,480 − 610 duplicates − 0 automation − 0 other = 1,870 records screened.
Title and abstract: 1,870 − 1,725 excluded = 145 reports sought.
Retrieval: 145 − 3 not retrieved = 142 reports assessed.
Full text: 32 + 21 + 27 + 9 + 13 = 102 excluded; 142 − 102 = 40 reports included, describing 38 studies.
Other methods: 771 + 7 + 2,448 = 3,226 records identified; 214 + 1,288 = 1,502 removed before screening; 3,226 − 1,502 = 1,724 records screened; 1,724 − 1,633 excluded = 91 reports sought (61 + 7 + 23); 91 − 2 = 89 assessed; 15 + 10 + 22 + 12 = 59 excluded; 89 − 59 = 30 included (26 grey-literature documents and 4 study reports).
Included: 38 + 4 = 42 studies; 40 + 4 = 44 reports of included studies; 26 grey-literature documents.
Two of these steps deserve attention. The first is the move from records to reports between title and abstract screening and retrieval. In Cedar Valley, each record that passed screening led to one report, so the number of reports sought equals the number of records that survived. In some reviews a single record leads to more than one report, or the team discovers at retrieval that two records point to the same document, and the diagram should then show the change with a note. The second is the move from reports to studies in the final box. The number of studies can never exceed the number of included reports, and when it is smaller the methods should say how reports were linked to studies.
| Box in the diagram | Where the number comes from |
|---|---|
| Records identified from databases and registers | The import log in Section 1, by database. |
| Records removed before screening | The de-duplication log: 571 automated, 27 manual and 12 found by the platform, totalling 610. |
| Records screened and excluded | The screening platform's counts at the title and abstract stage, checked against 1,870 − 145. |
| Reports sought, not retrieved and assessed | The retrieval log kept by the intern, listing every report requested and its source. |
| Reports excluded, with reasons | The full-text decisions, each with one reason from the ordered list. |
| Records identified from other methods | The grey-literature and citation-chasing logs from Lesson 5. |
| Studies and reports included | The list of included studies, with the reports linked to each. |
4.6 Adapting the Diagram for Scoping and Rapid Reviews
The Cedar Valley review is a rapid scoping review, and its reporting follows PRISMA-ScR, the extension for scoping reviews (Tricco et al., 2018). PRISMA-ScR asks for the same information about selection as PRISMA 2020, using the term "sources of evidence" because a scoping review may include documents that are not research studies. A scoping review can therefore use the PRISMA 2020 template with labels that fit its sources, as the Cedar Valley team did. A rapid review that used shortcuts, such as single screening of part of the records, should keep the same boxes and explain the shortcut in the methods; Lesson 10 discusses how rapid reviews report their shortcuts.
The environmental scan is reported separately. Its counts (18 programs identified, 11 responding to the survey and 7 taking part in key informant interviews) describe organizations and people, and Lesson 11 shows how to report a scan, including a simple flow of programs from identification to participation.
Tools for drawing the diagram
The PRISMA website provides the templates as editable documents. Covidence draws a PRISMA diagram from its stage counts, and the PRISMA2020 package for R and its accompanying Shiny web application produce compliant diagrams from a table of numbers (Haddaway et al., 2022). Any of these tools is acceptable. Whichever tool the team uses, the numbers must come from the team's own log, because a platform cannot know about duplicates removed in a reference manager before import or about records that were never imported.
The most common error is a box that does not equal the box above it minus the exclusions beside it. It usually arises when a diagram is drawn from memory or when screening continued after the diagram was drafted. Rechecking every subtraction against the log prevents it.
Some diagrams label the final box "articles included" and give a number that mixes reports and studies. The 2020 diagram asks for studies and reports separately, and the two numbers should match the review's tables.
A full-text exclusion box with a total and no reasons does not meet the template. Reasons are optional at the title and abstract stage and expected at full text.
Folding unobtainable reports into the exclusions makes the search look more complete than it was. The 2020 diagram has a separate box for them.
Records found by citation chasing that the database search had already found should be removed before they are counted in the right-hand column. Otherwise the diagram overstates what citation chasing added.
The abstract, the results text and the diagram should report the same numbers. Reviews that are revised after peer review often update one and forget the others.
The 2009 template lacks the boxes for reports not retrieved and for other methods. Journals that follow PRISMA 2020 expect the 2020 template.
A draft flow diagram for a review that searched only databases reports the following numbers: records identified 640; duplicates removed 152; records screened 488; records excluded 431; reports sought for retrieval 57; reports not retrieved 2; reports assessed for eligibility 57; reports excluded 44; studies included 13. Check each step and identify what is wrong.
The first three steps reconcile: 640 − 152 = 488 and 488 − 431 = 57. The error is at retrieval: if 57 reports were sought and 2 were not retrieved, only 55 can have been assessed. With 44 reports excluded, 55 − 44 = 11 reports were included, so the review cannot include 13 studies; it includes at most 11. The author of the diagram needs to check the log to see whether the number not retrieved, the number excluded or the number included is wrong, and should also report the number of reports of included studies alongside the number of studies.
The Cedar Valley team now knows which 42 studies and 26 grey-literature documents it will chart. Lesson 8 turns to data extraction and the appraisal of those studies.
Reflection
A review team drew a draft PRISMA 2020 flow diagram for a new review that searched databases and other sources. The draft reports: records identified from MEDLINE 180, CINAHL 140 and PsycINFO 92 (total 412); duplicates removed 96; records screened 326; records excluded 291; reports sought 35; reports not retrieved 1; reports assessed 34; reports excluded 26 (population 11, setting 6, outcome 9). In the other-methods column: websites 14 and citation searching 6, all 20 sought, none unretrieved, 20 assessed, 17 excluded with no reasons given, and 3 included. The final box reads "Studies included (n = 11)". The screening platform's export shows 281 records excluded at title and abstract. The review's tables list 8 studies from the database search, each described in one report, and from other methods 1 study (one report) and 2 grey-literature documents. Under PRISMA 2020, each box should equal the box above it minus the exclusions beside it, reasons are expected for full-text exclusions, and the final box reports studies and reports of included studies separately. Check each step, identify every error, and write the corrected numbers for the boxes that are wrong, including a corrected final box.
The identification box reconciles: 180 + 140 + 92 = 412. The next step does not: 412 − 96 = 316, so the records screened box should read 316 where the draft reports 326. The platform export confirms 281 exclusions at title and abstract, and 316 − 281 = 35 reports sought, which matches the draft. The draft's 326 and 291 are each 10 too high, which suggests that 10 records were counted twice, perhaps citation-chasing records added to the database column. The draft's own subtraction (326 − 291 = 35) hid the error, which is why each box must be checked against the box above it and against the log.
Retrieval and full text reconcile: 35 − 1 = 34 assessed; 11 + 6 + 9 = 26 excluded; 34 − 26 = 8 reports included, describing 8 studies. In the other-methods column, 20 − 17 = 3 included, which reconciles, but the 17 exclusions need reasons.
The final box is wrong. "Studies included (n = 11)" adds 8 studies, 1 study and 2 grey-literature documents as if all were studies. The corrected box reads: studies included in review (n = 9); reports of included studies (n = 9); grey-literature documents included (n = 2). The methods should explain that grey-literature documents are reported separately.
Minimum 20 characters required.
Question 1: What does item 16a of PRISMA 2020 ask authors to report?
Question 2: Which change did the PRISMA 2020 flow diagram introduce compared with the 2009 version?
Question 3: In the Cedar Valley diagram, 142 database reports were assessed and 102 were excluded, yet the final box credits 38 studies to the database search. Why does 142 − 102 = 40 differ from 38?
Question 4: A team with a fixed deadline screens a random sample of 300 of its 1,100 de-duplicated records and states the sampling in its methods. Where should the 800 unscreened records appear in its PRISMA 2020 flow diagram?
Final Assessment
Bringing It All Together
This lesson followed the fictional Cedar Valley evidence team from the moment its database searches finished to the moment it knew which studies it would chart. The team exported 2,480 records from five databases, imported them into Zotero by source and checked every batch, and removed 610 duplicates in automated, manual and platform passes, keeping a log that would later supply the first boxes of its flow diagram. Throughout, it kept the PRISMA 2020 distinction between records, reports and studies in view: duplicate records were merged, while different reports of one study were kept and linked at full text.
The team then set up a screening project in Covidence with a screening guide, an ordered list of exclusion reasons and blinded dual review, and compared that choice with Rayyan, which suits teams without a licence. A pilot of 200 records produced 93 percent agreement and a kappa of 0.72, and the pattern of disagreements, more than the kappa value itself, led the team to add two rules before screening the rest. Dual screening at both stages narrowed 1,870 records to 40 reports of 38 studies.
Finally, the team drew a complete PRISMA 2020 flow diagram, with a second column for grey literature, organizations and citation chasing, and checked every number against its log. The review includes 42 studies described in 44 reports, together with 26 grey-literature documents reported separately.
Key Takeaways from this lesson
- PRISMA 2020 distinguishes records (indexed titles and abstracts), reports (documents) and studies (investigations), and each stage of the review counts one of these units.
- De-duplication merges copies of the same report, while different reports of one study, such as a conference abstract and a later article, are kept and linked at full text.
- A false-positive merge removes evidence before screening, so automated matching should be cautious and uncertain pairs should be checked by a person.
- An import log that compares records exported with records imported for each source catches lost batches before screening begins.
- Screening platforms such as Covidence and Rayyan support blinded dual screening, detect conflicts and record reasons; Covidence imposes fixed stages, while Rayyan leaves more structure to the team.
- A screening guide with rules for borderline cases and an ordered list of exclusion reasons makes decisions consistent and the counts in the flow diagram reproducible.
- Title and abstract screening should be inclusive, and full-text screening applies every criterion strictly with one recorded reason for each exclusion.
- Percent agreement overstates agreement when most records are easy exclusions, and Cohen's kappa corrects for chance agreement but can be low when relevant records are rare.
- A pilot round, followed by discussion of every disagreement and revision of the guide, does more to improve screening than any single agreement statistic.
- A PRISMA 2020 flow diagram must reconcile box by box with the team's log, report studies and reports separately, and show reports not retrieved and other methods in their own boxes.
Core Concepts Reviewed
Section 1: records, reports and studies; reference managers and Zotero; RIS export and import logs; duplicate records, false-positive merges and manual checks; the audit trail.
Section 2: screening platforms; Covidence and Rayyan compared; blinding and conflicts; the screening guide; ordered exclusion reasons; keyword highlights; privacy and licences.
Section 3: title and abstract screening; full-text screening and reports not retrieved; pilot screening; dual independent screening and conflict resolution; percent agreement, chance-expected agreement and Cohen's kappa; the high agreement, low kappa paradox.
Section 4: PRISMA 2020 items 16a and 16b; changes from the 2009 diagram; choosing a template; the complete Cedar Valley diagram; reconciliation checks; PRISMA-ScR adaptations; common errors.
The final reflection asks you to turn the Cedar Valley numbers into the methods and results paragraphs that a published review would contain.
Reflection
Write the record management and study selection paragraphs for the methods and results of the fictional Cedar Valley rapid scoping review, in about 200 to 250 words of formal prose. Use these facts. The librarian searched MEDLINE (742 records), Embase (816), CINAHL (388), PsycINFO (296) and Web of Science (238). Records were imported into Zotero; duplicates were removed in an automated pass (571), a manual pass (27) and on import to Covidence (12). Two reviewers piloted the screening guide on a random sample of 200 records (186 agreements; kappa 0.72), after which rules on mixed-age samples and telephone delivery were added. Two reviewers then screened all records independently at both stages, resolving conflicts by discussion or by a third reviewer. At title and abstract, 1,725 records were excluded. Of 145 reports sought, 3 could not be retrieved; of 142 assessed, 102 were excluded (population 32, intervention or setting 21, no loneliness or isolation outcome 27, not a high-income country 9, publication type 13), and 40 reports of 38 studies were included. Other methods (557 grey-literature records and items screened, with 61 reports sought; 7 documents from organizations; 1,160 new citation-chasing records screened, with 23 reports sought) yielded 91 reports, of which 2 were not retrieved, 59 of 89 assessed were excluded, and 30 were included (26 grey-literature documents and 4 reports of 4 studies). Include the totals, refer to a PRISMA 2020 flow diagram, and end with one sentence on a limitation of the process.
Methods. Records from MEDLINE, Embase, CINAHL, PsycINFO and Web of Science were imported into Zotero, and duplicates were removed in an automated pass and a manual pass, with a final check on import to Covidence. Two reviewers piloted the screening guide on a random sample of 200 records (93 percent agreement, kappa 0.72) and revised it to clarify eligibility for mixed-age samples and telephone-delivered programs. Two reviewers then independently screened all titles and abstracts and all full texts, recording one reason for each full-text exclusion from an ordered list; disagreements were resolved by discussion or by a third reviewer.
Results. The database searches identified 2,480 records, of which 610 were duplicates. Of 1,870 records screened, 1,725 were excluded. Of 145 reports sought, 3 could not be retrieved, and 102 of the 142 assessed were excluded, most often because the population fell outside the criteria (32) or no loneliness or social isolation outcome was reported (27). Forty reports of 38 studies met the criteria. Grey-literature searches, organizations and citation chasing led to 91 further reports, of which 30 were included: 26 grey-literature documents and 4 reports of 4 studies. The review therefore includes 42 studies, described in 44 reports, and 26 grey-literature documents (Figure 1, PRISMA 2020 flow diagram). Three database reports and two reports from other sources could not be obtained, and their absence may have led the review to miss eligible studies.
Minimum 30 characters required.
Final Knowledge Assessment
Question 1: The Cedar Valley database searches retrieved 2,480 records, and 610 duplicates were removed. How many records were screened at title and abstract?
Question 2: Which pair of records should be merged during de-duplication?
Question 3: A team plans to import raw database exports directly into its screening platform instead of using a reference manager. Which consequence follows?
Question 4: In a pilot of 100 records, both screeners include 10, screener 1 alone includes 5, screener 2 alone includes 5, and both exclude 80. What is Cohen's kappa?
Question 5: Which statement about Cohen's kappa is correct?
Question 6: Where do records found by citation chasing appear in a PRISMA 2020 flow diagram?
Question 7: A team's screening guide has no rule for samples that mix ages, such as adults aged 60 and older. What is the most likely effect, and what is the remedy?
Question 8: Why are reasons for exclusion recorded at full text but usually not at title and abstract?
Question 9: Which setup fits a three-person student team with no budget and no institutional Covidence licence that wants blinded dual screening?
Question 10: A full-text report fails the population criterion and is also a commentary. Under an ordered list that places population first, how is it recorded, and why does the order matter?
Question 11: Which statement about the final box of the Cedar Valley flow diagram is correct?
Question 12: In the Cedar Valley pilot, the screeners agreed on 186 of 200 records (kappa 0.72), but of the 36 records that either wanted to include, they agreed on 22. What did the team do next?
Question 13: Why does the PRISMA 2020 flow diagram include a box for reports not retrieved?
Question 14: A team draws its flow diagram with the PRISMA2020 Shiny app. Where should the numbers come from?
Question 15: A rapid review team proposes that one person screen all records to save time. Which response reflects this lesson?
Glossary: Key Terms, People & Frameworks
📚 Reference page, available throughout the lesson
Definitions of the terms, tools and people introduced in this lesson on managing records and screening studies.