HSCI 241 · Lesson 7

Managing Records and Screening Studies

Finding & Synthesizing Health Evidence

Learning objectives for this lesson:

  • Explain why a review team gathers its search results in a reference manager and keeps a written log of counts from export to screening.
  • Distinguish records, reports and studies as PRISMA 2020 defines them, and explain why duplicate records are merged while multiple reports of one study are linked.
  • Import database exports into Zotero, check that every batch arrived, and remove duplicates with automated and manual passes.
  • Compare Covidence and Rayyan and set up a screening project with a screening guide, ordered exclusion reasons and blinded dual review.
  • Apply eligibility criteria at the title and abstract stage and the full-text stage, recording one ordered reason for each full-text exclusion.
  • Calculate and interpret percent agreement and Cohen's kappa for a pilot of dual screening, and explain why kappa can be low when agreement is high.
  • Draw a complete PRISMA 2020 flow diagram for a review that searched databases and other sources, and check that every number reconciles.

This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University, drawing on the Cochrane Handbook for Systematic Reviews of Interventions and the JBI Manual for Evidence Synthesis.

Lesson 7 · HSCI 241

Managing Records and Screening Studies

A short guided walkthrough before you work through the lesson at your own pace.

Finding and Synthesizing Health Evidence
Why this lesson matters

From search results to included studies

The steps between the search and the included studies decide whether the review keeps everything the search found.

Lost batches, mistaken merges and missing counts all weaken a review in ways that readers cannot see unless the team records each step.

The road map

Four sections

1 · Records and duplicates

Zotero, import logs and de-duplication without losing reports.

2 · Screening platforms

Covidence and Rayyan, screening guides and ordered exclusion reasons.

3 · Screening and agreement

Two screening stages, dual screening, percent agreement and kappa.

4 · PRISMA 2020 flow diagram

A complete diagram for the running case, checked number by number.

The running case

The fictional Cedar Valley review

2,480Records from five databases
1,870Records screened after removing 610 duplicates
42Studies included, plus 26 grey-literature documents
How to use this lesson

Read, practise and apply

  • Each section has a walkthrough, a reading, a reflection and a knowledge check.
  • The worked examples use the fictional Cedar Valley review throughout.
  • The Cedar Valley numbers are carried from the first import to the finished flow diagram.
Section 1 of 5

Reference Managers and De-duplication

⏱ Estimated reading time: 35 minutes
Section 1 of 5

Reference Managers and De-duplication

Records, reports and studies; import logs; finding duplicates; the audit trail.

Three units

Records, reports and studies

Record

A record is the indexed title or abstract of a report.

Report

A report is a document, such as an article or abstract, about a study.

Study

A study is the investigation, which may have several reports.

De-duplication merges records; full-text screening links reports to studies.

The tool

What a reference manager does

Zotero

Zotero is a free, open-source reference manager with group libraries and a Duplicate Items view.

RIS files

RIS is a plain-text format in which each field carries a two-letter tag such as TI for title.

One collection per database keeps each source’s count visible.

The import log

Check every batch

The log compares records exported with records imported for each database.

  • The first Embase import brought in 500 of 816 records because a second batch had been missed.
  • The log revealed the shortfall before anyone began screening.
  • All 2,480 records from five databases were accounted for.
Two kinds of error

Cautious matching

False-positive merge

Two different reports are merged, and one is lost before screening.

Missed duplicate

The same report is screened twice, which costs time and loses nothing.

Keep look-alike reports, such as an abstract and its later article, and link them at full text.

Worked example

Three passes in Cedar Valley

571Automated pass in Zotero
27Manual pass by title and author
12Flagged on import to Covidence
1,870Records left to screen

The total removed is 571 + 27 + 12 = 610 duplicates.

Carry forward

The audit trail feeds the flow diagram

  • Keep the original exports, the count log and a pre-merge copy of the library.
  • The counts become the first boxes of the PRISMA 2020 diagram.
  • Section 2 moves the 1,870 records into a screening platform.

Learning Objectives for this section

  • Explain why a review team gathers its search results in a reference manager and keeps a written record of counts from export to screening.
  • Distinguish records, reports and studies as PRISMA 2020 defines them, and explain why de-duplication works on records while the linking of reports to studies waits until full-text screening.
  • Export records from bibliographic databases in RIS format, import them into Zotero by source, and check that every batch arrived.
  • Identify and merge duplicate records with automated and manual checks, and recognize pairs of records that look alike but describe different reports.
  • Keep an audit trail of de-duplication that supplies the first numbers of the PRISMA 2020 flow diagram.

Introduction

By the end of Lessons 4 and 5, a review team holds the results of several database searches, documents from websites and organizations, and references found by citation chasing, in different file formats and with overlapping content. Before anyone can screen them, the team has to gather them in one place, count them, remove the copies, and move a clean set into a screening platform without losing a record. This work is clerical, and it is also part of the method: a review that cannot say how many records it found and how many duplicates it removed cannot complete its flow diagram, and a review that loses records during import has quietly narrowed its own search.

This section explains the units that the work counts, describes what a reference manager does, and walks through import and de-duplication with the running case. Sections 2 to 4 cover screening platforms, screening itself and the PRISMA 2020 flow diagram.

Case: the fictional Cedar Valley evidence review

The fictional Cedar Valley Health Authority in British Columbia serves about 210,000 residents, about 46,000 of them aged 65 and older, through 24 primary care clinics. Before launching a community connector (social prescribing) program for older adults, its planning team asked a small evidence team for a rapid scoping review and environmental scan within twelve weeks. The team is a health authority evidence officer, a university librarian who works with the team, and a student intern. The review asks which community-based interventions have been evaluated for reducing loneliness or social isolation among adults aged 65 and older, and with what outcomes.

The librarian has run the final searches in MEDLINE, Embase, CINAHL, PsycINFO and Web of Science, which together retrieved 2,480 records. The intern has been asked to bring these records together, remove duplicates, and prepare them for screening. All numbers about Cedar Valley in this lesson are illustrative.

1.1 Records, Reports and Studies

Three words recur throughout this lesson, and PRISMA 2020 gives each a specific meaning (Page et al., 2021). A record is the title, abstract or both of a report as it is indexed in a database or on a website. When MEDLINE returns a citation with its abstract, that citation is a record. A report is a document that supplies information about a study, such as a journal article, a preprint, a conference abstract, a thesis, a trial registry entry or a government report. A study is the investigation itself, for example a trial of a befriending program with a defined group of participants, an intervention and outcomes. One study can produce several reports, and one report can be indexed as a record in several databases.

These distinctions matter because different stages of the review count different units. De-duplication works on records: its job is to make sure that each report enters screening once, however many databases indexed it. Linking reports to studies happens later, usually at full-text screening, when the reviewer can read enough to see that a conference abstract from 2020 and a journal article from 2022 describe the same trial. If the team tried to merge reports of the same study during de-duplication, it would have to judge study identity from titles and abstracts alone, and it would risk discarding a report that holds information the full article leaves out.

Records Reports Study MEDLINE: article Embase: article CINAHL: article Embase: abstract PsycINFO: abstract Web of Science: abstract Journal article(full results, 2022) Conference abstract(early results, 2020) One study:a befriendingtrial De-duplication: 6 records become 2 Full-text stage: 2 reports, 1 study
A single befriending trial reported in a conference abstract and a journal article can appear as six records across five databases; de-duplication reduces the six records to two reports, and full-text screening links the two reports to one study.

1.2 What a Reference Manager Does

Background: Zotero basics

A reference manager is software that stores the details of each source as structured fields (authors, title, journal, year, abstract, DOI and so on), inserts citations as you write and builds the reference list in a chosen style. Zotero, the example in this lesson, is free, open-source software maintained by the Corporation for Digital Scholarship. Its basic parts are the library, which holds every item; collections, which are folders inside the library, with any item able to sit in several of them; the Zotero Connector, a browser extension that saves the article or page you are viewing, alongside an Add Item by Identifier button that retrieves an item's details from its DOI or ISBN; and the word-processor plug-in for Microsoft Word and LibreOffice (with a Zotero menu in Google Docs), which inserts citations and refreshes the bibliography. Imported details are often imperfect, so each entry should be checked. HSCI 207 Lesson 12, Section 4.4 (Writing Up Research), walks through installation and these steps in detail and is optional reading for anyone who has not used Zotero.

In a review, the same software does a different job: it serves as the team's central store of records between the search and the screening platform, holding each export in its own collection, finding duplicates and exporting a clean file for screening. Zotero's online group libraries let the whole team work from one shared library. EndNote (a commercial product from Clarivate) and Mendeley (from Elsevier) offer similar functions, and many librarians who support reviews work in EndNote. The principles in this section apply to any of them.

CollectionsClick to explore
Tags and notesClick to explore
Group librariesClick to explore
Duplicate Items viewClick to explore
Import, export and attachmentsClick to explore

Teams organize this stage in one of three ways, and the choice depends on the size of the search, the tools the institution supports and the preferences of the librarian.

The team imports every export into a reference manager such as Zotero, removes duplicates there, and exports one de-duplicated file to the screening platform. This route gives full control over de-duplication and a permanent library for later updates, at the cost of an extra transfer whose counts must be checked.

Many health sciences librarians de-duplicate in EndNote with a published sequence of field comparisons and manual checks, designed to remove most duplicates while keeping false merges rare (Bramer et al., 2016). Teams working with such a librarian often receive a de-duplicated file.

Covidence and Rayyan both look for duplicates on import, so a small team may import the raw exports directly. This route is quick, and it leaves the team dependent on the platform's matching rules.

1.3 Exporting and Importing Records

Each database interface exports records in its own way, and the team's first task is to export every record from each final search, with abstracts, in a format the reference manager can read. The RIS format is a plain-text format in which each field sits on its own line with a two-letter tag, such as TY for the type of reference, AU for an author, TI for the title, PY for the year and ER for the end of the record. A fictional record exported in RIS looks like this.

TY  - JOUR
AU  - Lindqvist, Maren
AU  - Osei, Kwame
TI  - A telephone befriending service for community-dwelling older adults: a pilot evaluation
JO  - Journal of Community Ageing
PY  - 2021
VL  - 14
SP  - 112
EP  - 124
DO  - 10.0000/fictional.2021.0112
AB  - Loneliness was measured with a three-item scale at baseline and twelve weeks...
ER  - 

Several practical rules prevent lost records. Some interfaces cap the number of records per export, so a large result set may need several batches whose sizes should add up to the search total. Each file should be named with the database and the date of the search (for example, CINAHL_2026-02-10.ris), kept unchanged as the original export, and imported into its own collection, and the number that arrives should be compared with the number exported. A shortfall usually means that a batch was skipped or that the import stopped at a malformed record.

DatabaseRecords exportedRecords in Zotero collectionCheck
MEDLINE742742Matches
Embase816816Matches (exported in two batches)
CINAHL388388Matches
PsycINFO296296Matches
Web of Science238238Matches
Total2,4802,480All records accounted for

The Cedar Valley import log above shows the check the intern completed before de-duplication began. The first import of the Embase file brought in only 500 records, because the librarian had exported the results in two batches and the second file had been saved in a different folder. The log made the shortfall obvious, and the second batch was imported before anyone started screening.

1.4 Finding and Removing Duplicates

A duplicate record is a second or later copy of the same report. Duplicates arise mainly because databases overlap: a gerontology article in a well-known journal is likely to be indexed in MEDLINE, Embase, CINAHL and PsycINFO, and to appear in Web of Science as well. Duplicates also arise within a single database when a report is indexed twice, for example once as an online-first version and again in its print issue.

De-duplication can go wrong in two directions. A false-positive merge joins two records that describe different reports, and the second report disappears from the review before anyone screens it. This error reduces recall, which Lesson 3 identified as the property that reviews protect most carefully. A false-negative leaves two copies of the same report in the set, so screeners see it twice. That error costs time and can produce inconsistent decisions on the same report, but it does not remove evidence, and it is usually caught later. For this reason, teams set automated matching to be cautious and send uncertain pairs to a person.

Matching on fields

Software decides whether two records match by comparing fields. The DOI (digital object identifier) is the strongest single identifier, but older records and conference abstracts often lack it. Titles differ across databases in punctuation, capitalization, spelling and the handling of subtitles, and author names appear as full names in one database and as initials in another. Good matching therefore normalizes fields (for example, by removing punctuation and converting text to lower case) and combines several fields, such as title, first author, year and journal.

The Cedar Valley de-duplication

The intern de-duplicated in three passes and recorded the number removed at each pass.

PassWhat was doneDuplicates removedRecords remaining
StartAll five database exports imported into Zotero.None2,480
1. AutomatedZotero's Duplicate Items view; each suggested group was checked and merged.5711,909
2. ManualRecords sorted by title and then by first author, and adjacent near-matches inspected.271,882
3. Platform checkDe-duplicated file imported into Covidence, which flagged further matches for confirmation.121,870
Total6101,870

The 27 duplicates found in the manual pass were mostly pairs in which one record lacked a DOI and the titles differed in punctuation or spelling (for example, "behaviour" in one database and "behavior" in another). The 12 found in Covidence were pairs whose author lists had been formatted differently. During the manual pass the intern also found nine pairs that looked alike but were different reports, including a conference abstract and its later journal article, and kept both records in each pair with a note for the full-text stage.

Conference abstract and later journal articlev

These are two reports of one study. Keep both records, link them to the same study at full text, and check the abstract for any results the article omits.

Preprint and published versionv

These are also two reports of one study. Keep both, note the relationship, and treat the published version as the primary report.

Protocol paper and results paperv

A protocol and the trial's results share authors and similar titles, and they are separate reports. The protocol is usually excluded at full text or kept as a supporting report, depending on the review's rules.

Correction noticev

A notice titled "Correction to" followed by the original title is a separate document. Most teams keep it with the article it corrects.

Same title, different reportsv

Series titles such as "part 1" and "part 2" can fool automated matching. Checking the year, pages and abstract usually separates them.

Try it: which of these are duplicates?

The intern's manual pass brought up the six fictional records below. Decide which pairs are duplicate records of one report, which are different reports of one study, and which are unrelated.

IDRecord
R1Okafor A, Byrne L. Intergenerational gardening and loneliness in older adults: a mixed-methods evaluation. Ageing Soc Res. 2023;9:41–55. DOI present.
R2Okafor, Adaeze; Byrne, Liam. Intergenerational gardening and loneliness in older adults - a mixed methods evaluation. 2023. No DOI.
R3Okafor A. Intergenerational gardening and loneliness: early findings. Conference abstract, 2021.
R4Chen M, Patel S. Loneliness in later life: part 1, measurement. 2019.
R5Chen M, Patel S. Loneliness in later life: part 2, interventions. 2019.
R6Correction to: Intergenerational gardening and loneliness in older adults. 2024.
Check your answerv

R1 and R2 are duplicate records of the same report: the authors, title and year match once punctuation and name formats are normalized, and R2 should be merged into R1, which carries the DOI. R3 is a different report (a conference abstract) of what is probably the same study, so it stays in the set and is linked at full text. R4 and R5 are different reports in a series and are unrelated as duplicates, even though the start of the title matches. R6 is a correction notice for R1; it is a separate document that most teams keep with R1 so that the full-text reviewer reads both.

1.5 Keeping an Audit Trail

An audit trail is a written record that would let another person repeat the steps and arrive at the same numbers. For record management it has three parts: the original export files, kept unchanged; a log of counts for each source and each de-duplication pass; and a copy of the library exported before merging, so that any merge can be checked later.

Counts to record before screening begins

The number of records retrieved from each database and register, and their total, become the first box of the PRISMA 2020 flow diagram. The number of duplicates removed becomes the "records removed before screening" box, together with any records removed by automation tools or for other reasons. The number of records sent to the screening platform becomes the "records screened" box. For Cedar Valley these are 2,480 records identified, 610 duplicates removed and 1,870 records screened, and the arithmetic 2,480 − 610 = 1,870 must hold.

The records found through grey-literature searching and citation chasing are kept in their own collections and counted separately, because PRISMA 2020 reports them in a second column of the flow diagram. Records from citation chasing should be checked against the database records before they are counted, so that a study the databases already found is not counted twice. Section 4 returns to these numbers.

With 1,870 unique records counted and logged, the Cedar Valley team is ready to move them into a screening platform. Section 2 describes the two platforms most often used for that step.

Reflection

You are de-duplicating 1,240 records exported for a review of walking programs for older adults: 520 from MEDLINE, 410 from CINAHL and 310 from PsycINFO. Under PRISMA 2020, a record is the indexed title or abstract of a report, a report is a document that supplies information about a study (such as an article, conference abstract or protocol), and a study is the investigation itself. Zotero's Duplicate Items view suggests four groups. Group 1: two records with the same authors, year, journal and pages, one with a DOI and one without, whose titles differ only by a hyphen. Group 2: a 2021 conference abstract and a 2023 journal article by the same first author with similar titles. Group 3: a 2022 trial protocol and a 2024 results paper with the same trial name in their titles. Group 4: two records titled "Walking for health: part 1" and "Walking for health: part 2" by the same authors in the same year. For each group, say whether you would merge the records and explain why, using the terms record, report and study. Then state the counts you would enter in your log if the automated pass removed 268 duplicates and the manual pass removed 14 more.

Model answer

Group 1 should be merged. The two items are duplicate records of one report: the authors, year, journal and pages match, and a hyphen is the kind of difference that databases introduce when they index the same title. I would keep the record that carries the DOI. Group 2 should not be merged. The conference abstract and the article are two different reports that probably describe one study, so both stay in the set with a note, and I would link them to the same study at full text, checking the abstract for results the article omits. Group 3 should also be kept apart. The protocol and the results paper are separate reports of one trial; the protocol will probably be excluded at full text for having no outcome data, or kept as a supporting report. Group 4 contains two different reports in a series, so the records are unrelated as duplicates despite the shared title.

For the log, I would record 520, 410 and 310 records by database, totalling 1,240 records identified; 268 duplicates removed in the automated pass and 14 in the manual pass, totalling 282 duplicates; and 1,240 − 282 = 958 records sent to screening. I would also keep the original export files and a copy of the library from before merging, so that any merge can be checked later.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: According to PRISMA 2020, how does a record differ from a report?

PRISMA 2020 defines a record as the title or abstract of a report as indexed in a database or website, and a report as a document (article, preprint, conference abstract, thesis or report) that supplies information about a study. The distinction concerns the unit counted, and it applies whatever the source; grey-literature documents are also reports.

Question 2: During de-duplication, the intern finds a 2020 conference abstract and a 2022 journal article that appear to describe the same befriending trial. What should she do?

The abstract and the article are different reports of one study. De-duplication removes copies of the same report; linking reports to studies happens at full text, where the reviewer can confirm that they describe the same trial and can check the abstract for results the article omits. Merging them at this stage risks losing information.

Question 3: Why do review teams set automated de-duplication to be cautious and send uncertain pairs to a person?

The two errors have different costs. Merging two different reports loses one of them before screening, which reduces recall. Leaving a duplicate in the set means it is screened twice, which wastes time but loses nothing. PRISMA 2020 allows automation tools and asks authors to report how they were used.

Question 4: The intern exported 816 Embase records in two batches, but only 500 arrived in the Zotero collection. What is the most likely explanation, and what should she do?

A shortfall after a batched export usually means that a batch was skipped. The import log exists to catch this, and the remedy is to import the missing file and confirm that 816 records arrive. Import does not remove duplicates in Zotero, and the log should never be altered to match a faulty import.
Section 2 of 5

Screening Platforms: Covidence and Rayyan

⏱ Estimated reading time: 30 minutes
Section 2 of 5

Screening Platforms: Covidence and Rayyan

Platform functions, two platforms compared, and setting up a screening project.

Core functions

What screening platforms do

Import and duplicate checks One record at a time Keyword highlights Blinded decisions Conflict lists Reasons and exports

Highlights direct attention, and the reviewer still reads the whole record.

Two platforms

Covidence and Rayyan

Covidence

Covidence is a subscription platform with fixed stages, required full-text reasons and PRISMA counts.

Rayyan

Rayyan has a free plan, one flexible screening space, a blind mode and team-defined rounds.

Both platforms change often, so check current features before choosing.

Setting up

Before the first record is screened

  • The team adds its screeners and names a third reviewer for conflicts.
  • The project requires two independent reviewers per record.
  • The imported count is checked against the log: 1,882 imported and 12 flagged.
  • The guide, exclusion reasons and keyword highlights are entered before screening.
The screening guide

Five questions for Cedar Valley

  • Participants are adults aged 65 and older who live in the community.
  • The intervention is delivered in the community, including by telephone or online.
  • Loneliness or social isolation is measured or explored as an outcome.
  • The study was conducted in a high-income country.
  • The report is a primary study or a program evaluation report.
Ordered reasons

One reason per excluded report

A report that fails several criteria receives the first reason in the list that applies.

Example

A commentary on a hospital program for adults aged 50 and older is recorded as population outside criteria.

Why it matters

A fixed order makes the counts reproducible, whoever screens the report.

Carry forward

Ready to screen

  • The team checks institutional policy and the licence terms for sharing full texts.
  • Decisions are exported after each stage and stored with the log.
  • Section 3 covers the pilot, both screening stages and agreement statistics.

Learning Objectives for this section

  • Describe the functions that screening platforms provide and explain why most teams prefer them to spreadsheets for dual screening.
  • Compare Covidence and Rayyan by cost model, workflow, blinding, conflict handling and recording of exclusion reasons.
  • Set up a screening project with a screening guide, an ordered list of exclusion reasons, keyword highlights and blinded dual review.
  • Write screening questions for the Cedar Valley review that turn the eligibility criteria into decisions that two reviewers can apply consistently.
  • Identify the privacy, licensing and record-keeping matters that a team should settle before screening begins.

Introduction

Screening is the process of deciding, record by record and then report by report, which items meet the review's eligibility criteria. With 1,870 records to screen, the fictional Cedar Valley team needs a place where two people can make independent decisions, where disagreements are found automatically, where reasons for exclusion are recorded, and where the numbers at each stage can be counted at the end. A spreadsheet can do some of this for a small review, but it makes blinding difficult, invites errors when people sort or overwrite rows, and leaves the team to count everything by hand. Dedicated screening platforms were built to solve these problems.

This section describes what screening platforms do, compares the two most widely used in health sciences, Covidence and Rayyan, and walks through setting up the Cedar Valley screening project. The decisions themselves, and how to measure whether two screeners agree, are the subject of Section 3.

2.1 What Screening Platforms Do

Most screening platforms share a core set of functions. They import records in RIS and similar formats and check for duplicates on the way in. They present one record at a time, with the title and abstract laid out for reading and with chosen keywords highlighted. They let each reviewer record a decision (usually include, exclude or maybe) without seeing the other reviewer's decision, a feature called blinding. They compare the two sets of decisions and list the conflicts, meaning the records on which the reviewers disagree, for resolution. They store reasons for exclusion and notes, and they export decisions so that the team can keep its own copy and report its numbers.

Title and abstract screening: record 214 of 1,870 A telephone befriending service for community-dwelling older adults Abstract: We tested a loneliness intervention in which volunteers called older adults weekly for twelve weeks. Unlike residential care programs, the service reached people at home. Loneliness fell on a three-item scale... Include Maybe Exclude Reviewer 2: vote hidden until both have voted Inclusion keyword Exclusion keyword Notes and labels
A schematic of a generic screening screen: the reviewer reads one record, sees inclusion and exclusion keywords highlighted, and records a decision without seeing the other reviewer's vote.

Keyword highlighting deserves a caution. Highlighting the word "residential" in red helps a reviewer notice that a study may have been set in long-term care, but in the record above the word appears in a sentence that contrasts the program with residential care. Highlights direct attention; they do not make decisions, and a reviewer who screens by colour alone will make errors.

2.2 Covidence

Covidence is a web-based platform produced by Veritas Health Innovation, a not-for-profit organization in Melbourne, Australia, and it is widely used by Cochrane review teams. It is a subscription service, and many universities and health organizations hold licences for their staff and students. Covidence organizes a review as a fixed sequence of stages: import, title and abstract screening, full-text review, and then data extraction and quality assessment, which Lesson 8 covers.

At the title and abstract stage, each reviewer votes yes, no or maybe on each record. By default Covidence asks for two votes on every record, keeps each reviewer's votes hidden from the other, and moves a record forward or out when the votes agree. Records with disagreeing votes go to a list of conflicts for the team to resolve. At the full-text stage, reviewers upload or attach the full text, and an exclusion must be accompanied by a reason chosen from a list that the team writes in advance. Covidence keeps a running count of records and reports at each stage and can produce a PRISMA flow diagram from those counts, which the team should still check against its own log. It also lets the team mark several reports as belonging to one study.

2.3 Rayyan

Rayyan was developed at the Qatar Computing Research Institute and described by Ouzzani and colleagues as a web and mobile application for systematic reviews (Ouzzani et al., 2016). It offers a free plan with limits on features and paid plans for larger teams. Rayyan presents all imported records in a single screening space, where each reviewer marks records as include, exclude or maybe and can add exclusion reasons, labels and notes. In blind mode, which collaborative reviews use during screening, no reviewer can see another's decisions; when the team turns blind mode off, the platform shows where decisions conflict, and the team resolves them.

Rayyan is more flexible than Covidence and less structured. It does not force a sequence of stages, so the team decides how to run full-text screening, for example as a second round in a separate review. Its filters, keyword highlights and labels make it easy to sort records, and it detects possible duplicates on import. The flexibility suits small teams and student projects, and it places more responsibility on the team to keep its stages and counts organized.

FeatureCovidenceRayyan
Cost modelPaid subscription, often through an institutional licence.Free plan with limits, and paid plans with more features.
WorkflowFixed stages from import to extraction.One flexible screening space; the team defines its rounds.
BlindingVotes are hidden from the other reviewer by default.Blind mode hides decisions until the team turns it off.
ConflictsListed automatically for resolution.Shown when blind mode is off, and filtered for resolution.
Exclusion reasonsRequired at full text, from the team's list.Added by reviewers; requiring them is a team rule.
Reports and studiesSeveral reports can be merged into one study.Relationships are recorded with labels and notes.
Flow diagram countsCounts kept by stage and drawn as a PRISMA diagram.Decisions can be exported; the team compiles or checks the counts.
Machine learningBoth platforms offer features that rank records or predict relevance from earlier decisions. Lesson 6 explains how such features work and how far to rely on them.

Both platforms change often, and the features above should be checked against the current version before a team commits to one. For the Cedar Valley review, the librarian's university holds a Covidence licence, and the team chose Covidence because its fixed stages and required exclusion reasons would make the PRISMA counts straightforward for a team that included a first-time screener. A student team without a licence could run the same process in Rayyan's free plan by setting its own rules for reasons and rounds.

EPPI-ReviewerClick to explore
DistillerSRClick to explore
ASReviewClick to explore
SpreadsheetsClick to explore

2.4 Setting Up a Screening Project

A screening project should be set up before anyone screens a record, because changing the rules part of the way through makes earlier decisions inconsistent with later ones. The steps below apply to any platform.

1. Create the project and add the teamv

Create one review, invite each screener, and decide who will resolve conflicts. In Cedar Valley, the evidence officer and the intern screen, and the librarian acts as third reviewer for conflicts that discussion does not settle.

2. Set the number of reviewers per recordv

Set two independent reviewers per record at both stages unless the protocol states otherwise. Section 3 discusses when a rapid review might reduce this.

3. Import the de-duplicated records and check the countv

Import the de-duplicated file and confirm that the platform shows the number in the log. Record any further duplicates the platform finds. For Cedar Valley, 1,882 records were imported, Covidence flagged 12 duplicates, and 1,870 records entered screening.

4. Write the screening guide into the projectv

Enter the eligibility criteria and the screening questions where screeners will see them. Covidence and Rayyan both allow the criteria to be displayed or attached to the project.

5. Set the exclusion reasons in orderv

Write a short list of exclusion reasons that match the criteria, and put them in a fixed order. A report that fails several criteria receives the first reason in the list that applies, which keeps the counts in the flow diagram consistent.

6. Add keyword highlightsv

Add inclusion terms such as lonel*, isolat*, befriend* and social prescri*, and exclusion terms such as nursing home, inpatient and adolescent. Check that the platform handles truncation in the way the team expects.

7. Plan the pilotv

Choose a random sample of records for both screeners to screen first, and agree to meet afterwards to compare decisions and revise the guide. Section 3 describes the pilot and how to measure agreement.

The Cedar Valley screening guide

A screening guide turns eligibility criteria into questions that a screener can answer from a title, an abstract or a full text. Each question should have a clear rule for the common borderline cases, because borderline cases cause most disagreements. The Cedar Valley guide below follows the review's Population, Concept and Context (PCC) question from Lesson 2: adults aged 65 and older who live in the community, community-based interventions evaluated for their effects on loneliness or social isolation, and studies conducted in high-income countries.

QuestionRule for borderline casesExclusion reason if the answer is no
1. Are the participants adults aged 65 and older who live in the community?Include if at least 80 percent of participants are 65 or older, if results are reported separately for this group, or, when the age distribution is not reported, if the mean or median age is 65 or older. Residents of long-term care and hospital inpatients are excluded.1. Population outside criteria
2. Is the intervention delivered in the community?Homes, community centres, libraries, faith settings, telephone and online delivery all count. Programs delivered only in hospitals or long-term care do not.2. Intervention or setting not community-based
3. Was loneliness or social isolation measured or explored as an outcome?Any validated scale, single question or qualitative exploration counts. A study of depression alone does not count.3. No loneliness or social isolation outcome
4. Was the study conducted in a high-income country?Use the World Bank income classification for the year of the study; multi-country studies count if results are reported for high-income countries.4. Not conducted in a high-income country
5. Is the report an eligible type of evidence source?Primary studies of any design and program evaluation reports are eligible. Commentaries, editorials and protocols without results are excluded. Reviews are excluded, and their reference lists are checked in citation chasing.5. Ineligible publication type

Why the order of reasons matters

A commentary about a hospital-based program for adults aged 50 and over fails questions 1, 2 and 5. Under the ordered list it is recorded once, as "population outside criteria", because that is the first reason that applies. Without a fixed order, one screener might record the population and another the publication type, and the counts in the flow diagram would depend on who screened which report. The order should be stated in the review's methods.

2.5 Privacy, Licences and Records

Bibliographic records do not contain information about study participants, so screening carries little privacy risk. Three related matters still need a decision before screening begins. The first is the institution's policy on cloud-based tools, which some health authorities restrict, and which may determine whether staff can use a given platform. The second is the licence under which full texts are obtained: PDFs supplied through a university library are licensed to that library's users, and the team should share them only among members entitled to access them. The third is the team's own copy of its decisions. Platforms can change their plans, and accounts can lapse, so the team should export the screening decisions and exclusion reasons at the end of each stage and store them with the audit trail described in Section 1.

Try it: draft three screening questions

A review of walking programs for older adults has these eligibility criteria: participants are adults aged 60 and older living in the community; the program is a structured walking program; physical activity or mobility is measured as an outcome; and the report is a primary study with results. Write three screening questions in the style of the Cedar Valley guide, each with a rule for one borderline case that screeners are likely to meet often, and an exclusion reason. Put the reasons in the order you would apply them, and write one sentence explaining the order.

With the platform set up and the guide written, the Cedar Valley screeners are ready to begin. Section 3 follows them through a pilot, through title and abstract screening and full-text screening, and through the calculation of their agreement.

Reflection

A three-person student team will screen about 600 records for a scoping review of school-based mental health programs. The team wants blinded dual screening, a recorded reason for every full-text exclusion, and accurate counts for a PRISMA 2020 flow diagram. Its university holds no Covidence licence, and the team has no budget. The two options are as follows. Covidence is a paid subscription service, usually accessed through an institutional licence; it moves records through fixed stages (title and abstract, full text, extraction), hides each reviewer's votes from the other by default, lists conflicts automatically, requires an exclusion reason from the team's list at full text, and produces PRISMA counts. Rayyan offers a free plan with limits and paid plans; it has one flexible screening space, a blind mode that hides decisions until the team turns it off, and exclusion reasons and labels that reviewers add, and the team compiles or checks its own counts from exported decisions. The review's criteria are: (1) participants are school-aged children or adolescents; (2) the program is delivered in schools; (3) a mental health outcome is measured; (4) the report is a primary study or a program evaluation report. Choose a platform and justify the choice, write an ordered list of exclusion reasons and explain why the order matters, and describe two team rules that would make the chosen platform's process consistent.

Model answer

The team should use Rayyan's free plan. Without a licence or budget, Covidence is not available, and Rayyan's blind mode supports independent dual screening of 600 records. Rayyan's flexibility means the team must supply some of the structure that Covidence would impose.

The ordered exclusion reasons would be: 1, population outside criteria (not school-aged); 2, setting outside criteria (program not delivered in schools); 3, no mental health outcome; 4, ineligible publication type. The order follows the screening questions, and it matters because a report that fails several criteria must receive one reason only, the first that applies. Without a fixed order, two screeners could record different reasons for the same kind of report and the counts in the flow diagram would depend on who screened it. The order would be stated in the methods.

The first rule is that every full-text exclusion must carry exactly one reason label from the list, which the team will check before closing the stage. The second rule is that full-text screening runs as a separate round, with blind mode on, and that decisions and reasons are exported at the end of each stage and stored with the team's log, so that the flow diagram is built from exported counts rather than from memory. A further rule worth adding is that maybe counts as include at the title and abstract stage.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: In Rayyan, what does blind mode do during collaborative screening?

Blind mode keeps each reviewer's decisions hidden from the others, so that the decisions are independent. When the team turns it off, the platform shows where decisions conflict. Masking author names is a different practice, and blind mode does not hide abstracts or records.

Question 2: Which feature distinguishes Covidence's workflow from Rayyan's, as this section describes them?

Covidence organizes a review as fixed stages (import, title and abstract, full text, extraction) and requires an exclusion reason from the team's list at full text. Rayyan presents records in one flexible screening space where the team defines its own rounds. Covidence is a paid subscription, Rayyan has a free plan, and both check for duplicates on import.

Question 3: A full text is a commentary about a hospital-based program for adults aged 50 and older. The Cedar Valley exclusion reasons are ordered 1 population, 2 intervention or setting, 3 outcome, 4 country, 5 publication type. Which reason is recorded?

Under an ordered list, each excluded report receives one reason, the first that applies. The report fails questions 1, 2 and 5, so it is recorded as population outside criteria. Recording several reasons would make the reasons add up to more than the number of reports excluded.

Question 4: Why should keyword highlighting be treated with caution during screening?

Highlights direct attention to words that often signal inclusion or exclusion, but a word such as "residential" can appear in a sentence contrasting the program with residential care. The reviewer still has to read the record. Highlights do not make decisions, and they are used at the title and abstract stage.
Section 3 of 5

Screening Studies: Stages, Dual Screening and Agreement

⏱ Estimated reading time: 40 minutes
Section 3 of 5

Screening Studies: Stages, Dual Screening and Agreement

Two stages, a pilot, dual screening, and percent agreement and kappa.

Two stages

Inclusive first, strict second

Title and abstract

Records move forward unless they clearly fail a criterion, and maybe counts as include.

Full text

Every criterion is applied, and each exclusion receives one ordered reason.

Cedar Valley at full text

From 145 reports to 38 studies

145Reports sought
3Not retrieved
102Excluded with reasons
40Reports of 38 studies
Pilot and dual screening

Calibrate, then screen independently

  • The pilot of 200 records produced 14 disagreements.
  • Nine involved mixed-age samples and three involved telephone programs.
  • The team added two rules and kept full dual screening at both stages.
  • Single screening misses more eligible studies (Waffenschmidt et al., 2019; Gartlehner et al., 2020).
Percent agreement

The pilot two-by-two table

Both include: 22

Both screeners included these records.

Intern only: 8

Only the intern included these records.

Officer only: 6

Only the officer included these records.

Both exclude: 164

Both screeners excluded these records.

Observed agreement is (22 + 164) ÷ 200 = 0.93.

Cohen’s kappa

Agreement beyond chance

Cedar Valley pilot
\[ \kappa = \frac{p_o - p_e}{1 - p_e} = \frac{0.93 - 0.752}{1 - 0.752} = \frac{0.178}{0.248} = 0.72 \]

Chance-expected agreement is (0.15 × 0.14) + (0.85 × 0.86) = 0.752.

A paradox

High agreement, low kappa

When relevant records are rare, chance agreement is high and kappa falls quickly.

0.955Observed agreement in the practice example
0.28Kappa in the same example

Feinstein and Cicchetti (1990) described this pattern.

Carry forward

Consistency is checked; now report it

  • Kappa measures consistency between screeners and does not measure accuracy.
  • The database column of the flow diagram is now complete.
  • Section 4 adds other methods and draws the full PRISMA 2020 diagram.

Learning Objectives for this section

  • Describe the purposes of title and abstract screening and full-text screening, and apply the decision rules appropriate to each stage.
  • Run a pilot screening round and use its results to revise a screening guide.
  • Explain why reviews use two independent screeners, and describe how conflicts are resolved and recorded.
  • Calculate percent agreement, chance-expected agreement and Cohen's kappa from a two-by-two table of screening decisions.
  • Interpret kappa with care, including the reason kappa can be low when percent agreement is high.

Introduction

Screening moves through two stages. At title and abstract screening, reviewers read the short record for each item and remove those that clearly fail the eligibility criteria. At full-text screening, reviewers obtain and read the full report for each item that survived the first stage, and decide whether it meets every criterion. The two stages exist because reading 1,870 full texts would take months, while most records can be excluded safely in a few seconds from the title and abstract. The cost of the two-stage design is that the first stage must be inclusive: a relevant study excluded on its abstract never reaches a reader who could see that it belonged.

This section follows the fictional Cedar Valley screeners through both stages, the pilot that came before them, and the agreement statistics they used to check that two people were applying the same rules.

3.1 Title and Abstract Screening

The working rule at the title and abstract stage is that a record moves forward unless it clearly fails at least one criterion. When the abstract is silent about a criterion, the record moves forward. When a record has no abstract, the screener decides from the title, and moves the record forward if the title leaves the question open. Most teams use three options, include, exclude and maybe, and treat maybe as include for the purpose of moving records to the next stage. The value of the maybe option is that it tells the full-text reviewer which records were uncertain and why.

Reasons for exclusion are usually not recorded at this stage. PRISMA 2020 asks only for the number of records excluded at title and abstract, and recording reasons for more than a thousand quick exclusions would slow the work without improving the report. Reasons become mandatory at full text.

Title: Effects of a peer-led walking group on loneliness and social participation among adults aged 70 and older in a mid-sized Canadian city (fictional).

Abstract summary: A controlled before-and-after evaluation of a twelve-week walking group run by a seniors' centre, with loneliness measured on a three-item scale.

Decision: Include. The participants are older adults in the community, the program runs in the community, loneliness is an outcome, and Canada is a high-income country. Every question in the guide is answered yes.

Title: Loneliness among hospital inpatients aged 65 and older: a cross-sectional survey (fictional).

Abstract summary: A survey of loneliness among older adults admitted to medical wards, with no intervention.

Decision: Exclude. Hospital inpatients fall outside the population, and the study evaluates no intervention. The abstract settles the question, so there is no need to read the full text.

Title: A community arts program for older people: outcomes from the first year (fictional).

Abstract: None available.

Decision: Maybe, which moves the record forward. The title suggests a community-based program for older adults with reported outcomes, but it does not say how old the participants were or whether loneliness was measured. Only the full text can answer those questions.

3.2 Full-Text Screening

At full text the team first has to obtain the reports. Most come through the university library; others come through interlibrary loan, from the authors, or from organization websites. Reports that cannot be obtained after reasonable effort are counted as reports not retrieved and listed in the review, which is more honest than leaving them out of the count. The Cedar Valley team sought 145 reports and could not retrieve three: a doctoral dissertation indexed in PsycINFO that was not available from the awarding university, and two articles from journals the library could not supply within the review's timeline after requests to the authors went unanswered. That left 142 reports to assess.

Full-text screening applies every criterion strictly. Each excluded report receives exactly one reason, the first reason in the ordered list from Section 2 that applies. This is also the stage at which reports are linked to studies: when two or more reports describe the same study, the reviewer links them so that the study is counted once. The Cedar Valley team included 40 reports, which described 38 studies, because two studies had each been reported in both a conference abstract and a journal article.

Exclusion reason (in the order applied)Reports excludedExample from the Cedar Valley full texts
1. Population outside criteria32Participants aged 50 and older, with no separate results for those aged 65 and older.
2. Intervention or setting not community-based21A visiting program delivered only to residents of long-term care homes.
3. No loneliness or social isolation outcome27A group exercise program that measured only physical function and falls.
4. Not conducted in a high-income country9A community program evaluated in a middle-income country.
5. Ineligible publication type13A protocol for a trial that has not yet reported results.
Total excluded102142 reports assessed − 102 excluded = 40 reports included.

3.3 Piloting the Screening Guide

A pilot (sometimes called a calibration exercise) is a small round of screening that both screeners complete before the main screening begins. Each screener works through the same random sample of records independently, and the team then compares the decisions, discusses every disagreement, and revises the guide. The pilot catches criteria that are ambiguous, rules that the guide leaves unstated, and differences in how screeners read the same words. Teams repeat the pilot with a fresh sample if agreement is poor or if the guide changed substantially.

Case: the Cedar Valley pilot

The evidence officer and the intern each screened the same random sample of 200 of the 1,870 records, with maybe counted as include. They agreed on 186 records and disagreed on 14. When they discussed the disagreements, they found that nine involved samples of adults aged 60 and older with no separate results for those aged 65 and older, three involved telephone befriending programs that one screener had not counted as community-based, and two were simple misreadings. The first version of the guide had shortened the protocol's age rule and left out its statement on telephone delivery, so the team restored the 80 percent rule with its mean or median age fallback and the statement that telephone and online delivery count as community-based, which are the rules shown in the Section 2 guide. The two screeners then screened the remaining 1,670 records independently with the revised guide.

3.4 Dual Independent Screening and Resolving Conflicts

In dual independent screening, two reviewers screen every record without seeing each other's decisions, and the team resolves each disagreement afterwards. The method rests on a simple observation: careful people make errors when they read hundreds of abstracts, and two people rarely make the same error on the same record. Methodological studies support the practice. A methodological systematic review comparing single with double screening found that single screening missed more eligible studies (Waffenschmidt et al., 2019), and a randomized trial of abstract screening reached the same conclusion for single-reviewer screening (Gartlehner et al., 2020). The Cochrane Handbook recommends that at least two people, working independently, decide whether each study meets the eligibility criteria (Higgins et al., 2023).

Conflicts are resolved in one of three ways. The two screeners can discuss the record and agree; a third reviewer can decide; or the team can use discussion first and refer anything unresolved to the third reviewer, which is the Cedar Valley approach. Whatever the method, the protocol should state it, and the platform should record the final decision. A record of conflicts is also useful in its own right, because a run of conflicts on one criterion signals that the guide needs another rule.

Rapid reviews sometimes reduce dual screening, for example by having one person screen most records while a second person screens a sample or checks all exclusions. These shortcuts save time at a known cost in missed studies, and Lesson 10 examines the guidance on when they are acceptable. The Cedar Valley protocol planned a shortcut of this kind: two reviewers would screen the first 20 percent of titles and abstracts (374 records, starting with the pilot sample), and one reviewer would then screen the rest while a second checked every exclusion. In the first 374 records, 345 (92 percent) were excluded, so checking every exclusion would have saved only about 3 percent of the work. The team amended its protocol on 20 February 2026, at the end of week 5, and screened every remaining record in duplicate, because the number of records was manageable and the planning team would rely on the review to choose a program model. Full texts were screened by two reviewers independently, as the protocol planned.

3.5 Measuring Agreement Between Screeners

Teams often summarize how well two screeners agreed, either for the pilot, to decide whether the guide is ready, or for each stage, to report in the methods. The two most common statistics are percent agreement and Cohen's kappa. Both start from a two-by-two table that cross-classifies the two screeners' decisions.

Cedar Valley pilot (200 records)Officer: includeOfficer: excludeIntern total
Intern: include22 (both include)830
Intern: exclude6164 (both exclude)170
Officer total28172200

Background: percent agreement

Percent agreement (also called observed agreement, po) is the proportion of items on which two people who classify the same items independently make the same decision. In the Cedar Valley pilot the screeners agreed on 22 includes and 164 excludes, so observed agreement is (22 + 164) ÷ 200 = 186 ÷ 200 = 0.93, or 93 percent. The statistic is easy to compute and explain, and it has two limits. First, some agreement arises by chance: when most records are obvious exclusions, two screeners will agree on most of them whatever their skill, so high percent agreement can hide poor agreement on the few records that matter. Second, agreement measures consistency and does not measure accuracy: two screeners who misread a criterion in the same way agree perfectly and are both wrong. HSCI 207 Lesson 8, Section 4.5 (Collecting Quantitative Data), teaches percent agreement on 30 pilot charts abstracted by two people, with a worked example of agreement that arises by chance, and is optional reading. The rest of this section deals with the first limit through Cohen's kappa; the screening guide, the pilot and the discussion of conflicts deal with the second.

Agreement expected by chance

Cohen's kappa corrects for this by asking how much agreement two screeners would reach if each made decisions at their own overall rate but independently of the content of the records. The intern included 30 of 200 records (0.15) and the officer included 28 of 200 (0.14). If their decisions were unrelated, the expected proportion of records that both include is 0.15 × 0.14 = 0.021, and the expected proportion that both exclude is 0.85 × 0.86 = 0.731. The chance-expected agreement is the sum, 0.021 + 0.731 = 0.752. In counts, this is 4.2 expected joint includes and 146.2 expected joint excludes, or 150.4 of 200 records.

Cohen's kappa

Cohen's kappa (Cohen, 1960) is the agreement achieved beyond chance, divided by the largest possible agreement beyond chance.

Formulas for agreement between two screeners

Observed agreement: po = (both include + both exclude) ÷ total records

Chance-expected agreement: pe = (proportion included by screener 1 × proportion included by screener 2) + (proportion excluded by screener 1 × proportion excluded by screener 2)

Cohen's kappa: κ = (po − pe) ÷ (1 − pe)

Cedar Valley pilot: κ = (0.93 − 0.752) ÷ (1 − 0.752) = 0.178 ÷ 0.248 = 0.72

Cedar Valley pilot: where the agreement of 0.93 comes from Room above chance: 1 − 0.752 = 0.248 Agreement expected by chance 0.752 0.178 0.07 00.250.500.751.00 Agreement beyond chance: 0.93 − 0.752 = 0.178 Disagreement: 14 of 200 records = 0.07 Kappa = 0.178 ÷ 0.248 = 0.72
Kappa expresses the agreement achieved beyond chance (0.178) as a share of the agreement that was possible beyond chance (0.248). The pilot's observed agreement of 0.93 corresponds to a kappa of 0.72.

Interpreting kappa

Kappa equals 1 when agreement is perfect, 0 when agreement is no better than chance, and falls below 0, to a minimum of −1, when screeners agree less often than chance would predict. The most often cited labels for values in between come from Landis and Koch (1977), who proposed them as convenient benchmarks for discussion. McHugh (2012) argued that these labels are too generous for health research, where a kappa in the 0.60s still leaves a substantial share of decisions in doubt. Labels are therefore a starting point for judgement.

KappaLabel proposed by Landis and Koch (1977)
Below 0Poor
0.00 to 0.20Slight
0.21 to 0.40Fair
0.41 to 0.60Moderate
0.61 to 0.80Substantial
0.81 to 1.00Almost perfect

The Cedar Valley pilot's kappa of 0.72 falls in the band Landis and Koch called substantial. The more useful reading comes from the table itself. Of the 36 records that at least one screener wanted to include (22 + 8 + 6), the screeners agreed on only 22, which is the agreement that matters most at this stage. That observation, together with the pattern in the disagreements, is what led the team to revise its guide before continuing.

High agreement, low kappa

Kappa depends on the proportion of records that are included. When very few records are relevant, chance-expected agreement is already close to 1, so even a small number of disagreements produces a low kappa. Feinstein and Cicchetti (1990) described this as a paradox of high agreement and low kappa. It is common in screening, where included records are usually a small minority: in Cedar Valley, the 40 included reports came from 1,870 screened records, about 2 percent. For this reason, a review should report the two-by-two table or the number of conflicts alongside kappa, and should judge agreement mainly by how the screeners did on the records that either of them included.

Try it: a review with few relevant records

Two screeners on another team each screen the same 200 records. Both include 2; screener 1 includes 5 that screener 2 excludes; screener 2 includes 4 that screener 1 excludes; both exclude 189. Calculate percent agreement, chance-expected agreement and kappa, and say what the results suggest.

Check your answerv

Observed agreement is (2 + 189) ÷ 200 = 0.955. Screener 1 included 7 of 200 (0.035) and screener 2 included 6 of 200 (0.03). Chance-expected agreement is (0.035 × 0.03) + (0.965 × 0.97) = 0.00105 + 0.93605 = 0.9371. Kappa is (0.955 − 0.9371) ÷ (1 − 0.9371) = 0.0179 ÷ 0.0629 = 0.28. Percent agreement looks excellent, but of the 11 records that either screener wanted to include, the screeners agreed on only 2. The low kappa reflects a real problem: the screeners are applying the criteria differently to exactly the records that matter, and the team should discuss the 9 conflicts and revise its guide before going further.

Reporting agreement in a reviewv

A clear methods statement names the stage, the number of records, the statistic and the way disagreements were resolved, for example: "Two reviewers independently screened a pilot sample of 200 records (agreement 93 percent, kappa 0.72) and, after revising the screening guide, screened all remaining records independently; disagreements were resolved by discussion or by a third reviewer."

More than two screenersv

When three or more screeners each rate the same records, Fleiss' kappa extends the same idea to several raters. When different pairs of screeners share the work, teams usually report agreement for each pair or for the pilot only.

A prevalence-adjusted alternativev

Some authors also report the prevalence-adjusted bias-adjusted kappa (PABAK), which equals 2 × po − 1 for two categories (Byrt et al., 1993). For the Cedar Valley pilot, PABAK is 2 × 0.93 − 1 = 0.86. It removes the effect of a low proportion of includes, which is useful for comparison, and for the same reason it can make agreement on the rare includes look better than it was.

Kappa does not measure accuracyv

Two screeners who make the same mistake agree perfectly. Kappa measures consistency between people, and the screening guide, the pilot and the discussion of conflicts are what bring the decisions in line with the criteria.

With 38 studies included from the database search, and the reports and exclusions counted at each stage, the Cedar Valley team has the numbers it needs for the left-hand column of its flow diagram. Section 4 adds the grey literature and citation chasing and draws the complete PRISMA 2020 flow diagram.

Reflection

Two reviewers independently assessed 142 full-text reports. Both included 36 reports; reviewer A included 4 that reviewer B excluded; reviewer B included 3 that reviewer A excluded; both excluded 99. Observed agreement is po = (both include + both exclude) ÷ total. Chance-expected agreement is pe = (proportion included by A × proportion included by B) + (proportion excluded by A × proportion excluded by B). Cohen's kappa is (po − pe) ÷ (1 − pe). Landis and Koch (1977) labelled kappa values of 0.61 to 0.80 substantial and 0.81 to 1.00 almost perfect. Calculate po, pe and kappa, showing your working. Interpret the result, commenting on agreement among the reports that at least one reviewer wanted to include. Explain what the team should do with the disagreements, state how many reports it would include if discussion led it to include 4 of the disputed reports, and write one sentence reporting this agreement in the methods.

Model answer

Observed agreement is (36 + 99) ÷ 142 = 135 ÷ 142 = 0.9507. Reviewer A included 40 of 142 reports (0.2817) and reviewer B included 39 (0.2746). Chance-expected agreement is (0.2817 × 0.2746) + (0.7183 × 0.7254) = 0.0774 + 0.5210 = 0.5984. Kappa is (0.9507 − 0.5984) ÷ (1 − 0.5984) = 0.3523 ÷ 0.4016 = 0.88.

A kappa of 0.88 falls in the band that Landis and Koch called almost perfect. Agreement is also good where it matters most: of the 43 reports that at least one reviewer wanted to include (36 + 4 + 3), the reviewers agreed on 36. Because more than a quarter of full texts were included, chance-expected agreement is much lower than at the title and abstract stage, and kappa gives a fair summary here.

The 7 disagreements should be resolved by discussion, with a third reviewer deciding any that remain, and each final decision and its reason should be recorded in the platform. If discussion led the team to include 4 of the 7, it would include 36 + 4 = 40 reports. A methods sentence could read: "Two reviewers independently assessed all 142 full-text reports (agreement 95 percent, kappa 0.88); the 7 disagreements were resolved by discussion or by a third reviewer."

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: Which rule best describes title and abstract screening?

The first stage must be inclusive, because a relevant study excluded on its abstract never reaches a full-text reader. Records whose abstracts are silent on a criterion move forward. Reasons are usually recorded only at full text, and a scoping review such as Cedar Valley includes many designs.

Question 2: In a pilot of 200 records, two screeners both include 22 records and both exclude 164. What is the observed agreement?

Observed agreement counts both kinds of agreement: (22 + 164) ÷ 200 = 186 ÷ 200 = 0.93. Kappa (0.72 in the Cedar Valley pilot) is a different statistic that corrects observed agreement for agreement expected by chance.

Question 3: Why can kappa be low when percent agreement is high in screening?

When almost every record is excluded, two screeners would agree on most records by chance alone, so chance-expected agreement is near 1 and the room above chance is small. A handful of disagreements then uses up much of that room. Kappa uses all four cells of the table, and the gap between it and percent agreement varies with prevalence.

Question 4: The Cedar Valley team sought 145 full-text reports and could not obtain 3 of them. How should those 3 be handled?

PRISMA 2020 has a box for reports not retrieved. Counting them separately, and listing them, tells readers that the team looked for them and could not assess them, which may matter for the review's conclusions. Treating them as exclusions or dropping them would misrepresent the process.
Section 4 of 5

Reporting Study Selection with the PRISMA 2020 Flow Diagram

⏱ Estimated reading time: 35 minutes
Section 4 of 5

Reporting Study Selection with the PRISMA 2020 Flow Diagram

Items 16a and 16b, the 2020 changes, templates and the Cedar Valley diagram.

What PRISMA asks

Items 16a and 16b

Item 16a

Authors report the search and selection results, ideally with a flow diagram.

Item 16b

Authors cite apparently eligible studies that were excluded, with reasons.

Page et al. (2021) published the statement and its templates.

What changed

The 2020 diagram

Records, reports and studies Records removed before screening Reports not retrieved A column for other methods Registers beside databases Templates for updates
Left-hand column

Databases and registers

2,480Records identified
1,870Records screened
142Reports assessed
40Reports of 38 studies

Each box equals the box above it minus the exclusions beside it.

Right-hand column and final box

Other methods and the totals

91Reports from grey literature, organizations and citation chasing
42Studies in 44 reports
26Grey-literature documents
Reconcile

Check every box against the log

  • Each box equals the box above it minus the exclusions beside it.
  • Reasons for exclusion add up to the number excluded.
  • The number of studies never exceeds the number of included reports.
  • The abstract, the text and the diagram report the same numbers.
Next steps

Reflection and assessment

  • The reflection below asks you to check a draft flow diagram.
  • The knowledge check and the final assessment follow.
  • Check each box against the one above it, and report studies and reports separately.

Learning Objectives for this section

  • Explain what the PRISMA 2020 flow diagram reports and which PRISMA 2020 items it supports.
  • Describe how the 2020 diagram differs from the 2009 version, including the distinction between records, reports and studies and the separate column for other methods.
  • Choose the PRISMA 2020 template that fits a new or updated review.
  • Draw a complete flow diagram for the Cedar Valley review and check that every number reconciles with the team's log.
  • Adapt the diagram for a scoping review and recognize the errors that most often appear in published flow diagrams.

Introduction

The PRISMA 2020 statement (Preferred Reporting Items for Systematic reviews and Meta-Analyses) is a 27-item checklist of what a systematic review report should contain, published with an explanation and elaboration paper and a set of flow diagram templates (Page et al., 2021). Item 16a asks authors to describe the results of the search and selection process, from the number of records identified to the number of studies included, ideally with a flow diagram. Item 16b asks authors to cite studies that might appear to meet the inclusion criteria but were excluded, and to explain why. The flow diagram answers item 16a in a single figure: it shows how many records the searches found, how many were removed at each stage and why, and how many studies and reports the review finally included.

The diagram is the public face of the record management and screening work in Sections 1 to 3. Every number in it should come straight from the team's log, and a reader should be able to add and subtract their way from the top of the diagram to the bottom. This section explains the structure of the 2020 diagram, draws the complete diagram for the fictional Cedar Valley review, and shows how to check it.

4.1 What the Flow Diagram Shows

A reader looks at a flow diagram to answer practical questions. How large was the search? How many records were duplicates? What share of records survived title and abstract screening? Were any reports impossible to obtain? Why were reports excluded at full text, and was any one reason dominant? How many studies did the review include, and how many reports described them? Answers to these questions help the reader judge whether the search was broad enough, whether the criteria were applied as described, and how the review's results might change if, for example, the unretrieved reports had been found.

The diagram does not describe the search strategies themselves. Those belong in the methods and appendices, reported according to PRISMA-S (Rethlefsen et al., 2021), which Lesson 4 introduced. The diagram also does not replace the list of excluded studies that item 16b asks for; it summarizes the reasons as counts, while the list names the specific studies that a reader might have expected to see included.

4.2 From PRISMA 2009 to PRISMA 2020

The original PRISMA statement (Moher et al., 2009) included a four-phase flow diagram labelled identification, screening, eligibility and included. The 2020 update kept the top-to-bottom design and changed its content in several ways that reflect how reviews are now conducted.

Records, reports and studiesClick to explore
Records removed before screeningClick to explore
Reports not retrievedClick to explore
A column for other methodsClick to explore
Registers beside databasesClick to explore
Templates for updated reviewsClick to explore

4.3 Choosing a Template

PRISMA 2020 offers templates for two kinds of review and two kinds of search. A new review starts from nothing; an updated review builds on a previous version and reports both the earlier and the new studies. A review that searched only databases and registers needs one column; a review that also searched websites, contacted organizations or chased citations needs the version with a second column for other methods.

TemplateUse it when
New review, databases and registers onlyThe review is new and every record came from bibliographic databases or study registers.
New review, databases, registers and other sourcesThe review is new and the team also searched websites, contacted organizations, chased citations or used other methods. The Cedar Valley review uses this template.
Updated review, databases and registers onlyThe review updates an earlier version and the new search used only databases and registers.
Updated review, databases, registers and other sourcesThe review updates an earlier version and the new search also used other methods.

4.4 The Cedar Valley Flow Diagram

The Cedar Valley review searched five databases (Sections 1 to 3), searched grey literature, trial registries and websites and contacted organizations (Lesson 5), and chased citations from the included studies and from relevant reviews (Lesson 5). It therefore uses the template for new reviews with other sources. The left-hand column summarizes Sections 1 to 3 of this lesson. The right-hand column summarizes the other methods. The grey-literature searches produced 771 records and items, of which 214 were duplicates; the remaining 557 were screened, and 61 were sought as full reports. Organizations contacted during the environmental scan sent 7 documents, all of which were sought. Citation chasing retrieved 2,448 records, of which 1,288 were duplicates or had already been screened; the remaining 1,160 were screened, and 23 were sought. In all, 91 reports were sought. Two could not be obtained (a program report that an organization had withdrawn from its website and an internal evaluation that its authors could not share), 89 were assessed, and 30 were included: 26 grey-literature documents, such as program evaluation reports and agency reports, and 4 reports of 4 further studies found by citation chasing. PRISMA 2020 places study registers beside databases in the left-hand column. The Cedar Valley team searched the trial registries in week 5 with its other grey-literature sources and screened their records in that set, so it counted them in the right-hand column and explained the adaptation in its methods; none of the registry records led to an included study.

Identification of studies viadatabases and registers Identification of studies viaother methods Records identified from databases (n = 2,480): MEDLINE (n = 742) Embase (n = 816) CINAHL (n = 388) PsycINFO (n = 296) Web of Science (n = 238) Registers: see note Records removed before screening: Duplicate records (n = 610) Marked ineligible by automation tools (n = 0) Removed for other reasons (n = 0) Records identified from: Grey-literature sources (n = 771) Organizations (n = 7) Citation searching (n = 2,448) Records removed before screening: Duplicates and records already screened (n = 1,502) Records screened (n = 1,724) Records excluded (n = 1,633) Records screened (n = 1,870) Records excluded (n = 1,725) Reports sought for retrieval (n = 145) Reports not retrieved (n = 3) Reports sought for retrieval (n = 91) Reports not retrieved (n = 2) Reports assessed for eligibility (n = 142) Reports excluded (n = 102): Population outside criteria (n = 32) Not community-based (n = 21) No loneliness or isolation outcome (n = 27) Not a high-income country (n = 9) Ineligible publication type (n = 13) Reports assessed for eligibility (n = 89) Reports excluded (n = 59): Population outside criteria (n = 15) Not community-based (n = 10) No loneliness or isolation outcome (n = 22) Program description only (n = 12) Studies included in review (n = 42) Reports of included studies (n = 44) Grey-literature documents included (n = 26) Identification Screening Included
The PRISMA 2020 flow diagram for the fictional Cedar Valley rapid scoping review, drawn on the template for new reviews that searched databases, registers and other sources. Note: trial registry records were screened with the other grey-literature records and are counted under other methods. The right-hand column is adapted to show the screening of grey-literature and citation records, and the final box reports grey-literature documents separately from studies.

The final box reports three numbers. The review includes 42 studies (38 from the databases and 4 from citation chasing), described in 44 reports (40 and 4). It also includes 26 grey-literature documents. The team decided at the protocol stage to chart grey-literature documents separately from research studies because their methods are often reported briefly, and it adapted the final box to show them. PRISMA 2020 presents its flow diagram as a template that authors can adapt, and an adaptation of this kind should be explained in the methods.

4.5 Checking That the Numbers Reconcile

Before a flow diagram goes into a report, someone other than the person who drew it should check every number against the log. The check is arithmetic: each box should equal the box above it minus the exclusions beside it, the reasons for exclusion should add up to their total, and the included counts should add up across columns.

Reconciliation checks for the Cedar Valley diagram

Database sources: 742 + 816 + 388 + 296 + 238 = 2,480 records identified.

Before screening: 2,480 − 610 duplicates − 0 automation − 0 other = 1,870 records screened.

Title and abstract: 1,870 − 1,725 excluded = 145 reports sought.

Retrieval: 145 − 3 not retrieved = 142 reports assessed.

Full text: 32 + 21 + 27 + 9 + 13 = 102 excluded; 142 − 102 = 40 reports included, describing 38 studies.

Other methods: 771 + 7 + 2,448 = 3,226 records identified; 214 + 1,288 = 1,502 removed before screening; 3,226 − 1,502 = 1,724 records screened; 1,724 − 1,633 excluded = 91 reports sought (61 + 7 + 23); 91 − 2 = 89 assessed; 15 + 10 + 22 + 12 = 59 excluded; 89 − 59 = 30 included (26 grey-literature documents and 4 study reports).

Included: 38 + 4 = 42 studies; 40 + 4 = 44 reports of included studies; 26 grey-literature documents.

Two of these steps deserve attention. The first is the move from records to reports between title and abstract screening and retrieval. In Cedar Valley, each record that passed screening led to one report, so the number of reports sought equals the number of records that survived. In some reviews a single record leads to more than one report, or the team discovers at retrieval that two records point to the same document, and the diagram should then show the change with a note. The second is the move from reports to studies in the final box. The number of studies can never exceed the number of included reports, and when it is smaller the methods should say how reports were linked to studies.

Box in the diagramWhere the number comes from
Records identified from databases and registersThe import log in Section 1, by database.
Records removed before screeningThe de-duplication log: 571 automated, 27 manual and 12 found by the platform, totalling 610.
Records screened and excludedThe screening platform's counts at the title and abstract stage, checked against 1,870 − 145.
Reports sought, not retrieved and assessedThe retrieval log kept by the intern, listing every report requested and its source.
Reports excluded, with reasonsThe full-text decisions, each with one reason from the ordered list.
Records identified from other methodsThe grey-literature and citation-chasing logs from Lesson 5.
Studies and reports includedThe list of included studies, with the reports linked to each.

4.6 Adapting the Diagram for Scoping and Rapid Reviews

The Cedar Valley review is a rapid scoping review, and its reporting follows PRISMA-ScR, the extension for scoping reviews (Tricco et al., 2018). PRISMA-ScR asks for the same information about selection as PRISMA 2020, using the term "sources of evidence" because a scoping review may include documents that are not research studies. A scoping review can therefore use the PRISMA 2020 template with labels that fit its sources, as the Cedar Valley team did. A rapid review that used shortcuts, such as single screening of part of the records, should keep the same boxes and explain the shortcut in the methods; Lesson 10 discusses how rapid reviews report their shortcuts.

The environmental scan is reported separately. Its counts (18 programs identified, 11 responding to the survey and 7 taking part in key informant interviews) describe organizations and people, and Lesson 11 shows how to report a scan, including a simple flow of programs from identification to participation.

Tools for drawing the diagram

The PRISMA website provides the templates as editable documents. Covidence draws a PRISMA diagram from its stage counts, and the PRISMA2020 package for R and its accompanying Shiny web application produce compliant diagrams from a table of numbers (Haddaway et al., 2022). Any of these tools is acceptable. Whichever tool the team uses, the numbers must come from the team's own log, because a platform cannot know about duplicates removed in a reference manager before import or about records that were never imported.

Numbers that do not add upv

The most common error is a box that does not equal the box above it minus the exclusions beside it. It usually arises when a diagram is drawn from memory or when screening continued after the diagram was drafted. Rechecking every subtraction against the log prevents it.

Mixing records, reports and studiesv

Some diagrams label the final box "articles included" and give a number that mixes reports and studies. The 2020 diagram asks for studies and reports separately, and the two numbers should match the review's tables.

Missing reasons at full textv

A full-text exclusion box with a total and no reasons does not meet the template. Reasons are optional at the title and abstract stage and expected at full text.

Hiding reports that were not retrievedv

Folding unobtainable reports into the exclusions makes the search look more complete than it was. The 2020 diagram has a separate box for them.

Counting citation-chasing records twicev

Records found by citation chasing that the database search had already found should be removed before they are counted in the right-hand column. Otherwise the diagram overstates what citation chasing added.

A diagram that contradicts the textv

The abstract, the results text and the diagram should report the same numbers. Reviews that are revised after peer review often update one and forget the others.

Using the 2009 templatev

The 2009 template lacks the boxes for reports not retrieved and for other methods. Journals that follow PRISMA 2020 expect the 2020 template.

Try it: find the error

A draft flow diagram for a review that searched only databases reports the following numbers: records identified 640; duplicates removed 152; records screened 488; records excluded 431; reports sought for retrieval 57; reports not retrieved 2; reports assessed for eligibility 57; reports excluded 44; studies included 13. Check each step and identify what is wrong.

Check your answerv

The first three steps reconcile: 640 − 152 = 488 and 488 − 431 = 57. The error is at retrieval: if 57 reports were sought and 2 were not retrieved, only 55 can have been assessed. With 44 reports excluded, 55 − 44 = 11 reports were included, so the review cannot include 13 studies; it includes at most 11. The author of the diagram needs to check the log to see whether the number not retrieved, the number excluded or the number included is wrong, and should also report the number of reports of included studies alongside the number of studies.

The Cedar Valley team now knows which 42 studies and 26 grey-literature documents it will chart. Lesson 8 turns to data extraction and the appraisal of those studies.

Reflection

A review team drew a draft PRISMA 2020 flow diagram for a new review that searched databases and other sources. The draft reports: records identified from MEDLINE 180, CINAHL 140 and PsycINFO 92 (total 412); duplicates removed 96; records screened 326; records excluded 291; reports sought 35; reports not retrieved 1; reports assessed 34; reports excluded 26 (population 11, setting 6, outcome 9). In the other-methods column: websites 14 and citation searching 6, all 20 sought, none unretrieved, 20 assessed, 17 excluded with no reasons given, and 3 included. The final box reads "Studies included (n = 11)". The screening platform's export shows 281 records excluded at title and abstract. The review's tables list 8 studies from the database search, each described in one report, and from other methods 1 study (one report) and 2 grey-literature documents. Under PRISMA 2020, each box should equal the box above it minus the exclusions beside it, reasons are expected for full-text exclusions, and the final box reports studies and reports of included studies separately. Check each step, identify every error, and write the corrected numbers for the boxes that are wrong, including a corrected final box.

Model answer

The identification box reconciles: 180 + 140 + 92 = 412. The next step does not: 412 − 96 = 316, so the records screened box should read 316 where the draft reports 326. The platform export confirms 281 exclusions at title and abstract, and 316 − 281 = 35 reports sought, which matches the draft. The draft's 326 and 291 are each 10 too high, which suggests that 10 records were counted twice, perhaps citation-chasing records added to the database column. The draft's own subtraction (326 − 291 = 35) hid the error, which is why each box must be checked against the box above it and against the log.

Retrieval and full text reconcile: 35 − 1 = 34 assessed; 11 + 6 + 9 = 26 excluded; 34 − 26 = 8 reports included, describing 8 studies. In the other-methods column, 20 − 17 = 3 included, which reconciles, but the 17 exclusions need reasons.

The final box is wrong. "Studies included (n = 11)" adds 8 studies, 1 study and 2 grey-literature documents as if all were studies. The corrected box reads: studies included in review (n = 9); reports of included studies (n = 9); grey-literature documents included (n = 2). The methods should explain that grey-literature documents are reported separately.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: What does item 16a of PRISMA 2020 ask authors to report?

Item 16a asks for the results of the search and selection process, ideally with a flow diagram. Item 16b asks for the list of studies that might appear eligible but were excluded. Search strategies are reported according to PRISMA-S, and PRISMA 2020 does not require kappa.

Question 2: Which change did the PRISMA 2020 flow diagram introduce compared with the 2009 version?

The 2020 template for reviews that used other sources has a second column for websites, organizations, citation searching and similar methods. It also added boxes for records removed before screening and for reports not retrieved, and it kept the two screening stages separate.

Question 3: In the Cedar Valley diagram, 142 database reports were assessed and 102 were excluded, yet the final box credits 38 studies to the database search. Why does 142 − 102 = 40 differ from 38?

PRISMA 2020 counts reports and studies separately. Two Cedar Valley studies were each reported in a conference abstract and a journal article, so 40 reports describe 38 studies. The number of studies can never exceed the number of included reports.

Question 4: A team with a fixed deadline screens a random sample of 300 of its 1,100 de-duplicated records and states the sampling in its methods. Where should the 800 unscreened records appear in its PRISMA 2020 flow diagram?

The records were removed before screening for a reason other than duplication, so they belong in "records removed for other reasons", with the sampling explained. Counting them as excluded would claim that someone screened them, and leaving them out would make the numbers fail to reconcile.
Section 5 of 5

Final Assessment

⏱ Estimated time: 25 minutes

Bringing It All Together

This lesson followed the fictional Cedar Valley evidence team from the moment its database searches finished to the moment it knew which studies it would chart. The team exported 2,480 records from five databases, imported them into Zotero by source and checked every batch, and removed 610 duplicates in automated, manual and platform passes, keeping a log that would later supply the first boxes of its flow diagram. Throughout, it kept the PRISMA 2020 distinction between records, reports and studies in view: duplicate records were merged, while different reports of one study were kept and linked at full text.

The team then set up a screening project in Covidence with a screening guide, an ordered list of exclusion reasons and blinded dual review, and compared that choice with Rayyan, which suits teams without a licence. A pilot of 200 records produced 93 percent agreement and a kappa of 0.72, and the pattern of disagreements, more than the kappa value itself, led the team to add two rules before screening the rest. Dual screening at both stages narrowed 1,870 records to 40 reports of 38 studies.

Finally, the team drew a complete PRISMA 2020 flow diagram, with a second column for grey literature, organizations and citation chasing, and checked every number against its log. The review includes 42 studies described in 44 reports, together with 26 grey-literature documents reported separately.

Key Takeaways from this lesson

  • PRISMA 2020 distinguishes records (indexed titles and abstracts), reports (documents) and studies (investigations), and each stage of the review counts one of these units.
  • De-duplication merges copies of the same report, while different reports of one study, such as a conference abstract and a later article, are kept and linked at full text.
  • A false-positive merge removes evidence before screening, so automated matching should be cautious and uncertain pairs should be checked by a person.
  • An import log that compares records exported with records imported for each source catches lost batches before screening begins.
  • Screening platforms such as Covidence and Rayyan support blinded dual screening, detect conflicts and record reasons; Covidence imposes fixed stages, while Rayyan leaves more structure to the team.
  • A screening guide with rules for borderline cases and an ordered list of exclusion reasons makes decisions consistent and the counts in the flow diagram reproducible.
  • Title and abstract screening should be inclusive, and full-text screening applies every criterion strictly with one recorded reason for each exclusion.
  • Percent agreement overstates agreement when most records are easy exclusions, and Cohen's kappa corrects for chance agreement but can be low when relevant records are rare.
  • A pilot round, followed by discussion of every disagreement and revision of the guide, does more to improve screening than any single agreement statistic.
  • A PRISMA 2020 flow diagram must reconcile box by box with the team's log, report studies and reports separately, and show reports not retrieved and other methods in their own boxes.

Core Concepts Reviewed

Section 1: records, reports and studies; reference managers and Zotero; RIS export and import logs; duplicate records, false-positive merges and manual checks; the audit trail.

Section 2: screening platforms; Covidence and Rayyan compared; blinding and conflicts; the screening guide; ordered exclusion reasons; keyword highlights; privacy and licences.

Section 3: title and abstract screening; full-text screening and reports not retrieved; pilot screening; dual independent screening and conflict resolution; percent agreement, chance-expected agreement and Cohen's kappa; the high agreement, low kappa paradox.

Section 4: PRISMA 2020 items 16a and 16b; changes from the 2009 diagram; choosing a template; the complete Cedar Valley diagram; reconciliation checks; PRISMA-ScR adaptations; common errors.

The final reflection asks you to turn the Cedar Valley numbers into the methods and results paragraphs that a published review would contain.

Reflection

Write the record management and study selection paragraphs for the methods and results of the fictional Cedar Valley rapid scoping review, in about 200 to 250 words of formal prose. Use these facts. The librarian searched MEDLINE (742 records), Embase (816), CINAHL (388), PsycINFO (296) and Web of Science (238). Records were imported into Zotero; duplicates were removed in an automated pass (571), a manual pass (27) and on import to Covidence (12). Two reviewers piloted the screening guide on a random sample of 200 records (186 agreements; kappa 0.72), after which rules on mixed-age samples and telephone delivery were added. Two reviewers then screened all records independently at both stages, resolving conflicts by discussion or by a third reviewer. At title and abstract, 1,725 records were excluded. Of 145 reports sought, 3 could not be retrieved; of 142 assessed, 102 were excluded (population 32, intervention or setting 21, no loneliness or isolation outcome 27, not a high-income country 9, publication type 13), and 40 reports of 38 studies were included. Other methods (557 grey-literature records and items screened, with 61 reports sought; 7 documents from organizations; 1,160 new citation-chasing records screened, with 23 reports sought) yielded 91 reports, of which 2 were not retrieved, 59 of 89 assessed were excluded, and 30 were included (26 grey-literature documents and 4 reports of 4 studies). Include the totals, refer to a PRISMA 2020 flow diagram, and end with one sentence on a limitation of the process.

Model answer

Methods. Records from MEDLINE, Embase, CINAHL, PsycINFO and Web of Science were imported into Zotero, and duplicates were removed in an automated pass and a manual pass, with a final check on import to Covidence. Two reviewers piloted the screening guide on a random sample of 200 records (93 percent agreement, kappa 0.72) and revised it to clarify eligibility for mixed-age samples and telephone-delivered programs. Two reviewers then independently screened all titles and abstracts and all full texts, recording one reason for each full-text exclusion from an ordered list; disagreements were resolved by discussion or by a third reviewer.

Results. The database searches identified 2,480 records, of which 610 were duplicates. Of 1,870 records screened, 1,725 were excluded. Of 145 reports sought, 3 could not be retrieved, and 102 of the 142 assessed were excluded, most often because the population fell outside the criteria (32) or no loneliness or social isolation outcome was reported (27). Forty reports of 38 studies met the criteria. Grey-literature searches, organizations and citation chasing led to 91 further reports, of which 30 were included: 26 grey-literature documents and 4 reports of 4 studies. The review therefore includes 42 studies, described in 44 reports, and 26 grey-literature documents (Figure 1, PRISMA 2020 flow diagram). Three database reports and two reports from other sources could not be obtained, and their absence may have led the review to miss eligible studies.

Minimum 30 characters required.

✓ Reflection saved

Final Knowledge Assessment

Final Assessment, this lesson: Managing Records and Screening Studies (15 Questions)

Question 1: The Cedar Valley database searches retrieved 2,480 records, and 610 duplicates were removed. How many records were screened at title and abstract?

Records screened equal records identified minus records removed before screening: 2,480 − 610 = 1,870. The 1,725 figure is the number excluded at title and abstract, which leaves 145 reports sought for retrieval.

Question 2: Which pair of records should be merged during de-duplication?

Two records of the same article are duplicate records of one report, so they are merged, keeping the record with the DOI. A protocol and a results paper, or a preprint and its published version, are different reports of one study and are linked at full text. Parts of a series are different reports.

Question 3: A team plans to import raw database exports directly into its screening platform instead of using a reference manager. Which consequence follows?

Covidence and Rayyan both check for duplicates on import, so this route is quick, and the platform reports how many it found. The trade-off is dependence on the platform's matching rules and less opportunity to inspect borderline pairs before screening. The PRISMA template does not depend on the tools used.

Question 4: In a pilot of 100 records, both screeners include 10, screener 1 alone includes 5, screener 2 alone includes 5, and both exclude 80. What is Cohen's kappa?

Observed agreement is (10 + 80) ÷ 100 = 0.90. Each screener included 15 of 100 (0.15), so chance-expected agreement is (0.15 × 0.15) + (0.85 × 0.85) = 0.0225 + 0.7225 = 0.745. Kappa is (0.90 − 0.745) ÷ (1 − 0.745) = 0.155 ÷ 0.255 = 0.61. The value 0.90 is the observed agreement, and 0.745 is the chance-expected agreement.

Question 5: Which statement about Cohen's kappa is correct?

Kappa measures agreement beyond chance between two people, so two screeners who misread a criterion in the same way agree perfectly. Accuracy depends on the screening guide, the pilot and the discussion of conflicts. Maybe decisions are usually counted as include when kappa is calculated.

Question 6: Where do records found by citation chasing appear in a PRISMA 2020 flow diagram?

The 2020 template for reviews that used other sources has a column for websites, organizations and citation searching. Records the databases had already found should be removed before counting, so that the column shows what citation chasing added; Cedar Valley screened 1,160 such records and sought 23 reports.

Question 7: A team's screening guide has no rule for samples that mix ages, such as adults aged 60 and older. What is the most likely effect, and what is the remedy?

Unstated rules produce disagreements on borderline records. In the Cedar Valley pilot, nine of fourteen disagreements involved mixed-age samples, and the team added a rule about mean or median age before continuing. Platforms do not resolve such questions, and the diagram has no special box for them.

Question 8: Why are reasons for exclusion recorded at full text but usually not at title and abstract?

The flow diagram reports only the number of records excluded at title and abstract, while full-text exclusions are reported with reasons. Recording a reason for each of more than a thousand quick exclusions would slow screening without improving the report. Platforms do allow notes at the first stage, and exclusions there are counted.

Question 9: Which setup fits a three-person student team with no budget and no institutional Covidence licence that wants blinded dual screening?

Rayyan's free plan and blind mode support independent dual screening, and the team must add rules for reasons and stages that Covidence would impose. Covidence normally requires a subscription. A single shared spreadsheet breaks blinding, and a reference manager has no conflict detection.

Question 10: A full-text report fails the population criterion and is also a commentary. Under an ordered list that places population first, how is it recorded, and why does the order matter?

An ordered list assigns the first reason that applies, which gives every excluded report one reason and makes the counts reproducible. Without an order, different screeners could record different reasons for similar reports, and the diagram's reasons would not add up to the total.

Question 11: Which statement about the final box of the Cedar Valley flow diagram is correct?

The databases contributed 38 studies in 40 reports and citation chasing 4 studies in 4 reports, giving 42 studies in 44 reports. The team charts the 26 grey-literature documents separately and adapted the final box to show them, explaining the adaptation in its methods.

Question 12: In the Cedar Valley pilot, the screeners agreed on 186 of 200 records (kappa 0.72), but of the 36 records that either wanted to include, they agreed on 22. What did the team do next?

The disagreements showed two gaps in the guide (mixed-age samples and telephone programs), and two were simple misreadings. The team added rules and continued with full dual screening. Kappa labels are a starting point for judgement, and repeating the same sample would test memory rather than the guide.

Question 13: Why does the PRISMA 2020 flow diagram include a box for reports not retrieved?

Reports that could not be obtained were never assessed, and readers need to know how many there were because they may contain eligible studies. Folding them into exclusions would make the search look more complete than it was.

Question 14: A team draws its flow diagram with the PRISMA2020 Shiny app. Where should the numbers come from?

Drawing tools only lay out the numbers they are given. A screening platform does not know about duplicates removed in a reference manager before import, so the team's own log is the source for every box, and someone other than the person who drew the diagram should check it.

Question 15: A rapid review team proposes that one person screen all records to save time. Which response reflects this lesson?

A methodological review (Waffenschmidt et al., 2019) and a randomized trial (Gartlehner et al., 2020) found that single screening misses more eligible studies than dual screening. Rapid reviews sometimes accept this cost, and Lesson 10 examines the guidance; the shortcut should be justified and reported. PRISMA 2020 is a reporting guideline and does not prohibit methods.
✦ Complete the final reflection above before submitting

Congratulations!

You have successfully completed this lesson: Managing Records and Screening Studies.

You can now gather and log search results in a reference manager, remove duplicate records without losing distinct reports, set up a screening project with a guide and ordered exclusion reasons, run and evaluate a pilot with percent agreement and Cohen's kappa, carry out dual screening at both stages, and report the whole process in a PRISMA 2020 flow diagram that reconciles with your log.

Lesson 8, Data Extraction, Risk of Bias and Certainty of Evidence, takes the studies that survive screening and shows how to design and pilot an extraction form, choose and apply risk-of-bias tools matched to study design, apply them consistently between two reviewers, and rate the certainty of a body of evidence with GRADE.

Continue to Lesson 8 →
Reference

Glossary: Key Terms, People & Frameworks

📚 Reference page, available throughout the lesson

Definitions of the terms, tools and people introduced in this lesson on managing records and screening studies.

Core Concepts
Record The title, abstract or both of a report as indexed in a database or on a website. One report can appear as several records in different databases.
Report A document that supplies information about a study, such as a journal article, preprint, conference abstract, thesis, registry entry or agency report.
Study An investigation, such as a trial or an evaluation, with a defined group of participants, an intervention or exposure and outcomes. One study can have several reports.
Duplicate record A second or later copy of the same report, usually because two databases index the same article.
De-duplication The removal of duplicate records before screening, usually by an automated pass followed by manual checks.
False-positive merge An error in de-duplication in which two records describing different reports are merged, so that one report is lost before screening.
Audit trail A written record of files, counts and decisions that would let another person repeat the process and obtain the same numbers.
Screening The process of deciding which records and reports meet the review's eligibility criteria.
Title and abstract screening The first stage of screening, in which reviewers exclude records that clearly fail the criteria and move all others forward.
Full-text screening The second stage of screening, in which reviewers read the full report and decide whether it meets every criterion, recording one reason for each exclusion.
Screening guide A set of questions that turns the eligibility criteria into decisions, with rules for common borderline cases.
Ordered exclusion reasons A fixed list of exclusion reasons applied in sequence, so that each excluded report receives the first reason that applies.
Pilot screening A small round in which all screeners screen the same sample, compare decisions and revise the guide before the main screening; also called calibration.
Dual independent screening Screening in which two reviewers assess every record or report without seeing each other's decisions, with disagreements resolved afterwards.
Blinding (in screening) Keeping each reviewer's decisions hidden from the other reviewers until both have screened a record.
Conflict A record or report on which reviewers made different decisions, resolved by discussion or by a third reviewer.
Reports not retrieved Reports that a team sought at the full-text stage but could not obtain; PRISMA 2020 reports them in their own box.
Percent agreement The proportion of records on which two screeners made the same decision; also called observed agreement.
Chance-expected agreement The agreement two screeners would reach if each decided at their own overall rate but independently of the records' content.
Cohen's kappa Agreement beyond chance divided by the maximum possible agreement beyond chance: (po − pe) ÷ (1 − pe). It ranges from −1 to 1, with 0 for agreement no better than chance.
High agreement, low kappa paradox The pattern in which kappa is low despite high percent agreement because almost all records fall in one category, as in screening.
Frameworks & Tools
Reference manager Software that stores bibliographic records as structured fields, organizes them, finds duplicates and exports them to other tools.
Zotero A free, open-source reference manager with collections, group libraries and a Duplicate Items view.
RIS format A plain-text format for bibliographic records in which each field has a two-letter tag, used to move records between databases and tools.
Covidence A subscription screening platform that moves records through fixed stages, supports blinded dual review and produces PRISMA counts.
Rayyan A screening platform with a free plan, a flexible screening space, blind mode, labels and exclusion reasons (Ouzzani et al., 2016).
PRISMA 2020 statement A 27-item reporting guideline for systematic reviews, with an explanation and elaboration paper and flow diagram templates (Page et al., 2021).
PRISMA 2020 flow diagram A figure that reports the number of records, reports and studies at each stage of identification, screening and inclusion, with reasons for full-text exclusions.
PRISMA-ScR The PRISMA extension for scoping reviews, which uses the term sources of evidence (Tricco et al., 2018).
PRISMA2020 Shiny app An R package and web application that draws PRISMA 2020 flow diagrams from a table of numbers (Haddaway et al., 2022).
Landis and Koch benchmarks Labels for kappa values, from slight to almost perfect, proposed in 1977 as conventions for discussion.
PABAK The prevalence-adjusted bias-adjusted kappa, equal to 2 × po − 1 for two categories (Byrt et al., 1993).
Key People
Jacob Cohen An American psychologist and statistician who introduced the kappa coefficient of agreement in 1960 and wrote widely on effect sizes and statistical power.
J. Richard Landis and Gary G. Koch Biostatisticians whose 1977 paper on observer agreement for categorical data proposed the widely cited labels for kappa values.
Alvan R. Feinstein and Domenic V. Cicchetti Authors of the 1990 papers describing the paradoxes of high agreement and low kappa.
Matthew J. Page A research methodologist in Australia and lead author of the PRISMA 2020 statement and its explanation and elaboration.
David Moher A Canadian epidemiologist at the Ottawa Hospital Research Institute who led the original 2009 PRISMA statement and co-authored PRISMA 2020.
Mourad Ouzzani A computer scientist at the Qatar Computing Research Institute and lead author of the 2016 paper describing Rayyan.
Wichor M. Bramer A biomedical information specialist in the Netherlands whose team published a step-by-step method for de-duplicating search results in EndNote (Bramer et al., 2016).
Andrea C. Tricco A Canadian epidemiologist in Toronto who led the development of PRISMA-ScR, the reporting guideline for scoping reviews.
No matching entries. Try a different search term.