HSCI 241 · Lesson 6

Artificial Intelligence in Evidence Retrieval and Synthesis

Finding & Synthesizing Health Evidence

Learning objectives for this lesson:

  • Describe how a large language model produces text by predicting one token after another, and explain why this process can generate fabricated or mis-attributed citations.
  • Explain how retrieval-augmented generation works and why it reduces some citation errors while leaving others in place.
  • Distinguish conversational assistants, AI search engines and research assistants, citation-context tools and AI-assisted screening tools by where their answers come from and by what each can and cannot do in a review.
  • Evaluate an unfamiliar AI research tool by asking about its index, its ranking method, its evaluation, its documentation and its handling of data.
  • Verify AI-suggested citations with a four-step procedure covering existence, bibliographic details, claim and fit, and record the results in a verification log.
  • Document an AI-assisted search, identify privacy, confidentiality and copyright risks, and write a disclosure statement that follows emerging guidance such as the RAISE recommendations.
  • Explain how active-learning screening works and interpret recall, work saved over sampling and stopping rules from a screening simulation.
  • Judge when AI assistance is acceptable at each stage of a review, and prepare an AI-assisted search log, a verification log and a disclosure statement for a review.

This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University, drawing on the Cochrane Handbook for Systematic Reviews of Interventions and the JBI Manual for Evidence Synthesis.

Lesson 6 · HSCI 241

Artificial Intelligence in Evidence Retrieval and Synthesis

A short guided orientation before you work through the lesson at your own pace.

Finding & Synthesizing Health Evidence
The running case

The Cedar Valley evidence review (fictional)

A planning team asks a small evidence team for a rapid scoping review and environmental scan on community programs for loneliness in adults aged 65 and older.

2,480 records

Five databases were searched, and 1,870 records remain after 610 duplicates are removed.

An intern's idea

The intern proposes asking an AI assistant for the key studies.

Three conditions

Every reference is verified, every use is documented, and no scan data enter any tool.

The road map

Four sections

1 · How AI tools answer

Next-token prediction, retrieval-augmented generation and why fabricated citations occur.

2 · Categories of tools

Assistants, AI search engines, citation-context tools and AI-assisted screening.

3 · Verification and disclosure

A four-step check, a worked citation exercise, privacy and the RAISE recommendations.

4 · Active-learning screening

How it works, how it is measured, when to stop, and a decision table for every stage.

What you will learn

Four documents that record AI use

  • An AI-assisted search log records the tool, version, date, settings, prompt and output.
  • A verification log records four checks for every source the tool suggests.
  • A marked decision table shows where the team will and will not use AI in the review.
  • A disclosure statement tells readers what was used and how it was checked.
A note on products

Categories outlast product names

Products named in this lesson are examples available in 2026. They change quickly, so the lesson teaches you to evaluate categories of tools.

Read each section, try the tasks before opening the answers, complete the quizzes and write the reflections.

Section 1 of 5

How Large Language Models and Retrieval-Augmented Tools Generate Answers

⏱ Estimated reading time: 35 minutes
Section 1 of 5

How Large Language Models and Retrieval-Augmented Tools Generate Answers

Next-token prediction, the retrieval step, and why fabricated citations occur.

The mechanism

Predicting the next token

Tokens

Words, parts of words and punctuation marks are the units the model reads and writes.

Training

The model learns to predict the next token from a very large collection of text.

Generation

The model chooses one token, adds it to the text and predicts again.

A temperature setting controls how much chance enters each choice (Vaswani et al., 2017, introduced the transformer architecture).

Four consequences

What next-token prediction means for a reviewer

  • The model has no catalogue of sources, so it writes references from patterns.
  • Its information stops at a knowledge cutoff unless it can search.
  • The same prompt can give a different answer on another day.
  • Fluent, confident text carries no information about accuracy.
Adding a search step

Retrieval-augmented generation

Prompt
Ranked search of an index
Top passages in context
Answer with citations

Retrieval ties each citation to a real document (Lewis et al., 2020). The summary can still misstate the source, and the retriever returns only a top set.

Why references are fabricated

Highly patterned text, assembled from plausible parts

Fabricated

A complete reference to a work that does not exist.

Conflated

Parts of two real works merged into one reference.

Wrong details

A real work with an incorrect year, journal or DOI.

Unsupported claim

A real work attached to a finding it does not report.

The evidence

Fabrication fell with a newer model and did not disappear

55%of GPT-3.5 references were fabricated
18%of GPT-4 references were fabricated

Walters and Wilder (2023) checked 636 references in 84 short literature reviews. The figures describe the models tested at that time.

Carry forward

The team's working rule

Any source suggested by an AI tool is a lead. It enters the review only after it has been found in an independent system, its details confirmed, its claim checked and the checks recorded.

Section 2 asks how this rule applies to each category of AI research tool.

Learning Objectives for this section

  • Describe in plain terms how a large language model produces text by predicting one token after another from patterns learned in training.
  • Explain how retrieval-augmented generation adds a search step to a language model and why that step reduces some errors while leaving others in place.
  • Explain why fabricated and mis-attributed citations occur, and distinguish the main types of citation error.
  • Relate the behaviour of AI tools to the retrieval ideas from Lesson 3, including relevance ranking, recall and reproducibility.

1.1 Why a Review Team Needs to Understand AI Tools

By the middle of a review, a team has a question, a protocol, a database search and a plan for grey literature, and it has probably been told that an artificial intelligence (AI) tool could do much of the remaining work. Some of these tools help with parts of a review, and some of their outputs are wrong in ways that are hard to see. A reviewer who understands how the tools produce their answers can decide where they help and how to check them.

The running example in this course is the Cedar Valley evidence review, a fictional project. Before launching a community connector (social prescribing) program for older adults, the planning team of the fictional Cedar Valley Health Authority in British Columbia asked its small evidence team, made up of an evidence officer, a university librarian and a student intern, for a rapid scoping review and environmental scan within twelve weeks. The review asks which community-based interventions have been evaluated for reducing loneliness or social isolation among adults aged 65 and older, and with what outcomes. By this lesson the team has searched MEDLINE, Embase, CINAHL, PsycINFO and Web of Science (Lesson 4), retrieving 2,480 records, of which 1,870 remain after 610 duplicates are removed, and it has planned its grey-literature and web searching (Lesson 5).

Case: The intern's proposal

The intern has used a conversational AI assistant for coursework and suggests that the team ask it for the key studies on loneliness interventions for older adults. The librarian agrees to a test on two conditions. Every reference the tool produces will be checked against a primary source and recorded in a verification log, and the test will be documented in enough detail that a reader of the final report can see what was done. The evidence officer adds a third condition: nothing from the environmental scan, such as interview notes, will be entered into any tool. Section 3 works through the intern's results, and this section explains why the librarian's conditions are needed.

Artificial intelligence is a broad label for computer systems that perform tasks usually associated with human judgement. Machine learning is the part of AI in which a system learns patterns from examples, as the screening tools in Section 4 do. A large language model is a generative AI system that produces text. HSCI 841 Lesson 12 complements this lesson by examining the use of large language models to code qualitative data.

1.2 How a Large Language Model Produces Text

Tokens and next-token prediction

A large language model works with tokens, which are short pieces of text such as a whole word, part of a word or a punctuation mark. A long or unusual word, such as a surname or a technical term, is usually split into several tokens. In training, the model reads a very large collection of text and learns to predict the next token in a passage from the tokens before it, storing what it learns as billions of numerical settings called parameters. Most current models use the transformer architecture introduced by Vaswani and colleagues (2017), which lets the model weigh every earlier token when it predicts the next one.

Developers then train the model further on examples of instructions and on human ratings of its answers, so that it follows requests and holds a conversation. When it answers, it repeats one step many times: it calculates a probability for every possible next token, selects one, adds it to the text and calculates again. A setting often called temperature controls how much randomness enters the selection, which is one reason the same prompt can produce different answers on different occasions.

TokenClick to explore
ParametersClick to explore
Knowledge cutoffClick to explore
Context windowClick to explore
TemperatureClick to explore
HallucinationClick to explore

What next-token prediction means for a reviewer

Four consequences follow for evidence work. First, a language model on its own has no catalogue of sources. What it knows about the literature is spread across its parameters as patterns, so when it writes a reference it produces text that resembles references it has seen. Second, the model's information stops at its knowledge cutoff, so recent studies are missing unless the tool can search. Third, because selection involves chance, a repeated prompt can give a different list of studies, which makes the output difficult to reproduce. Fourth, fluency and confidence in the text carry no information about accuracy, because the model produces fluent text whether or not the content is correct.

Bender and colleagues (2021) described large language models as systems that assemble plausible sequences of language without grounding in meaning, and argued that this creates risks when people treat the output as reliable. Reviewers need only recognize that a model's output is shaped by its training text and by the wording of the prompt. A prompt that asks for "studies showing that befriending reduces loneliness" invites a list of supportive studies, whereas a neutral prompt that asks what has been evaluated and with what results gives the model less to agree with.

1.3 Retrieval-Augmented Generation

Many AI research tools add a search step to the language model. Lewis and colleagues (2020) introduced the term retrieval-augmented generation for systems that combine a retriever, which finds relevant passages in a collection of documents, with a generator, which writes an answer using those passages. In a typical tool, the user's question is turned into a search, the retriever returns the passages it ranks highest from its index (the web, or a database of scholarly abstracts), those passages are placed in the model's context window, and the model writes an answer with links to the passages it drew on. Many conversational assistants available in 2026 can also run a web search during a conversation, which turns them into retrieval-augmented systems for that response.

A. Language model alone Prompt from the user The model predicts the next token again and again from patterns learned in training, with no lookup of sources Answer and references written from patterns, so any of them may be false B. Retrieval-augmented generation Prompt from the user A retriever runs a ranked search over an index of web or scholarly text Top-ranked passages go into the context window The model writes an answer citing those passages; its claims still need checks Retrieval ties each citation to a real document. The retriever still returns only a ranked top set, and the summary can misstate what a source says.
Figure 6.1. A language model used alone writes references from learned patterns, whereas a retrieval-augmented tool first retrieves a ranked set of passages and then writes an answer that cites them.

What retrieval changes and what it leaves in place

Because each citation points to a document the retriever actually returned, a retrieval-augmented tool is much less likely to invent a reference outright. Other errors remain. The model can attach a claim to a source that does not support it, for example by turning a cautious conclusion into a confident one or by combining findings from two passages into a statement neither makes. The retriever can return weak, outdated or off-topic documents, and the model will summarize them with the same fluency as strong ones.

That last limit connects directly to Lesson 3. A bibliographic database running a Boolean search returns every record that matches the search, which is why reviews depend on it for high recall, the proportion of all relevant records that a search finds. A retrieval-augmented tool usually converts the question into an embedding, a list of numbers that represents its meaning, and returns the documents whose embeddings are closest. This approach, often called semantic search, can find relevant documents that use different words from the question, which is useful. It is still relevance ranking: the tool returns a top set of perhaps ten to a few dozen sources and has no notion of retrieving every eligible study. Its index also limits what it can find: many scholarly AI tools index abstracts from open sources and hold little of the content of Embase or CINAHL and little Canadian grey literature.

1.4 Why Fabricated and Mis-Attributed Citations Occur

References are among the most patterned text that exists. A reference list follows a predictable shape: surnames and initials, a year, a title with a colon, a journal name, a volume, an issue, a page range and a digital object identifier (DOI). A model trained on millions of reference lists learns that shape very well. Asked for sources without a retrieval step, it produces text in that shape from patterns associated with the topic. The surnames may belong to real researchers, the journal may be real and the title may read like a typical title, while the combination corresponds to no published work. Ji and colleagues (2023), in a survey of the problem across text-generation systems, describe this kind of output as hallucination: content that is fluent but unsupported by the source or by the facts.

Studies that tested chatbots in 2023 documented the problem. Walters and Wilder (2023) asked ChatGPT to write short literature reviews on 42 topics and checked the 636 references in the 84 reviews it produced. They found that 55 percent of the references produced by GPT-3.5 and 18 percent of those produced by GPT-4 were fabricated, and that 43 percent and 24 percent of the real references, respectively, contained substantive errors. Chelli and colleagues (2024) gave ChatGPT and Bard (later renamed Gemini) the inclusion criteria of eleven published systematic reviews on shoulder rotator cuff conditions and compared the references the tools returned with the references in the reviews. They classified 39.6 percent of the GPT-3.5 references, 28.6 percent of the GPT-4 references and 91.4 percent of the Bard references as hallucinated, and the GPT models found fewer than one in seven of the studies the human reviews had included. These figures describe the model versions tested at that time, and newer tools may perform differently. They show that fabrication fell with a newer model without disappearing.

Type of citation errorWhat it looks likeHow a reviewer detects it
Fabricated referenceA complete, plausible reference to an article, report or book that does not exist. The DOI, if given, fails to resolve or leads to a different work.No match for the exact title in a bibliographic database, Google Scholar or Crossref, and no match for the DOI.
Conflated referenceParts of two or more real works merged into one reference, such as the authors and year of one article with the title of another.The title search finds a real article, but its authors, year or journal differ from those given.
Real work, wrong detailsA real article with an incorrect year, journal, volume, page range or DOI.The publisher's record or the database record shows different bibliographic details.
Real work, unsupported claimA correct reference attached to a statement that the source does not make, or that overstates what it found.Reading the abstract and the relevant part of the full text shows a different finding or conclusion.
Real work, wrong fitAn accurate reference and summary for a source that does not meet the review's eligibility criteria, or that has been retracted.Comparing the source with the review's question and criteria, and checking for a retraction notice.

Two features of these errors make them hard to see. They look exactly like correct references, so a reader cannot tell them apart by inspection, and they are often mixed in with real references, so a list that contains several familiar studies earns trust it has not earned. The only reliable check is to look each source up in an independent system and to read what it says, which is the procedure Section 3 sets out.

Try it: Name the error

Each statement describes what a reviewer found when checking a reference that an AI tool suggested. Name the type of error from the table above before opening the answers below. (a) The title matches a real 2015 article, but the reference gives the authors and year of a different 2010 article by the same lead author. (b) The article exists and the details are correct, but the tool's summary says it found a large reduction in loneliness, and the abstract reports that most included evaluations were at high risk of bias and that effectiveness could not be judged. (c) No database, search engine or DOI registry has any record of the title, and the DOI returns an error.

Answers to the error-naming taskv

Statement (a) describes a conflated reference, because elements of two real works have been merged. Statement (b) describes a real work with an unsupported claim, because the reference is correct but the summary misstates the source. Statement (c) describes a fabricated reference, because nothing matching it can be found anywhere and the identifier does not resolve. Section 3 meets all three of these errors in the Cedar Valley verification exercise.

A common belief: a reference with a DOI must be realv

A DOI is a string of characters with a predictable format, beginning with 10 and a publisher prefix, and a language model can produce one as easily as it produces a title. A DOI shows that a reference is real only when it resolves at doi.org to the same work that the reference describes.

A common belief: newer models have solved the problemv

Walters and Wilder (2023) found far fewer fabricated references from GPT-4 than from GPT-3.5, which shows that newer models can improve. Their GPT-4 results still included fabricated references and errors in real ones, and no published evaluation that this course is aware of shows that any tool has eliminated them. Every reference therefore still needs checking, whichever tool produced it.

A common belief: a tool that searches the web is always groundedv

A retrieval-augmented tool links its claims to retrieved documents, which reduces fabrication. The claims can still misstate the documents, and the documents can still be weak or irrelevant. Reviewers check the claim against the source in the same way for these tools as for any other.

1.5 From Mechanism to Practice

The mechanism explains the rules the rest of the lesson applies. A generated reference is a lead to be checked. Retrieval-augmented tools rank and summarize a small top set, so they can supplement a documented database search without replacing it (Section 2). Outputs vary with the prompt, the model version and the date, so every use has to be recorded (Section 3). Machine learning can rank records well without making inclusion decisions, which is why ordering records for screening is its most established use in reviews (Section 4).

The Cedar Valley team's working rule

After discussing how these tools work, the Cedar Valley team adopts a working rule for the rest of the review. Any source suggested by an AI tool is treated as a lead. It enters the review only after a team member has found it in an independent system, confirmed its details, read enough of it to confirm the claim, and recorded these checks in the verification log.

Reflection

A colleague working on a review of school-based physical activity programs asks a conversational AI assistant, with its web search turned off, for ten key studies. The assistant returns ten complete references in APA style, each with a digital object identifier (DOI) and a one-sentence summary, and several of the authors are well-known researchers in the field. The colleague wants to paste the list straight into the reference list. A second colleague suggests switching to a retrieval-augmented tool, which first searches an index of documents and then writes an answer that cites the passages it retrieved, and says that this will make every claim correct. Write a reply to both colleagues that (a) explains how a large language model without a search step produces a reference and why a DOI and familiar author names do not show that a reference is real, (b) explains what a retrieval-augmented tool changes, (c) names two kinds of error that can remain when retrieval is used, and (d) ends with one practical rule for the review.

Model answer

A language model without search has no catalogue of articles to consult. It produces a reference one token at a time, following the patterns of the millions of references in its training text, so the result has the right shape (authors, year, title, journal and DOI) whether or not the work exists. Familiar author names show only that those researchers are associated with the topic in the training text, and a DOI is a predictable string that the model can generate as easily as a title. A DOI shows that a reference is real only when it resolves to the same work at doi.org.

A retrieval-augmented tool first retrieves real documents and then writes an answer citing them, so it is much less likely to invent a source. Two errors remain. The tool can attach a claim to a real source that does not support it, for example by turning a cautious conclusion into a confident one, and it can retrieve weak or off-topic documents and summarize them as fluently as strong ones. Because it returns a ranked top set, it also misses relevant studies.

The practical rule is that every AI-suggested source is a lead: it enters the review only after someone has found it in an independent database, checked its details, read enough of it to confirm the claim and recorded those checks in a verification log.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: When a large language model without a search tool writes a reference list, what is it doing?

A language model without retrieval has no catalogue of sources. It writes each reference token by token from the patterns of references it saw in training, which is why a reference can look correct and still refer to no real work. Option (a) is the most tempting distractor, but the parameters store patterns, not a list of documents.

Question 2: Which error is a retrieval-augmented tool still likely to make?

Retrieval ties each citation to a document the tool actually returned, so outright fabrication becomes much less likely. The tool can still misstate what a source says, for example by turning a cautious conclusion into a confident one, so claims still need checking against the source.

Question 3: A reference gives the title of a 2015 article but the authors, year and journal of a 2010 article by the same lead author. What type of error is this?

Both articles exist, and the reference merges parts of them, which is a conflated reference. It differs from a fabricated reference, because a title search finds a real article, but the other details belong to a different one.

Question 4: What did Walters and Wilder (2023) find when they compared references produced by GPT-3.5 and GPT-4?

Walters and Wilder (2023) found that 55 percent of GPT-3.5 references and 18 percent of GPT-4 references were fabricated. The newer model improved, but fabrication did not disappear, which is why every reference still needs checking.
Section 2 of 5

Categories of AI Research Tools and What Each Can and Cannot Do

⏱ Estimated reading time: 30 minutes
Section 2 of 5

Categories of AI Research Tools

Where each kind of tool gets its answers, and what it can and cannot do in a review.

Two sorting questions

Where does the answer come from, and what does the tool produce?

Source of the answer

Training data only, a search index, or the team's own records.

Type of output

Written text, or a ranking or classification of records.

Four categories

Examples available in 2026

Conversational assistants

ChatGPT, Claude, Gemini and Microsoft Copilot suit synonyms, explanations and editing.

AI search engines

Perplexity, Elicit, Consensus and Scopus AI find a ranked top set with summaries.

Citation-context tools

scite and Semantic Scholar show how later papers cite a study.

AI-assisted screening

ASReview, Rayyan, EPPI-Reviewer and DistillerSR order records for screening.

Purpose matters

Why an AI search cannot replace the database search

Database search

It returns every matching record, can be published in full and can be run again.

AI search

It ranks, selects and summarises, so recall and reproducibility both fall.

Studies an AI tool finds enter PRISMA 2020 as records identified by other methods.

The Cedar Valley test

Twenty sources, no new eligible study

20sources returned for the review's question
16already among the 1,870 database records
4outside the review's scope

The result offers some reassurance about the search. It cannot show what the tool missed.

Evaluating a new tool

Six questions for any tool

  • The documentation should state what the tool searches and what its index covers.
  • It should explain how the tool decides what to show.
  • An independent evaluation for a similar task carries the most weight.
  • The team must be able to save the prompt, version, date and output.
  • The terms should state what happens to information that is entered.
  • Funding and commercial interests should be known and reported.
Carry forward

From categories to checks

Every category produces output that a reviewer must check, record and disclose.

Section 3 works through the eight references the intern's assistant produced.

Learning Objectives for this section

  • Distinguish four categories of AI research tools: conversational assistants, AI search engines and research assistants, citation-context tools, and AI-assisted screening and extraction tools.
  • Explain, for each category, where its answers come from and which review tasks it can support.
  • Identify the limitations of each category that matter for a systematic or scoping review, including coverage, recall, reproducibility and transparency.
  • Evaluate an unfamiliar AI tool with a short set of questions about its index, its method, its evaluation and its handling of data.

2.1 Categories Last Longer Than Products

AI research tools change quickly. Products are launched, renamed, merged, given new features and withdrawn within the span of a single degree program, and a guide that ranked individual products would be out of date before the term ended. This section therefore organizes tools into four categories by how they work. Product names appear only as examples available in 2026, and their presence in the text is no endorsement. A student who understands the categories can place a new tool into one of them, predict its strengths and weaknesses, and ask the right questions about it.

Two questions sort most tools. The first asks where the answer comes from: from patterns the model learned in training, from a search of an index of documents, or from the review team's own records and decisions. The second asks what the tool produces: written text, such as an answer or a summary, or a ranking or classification of records. Figure 6.2 places the four categories on these two dimensions.

Where the answer comes from Training data only A search index The team's own records Writes text Ranks or classifies Conversational assistant without search Highest risk of fabricated references AI search engines and assistants with search Real sources, ranked top set, claims to check Summaries and extraction from uploaded texts Emerging; every value needs checking Few review tools work this way Citation-context tools Classify how later papers cite a study Active-learning screening Orders records; the human still decides
Figure 6.2. Four categories of AI research tools placed by where their answers come from and by what they produce. Red marks the highest risk of fabricated content, amber marks real sources with claims that need checking, and teal marks tools that order or classify material without writing claims.

2.2 The Four Categories

The tabs below describe each category in turn: how it works, examples available in 2026, the review tasks it can support and the limits that matter in a review.

How they work. A conversational assistant is a general-purpose large language model with a chat interface. Without a search tool, it answers from patterns learned in training, as Section 1 described. Many assistants can now run a web search during a conversation, and several offer extended research modes that run a series of searches and compile a report with links. Examples available in 2026 include ChatGPT (OpenAI), Claude (Anthropic), Gemini (Google) and Microsoft Copilot.

What they can support. Assistants are good at work with words. They can suggest synonyms, spelling variants and related terms for a concept table, which the librarian then tests against the databases (Lesson 4). They can explain an unfamiliar method in plain language, draft a plain-language summary from text the team has written, check a protocol section for clarity, or suggest how a search might be translated from one database's syntax to another, which a librarian must then check line by line.

What they cannot do. An assistant without search cannot be trusted to name sources, because it generates references from patterns. An assistant with search returns a small, ranked set of web pages chosen by criteria the user cannot see. Neither produces a search that another team could repeat, and neither has access to subscription databases such as Embase unless an institution has connected them.

How they work. AI search engines and research assistants are retrieval-augmented systems built for finding information. General tools search the web, and scholarly tools search large collections of article records and abstracts. They typically return a written answer with citations, a list of sources, or a table summarizing each source. Examples available in 2026 include Perplexity, which searches the web, and Elicit, Consensus and Scopus AI, which search scholarly literature; Scopus AI draws on the Scopus database.

What they can support. These tools can find a handful of relevant papers quickly, which helps when a team is scoping a topic or deciding whether a review already exists. They can supply seed studies for citation chasing (Lesson 5), and they can serve as a check on a database search: if a tool finds an eligible study that the database search missed, the librarian can ask why and revise the search.

What they cannot do. They return a ranked top set, so their recall is low by design and unknown in any particular case. Their index may hold only abstracts, may favour open-access and English-language work, and may include little grey literature. The ranking method is usually undisclosed, results change as the index and model change, and the written summaries can misstate the sources they cite.

How they work. Citation-context tools analyze the sentences in which later papers cite an earlier one. Some use machine learning to classify each citing statement, for example as supporting, contrasting or simply mentioning the cited work. Examples available in 2026 include scite, whose Smart Citations classify citing statements in this way, and Semantic Scholar, which shows the sentences in which a paper is cited and flags citations it judges influential. The citation-network tools introduced in Lesson 5, such as Connected Papers, ResearchRabbit and Litmaps, recommend related papers from patterns of citation and sit close to this category.

What they can support. They show quickly how a study has been received, whether later work questioned it, and which later papers built on it. This helps a reviewer interpret an influential study and can surface related work for citation chasing.

What they cannot do. A classification describes what a citing sentence says about a paper. It says nothing about the quality of the cited study, and an automated classifier can misread a sentence. Coverage depends on the full texts the tool can access, so citation counts and classifications are incomplete.

How they work. Screening tools use machine learning on the review team's own records. In active learning, a model learns from the reviewer's decisions on a few records and then shows the records it predicts are most likely to be relevant, retraining as each decision is made. Classifiers trained on large labelled collections can identify a particular kind of record, such as reports of randomized trials. Large language models are also being tested for screening and for extracting data from full texts. Examples available in 2026 include ASReview, an open-source active-learning tool, and the machine-learning features of screening platforms such as Rayyan, EPPI-Reviewer and DistillerSR.

What they can support. Ordering records so that relevant ones appear early is the most established use of machine learning in reviews. It lets a team begin full-text work sooner and supports decisions about when screening can safely end, which Section 4 examines.

What they cannot do. The model only predicts. A human reviewer still decides on every record the model shows, and a decision to stop screening before every record has been seen carries a risk of missing studies that the team must estimate and report. Extraction by a large language model produces values that each need checking against the full text.

The categories side by side

CategoryWhere answers come fromReview tasks it can supportMain limits in a review
Conversational assistantsPatterns learned in training, plus web search when that feature is onSynonyms for a concept table, plain-language explanations, drafting and editing the team's own textFabricated references without search, no reproducible search, no access to subscription databases
AI search engines and research assistantsA ranked search of a web or scholarly index, summarized by a language modelScoping a topic, finding seed studies, checking a database search for missed studiesLow and unknown recall, partial coverage, undisclosed ranking, summaries that need checking
Citation-context toolsCiting sentences in later papers, classified by machine learningInterpreting how a study was received and finding related workIncomplete coverage, misclassified sentences, no judgement of study quality
AI-assisted screening and extractionThe team's own records and screening decisionsOrdering records for screening, flagging study designs, drafting extraction for checkingUncertain stopping point, performance that varies by topic, extraction errors

2.3 Why an AI Search Cannot Replace the Database Search

The difference between these tools and the search built in Lesson 4 is a difference of purpose. A systematic or scoping review promises its readers that the team has made a documented, reproducible attempt to find all of the eligible evidence. The database search keeps that promise because it is exhaustive within each database, because its full strategy can be published, and because another team can run it again and obtain essentially the same records from the same databases on the same date. An AI search engine is designed to give a busy user a good answer quickly, so it ranks, selects and summarizes. Each of those steps reduces recall and reproducibility. For this reason the AI tools in this lesson can add to a review's search and can check it, and the database search remains the backbone of the review.

The same reasoning explains where AI-found studies belong in reporting. PRISMA 2020 separates records identified from databases and registers from records identified by other methods, such as websites, organizations and citation searching. A study that an AI search engine finds, and that the team then confirms and screens, enters the review through the other-methods route, and the search that found it is documented in the same way as a web search (Lesson 5 and Section 3).

Case: The Cedar Valley team tests three tools

The team runs three short tests during one week, recording each in its search log. First, the intern asks a conversational assistant, with web search turned off, for key studies on community-based interventions for loneliness in older adults; Section 3 verifies the eight references it produces. Second, the librarian enters the review's PCC question (population, concept and context) into a scholarly AI research assistant and exports the 20 sources it returns. After checking them against the team's reference library, the librarian finds that 16 were already among the 1,870 database records and that the other 4 are outside the review's scope: two studied adults younger than 65, one was a commentary, and one evaluated a hospital discharge program. Third, the evidence officer uses a citation-context tool to see how later papers cited Bickerdike and colleagues' (2017) systematic review of social prescribing, and reads several of the citing sentences to confirm the tool's labels.

The team draws a measured conclusion. The AI research assistant found no eligible study that the database search had missed, which offers some reassurance about the search. It cannot show what the tool itself missed, because a ranked top set of 20 says nothing about the hundreds of other relevant records in the index. The citation-context tool helped the team understand the debate about the evidence for social prescribing, and two of the labels it assigned did not match the citing sentences when the evidence officer read them.

2.4 Evaluating a Tool You Have Not Seen Before

A new tool will appear during this term, and another will appear next year. The questions below turn the categories into an evaluation that a student can complete in a few minutes using the tool's own documentation. They echo the expectations that the 2025 joint position statement on AI in evidence synthesis, discussed in Section 3, sets for tool developers, which include publishing how a tool works, making evaluations of it available and stating its limitations and potential biases.

What does the tool search, and what does its index cover?v

A tool can only return what is in its index. The documentation should state the source of its records (for example the open web, a named scholarly database or a set of open abstracts), the date range, and whether it holds full texts or abstracts only. A tool that indexes open abstracts in English will underrepresent some journals, languages and grey literature.

How does it decide what to show?v

Relevance ranking, semantic search, filters for study design, and limits on the number of results all shape the output. If the documentation does not explain the ranking, the team should treat the results as an unrepresentative sample of the relevant literature.

Has it been evaluated for this task, and by whom?v

An evaluation published by independent methodologists for a task like the team's, such as identifying intervention studies in health services research, carries more weight than a developer's general claims. A team that finds no evaluation can pilot the tool on a small set of studies it already knows are eligible.

Can the use be recorded and repeated?v

The team needs to save the prompt, settings, date, version and full output, and to export the sources in a format a reference manager can import. A tool that allows none of this cannot be used in a way that meets the reporting expectations in Section 3.

What happens to the information entered?v

The terms of use should state whether inputs are stored, whether they are used to train future models, and where the data are held. These terms decide whether anything beyond published material can be entered, a question Section 3 takes up for privacy and copyright.

Who provides it, and what are their interests?v

Funding, ownership and commercial relationships with publishers can affect what a tool indexes and promotes. The joint position statement asks review authors to report financial and non-financial interests related to the AI tools they use.

Try it: Classify and question a tool

Choose one AI tool that you have used or that your library offers. Using its documentation, place it in one of the four categories, state where its answers come from, and answer the six questions above in one or two sentences each. Note any question that the documentation does not answer, because a missing answer is itself useful information about whether the tool belongs in a review. Keep your notes, since an AI-assisted search log (Section 3.4) asks for the same details.

A note on product names

The examples named in this section were available in 2026. Products in this field are renamed, merged, given new features and withdrawn often. When a product named here has changed, the category and the questions still apply, and the tool's current documentation is the source to consult.

Reflection

A health authority manager proposes to replace the database search for a scoping review of community transportation programs for older adults with an AI research assistant. The tool returned 25 relevant-looking papers with summaries in ten minutes, whereas the librarian's Boolean search of MEDLINE, Embase, CINAHL, PsycINFO and Web of Science retrieved 3,100 records that will take weeks to screen. The tool's documentation states that it searches a large collection of open scholarly abstracts, returns its 25 highest-ranked results by relevance, and does not describe how it ranks them. Recall is the proportion of all relevant records that a search finds. Write a reply to the manager that (a) explains the difference between exhaustive Boolean retrieval and relevance ranking in terms of recall and reproducibility, (b) identifies two limits of this tool's index or method that matter for this review, and (c) proposes a role for the tool that the review could report transparently.

Model answer

A Boolean search in a bibliographic database returns every record that matches the search, which is why it can achieve high recall, and because the full strategy can be published, another team can run it again in the same databases and obtain essentially the same records. The AI research assistant ranks documents by predicted relevance and returns its top 25. Its recall for this question is therefore low by design and unknown, because the 25 results say nothing about the relevant records it did not show. Its ranking is undisclosed and its index and model change over time, so the search cannot be repeated.

Two limits of this tool matter here. Its index holds open scholarly abstracts, so it is likely to underrepresent subscription content from Embase and CINAHL and to hold little grey literature, such as program evaluations published by health authorities and transit agencies, which a scoping review of community programs needs. It also gives no information about how it ranks results, so the team cannot judge which kinds of studies it favours.

An appropriate role is as a supplementary search. The librarian can run the tool once, save the prompt, date, version and full output in an AI-assisted search log, check each source against the database results, and screen any new source through the usual process. New studies found this way are reported as records identified by other methods, and a new eligible study is also a signal to revisit the database search.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: Why does an AI search engine have low and unknown recall for a review question?

AI search engines rank documents by predicted relevance and return a top set, so they are not designed to retrieve every eligible study, and the user cannot tell how many relevant studies were left out. A Boolean database search returns every matching record, which is why reviews rely on it for recall.

Question 2: Which task is best suited to a citation-context tool such as scite?

Citation-context tools classify the sentences in which later papers cite a study, which helps a reviewer see how the study was received. The classification describes the citing sentence and says nothing about the quality of the cited study, so it cannot replace risk-of-bias appraisal.

Question 3: Of 20 sources from an AI research assistant, 16 were already in the Cedar Valley database results and the other 4 were out of scope. What is the most defensible conclusion?

Finding no new eligible study is some reassurance about the database search. A top set of 20 says nothing about the relevant records the tool did not show, so the test cannot prove the search complete, and the grey-literature plan still applies.

Question 4: If a tool's documentation does not explain how it ranks results, how should a review team treat those results?

Without knowing how a tool ranks results, the team cannot tell which kinds of studies it favours, so the results are best treated as an unrepresentative sample. Phrasing the prompt as a PCC question does not change how the tool ranks or what its index holds.
Section 3 of 5

Verifying AI Output: Reproducibility, Privacy and Disclosure

⏱ Estimated reading time: 45 minutes
Section 3 of 5

Verifying AI Output: Reproducibility, Privacy and Disclosure

A four-step check, a worked citation exercise, the search log and the disclosure statement.

The principle

The reviewer remains responsible

Evidence synthesists remain ultimately responsible for their work, including the decision to use AI.Paraphrase of the 2025 joint position statement (Flemyng et al., 2025)

The Committee on Publication Ethics stated in 2023 that AI tools cannot be authors because they cannot take responsibility for the work.

The procedure

Four checks for every AI-suggested source

1 · Existence

Search the exact title and resolve the DOI.

2 · Details

Match authors, year, journal and pages.

3 · Claim

Read the source and confirm the finding.

4 · Fit

Check eligibility and retraction status.

Worked exercise

Eight invented AI suggestions, checked

3correct in every detail and claim
3real but conflated, wrong in detail or misstated
2not found and treated as fabricated

None of the eight was an eligible primary study. The output was invented for teaching; the real sources are genuine publications.

Reproducibility

The AI-assisted search log

Why exact repetition fails

Outputs are sampled, models are updated and retired, and indexes change.

What to record

Tool, version, date, person, settings, exact prompt, full output and handling.

What stays out

Privacy, confidentiality and copyright

  • Interview notes and survey responses from the environmental scan stay out of consumer tools.
  • Unpublished manuscripts, grant applications and internal drafts are confidential.
  • Library licences may restrict uploading full texts, so the team asks the librarian first.
Disclosure

What a disclosure statement reports

The RAISE recommendations are emerging guidance, and the 2025 joint position statement asks authors to report AI use that makes or suggests judgements.

Name, version, dates
Purpose and stages
Justification
Interests
Limitations
Carry forward

From verification to screening

Machine learning can rank records well without making inclusion decisions.

Section 4 examines active-learning screening, the most established use of AI in reviews.

Learning Objectives for this section

  • Apply a four-step verification procedure (existence, bibliographic details, claim and fit) to AI-suggested citations and record the results in a verification log.
  • Explain why AI-assisted searches are difficult to reproduce, and document an AI-assisted search so that a reader can see exactly what was done.
  • Identify the privacy, confidentiality and copyright risks of entering material into AI tools.
  • Write a disclosure statement for AI use that follows emerging guidance, including the RAISE recommendations and the 2025 joint position statement on AI in evidence synthesis.

3.1 The Reviewer Remains Responsible

Every use of AI in a review rests on one principle: the authors of the review are responsible for its content, whatever tools they used. The 2025 position statement on AI use in evidence synthesis issued jointly by Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence (Flemyng et al., 2025) states that evidence synthesists remain ultimately responsible for their work, including the decision to use AI, and that AI should be used with human oversight. The Committee on Publication Ethics (COPE) stated in 2023 that AI tools cannot be authors because they cannot take responsibility for the work. In practice, every source an AI tool suggests is checked against a primary source, and every use is recorded and reported.

3.2 A Four-Step Verification Procedure

A verification log is a table in which a reviewer records, for each AI-suggested source, the checks performed, the result and the action taken. The procedure below fills one row of the log for each source and takes a few minutes per reference.

1. Does it exist? Search the title, resolve the DOI 2. Details match? Authors, year, journal, pages 3. Claim supported? Read the abstract and full text 4. Eligible, current? Criteria and retraction check no no no no Not found: treat as fabricated and remove it Correct it from the publisher's record Rewrite the summary from the source itself Use as background or for citation chasing; exclude retracted work Every check, verdict and action is recorded in one row of the verification log.
Figure 6.3. The four-step verification procedure. A source passes to the next step only when it passes the current one, and every outcome is recorded in the verification log.

Step 1, existence. Search for the exact title, in quotation marks, in a bibliographic database such as MEDLINE, in Google Scholar and in the library's discovery tool, and enter any DOI at doi.org or look it up in Crossref, the agency that registers most journal DOIs. A source that none of these systems can find is treated as fabricated. Step 2, details. Compare the authors, year, title, journal, volume, issue, pages and DOI with the publisher's record. Many articles appear online a year or more before they are assigned to an issue, so a difference between the two dates is common. Step 3, claim. Read the abstract and the relevant part of the full text, and decide whether the source says what the AI output says it says. Step 4, fit. Compare the source with the eligibility criteria, and check for a retraction or correction on the publisher's page or in the Retraction Watch database, which Crossref now makes openly available.

3.3 Worked Verification Exercise: The Cedar Valley Citation Check

This exercise uses AI output invented for teaching, imitating what a conversational assistant without web search can produce. It mixes real publications with fabricated and mis-attributed entries, which the answers identify. Do not cite the fabricated entries.

Invented AI output for teaching (this is not output from any real tool)

Prompt entered by the intern: "List eight peer-reviewed sources on community-based interventions to reduce loneliness or social isolation among adults aged 65 and older. Give a full APA reference and a one-sentence summary of the main finding for each."

Invented response:

  1. Dickens, A. P., Richards, S. H., Greaves, C. J., & Campbell, J. L. (2011). Interventions targeting social isolation in older people: A systematic review. BMC Public Health, 11, 647. https://doi.org/10.1186/1471-2458-11-647. Summary given: group-based and participatory interventions were more often associated with improved outcomes than one-to-one and non-participatory interventions.
  2. Holt-Lunstad, J., Smith, T. B., & Layton, J. B. (2010). Loneliness and social isolation as risk factors for mortality: A meta-analytic review. PLoS Medicine, 7(7), e1000316. Summary given: loneliness and social isolation are associated with an increased risk of death.
  3. Thompson, R., Nguyen, L., & Bartlett, K. (2021). Community connector programs and loneliness among rural Canadian seniors: A cluster randomized trial. Canadian Journal on Aging, 40(3), 412–425. https://doi.org/10.1017/S0714980821000187. Summary given: rural seniors referred to community connectors reported lower loneliness at six months than those receiving usual care.
  4. Gardiner, C., Geldenhuys, G., & Gott, M. (2018). Interventions to reduce social isolation and loneliness among older people: An integrative review. Health & Social Care in the Community, 26(2), 147–157. https://doi.org/10.1111/hsc.12367. Summary given: adaptability, a community development approach and productive engagement were features of the most effective interventions, although the quality of the evidence was generally weak.
  5. Bickerdike, L., Booth, A., Wilson, P. M., Farley, K., & Wright, K. (2017). Social prescribing: Less rhetoric and more reality. A systematic review of the evidence. BMJ Open, 7(4), e013384. https://doi.org/10.1136/bmjopen-2016-013384. Summary given: the review found consistent evidence that social prescribing through link workers reduces loneliness among older adults in primary care.
  6. Cattan, M., White, M., Bond, J., & Learmouth, A. (2008). Preventing social isolation and loneliness among older people: A systematic review of health promotion interventions. Journal of Aging and Health, 20(1), 41–67. Summary given: group activities with an educational or support component were the most effective interventions, and the effectiveness of home visiting and befriending remained unclear.
  7. Okafor, M., Lindqvist, E., & Harland, J. (2023). Social prescribing link workers and loneliness in adults over 65: A systematic review and meta-analysis. Journal of Applied Gerontology, 42(9), 1874–1889. Summary given: link worker programs produced a moderate reduction in loneliness, with larger effects in longer programs.
  8. National Academies of Sciences, Engineering, and Medicine. (2020). Social isolation and loneliness in older adults: Opportunities for the health care system. The National Academies Press. https://doi.org/10.17226/25663. Summary given: a consensus report on how the health care system can identify and respond to social isolation and loneliness in older adults.
Try it: Verify before you read the answers

Choose at least four of the eight entries, including entries 2, 5 and 7, and apply the four steps using PubMed, Google Scholar, doi.org or Crossref, and your library's discovery tool. Record your result for each step, then compare your results with the team's results below.

Entry 1: Dickens and colleagues (2011)v

Exists: yes; the title, the DOI and PubMed lead to the same article. Details: correct. Claim: supported. The abstract reports that 79 percent of group-based and 55 percent of one-to-one interventions reported at least one improved outcome, and over 80 percent of participatory compared with 44 percent of non-participatory interventions. These figures come from vote counting across studies at medium to high risk of bias, a method whose limits Lesson 9 explains. Fit: a systematic review, used for citation chasing. Verdict: verified as given.

Entry 2: Holt-Lunstad and colleagues (2010)v

Exists: the title search finds a real article, but it is a 2015 article by Holt-Lunstad, Smith, Baker, Harris and Stephenson in Perspectives on Psychological Science, 10(2), 227–237. Details: the authors, year, journal, volume and article number given belong to a different real article, Holt-Lunstad, Smith and Layton (2010), "Social relationships and mortality risk: A meta-analytic review," PLoS Medicine, 7(7), e1000316. Claim: both articles concern social relationships, loneliness or isolation as risk factors for mortality, so the summary fits the topic; the team cites whichever article it reads. Fit: neither evaluates an intervention, so both are background. Verdict: conflated reference, corrected.

Entry 3: Thompson and colleagues (2021)v

Exists: no. No record matching the title or authors appears in MEDLINE, Google Scholar or Crossref, and the DOI returns a "DOI not found" message at doi.org. The Canadian Journal on Aging is a real journal, which helps the entry look credible. Verdict: not found, treated as fabricated, removed and recorded in the log. This entry was invented for this exercise.

Entry 4: Gardiner and colleagues (2018)v

Exists: yes. Details: correct for the 2018 issue, 26(2), 147–157; the article appeared online in 2016, which is why some databases list that year. Claim: supported. The abstract reports that adaptability, a community development approach and productive engagement were associated with the most effective interventions, and that the evidence was generally weak. Fit: an integrative review, used for citation chasing. Verdict: verified as given.

Entry 5: Bickerdike and colleagues (2017)v

Exists: yes. Details: correct. Claim: unsupported. The review included 15 evaluations of United Kingdom programs that referred primary care patients, without an age restriction, to a link worker, and rated all 15 at high risk of bias. The authors noted that most evaluations presented positive conclusions despite clear methodological shortcomings, and concluded that the evidence was insufficient to judge either success or value for money. The AI summary turned a cautious conclusion into a confident one and added an older-adult focus. Verdict: real source with an unsupported claim; the team rewrites the summary from the abstract and uses the review for citation chasing.

Entry 6: Cattan and colleagues (2008)v

Exists: yes, under the same title and authors. Details: wrong. The article was published in Ageing & Society, 25(1), 41–67, in 2005, and the journal, year and volume given are incorrect. Claim: supported; the abstract reports that nine of the ten effective interventions were group activities with an educational or support input, and that the effectiveness of home visiting and befriending remained unclear. Verdict: real source with wrong details, corrected and used for citation chasing.

Entry 7: Okafor and colleagues (2023)v

Exists: no. No record matching the title or the authors appears in any of the systems searched, and the entry has no DOI, which is unusual for a 2023 journal article. Verdict: not found, treated as fabricated and removed. This entry was invented for this exercise. It is the most tempting entry because it appears to answer the planning team's question directly, and a fabricated source that fits the reader's question closely is the kind most likely to survive an unverified reading.

Entry 8: National Academies of Sciences, Engineering, and Medicine (2020)v

Exists: yes; the DOI resolves to the report. Details: correct. Claim: the summary restates the report's stated scope, which the title confirms. Fit: the prompt asked for peer-reviewed sources, and this is a consensus study report, so the team records it as a report, a form of grey literature, and uses it as background. Verdict: verified, with the source type corrected.

The team's verification log

No.Source as givenExistsDetailsClaimVerdict and action
1Dickens et al., 2011YesCorrectSupportedVerified; citation chasing
2Holt-Lunstad et al., 2010YesTwo articles mergedFits topicConflated; corrected; background
3Thompson et al., 2021NoNot applicableNot applicableFabricated; removed
4Gardiner et al., 2018YesCorrectSupportedVerified; citation chasing
5Bickerdike et al., 2017YesCorrectUnsupportedSummary rewritten; citation chasing
6Cattan et al., 2008YesJournal, year, volume wrongSupportedCorrected to 2005; citation chasing
7Okafor et al., 2023NoNot applicableNot applicableFabricated; removed
8National Academies, 2020YesCorrectSupportedVerified; recorded as report

Six of the eight sources exist, only three were correct in every detail and claim as given, and none is an eligible primary study for the scoping review. The tool did name several well-cited reviews that are useful for citation chasing. A reviewer who copied the list unchecked would have cited two articles that do not exist, attributed a confident finding to a review that judged the evidence insufficient, and given wrong details for two more.

3.4 Reproducibility and the AI-Assisted Search Log

Lesson 3 showed that web search engines personalize and re-rank results. AI tools add further variation: the model samples its output, providers update and retire models, the index behind a retrieval tool changes, and small changes in a prompt's wording can change the results. Because an AI-assisted search cannot be repeated exactly, the goal is a record full enough that a reader can see what was done, judge its influence and repeat it approximately.

Field in the AI-assisted search logWhat to recordCedar Valley entry for the second test
Tool and providerProduct name, provider, and type of account or licenceA scholarly AI research assistant, recorded by name; university licence
Model or versionThe model or version shown, or a note that none was shownThe version shown on the settings page
Date and personThe date of the search and the team member who ran itThe date of the test; the librarian
SettingsAny filters, modes, date limits or source options selectedDefault settings, with results limited to journal articles
PromptThe exact wording entered, including any follow-up promptsThe review's PCC question as written in the protocol
OutputThe full output saved as a file, with its file nameExported list of 20 sources and a saved copy of the answer page
HandlingHow the output was verified, de-duplicated and screened16 duplicates of database records; 4 outside the review's scope; none added

This log sits with the web search documentation from Lesson 5 and the PRISMA-S log from Lesson 4. PRISMA 2020 also asks authors to report automation tools used in selecting studies and collecting data, which covers the tools in Section 4.

3.5 Privacy, Confidentiality and Copyright

Whatever is entered into an AI tool leaves the team's control to some degree. Depending on the provider's terms and the account, inputs may be stored, reviewed or used to train future models, and may be held in another country. Public bodies in British Columbia, including health authorities and universities, are subject to the Freedom of Information and Protection of Privacy Act and have policies on which AI tools staff and students may use. The team checks both institutions' policies before entering anything other than published material.

Interview notes, survey responses and any other information about identifiable people stay out of consumer AI tools. For the Cedar Valley environmental scan, this covers the responses from the 11 programs that complete the survey and the notes from the 7 key informant interviews (Lesson 11). Data that seem anonymous can still identify a person or a small organization.

Unpublished manuscripts, grant applications, internal health authority documents and draft reports belong to their authors or organizations. Many journals and funders prohibit peer reviewers from entering confidential material into generative AI tools; the United States National Institutes of Health, for example, prohibits its peer reviewers from using them to analyze or critique grant applications.

Published articles are protected by copyright, and the library's licences with publishers set terms for subscribed content, which in some cases restrict uploading full texts into AI tools. Before uploading full texts, the team asks the librarian whether the licence allows it.

3.6 Disclosure

A disclosure statement tells readers which AI tools were used, for what purpose, at which stages, and how their output was checked. The RAISE recommendations (Responsible use of AI in evidence SynthEsis), developed through work involving Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence, are emerging guidance on responsible AI use in evidence synthesis that was still being revised in 2025. The joint position statement that endorses them (Flemyng et al., 2025) asks authors to report fully any AI use that makes or suggests judgements, such as eligibility, risk of bias, extraction or synthesis decisions, giving each tool's name, version and dates of use, its purpose and the stages affected, the justification for using it, any related interests, and its limitations. Editing limited to spelling and grammar generally need not be reported, although journal policies vary. The template below covers these elements.

Copy the template and replace each bracketed field.

We used [tool name, provider and version or model] between [dates] to [purpose] at the [review stage or stages] stage. [The tool was used with the following settings: settings.] The prompts and full outputs are provided in [appendix or repository]. We chose this tool because [justification, including any evaluation of the tool or the team's own pilot]. Every output was [verification method, for example checked against the primary source by one reviewer and confirmed by a second], and [number] of [number] AI-suggested sources were retained after verification. The tool did not [list the judgements it did not make, such as decisions to include or exclude studies]. The authors declare [interests related to the tool, or no such interests]. Limitations of this use include [limitations and their likely effect on the review]. The authors take full responsibility for the content of this review.

During week five of the review we used a general-purpose conversational assistant, with web search turned off, and a scholarly AI research assistant, each recorded by name and version with the prompts and full outputs in Appendix C, to identify sources for citation chasing and to check the sensitivity of the database search. These tools supplemented a peer-reviewed database search. The student intern checked every AI-suggested source for existence, bibliographic accuracy and support for the stated finding, and the librarian confirmed the checks. Six of the eight sources suggested by the conversational assistant existed; two could not be found and were removed, three required corrections to their details or summaries, and none was an eligible primary study. None of the 20 sources from the research assistant was a new eligible study. No AI tool made screening, eligibility, charting or synthesis decisions, and no environmental scan data were entered into any AI tool. The authors declare no interests related to these tools. These searches cannot be reproduced exactly because AI outputs vary over time. The authors take full responsibility for the content of this review.

I used [tool and version] on [date] to [purpose]. My prompts and the full output are in [appendix]. I checked every source it suggested against [databases or systems] and recorded the results in my verification log; [number] of [number] sources were verified and kept. I made all screening, eligibility and synthesis decisions myself, and I did not enter any personal or confidential information into the tool.

Where the disclosure goes

Place the statement where the journal or funder specifies, usually the methods section, and put the prompts, outputs and verification log in an appendix or repository.

Reflection

You are checking AI-suggested references with a four-step procedure: (1) existence: can the exact title be found in a bibliographic database, Google Scholar or Crossref, and does the DOI resolve at doi.org; (2) details: do the authors, year, journal, volume and pages match the publisher's record; (3) claim: does the source say what the AI summary says; and (4) fit: does the source meet the review's eligibility criteria and has it been retracted. Verdicts are: verified as given; real with wrong details (correct it); real with an unsupported claim (rewrite the summary); conflated (corrected); or not found (treat as fabricated and remove). One AI-suggested reference describes a 2019 randomized trial of volunteer befriending for older adults by two named authors, with the summary "befriending halved loneliness scores over twelve months." Suppose that your checks find that the exact title appears in no database, that the DOI returns "DOI not found", and that Google Scholar shows a 2019 trial of volunteer befriending with a similar title, by different authors in a different journal, which reported a small reduction in loneliness that was not statistically significant. Write the verification-log entry for the reference (existence, details, claim, verdict and action), explain whether the similar trial can replace it and on what conditions, and write a two-sentence disclosure statement for this use of the AI tool.

Model answer

Log entry. Existence: not found; the title is absent from MEDLINE, Google Scholar and Crossref, and the DOI does not resolve. Details and claim: not applicable, because there is no source to compare. Verdict: not found, treated as fabricated. Action: removed from the reference list, with the entry kept in the log so that the removal is documented.

The similar trial. The real trial is a different source, by different authors in a different journal, so it cannot be substituted as a correction of the AI reference. It can enter the review only as a new lead found while verifying: I would retrieve it, confirm its details from the publisher's record, read it to describe its finding accurately (a small, non-significant reduction in loneliness, which contradicts the AI summary), check it against the eligibility criteria, and screen it through the usual process. If it is included, it is reported as a record identified by other methods.

Disclosure. I used [tool name and version] on [date] to suggest sources on volunteer befriending, and the prompt and full output are in Appendix B. I checked every suggested source against MEDLINE, Google Scholar and Crossref, removed one reference that could not be found, and made all eligibility decisions myself.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: In the four-step verification procedure, what is checked in step 3?

Step 3 checks the claim by reading the abstract and the relevant part of the full text. Step 1 checks existence, step 2 checks bibliographic details, and step 4 checks fit with the eligibility criteria and retraction status.

Question 2: In the Cedar Valley exercise, entry 5 (Bickerdike and colleagues, 2017) existed and had correct details. Why was it flagged?

The review rated all 15 included evaluations at high risk of bias and concluded that the evidence was insufficient to judge success or value for money. The AI summary turned this cautious conclusion into a claim of consistent benefit for older adults, so the team kept the reference and rewrote the summary from the source.

Question 3: Why can an AI-assisted search not be reproduced exactly?

A model samples its output, providers update and retire models, and the index behind a retrieval tool changes. The reasonable goal is a record full enough that a reader can see what was done and repeat it approximately, which is what the AI-assisted search log provides.

Question 4: Under the 2025 joint position statement on AI in evidence synthesis, which use of AI must be fully reported?

The statement asks authors to report fully any AI use that makes or suggests judgements, including eligibility, risk of bias, data extraction and synthesis, with each tool's name, version, dates, purpose, justification, interests and limitations. Editing limited to spelling and grammar generally need not be reported.
Section 4 of 5

Active-Learning Screening and the Evidence on Its Performance

⏱ Estimated reading time: 40 minutes
Section 4 of 5

Active-Learning Screening and the Evidence on Its Performance

Ordering records, measuring performance, deciding when to stop, and judging AI use at every stage.

Three uses

Ordering, second screening and stopping early

Prioritized screening

Every record is screened in the model's order.

Second screener

The model replaces one of two human reviewers.

Truncation

Screening stops before every record is read.

O'Mara-Eves et al. (2015) judged prioritization safe and ready for live reviews.

The cycle

How active learning orders records

Reviewer labels a few records
Model learns from every decision
Model shows its top-ranked record
Reviewer decides, and the cycle repeats

ASReview, from Utrecht University, implements this cycle and has a simulation mode for completed reviews (van de Schoot et al., 2021).

Measuring performance

An illustrative Cedar Valley simulation

540records screened to reach 95% recall
28.9%of the 1,870 records
66%work saved over sampling at 95% recall
Work saved over sampling
\[ \text{WSS@95} = \frac{N - n_{95}}{N} - 0.05 \]
The stopping problem

In a live review, nobody knows the total

Heuristic rule

Stopping after 100 consecutive irrelevant records would have ended at record 710, with 138 of 142 found.

Statistical rule

A random sample of the remaining records tests whether a recall target has been met (Callaghan & Müller-Hansen, 2020).

What the evidence shows

Savings are real, and conditions vary

  • Most studies reviewed by O'Mara-Eves et al. (2015) suggested savings of 30 to 70 percent.
  • Single-reviewer abstract screening missed 13 percent of relevant studies (Gartlehner et al., 2020).
  • GPT-4's screening performance fell after adjustment for chance agreement (Khraisha et al., 2024).
Bringing it together

A decision table for every stage

Generally acceptable

Synonyms for the concept table, ordering records while all are screened, editing the team's own text.

Only with safeguards

Supplementary AI searches, a model as second screener, model-drafted extraction checked value by value.

Not acceptable

Replacing the database search, unchecked exclusions, and scan data in consumer tools.

What comes next

Reflection, assessment and Lesson 7

The final reflection asks you to write an AI-use plan and disclosure statement for a new Cedar Valley rapid review.

  • Complete the reflection before submitting the final assessment.
  • The disclosure statement goes in the methods, and the logs go in an appendix.
  • Lesson 7 imports, de-duplicates and screens the Cedar Valley records.

Learning Objectives for this section

  • Explain how active learning orders records during title and abstract screening, and how prioritized screening differs from using a model as a second screener or as a basis for stopping early.
  • Interpret recall, the proportion of records screened and work saved over sampling from a screening simulation.
  • Describe the stopping problem and compare heuristic and statistical stopping rules.
  • Summarize what published evaluations show about active learning and large language model screening, and the conditions that limit those findings.
  • Use a decision table to judge when AI assistance is acceptable at each stage of a review, and apply it to the Cedar Valley review.

4.1 The Screening Workload

Title and abstract screening is usually the most time-consuming stage of a review. In the fictional Cedar Valley review, 1,870 records remain after de-duplication, and Lesson 7 describes how the evidence officer and the intern screen each one independently, which amounts to 3,740 decisions before any full text is read. Machine learning has been applied to this stage for longer than to any other, and its use here has the most evidence behind it.

Screening tools can use machine learning in three ways, and the distinction matters more than the choice of product. In prioritized screening, the tool changes the order in which records are shown, and the team still screens every record. With a model acting as a second screener, the model's prediction takes the place of one of two human reviewers. In screening truncation, the team stops screening before it has seen every record, because the model's ranking suggests that few relevant records remain. O'Mara-Eves and colleagues (2015), in a systematic review of text mining for identifying studies, concluded that using text mining to prioritize the order of screening should be considered safe and ready for use in live reviews, that its use as a second screener could be considered with caution, and that using it to eliminate studies automatically was promising but not yet fully proven.

4.2 How Active Learning Works

Active learning is a form of machine learning in which the model chooses which unlabelled record a human should label next, learns from that label, and chooses again. In screening, the reviewer first labels a few records as relevant or irrelevant. The tool converts each title and abstract into numerical features, such as word counts weighted by how distinctive the words are, or embeddings of the kind described in Section 1. A classifier, such as naive Bayes or logistic regression, learns from the labelled records which features predict relevance, and then scores every unscreened record. The tool shows the reviewer the record with the highest score. The reviewer decides, the model retrains on the larger set of decisions, and the cycle repeats.

Start The reviewer labels a few relevant and irrelevant records The model learns from every decision made so far The model ranks the unscreened records and shows the top one The reviewer decides after reading the title and abstract The reviewer makes every decision. The model changes only the order.
Figure 6.4. The active-learning cycle in title and abstract screening. After a few starting labels, the model retrains after each decision and shows the record it now ranks highest.

ASReview, described by van de Schoot and colleagues (2021), is an open-source tool developed at Utrecht University that implements this cycle and lets the user choose among several feature extractors and classifiers. Its simulation mode replays a completed review whose decisions are known, showing how early the relevant records would have appeared. Screening platforms such as Rayyan, EPPI-Reviewer and DistillerSR offer comparable machine-learning features within their own workflows. A different kind of tool, the classifier trained once on a large labelled collection, is used for specific tasks. Thomas and colleagues (2021), for example, developed and evaluated a classifier that identifies reports of randomized controlled trials for Cochrane Reviews.

4.3 Measuring Performance: A Worked Cedar Valley Simulation

Three measures describe how well a screening tool performs. Recall, also called sensitivity, is the proportion of all relevant records that have been found at a given point. The proportion screened is the share of all records the reviewer has read at that point. Work saved over sampling, introduced by Cohen and colleagues (2006), compares the work needed to reach a target recall with the work needed to reach the same recall by screening in random order.

Work saved over sampling at 95 percent recall

WSS@95 = (N − n95) ÷ N − 0.05

Here N is the total number of records and n95 is the number screened when 95 percent of the relevant records have been found. In random order, a reviewer must screen about 95 percent of the records to find 95 percent of the relevant ones, so subtracting 0.05 removes the saving that random order would give.

Case: Looking ahead to a screening simulation (illustrative figures)

Lesson 7 describes how the team screens all 1,870 records in duplicate and sends 145 forward for full-text retrieval, of which 142 full texts are obtained and assessed. To learn what active learning might offer when the review is updated, the intern replays those completed decisions in ASReview's simulation mode, treating the 142 records whose full texts were assessed as the relevant records. The figures that follow are illustrative and were constructed for teaching. In the simulation, 88 of the 142 relevant records appear within the first 187 records screened, which is the first 10 percent. The 135th relevant record, which takes recall past 95 percent (135 ÷ 142 = 95.1 percent), appears at record 540. The last relevant record appears at record 1,395.

Point in the simulationRecords screenedProportion screenedRelevant foundRecall
First 10 percent of records18710.0%8862.0%
95 percent recall reached54028.9%13595.1%
Stopping heuristic triggered71038.0%13897.2%
Last relevant record found1,39574.6%142100%
All records screened1,870100%142100%

The work saved over sampling at 95 percent recall is (1,870 − 540) ÷ 1,870 − 0.05 = 0.711 − 0.05 = 0.661, or about 66 percent. In words, prioritized screening reached 95 percent recall after the reviewer had read 540 records, whereas random order would have required about 1,777 records (95 percent of 1,870) to reach the same recall on average. Figure 6.5 shows the same simulation as a curve.

0 50 100 142 0 500 1,000 1,500 1,870 Records screened Relevant records found 1 2 3 1: 95 percent recall at 540 2: stopping rule at 710, with 138 of 142 found 3: last relevant record at 1,395 Active-learning order Expected in random order
Figure 6.5. Illustrative recall curve for a simulation of the Cedar Valley screening. The curve rises steeply and then flattens, and the last few relevant records appear long after most have been found.

4.4 The Stopping Problem

The simulation measures are calculated after the fact, when the team knows that 142 records are relevant. In a live review nobody knows the total, so nobody knows the recall at any moment. Prioritized screening that continues to the last record is safe, and the time saving from stopping early is where the risk lies. A stopping rule is a pre-specified criterion for ending screening before every record has been read.

Heuristic stopping rules are simple to apply. The most common stops after a fixed number of consecutive irrelevant records, such as 50 or 100, on the reasoning that a long run without a relevant record means few remain. Boetje and van de Schoot (2024) proposed the SAFE procedure, a practical and deliberately conservative combination of stopping heuristics for active-learning screening in tools such as ASReview. Statistical stopping rules take a different approach. Callaghan and Müller-Hansen (2020) proposed drawing a random sample from the records that remain unscreened and using the number of relevant records in the sample to test whether a recall target, such as 95 percent, has been reached with a stated level of confidence. Callaghan and colleagues (2024) later argued that stopping criteria for computer-assisted screening need to be well evaluated and transparently reported before reviews rely on them.

In the Cedar Valley simulation, a heuristic of 100 consecutive irrelevant records would have stopped screening at record 710, after the 138th relevant record had appeared at record 610. Stopping there would have saved reading 1,160 records, or 62.0 percent of the total, and would have missed 4 of the 142 relevant records (recall 97.2 percent). Suppose that one of the four missed records was a qualitative evaluation of a men's shed program whose abstract never used the words loneliness or isolation, and that it was later among the included studies. A missed record can matter even at high recall. Hard-to-find records usually describe the concept in unexpected language or have weak abstracts, and a scoping review of a varied public health literature most needs to catch them.

Prioritized screeningClick to explore
Model as second screenerClick to explore
Heuristic stopping ruleClick to explore
Statistical stopping ruleClick to explore

4.5 What the Evidence Shows

The evidence on active learning comes mainly from simulations on completed reviews. O'Mara-Eves and colleagues (2015) found that most of the studies in their systematic review suggested workload savings of between 30 and 70 percent, sometimes accompanied by the loss of 5 percent of relevant studies, which is 95 percent recall. They noted that studies rarely replicated each other, and that more evaluation was needed outside highly technical and clinical areas. Van de Schoot and colleagues (2021) reported simulation studies with ASReview in which relevant records were found after screening a fraction of the total. Hamel and colleagues (2021) evaluated active machine learning across ten completed systematic reviews and developed a seven-step framework for teams adopting it, which ends with guidance on truncating screening.

Several conditions limit how far these findings transfer to a new review. Performance varies with the topic, the share of records that are relevant, the clarity of the abstracts and the size of the dataset. Simulations treat the original human decisions as the truth, although human screeners also miss studies; Gartlehner and colleagues (2020) found in a randomized trial that single-reviewer abstract screening missed 13 percent of relevant studies. The joint position statement on AI in evidence synthesis (Flemyng et al., 2025) cites that finding as an example of the trade-offs a team should weigh, since in a rapid review an AI tool used as a second reviewer may catch studies that a single human reviewer would miss.

Evaluations of large language models for screening are newer and less settled. Khraisha and colleagues (2024) tested GPT-4 on title and abstract screening, full-text screening and data extraction across several literature types and languages. GPT-4's accuracy matched human performance on some tasks, but performance fell across all stages once the authors adjusted for chance agreement and for the imbalance between included and excluded records. The authors concluded that substantial caution was warranted, while noting that for certain tasks under specific conditions the model approached human performance.

Questions to ask of a performance claimv

A reader can test a claim that a tool saves a stated share of screening work by asking whether the result came from a simulation or a live review, at what recall the saving was measured, how similar the evaluation datasets were to the reader's topic, whether the stopping point was chosen in advance, and whether the evaluators had an interest in the tool.

What to report when active learning is usedv

PRISMA 2020 asks authors to describe any automation tools used in selecting studies. A full description names the tool and version, the model settings, the starting labels, whether the tool prioritized, acted as a second screener or supported stopping early, the stopping rule and when it was specified, the number of records left unscreened, and any estimate of recall.

4.6 When AI Assistance Is Acceptable: A Decision Table

The table below brings the lesson together. It is a teaching tool for this course that applies three principles found in the joint position statement and in the earlier sections: human oversight of every judgement, verification of every output against primary sources, and full disclosure of any AI use that makes or suggests judgements. Journals, funders, the university and the health authority may set stricter rules, and those rules take precedence.

Review stageGenerally acceptable, with human checkingAcceptable only with added safeguardsNot acceptable
Question, scope and protocol (Lessons 1 and 2)Explaining question frameworks; checking the clarity of the team's own draftSuggesting options for scope, with the team and interest holders decidingLetting a tool set the question or the eligibility criteria
Search development (Lessons 3 and 4)Suggesting synonyms and spelling variants for the concept tableTranslating a search between databases, with the librarian checking every line and a PRESS peer reviewRunning an AI-generated search string without librarian review and testing
Searching (Lessons 4 and 5)Suggesting organizations and websites to searchAI search engines as a supplementary, documented search, with every source verifiedReplacing the database search with an AI search; citing unverified AI-suggested sources
Citation chasing (Lesson 5)Citation-network and citation-context tools to find related work, with records screened as usualRelying on a tool's recommendations as the only form of citation chasing, with the limits reportedTreating a "supporting" citation label as evidence of a study's quality
Title and abstract screening (Lesson 7)Active learning to order records while every record is screenedA model as second screener, or stopping early under a pre-specified, evaluated stopping rule, with recall estimated and reportedExcluding records by a tool alone, with no human check and no estimate of recall
Full-text screening (Lesson 7)Tools that find full texts or highlight passagesModel-suggested decisions, with a human deciding on every recordA model making final eligibility decisions
Data extraction and charting (Lessons 8 and 10)Formatting a charting table the team designedModel-drafted values, with every value checked against the full textUnchecked model output entering the charting table
Appraisal and certainty (Lesson 8)Explaining the items of an appraisal toolModel suggestions as one input, with two human reviewers decidingReporting model judgements as the reviewers' own
Synthesis and reporting (Lessons 9 and 12)Editing the grammar and clarity of the team's textA model-drafted plain-language summary, checked line by line against the findings and disclosedModel-written findings, or text containing unverified references or claims
Environmental scan (Lesson 11)Suggesting survey wording for the team to reviseTranscription or analysis only with tools approved for personal information and with participants' consentEntering interview notes or survey responses into consumer AI tools
Case: The Cedar Valley team's AI plan

The team records its decisions in the protocol's amendment log. A conversational assistant may suggest synonyms, which the librarian tests in each database. It has run the two supplementary AI searches described in Sections 2 and 3, with every source verified. It will screen all 1,870 records in duplicate, as Lesson 7 describes, and will use active learning only in a simulation to plan future updates. No AI tool will make eligibility, charting or synthesis decisions, and no information from the 18 programs in the environmental scan will be entered into any AI tool. The final report will include the disclosure statement from Section 3, and the AI-assisted search log and verification log will appear in an appendix. Lesson 10 returns to these choices when it considers the shortcuts that rapid reviews take.

Putting the lesson into practice

A review team that plans to use AI assistance usually works through four steps. Using a tool its institution permits, it runs one AI-assisted search on its question and records the tool, version, date, settings, exact prompt and full output in an AI-assisted search log. It verifies every source the tool suggests with the four-step procedure from Section 3 and records each check in a verification log, noting which sources exist, which needed corrections and which, if any, are eligible for the review. It marks, in a copy of the decision table, the stages at which it plans to use AI assistance and the safeguard it will apply at each one. And it writes a disclosure statement from the template in Section 3.6, entering no personal or confidential information into any tool.

Reflection

A team replays a completed review in an active-learning tool's simulation mode. The review had 2,000 records after de-duplication, of which 80 were relevant at the title and abstract stage. In the simulation, 95 percent recall (76 of 80 relevant records) was reached after 600 records had been screened. A heuristic stopping rule of 100 consecutive irrelevant records would have stopped screening at record 750, when 77 of the 80 relevant records had been found. The last relevant record appeared at record 1,420. Recall is the proportion of all relevant records found so far. Work saved over sampling at 95 percent recall is WSS@95 = (N − n95) ÷ N − 0.05, where N is the total number of records and n95 is the number screened when 95 percent recall is reached. (a) Calculate the recall and the proportion of records screened at the stopping point, and WSS@95. (b) Explain why these figures would not be available during a live review. (c) Recommend whether a rapid review for a health authority should stop screening at the heuristic's stopping point, and state what the review would need to report.

Model answer

(a) At the stopping point, recall is 77 ÷ 80 = 96.25 percent, and the proportion screened is 750 ÷ 2,000 = 37.5 percent. WSS@95 = (2,000 − 600) ÷ 2,000 − 0.05 = 0.70 − 0.05 = 0.65, so prioritized screening saved about 65 percent of the work that random order would need to reach the same recall.

(b) Recall and WSS depend on knowing the total number of relevant records, which is known in a simulation only because every record was screened. In a live review the total is unknown, so the team cannot tell how many relevant records remain when the heuristic triggers. The last relevant record here appeared at record 1,420, long after the stopping point.

(c) For a rapid review, stopping at the heuristic's point can be defensible if the stopping rule was specified in the protocol before screening began, if the topic resembles those on which the method has been evaluated, and if the team accepts a small risk of missing studies. I would prefer to add a statistical check, such as screening a random sample of the remaining records, to estimate recall. The review would need to report the tool and version, the starting labels, the stopping rule and when it was chosen, the number of records left unscreened, any estimate of recall, and the possibility that relevant studies were missed. If the review must find every eligible study, the team should screen all records and use the tool only to order them.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: What distinguishes prioritized screening from screening truncation?

In prioritized screening the model changes only the order, and every record is still screened, which carries little risk. Truncation stops before every record has been read, which saves more time and carries a risk of missing relevant records.

Question 2: In the Cedar Valley simulation, 95 percent recall was reached after 540 of 1,870 records. What is the work saved over sampling at 95 percent recall?

WSS@95 = (1,870 − 540) ÷ 1,870 − 0.05 = 0.711 − 0.05 = 0.661, or about 66 percent. Option (b) omits the subtraction of 0.05, and option (a) is the proportion of records screened.

Question 3: What does a statistical stopping rule such as the one proposed by Callaghan and Müller-Hansen (2020) do?

The statistical approach screens a random sample of the records that remain and uses the result to test whether a recall target has been reached at a stated confidence level. Option (a) describes a heuristic rule, whose recall in a particular review is unknown.

Question 4: Which conclusion is consistent with Khraisha and colleagues' (2024) evaluation of GPT-4 for screening and extraction?

GPT-4 matched human accuracy on some tasks, but performance fell across all stages once the authors adjusted for chance agreement and dataset imbalance, and they concluded that substantial caution was warranted. The study tested several languages, and it did not recommend routine replacement of human screeners.
Section 5 of 5

Final Assessment

⏱ Estimated time: 25 minutes

Bringing It All Together

This lesson examined how artificial intelligence tools produce answers and what that means for a review team. Section 1 showed that a large language model writes text, including references, by predicting one token after another from patterns learned in training, which explains why fabricated and conflated references occur. Retrieval-augmented generation ties citations to retrieved documents and reduces fabrication, but its claims can still misstate their sources, and its relevance ranking returns a top set with unknown recall. Section 2 sorted AI research tools into conversational assistants, AI search engines and research assistants, citation-context tools and AI-assisted screening tools, and set out questions for evaluating any new tool.

Section 3 turned these ideas into practice. The fictional Cedar Valley team verified eight AI-suggested references with a four-step procedure and found that six existed, three were correct as given and two had been fabricated. The section introduced the AI-assisted search log, the privacy, confidentiality and copyright limits on what may be entered into a tool, and a disclosure statement built on the RAISE recommendations and the 2025 joint position statement. Section 4 explained active-learning screening, the measures used to evaluate it, the stopping problem and the evidence on performance, and ended with a decision table for judging AI use at each stage of a review.

Key Takeaways from this lesson

  • A large language model generates references from learned patterns, so a reference produced without retrieval is a lead to be checked and never evidence that a source exists.
  • A DOI and familiar author names do not show that a reference is real; a DOI counts only when it resolves to the same work at doi.org.
  • Retrieval-augmented tools reduce fabricated references, but their summaries can misstate sources and their relevance ranking returns a top set with low and unknown recall.
  • Conversational assistants, AI search engines, citation-context tools and AI-assisted screening tools differ in where their answers come from, and each supports different review tasks.
  • AI search tools can supplement and check a documented database search, and studies they find enter the review as records identified by other methods.
  • Every AI-suggested source passes four checks (existence, bibliographic details, claim and fit), and each check is recorded in a verification log.
  • An AI-assisted search cannot be repeated exactly, so the tool, version, date, settings, prompt and full output are recorded in a search log.
  • Personal information, confidential documents and licensed full texts stay out of AI tools unless institutional policy and licence terms allow them.
  • Disclosure follows emerging guidance such as the RAISE recommendations, and reports each tool's name, version, dates, purpose, stages, justification, interests and limitations.
  • Active learning safely orders records for screening, while stopping early saves more time at the cost of a recall that is unknown unless it is estimated.

Core Concepts Reviewed

Section 1: tokens, next-token prediction, parameters, knowledge cutoff, context window, temperature, retrieval-augmented generation, embeddings and semantic search, hallucination, and fabricated, conflated and mis-attributed references.

Section 2: conversational assistants, AI search engines and research assistants, citation-context tools, AI-assisted screening and extraction, relevance ranking and recall, and questions for evaluating an unfamiliar tool.

Section 3: reviewer responsibility, the four-step verification procedure, the verification log, the AI-assisted search log, privacy, confidentiality and copyright, the RAISE recommendations, the 2025 joint position statement and the disclosure statement.

Section 4: prioritized screening, the model as second screener, screening truncation, active learning, recall, work saved over sampling, heuristic and statistical stopping rules, evidence on performance and the decision table for AI use.

The final reflection asks you to write an AI-use plan and disclosure statement for a new rapid review by the Cedar Valley evidence team.

Reflection

The fictional Cedar Valley Health Authority asks its three-person evidence team (an evidence officer, a university librarian and a student intern) for a rapid review, due in eight weeks, of community-based falls prevention programs for adults aged 65 and older. The librarian expects the database search to yield about 3,000 records after de-duplication. The team has access to a conversational AI assistant, a scholarly AI research assistant and an active-learning screening tool. The health authority's policy forbids entering personal information into consumer AI tools, and the team also plans to interview staff from five falls prevention programs. Write an AI-use plan of 250 to 400 words that states, for (1) searching, (2) title and abstract screening, (3) data extraction and (4) the program interviews, whether and how AI will be used and what safeguard applies at each stage; describes how AI-suggested sources will be verified, using the four steps of existence, bibliographic details, claim and fit; and ends with a disclosure statement of three or four sentences that names the tools, purposes and stages, the verification method, and the judgements the tools did not make.

Model answer

Searching. The librarian will build and peer-review the database search. The conversational assistant may suggest synonyms, which the librarian will test in each database. The research assistant will be run once as a supplementary search, with the prompt, version, date, settings and full output saved in an AI-assisted search log; new sources will be screened as records identified by other methods.

Verification. The intern will check every AI-suggested source for existence, bibliographic details, claim and fit, record each check in a verification log for the librarian to confirm, and remove any source that cannot be found.

Screening. One reviewer will screen every record in the order set by the active-learning tool, which will act as a second screener. Any early stop will follow a stopping rule set in the protocol, with recall estimated from a random sample of the remaining records.

Extraction and interviews. Two reviewers will extract data by hand, and no interview notes will be entered into any AI tool.

Disclosure. We used [assistant, version] to suggest search terms, [research assistant, version] for one supplementary search and [screening tool, version] to prioritize and second-screen titles and abstracts between [dates]. Prompts, outputs and verification logs are in Appendix C. No AI tool made eligibility, extraction or synthesis decisions, and the stopping rule and estimated recall are reported in the methods. The authors take full responsibility for the review.

Minimum 30 characters required.

✓ Reflection saved

Final Knowledge Assessment

Final Assessment, this lesson: Artificial Intelligence in Evidence Retrieval and Synthesis (15 Questions)

Question 1: The intern asks a conversational assistant with web search turned off for studies published last month. What is the most likely problem?

Without search, a model knows nothing published after its knowledge cutoff, and if pressed for sources it can generate plausible references that do not exist. Recent studies require a search of a current index.

Question 2: Which pairing of a tool category with a review task is most appropriate?

Suggesting synonyms is work with words, which assistants do well, and the librarian's testing in each database supplies the check. The other pairings ask a tool to do something it cannot do reliably or that requires human judgement.

Question 3: An AI-suggested reference has a DOI that returns "DOI not found", and no database or search engine has the title. What belongs in the verification log?

A source that no independent system can find is treated as fabricated. The entry stays in the log so that the removal is documented, and a similar real article can enter the review only as a new, separately verified lead.

Question 4: How does the semantic search used in many retrieval-augmented tools differ from Boolean retrieval in a bibliographic database?

Semantic search compares embeddings of the question and the documents and returns the closest matches, ranked. This can find documents that use different words, but it remains relevance ranking, so it does not return every eligible record.

Question 5: In the Cedar Valley log, entry 6 (Cattan and colleagues) had the right title and authors but the wrong journal, year and volume. What is the correct action?

The article exists in Ageing & Society (2005), so it is a real source with wrong details. The team corrects the reference from the publisher's record and, because the abstract supports the summary, keeps it for citation chasing.

Question 6: Which element belongs in a disclosure statement under the 2025 joint position statement?

The statement asks for each tool's name, version and dates of use, its purpose and the stages it affected, the justification for its use, any related interests and its limitations. No provider can guarantee error-free output, which is why verification is reported.

Question 7: Which material could the Cedar Valley team enter into a consumer AI tool, following the lesson's guidance?

Published material carries the lowest risk, subject to licence terms. Interview notes and survey responses contain information about identifiable people and organizations, and unreleased internal drafts are confidential, so none of them belongs in a consumer tool.

Question 8: Why does the formula for work saved over sampling at 95 percent recall subtract 0.05?

In random order, a reviewer must screen about 95 percent of the records to find 95 percent of the relevant ones, saving about 5 percent. Subtracting 0.05 credits the tool only with the saving beyond what random order would give.

Question 9: A team stops screening after 100 consecutive irrelevant records. Which description of this rule is accurate?

A run of irrelevant records is a heuristic signal that few relevant records remain, but in a live review the total is unknown, so the recall achieved is unknown. In the Cedar Valley simulation such a rule missed 4 of 142 relevant records.

Question 10: According to the lesson's decision table, which use of AI at title and abstract screening is generally acceptable with human checking?

Ordering records while every record is screened leaves every decision with a human and carries little risk. Stopping early is acceptable only under a pre-specified, evaluated rule with recall estimated and reported, and automated exclusion without a human check is not acceptable.

Question 11: Why was entry 7 (Okafor and colleagues, 2023) described as the most tempting entry in the Cedar Valley exercise?

The fabricated entry described a meta-analysis of link worker programs for adults over 65, which is close to the planning team's question. Fabricated sources that fit the reader's question closely are the ones most likely to survive an unverified reading.

Question 12: Which statement about active learning in a tool such as ASReview is accurate?

After a few starting labels, the model scores the unscreened records, shows the highest-ranked one, and retrains when the reviewer decides. The reviewer makes every inclusion decision, and the model changes only the order.

Question 13: What did O'Mara-Eves and colleagues (2015) conclude about text mining for screening?

They concluded that prioritizing the screening order was safe and ready for use, that a model as second screener could be used cautiously, and that automatic elimination was promising but not yet fully proven. Most studies suggested workload savings of 30 to 70 percent.

Question 14: The 2025 joint position statement cites the finding that single-reviewer abstract screening missed 13 percent of relevant studies. How does it use this finding?

The statement uses the finding to illustrate the trade-offs a team weighs: where a rapid review would otherwise use a single human screener, an AI tool acting as a second reviewer may catch studies that the human would miss.

Question 15: A student writes: "The AI tool cited twelve papers on my topic, so my search is done." Which response best applies the lesson?

AI-suggested sources are leads that need independent verification, and an AI search can supplement a documented database search without replacing it. Familiar authors, repeated outputs and the tool's own confirmation are all produced by the same process that can fabricate references.
✦ Complete the final reflection above before submitting

Congratulations!

You have successfully completed this lesson: Artificial Intelligence in Evidence Retrieval and Synthesis.

You can now explain how large language models and retrieval-augmented tools produce answers and why they fabricate or mis-attribute citations, place an AI research tool in its category and evaluate it, verify AI-suggested sources with a documented four-step procedure, record an AI-assisted search, protect personal and confidential information, write a disclosure statement, interpret the performance of active-learning screening, and decide when AI assistance is acceptable at each stage of a review.

Lesson 7, Managing Records and Screening Studies, takes the Cedar Valley team's 2,480 database records into a reference manager, removes the 610 duplicates, and screens the remaining 1,870 titles and abstracts in a screening platform with two independent reviewers. The decisions about AI-assisted screening that you made in this lesson's decision table will apply there.

Continue to Lesson 7 →
Reference

Glossary: Key Terms, People & Frameworks

📚 Reference page, available throughout the lesson

Definitions of the terms, tools, frameworks and people introduced in this lesson.

Core Concepts
Active learning A form of machine learning in which the model chooses which record a human should label next, learns from the label and chooses again; in screening, it shows the record it ranks most likely to be relevant.
Artificial intelligence (AI) A broad label for computer systems that perform tasks usually associated with human judgement, such as classifying text or answering questions.
Conflated reference A reference that merges elements of two or more real works, such as the title of one article with the authors and year of another.
Context window The amount of text a language model can consider at once, including the prompt, any documents or retrieved passages and its own reply.
Disclosure statement A statement in a review that reports which AI tools were used, for what purpose and at which stages, how their output was checked, and their limitations.
Embedding A list of numbers that represents the meaning of a piece of text, used in semantic search to find documents whose meaning is close to a question.
Fabricated reference A complete, plausible reference to a work that does not exist, typically produced by a language model writing from learned patterns.
Hallucination The common name for generated content that is fluent but not supported by the model's sources or by the facts.
Knowledge cutoff The date after which no text was included in a model's training data, so that the model has no information about later events or publications.
Large language model A generative AI system trained on very large collections of text to predict the next token, which allows it to answer questions and write fluent prose.
Machine learning The part of artificial intelligence in which a system learns patterns from examples instead of following rules written by a programmer.
Prioritized screening Screening in which a model sets the order in which records are shown while every record is still screened by a human.
Recall The proportion of all relevant records that a search or screening process has found; also called sensitivity.
Retrieval-augmented generation An approach that combines a retriever, which finds relevant passages in a collection of documents, with a language model that writes an answer citing those passages.
Screening truncation Stopping title and abstract screening before every record has been read, because a model's ranking suggests that few relevant records remain.
Stopping rule A pre-specified criterion for ending screening before every record has been read, either a heuristic signal or a statistical test of recall.
Temperature A setting that controls how much randomness enters a language model's choice of the next token.
Token A short piece of text, such as a word, part of a word or a punctuation mark, that a language model reads and writes.
Verification log A table recording, for each AI-suggested source, the checks performed (existence, details, claim and fit), the result and the action taken.
Work saved over sampling A measure of screening efficiency that compares the records screened to reach a target recall with the number random order would need; WSS@95 uses a 95 percent recall target.
Frameworks & Tools
AI search engine A retrieval-augmented tool that searches a web or scholarly index and returns a written answer, a ranked list of sources or a summary table.
ASReview Open-source software developed at Utrecht University that applies active learning to title and abstract screening and includes a simulation mode for completed reviews.
Citation-context tool A tool that analyses and classifies the sentences in which later papers cite a study, for example as supporting, contrasting or mentioning it.
Conversational assistant A general-purpose large language model with a chat interface, sometimes able to run a web search during a conversation.
Crossref The registration agency that holds metadata for most journal DOIs and that makes the Retraction Watch database openly available.
Digital object identifier (DOI) A persistent identifier for a published work, beginning with 10 and a publisher prefix, that resolves at doi.org to the work it identifies.
Joint position statement on AI in evidence synthesis (2025) A statement by Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence (Flemyng et al., 2025) on the responsible use and reporting of AI in evidence synthesis.
RAISE recommendations Responsible use of AI in evidence SynthEsis: emerging guidance, still being revised in 2025, on the responsible use of AI and automation across the roles in evidence synthesis.
SAFE procedure A practical and conservative set of stopping heuristics for active-learning screening proposed by Boetje and van de Schoot (2024).
Retraction Watch database A database of retracted and corrected publications, now openly available through Crossref, used to check whether a source has been retracted.
Key People
Ashish Vaswani Computer scientist and first author of the 2017 paper that introduced the transformer architecture used in most large language models.
Patrick Lewis First author of the 2020 paper that introduced the term retrieval-augmented generation for systems combining a retriever with a text generator.
Emily M. Bender Computational linguist at the University of Washington and lead author of a 2021 paper arguing that language models produce fluent text without grounding in meaning.
Aaron M. Cohen Biomedical informatics researcher at Oregon Health & Science University whose 2006 study with colleagues introduced work saved over sampling.
Rens van de Schoot Methodologist at Utrecht University who leads the development of ASReview and co-authored the SAFE procedure.
James Thomas Researcher at the EPPI-Centre, University College London, known for work on text mining and automation in systematic reviews and for co-developing the RAISE recommendations.
Alison O'Mara-Eves Researcher at the EPPI-Centre, University College London, and lead author of a 2015 systematic review of text mining for identifying studies.
Max Callaghan Researcher who, with Finn Müller-Hansen, proposed statistical stopping criteria for computer-assisted screening in 2020.
No matching entries. Try a different search term.