HSCI 241 · Lesson 5

Grey Literature, Web Searching and Citation Chasing

Finding & Synthesizing Health Evidence

Learning objectives for this lesson:

  • Define grey literature and explain how publication bias, time-lag bias and practice-based evaluation make it necessary to search beyond bibliographic databases.
  • Identify government and agency reports, program evaluations, theses, conference abstracts, preprints and trial registry records, and name Canadian and British Columbian sources for each.
  • Judge the credibility of a grey-literature document with the AACODS checklist.
  • Plan targeted website searches and write correct Google and Google Scholar queries using quotation marks, OR, the minus sign and the site:, filetype: and intitle: operators.
  • Carry out backward and forward citation searching from a set of seed references, using the terminology and recommendations of the TARCiS statement.
  • Explain how co-citation and bibliographic coupling underlie citation-network tools and judge what those tools can and cannot contribute to a review.
  • Keep a web search log and preserve the documents found so that a grey-literature search is transparent and as reproducible as its sources allow.
  • Report grey-literature, web and citation searches with PRISMA-S and place their results in the PRISMA 2020 flow diagram.
  • Draft a grey-literature and web search plan, a web search log and a pilot citation search for a review question.

This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University, drawing on the Cochrane Handbook for Systematic Reviews of Interventions and the JBI Manual for Evidence Synthesis.

Lesson 5 · HSCI 241

Grey Literature, Web Searching and Citation Chasing

A short guided orientation before you work through the lesson at your own pace.

Finding & Synthesizing Health Evidence
Why this lesson

Databases see only part of the evidence

Program evaluations, theses, conference abstracts and unpublished trials often sit outside journals and outside the major indexes.

A review that searches only journals can give a partial picture, and publication bias can make interventions look more effective than they are.

The road map

Four sections

1 · Kinds of grey literature

This section defines grey literature, explains why reviews need it and introduces Canadian sources and the AACODS checklist.

2 · Websites and operators

This section covers targeted website searching, Google operators and Google Scholar.

3 · Citation chasing

This section teaches backward and forward citation searching and citation-network tools.

4 · Documenting searches

This section shows how to plan, log, preserve and report searches.

Running case

The Cedar Valley evidence review (fictional)

The question

Which community-based interventions have been evaluated for reducing loneliness or social isolation among adults aged 65 and older, and with what outcomes?

In this lesson

The database searches have retrieved 2,480 records. The team now plans and logs its grey-literature and web searches.

What you will learn

Searching beyond the databases

  • A grey-literature search plan sets out sources, methods and stopping rules.
  • Targeted Google searches are recorded in a web search log.
  • Backward and forward citation searches start from seed articles.
  • Lesson 6 adds the log for AI-assisted searches.
Section 1 of 5

Kinds of Grey Literature and Why Reviews Need Them

⏱ Estimated reading time: 35 minutes
Section 1 of 5

Kinds of Grey Literature

What grey literature is, why reviews need it, where to find it in Canada, and how to judge it.

Definition

Grey literature is defined by who publishes it

Material produced by government, academia, business and industry whose publication is not controlled by commercial publishers.

The definition says nothing about quality. A grey document may be rigorous or weak, so each one is judged separately.

Why search beyond journals

Four reasons

Publication bias

Published trials tended to show larger effects than grey trials (Hopewell et al., 2007).

Time-lag bias

Favourable results tend to be published sooner than null results.

Practice-based evidence

Community programs are often evaluated in reports written for funders.

Context

Reports describe who runs and funds programs, which also feeds the scan.

Kinds

The main kinds of grey literature

Government and agency reports
Program evaluation reports
Theses and dissertations
Conference abstracts
Preprints
Trial registry records
Guidelines and briefs
Statistical reports
Trial registries

Finding trials that were never published

Since 2005, journals that follow the International Committee of Medical Journal Editors have required prospective registration.

A completed trial with no publication is a signal that results may be unpublished.

Main sources

The main registries are ClinicalTrials.gov, the ISRCTN registry and the WHO ICTRP search portal.

Canadian sources

Where the Cedar Valley team looks

Federal

The team searches the Public Health Agency of Canada, the National Seniors Council and Statistics Canada.

British Columbia

The team searches the Ministry of Health, the Office of the Seniors Advocate and the regional health authorities.

Community and national

The team searches United Way British Columbia, the Canadian Institute for Social Prescribing and the National Institute on Ageing.

Theses

The team searches Theses Canada, Summit at SFU, cIRcle at UBC and ProQuest Dissertations & Theses Global.

Judging credibility

The AACODS checklist

Authority

Who produced it, with what expertise?

Accuracy

Are methods and data stated?

Coverage

Are its scope and limits clear?

Objectivity

Could the producer's interests shape it?

Date

Is it dated and current?

Significance

Is it relevant to the question?

Carry forward

From what to look for to how to search

  • Grey literature is defined by who controls its publication.
  • Searching it reduces publication bias and finds practice-based evaluations.
  • Canadian sources span federal, provincial, community and academic organizations.
  • Section 2 shows how to search websites and search engines systematically.

Learning Objectives for this section

  • Define grey literature and explain why evidence reviews search for it in addition to bibliographic databases.
  • Describe how publication bias and time-lag bias can distort a review that relies only on journal articles.
  • Identify the main kinds of grey literature, including government and agency reports, program evaluations, theses, conference abstracts, preprints and trial registry records, and state where each is usually found.
  • Name Canadian and British Columbian organizations whose publications are relevant to a public health review.
  • Judge the credibility of a grey-literature document with the AACODS checklist.

1.1 What Grey Literature Is

In Lesson 4 you built a search for bibliographic databases such as MEDLINE, Embase and CINAHL, which index articles published in journals. A large share of the evidence on health programs, however, never appears in a journal. It sits in reports posted on government websites, in theses held by university libraries, in abstracts presented at conferences and in records of trials that were registered but never published. This material is called grey literature (spelled "gray" in American usage), and a review that ignores it may give decision-makers a partial and sometimes misleading picture.

The most widely cited definition was agreed at the International Conference on Grey Literature in Luxembourg in 1997 and extended in New York in 2004. In paraphrase, grey literature is material produced at all levels of government, academia, business and industry, in print and electronic formats, whose publication is not controlled by commercial publishers, meaning that publishing is not the main activity of the organization that produces it. The definition concerns who controls publication. A grey document can be a careful evaluation with a large sample and clear methods, or it can be a promotional summary with no methods at all, so its quality has to be judged separately (Section 1.5).

Some writers contrast grey literature with commercially published, indexed journal articles on one side and with unpublished material, such as an internal evaluation that was never posted, on the other. Unpublished material can be reached only by asking the people who hold it. The figure shows how each zone is reached.

Where the evidence on a review question can sit Indexed journal articles Peer-reviewed papers indexed in MEDLINE, Embase, CINAHL and similar databases Grey literature Government and agency reports, evaluations, theses, abstracts, preprints and trial registry records Unpublished material Internal evaluations, unreported trial results and program data held by organizations Reached by database searches (Lesson 4) Reached by grey-literature databases, websites, search engines and citations Reached only by contacting people
Evidence on a question is spread across indexed journal articles, publicly available grey literature and unpublished material, and each zone requires a different search method.
Case: grey literature for the Cedar Valley review

The Cedar Valley Health Authority is a fictional health authority in British Columbia that serves about 210,000 residents, of whom about 46,000 are aged 65 and older. Before launching a community connector (social prescribing) program, its planning team has asked a small evidence team, made up of an evidence officer, a university librarian and a student intern, for a rapid scoping review and an environmental scan within twelve weeks. The review asks which community-based interventions have been evaluated for reducing loneliness or social isolation among adults aged 65 and older, and with what outcomes.

The database searches built in Lesson 4 have now been run in MEDLINE, Embase, CINAHL, PsycINFO and Web of Science and retrieved 2,480 records. The librarian expects that many evaluations of community programs, such as a seniors' centre's report on its telephone check-in service, will not be among them.

1.2 Why Reviews Search Beyond Journals

Four reasons to search for grey literature apply with particular force to questions about community programs.

Publication bias

Publication bias occurs when the results of a study influence whether, where and how quickly it is published. Studies that find a statistically significant or favourable effect are more likely to be submitted and accepted than studies that find no effect. A review that includes only published studies will then overestimate how well an intervention works. Hopewell and colleagues (2007) reviewed methodological studies for the Cochrane Collaboration and found that trials published in journals tended to show larger intervention effects than trials found only in the grey literature. Searching grey literature is one of the main ways a review team reduces this bias at the search stage. Statistical methods that look for signs of publication bias after studies have been collected, such as funnel plots, are taught in HSCI 230 Lesson 2.

Time-lag bias

Time-lag bias is a related problem: studies with favourable results tend to be published sooner, while studies with null results can take years to appear or never appear. Conference abstracts, preprints and trial registry records show recent or ongoing work that may change a review's conclusions.

Practice-based evidence

Many community programs are evaluated by the organizations that deliver them or by the agencies that fund them. A non-profit that runs a friendly-visiting program may commission an evaluation for its funder and post the report on its website. A health authority may evaluate a pilot and present the results to its board. These evaluations are rarely submitted to journals, yet they often describe exactly the kinds of programs a planning team is considering. For the Cedar Valley question, the team expects much of the evidence on Canadian community programs to be found in this form.

Context for decisions and for the scan

Grey literature also describes who runs programs, how they are funded and how they are delivered. These details matter to decision-makers even when a document contains no outcome data, and they feed the environmental scan, which Lesson 11 teaches in full.

The cost of grey-literature searching

Mahood, Van Eerd and Irvin (2014) found grey-literature searching for a systematic review to be time-consuming and difficult to document, and the documents it produced varied widely in format and quality. A team with twelve weeks cannot search everywhere, so it plans the search in advance, focuses on the sources most likely to hold relevant documents, limits how much of each source it examines and records every step (Sections 2 and 4).

1.3 Kinds of Grey Literature

The cards describe the kinds of grey literature that matter most to health evidence reviews, with a Cedar Valley example of each. Select a card to read more.

Government and agency reportsClick to explore
Program evaluation reportsClick to explore
Theses and dissertationsClick to explore
Conference abstracts and proceedingsClick to explore
PreprintsClick to explore
Trial registry recordsClick to explore
Guidelines, briefs and working papersClick to explore
Statistical and survey reportsClick to explore

Trial registries in more detail

Trial registries deserve particular attention because they can reveal studies that would otherwise be invisible. Since 2005, the International Committee of Medical Journal Editors has required that trials be registered at or before the enrolment of the first participant as a condition of consideration for publication in the journals that follow its recommendations. A registry record therefore exists for many trials whose results were never published, and comparing a registered protocol with the later publication can show whether outcomes were changed after the data were seen (Lesson 8). Registration may be less consistent for trials of social and community interventions than for drug trials, so a registry search for Cedar Valley will be incomplete, but it remains worthwhile.

1.4 Canadian and British Columbian Sources

A grey-literature search is organized around the organizations that produce relevant documents. The table lists sources that a team working on a British Columbia question about older adults would consider, as a starting point to which each team adds organizations discovered while searching.

LevelExample sourcesWhat the Cedar Valley team might find
FederalPublic Health Agency of Canada (publications on Canada.ca); the National Seniors Council; Employment and Social Development Canada, which runs the New Horizons for Seniors Program; Statistics Canada; the Canadian Institute for Health InformationReports on the social isolation of seniors, descriptions of funded projects and national data
Provincial and regional (British Columbia)BC Ministry of Health; the Office of the Seniors Advocate; the BC Centre for Disease Control; the regional health authorities (Fraser Health, Interior Health, Island Health, Northern Health and Vancouver Coastal Health), the Provincial Health Services Authority and the First Nations Health AuthorityProgram descriptions, pilot evaluations and reports on older adults in the province
Community sectorUnited Way British Columbia, which manages the provincially funded Better at Home program and supports community-based seniors' services; local seniors' centres and non-profitsProgram evaluations and descriptions of who runs and funds programs
National non-governmental and academicCanadian Institute for Social Prescribing; the National Institute on Ageing at Toronto Metropolitan University; the National Collaborating Centre for Methods and Tools and its Health Evidence repositoryReports on social prescribing in Canada, policy papers on ageing and quality-rated public health reviews
ThesesTheses Canada (a program of Library and Archives Canada and Canadian universities, begun in 1965); university repositories such as Summit at Simon Fraser University and cIRcle at the University of British Columbia; ProQuest Dissertations & Theses GlobalGraduate evaluations of local programs
Health technology assessment and search toolsCanada's Drug Agency, which CADTH (the Canadian Agency for Drugs and Technologies in Health) became in 2024, and its Grey Matters checklist of grey-literature sourcesA structured list of sources to check and record, described in Section 2
Other provinces and internationalPublic Health Ontario; the Institut national de santé publique du Québec (for French-language documents, if the protocol includes them); the World Health OrganizationPrograms and evaluations from other jurisdictions

Theses Canada does not hold every recent Canadian thesis, because deposit practices vary by university, so searchers also check institutional repositories directly. When the search reaches documents produced by or about First Nations, Métis or Inuit communities, the team should respect the community's ownership of its information. The First Nations principles of OCAP (ownership, control, access and possession) describe how First Nations data and information should be governed, and the First Nations Health Authority is the appropriate starting point for First Nations health information in British Columbia.

Try it: classify what the intern found

During a first afternoon of searching, the intern saved six items: (1) a PDF titled "Seniors Connect pilot: year one evaluation" from a regional health authority's website; (2) a 2024 doctoral dissertation from a Canadian university on loneliness among rural older adults; (3) a ClinicalTrials.gov record for a completed trial of group exercise for isolated older adults, with no results posted; (4) a medRxiv manuscript on a digital befriending service; (5) a two-paragraph abstract from a gerontology conference; and (6) a Statistics Canada report on loneliness in the population. For each item, name its type and say whether it is likely to contribute outcome evidence, context, or a lead to follow up. Then open the accordion below to compare your answers.

Suggested answersv

Item 1 is a program evaluation report that may give outcome evidence. Item 2 is a dissertation, which gives outcome evidence only if it evaluates an intervention. Item 3 is a registry record and a lead: look for posted results or a publication, and contact the investigators. Item 4 is a preprint that may give outcome evidence once the team checks for a peer-reviewed version. Item 5 is a conference abstract, usually a lead to a fuller report. Item 6 is a statistical report that gives context.

1.5 Judging Grey Literature with AACODS

Grey documents have not passed through a journal's peer review, so the team needs a quick, consistent way to judge whether a document is credible enough to include and how much weight to give it. Tyndall (2010), at Flinders University in Australia, developed the AACODS checklist for this purpose. Its name lists six criteria: authority, accuracy, coverage, objectivity, date and significance. The accordion shows the questions each criterion asks, applied to a fictional evaluation report that the Cedar Valley team found on a non-profit's website.

Authorityv

Who produced the document, and with what expertise? The example report names a university-based evaluator and the program director.

Accuracyv

Are the methods, data sources, sample sizes and measures stated clearly enough to judge the findings? The example report describes a pre-post survey with a named loneliness scale and states how many participants completed both surveys.

Coveragev

Are the population, setting, period and limits of the document stated? The example report covers one program in one city over eighteen months and says so.

Objectivityv

Could the producer's interests have shaped the findings, and are negative findings reported? The example report was funded by the program's main donor, which the team records as a possible conflict of interest.

Datev

Is the document dated and current? The team records any publication date and the date it accessed the document. The example report is dated.

Significancev

Is the document relevant to the review question, and does it add to the evidence? The example report evaluates a community-based intervention for adults aged 65 and older and reports a loneliness outcome.

AACODS is a credibility screen, and a document that passes it may still have a weak design; Lesson 8 introduces the risk-of-bias tools. Formal critical appraisal is often optional in a scoping review (Lesson 10), but recording an AACODS judgement helps the team describe the grey evidence in its report.

Summary of Section 1

Grey literature is defined by who controls its publication, and its quality varies. Reviews search for it to reduce publication bias and time-lag bias, to find practice-based evaluations of community programs and to gather context for decisions. Canadian sources span federal agencies, provincial ministries and health authorities, community organizations and national centres, and the AACODS checklist gives a consistent first judgement of each document's credibility.

Reflection

A health authority evidence team in British Columbia is conducting a rapid scoping review of community-based interventions that aim to reduce loneliness or social isolation among adults aged 65 and older. Grey literature is material produced by government, academia, business and industry whose publication is not controlled by commercial publishers. The team has found three items. Item A is a 2019 PDF report on a non-profit's website that evaluates the organization's volunteer visiting program; it names no authors, reports a pre-post loneliness survey completed by 41 participants, and was funded by the program's main donor. Item B is a 2023 master's thesis from a Canadian university that reports a mixed methods evaluation of a community connector pilot in a small British Columbia town. Item C is a 2024 conference abstract reporting preliminary results of a randomized trial of telephone befriending; the trial's ClinicalTrials.gov record lists it as completed, with no results posted and no linked publication. (1) Name the type of grey literature for each item. (2) For each item, give one reason it could strengthen the review and one concern about using it. (3) Apply two criteria from the AACODS checklist to Item A: authority (who produced the document and with what expertise) and objectivity (whether the producer's interests could have shaped the findings). (4) State what the team should do next about Item C, and why.

Model answer

Item A is a program evaluation report, Item B is a thesis, and Item C is a conference abstract linked to a trial registry record.

Item A could strengthen the review because it evaluates exactly the kind of community program the review is about, and such evaluations are rarely published in journals. The concern is that its design is weak, with no comparison group and only 41 participants. Item B is likely to report full methods and complete results, including findings a journal article might omit, but it describes one pilot in one town, so its findings may not transfer. Item C reports a randomized trial, the strongest design among the three, but its results are preliminary and may change.

On authority, Item A names no authors, so the team cannot judge the evaluators' expertise; it should check whether the report names an external evaluator anywhere in the document. On objectivity, funding by the program's main donor creates a possible conflict of interest, which the team should record in its charting table.

For Item C, the team should search the registry and the literature again for posted results or a publication, and contact the investigators to ask whether results are available. A completed trial with no published results may have found no effect, and leaving it out could contribute to publication bias.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: Under the widely cited definition agreed in Luxembourg in 1997 and extended in 2004, what makes a document grey literature?

The definition turns on who controls publication: grey literature is produced by government, academia, business and industry where publishing is not the producer's main activity. Many grey documents lack peer review, but that is a common feature and the definition does not depend on it. The definition explicitly includes electronic formats, and government is only one of several kinds of producer.

Question 2: Hopewell and colleagues (2007) compared trials published in journals with trials found only in the grey literature. What did they find, and what does it imply for a review?

Trials published in journals tended to show larger intervention effects than grey trials, which is a sign of publication bias. A review that searches only journals can therefore overestimate how well an intervention works. The finding does not show that grey trials are more biased, and it does not show that the two groups had similar effects.

Question 3: The Cedar Valley intern finds a ClinicalTrials.gov record for a completed trial of telephone befriending for older adults, with no linked publication. What should the team do next?

A completed trial with no publication may have unpublished results, which is exactly what registry searching is meant to reveal. The team should look for posted results and contact the investigators. Excluding the trial would contribute to publication bias, and assuming null results would replace evidence with a guess. Registry records can be cited, and the trial may never appear in MEDLINE.

Question 4: A non-profit's evaluation report found on its website names no authors, has no date and was funded by the program it evaluates. Which AACODS criteria does this most directly raise concerns about?

Missing authors weaken the judgement of authority, a missing date is a problem under the date criterion, and funding by the program being evaluated is a possible conflict of interest under objectivity. Coverage concerns the stated population, setting and limits, and significance concerns relevance to the question; neither is directly affected by the features described.
Section 2 of 5

Targeted Website Searching and Search-Engine Operators

⏱ Estimated reading time: 40 minutes
Section 2 of 5

Targeted Website Searching and Search-Engine Operators

Four strategies, website searching, Google operators and Google Scholar.

Four strategies

Combining strategies finds more

14of 15 found by targeted websites
9of 15 identified by content experts
1of 15 found by grey-literature databases

The figures come from Godin et al. (2015), a case study of Canadian school breakfast program guidelines.

Targeted websites

Searching an organization's website

  • The team lists organizations that fund, deliver, evaluate or set policy for the programs.
  • The searcher browses publication and report pages and scans every relevant item.
  • The searcher uses the site's own search box with a few simple terms.
  • The searcher runs a Google query restricted to the site with site:.
Google operators

The operators that matter most

"..."

It matches an exact phrase.

OR

It offers alternatives and is written in capitals.

-word

It excludes a word.

site:

It limits results to a domain.

filetype:

It limits results to a file format.

intitle:

It requires a word in the title.

Cedar Valley

A query for provincial reports

loneliness OR "social isolation" seniors evaluation site:gov.bc.ca filetype:pdf

Federal content is split across two domains, so federal queries use site:canada.ca OR site:gc.ca.

Each database concept becomes part of a short, focused query.

Common errors

Four errors to avoid

Space after the colon

site: canada.ca fails; site:canada.ca works.

Lower-case or

Write OR in capitals.

Truncation

isolat* does not truncate in Google.

Pasted database strings

Google limits queries to 32 words.

Google Scholar and stopping rules

Sampling a ranked list on purpose

Google Scholar

Google Scholar is a useful supplement, but its ranking is opaque, it displays at most 1,000 results and it has no bulk export.

Stopping rule

The rule is set in advance. For Cedar Valley it is the first 100 Google results and the first 200 Google Scholar results.

Carry forward

From words to citations

  • Combining strategies finds documents that any single strategy misses.
  • Websites are searched through publication pages, site search and site-restricted queries.
  • Short queries with correct operators and a preset stopping rule make web searching systematic.
  • Section 3 follows citation links from studies the team has already found.

Learning Objectives for this section

  • Combine grey-literature databases, targeted websites, search engines and contacts with experts into one search, following the approach of Godin and colleagues (2015).
  • Choose organizational websites for a review question and search each one systematically through its publication pages, its own search box and a site-restricted search-engine query.
  • Write correct Google queries with quotation marks, OR, the minus sign and the site:, filetype: and intitle: operators, and recognize common syntax errors.
  • Explain what Google and Google Scholar queries cannot do that database queries can, and set a stopping rule for each search.
  • Use the Grey Matters checklist and programmable search engines to make web searching more systematic.

2.1 Four Complementary Strategies

There is no single index of grey literature, so a grey-literature search combines several strategies. Godin and colleagues (2015) set out a widely cited model in a case study that looked for Canadian guidelines on school breakfast programs. They used four strategies: searching grey-literature databases, searching customized Google search engines, searching the websites of targeted organizations, and asking content experts by email. For each Google search, they decided in advance to screen the first ten pages of results, about 100 hits.

Their results show why a combination is needed. The strategies together produced 302 potentially relevant items, of which 15 publications were included. Targeted website searching found 14 of the 15, content experts identified 9, and the grey-literature databases found only 1. Eleven of the 15 publications were found by more than one strategy, but four were found by only one, so dropping any single strategy would have produced a less complete result. The yields will differ for other topics. For a question about community programs, such as the Cedar Valley question, the case study suggests that organizational websites and contacts with the people who run programs are likely to be productive.

Recall from Lesson 3

Lesson 3 explained that a bibliographic database returns every record that matches a Boolean query, while a web search engine ranks pages by estimated relevance and shows only the top of that ranking. Searching the web for grey literature therefore means examining a sample of ranked results, and the person searching has to decide how far down the ranking to go. Lesson 3 also explained that results vary with location, account history and time. Section 4 of this lesson shows how to record searches so that these limits are visible to readers.

2.2 Targeted Website Searching

Targeted website searching means identifying, in advance, the organizations most likely to publish relevant documents and searching each of their websites in a planned way. Stansfield, Dickson and Bangpan (2016) describe the problems this raises for systematic reviews: site search functions vary in quality, the volume of material can be large, and the steps taken are easy to forget. Their advice, which this lesson follows, is to plan which sites to search, how to search each one and how far to go, and to keep a record of each decision.

Choosing the websites

The list of websites follows from the review question. For Cedar Valley, the question concerns community-based interventions for loneliness and social isolation among adults aged 65 and older, so the team asks which organizations fund, deliver, evaluate or set policy for such programs. The librarian proposes candidates from the Canadian sources listed in Section 1.4, the intern adds organizations named in the documents found so far, and the evidence officer adds organizations suggested by members of the planning team. The final list is recorded in the search plan with a short reason for each site.

Searching a website

Each website is searched with up to three methods, and the log records which were used.

  • The searcher browses the pages where documents are listed, such as "Publications", "Reports", "Resources" or "Research and evaluation", and scans every item listed under relevant headings.
  • The searcher uses the site's own search box with a few simple terms, since most site search boxes do not support Boolean operators, phrases or truncation reliably.
  • The searcher runs a Google query restricted to the site with the site: operator, which often finds PDF reports that the site's own search misses.

Health authority and government websites are reorganized often, so documents may move or disappear. Section 4 explains how to capture what is found.

The Grey Matters checklist

Grey Matters: A Practical Tool for Searching Health-Related Grey Literature was developed by CADTH (the Canadian Agency for Drugs and Technologies in Health) and is now published by Canada's Drug Agency. It lists sources of health-related grey literature by type and by jurisdiction, including health technology assessment agencies, clinical practice guideline sources, drug and device information, trial registries and health statistics, and it is laid out as a checklist on which a searcher can record which sources were checked. Its emphasis on drugs and health technologies means that it covers only part of what a community-program review needs, but it is a useful way to make sure that no major agency is overlooked, and the completed checklist becomes part of the search record.

2.3 Google Advanced Operators

A search operator is a word or symbol that tells a search engine how to treat part of a query. Google supports a small set of operators that make web searching more precise. The table lists those most useful for grey literature. Operators are typed with no space between the operator, its colon and the term, so site:gov.bc.ca works and site: gov.bc.ca does not.

OperatorWhat it doesCedar Valley exampleNotes
"..."Matches an exact phrase"community connector"Use for multi-word concepts. The Verbatim setting in Google's Tools menu also turns off automatic synonyms and spelling corrections.
ORMatches either termloneliness OR "social isolation"Must be written in capitals. In practice Google applies it to the terms immediately on either side.
-Excludes a term or an operator-jobs -careersPlace the minus sign directly before the word, with no space.
site:Limits results to a domain, a subdomain, a top-level domain or a URL prefixsite:gov.bc.ca, site:ca, site:canada.ca/en/public-healthA domain includes its subdomains. -site: excludes a site.
filetype:Limits results to one file formatfiletype:pdfReports are often posted as PDF files. ext: works the same way. Other formats include docx, xlsx and pptx.
intitle:Requires the next word to appear in the page titleintitle:lonelinessallintitle: requires all the following words to appear in the title.
inurl:Requires the next word to appear in the web addressinurl:evaluationUseful when an organization files reports under a recognizable folder name.
before: and after:Limit results to pages last updated before or after a dateafter:2018Google's dating of pages is approximate, so these are a rough filter.
*Stands for any whole word inside a phrase"social * program"The asterisk is a word placeholder and does not truncate.

Several features that searchers learn in databases do not carry over to Google. Google has no truncation, so isolat* does not retrieve "isolated" and "isolation"; instead Google matches some word variants automatically. Google combines terms with an implied AND, ignores most punctuation and symbols, and limits a query to 32 words. It has also retired operators over time, including the plus sign in 2011 and the tilde for synonyms in 2013, and in 2024 it removed its links to cached copies of pages. Because of these differences, a database search string from Lesson 4 cannot be pasted into Google. It has to be rewritten as several short queries, each combining one or two key concepts with operators.

Anatomy of an advanced Google query "community connector" OR "social prescribing" seniors Exact phrase Either one Exact phrase Required word site:gov.bc.ca filetype:pdf -jobs BC government domain only PDF files only Exclude job postings Google joins the parts with an implied AND. The query is typed on one line, with no space after any colon and no space after the minus sign.
Each part of an advanced query has one job: phrases fix wording, OR offers alternatives, operators restrict the domain and format, and the minus sign removes noise.

Building queries for Cedar Valley

The intern drafts a set of short queries, each aimed at one group of sources. The tabs show four of them, what each is meant to find and how the team refined it after looking at the first page of results.

Query: "social prescribing" OR "community connector" loneliness seniors site:canada.ca

This query looks for federal documents on social prescribing or community connectors that mention loneliness and seniors. Government of Canada content is split between the canada.ca domain and older gc.ca domains, such as the one used by Statistics Canada, so the team runs a second version with site:gc.ca. Combining them in one query with site:canada.ca OR site:gc.ca also works.

Query: loneliness OR "social isolation" seniors evaluation site:gov.bc.ca filetype:pdf

This query looks for PDF reports on the British Columbia government domain. A second query covers the regional health authorities by joining their domains with OR: loneliness seniors program site:fraserhealth.ca OR site:interiorhealth.ca OR site:islandhealth.ca OR site:northernhealth.ca OR site:vch.ca OR site:fnha.ca. When the first page showed mostly event listings, the team added evaluation to the first query and kept the second unchanged so that program descriptions would still appear.

Query: intitle:loneliness seniors program evaluation site:ca filetype:pdf -jobs

Restricting to the .ca top-level domain, to PDF files and to pages with "loneliness" in the title brings evaluation reports from non-profits, municipalities and universities across Canada to the top. The minus sign removes job postings for program coordinators, which crowded the first results page in a trial run.

Query: "social isolation" "older people" intervention evaluation site:who.int

This query looks for World Health Organization documents on interventions. The phrase "older people" is used because the World Health Organization uses it more often than "seniors". Matching the vocabulary of each organization is as important on the web as matching controlled vocabulary in a database.

Error 1: a space after the colonv

site: canada.ca is read as the word "site" followed by the word "canada.ca". The operator works only when written as site:canada.ca.

Error 2: a lower-case orv

loneliness or isolation treats "or" as an ordinary word. Write loneliness OR isolation.

Error 3: truncation symbolsv

isolat* does not truncate in Google. Write the forms you need, such as isolated OR isolation, or rely on Google's automatic matching of variants.

Error 4: pasting a database stringv

A fifty-term Boolean string with field tags and nested parentheses exceeds Google's 32-word limit and uses syntax that Google ignores. Split the concepts into several short queries.

Error 5: trusting the results countv

The number of results Google reports is an estimate that changes from day to day and is not a count of documents that can be screened. Record it if you wish, but base the stopping rule on the number of results actually examined.

2.4 Google Scholar

Google Scholar indexes scholarly material found on the web, including journal articles, theses, preprints, books and reports posted on academic and organizational sites. It supports quotation marks, OR, the minus sign, intitle: and allintitle:, author: and source:, and an advanced search form that restricts terms to the title, an author, a publication or a range of years. It does not support truncation or nested Boolean logic reliably, it ranks results by an undisclosed method, it displays at most 1,000 results for any query, and it cannot export a full result set in one step.

Two studies by Haddaway and colleagues help define its role. Haddaway, Collins, Coughlin and Kirk (2015) tested Google Scholar against case studies of environmental science systematic reviews. They found that its results contained moderate amounts of grey literature, that much of it appeared well beyond the first pages of results, and that Google Scholar missed important studies in five of six case studies. They recommended that searches of article titles for grey literature focus on the first 200 to 300 results and concluded that Google Scholar should not be used alone. Gusenbauer and Haddaway (2020) compared 28 academic search systems and concluded that Google Scholar is inadequate as the principal search system for a systematic review, chiefly because its queries cannot be controlled or reproduced precisely. The Cedar Valley team therefore uses Google Scholar as a supplementary source with a fixed stopping rule of 200 results, and it uses Google Scholar's "Cited by" links for forward citation chasing, which Section 3 describes.

2.5 Making Web Searching More Systematic

Two further tools reduce the haphazard quality of web searching. Google's Programmable Search Engine (formerly called Custom Search Engine) lets a team build a search box that searches only a chosen list of websites. Godin and colleagues searched two customized engines of this kind, one for Canadian public health information and one for Canadian government documents. The Cedar Valley team could build one engine covering the provincial government, the five regional health authorities, the First Nations Health Authority, the Office of the Seniors Advocate and United Way British Columbia, and then run the same queries across all of them at once. The team must still record the list of sites included in the engine.

The second tool is a stopping rule, a decision made before searching about how many ranked results will be examined for each query. A fixed number, such as the first 100 Google results or the first 200 Google Scholar results, is the most common rule, and it can be combined with an early stop when, for example, two consecutive pages contain nothing relevant. Setting the rule in advance keeps effort consistent across queries and lets the team report exactly how much of each ranking was examined.

If the protocol includes French-language documents, each query is also written in French, with terms such as "isolement social", solitude, aînés and "prescription sociale", and French-language sources such as the Institut national de santé publique du Québec are added to the website list.

Try it: write three queries

Write a Google query for each task. (1) Find PDF documents on the Office of the Seniors Advocate's website, seniorsadvocatebc.ca, that mention social isolation. (2) Find pages on any .ca website that have "befriending" in the title and mention older adults. (3) Find documents on Government of Canada sites about the New Horizons for Seniors Program and social isolation, excluding pages about how to apply for funding. Then compare your answers with the suggested queries below.

Suggested queriesv

(1) "social isolation" site:seniorsadvocatebc.ca filetype:pdf. (2) intitle:befriending "older adults" OR seniors site:ca. (3) "New Horizons for Seniors" "social isolation" site:canada.ca OR site:gc.ca -apply. Other correct versions are possible. In each case, check that there is no space after any colon, that OR is in capitals and that the minus sign is attached to the word it excludes.

Summary of Section 2

A grey-literature search combines grey-literature databases, targeted websites, search engines and contacts, because each finds documents the others miss. Targeted website searching starts from a reasoned list of organizations and uses publication pages, site search and site-restricted Google queries. Google operators such as quotation marks, OR, the minus sign, site:, filetype: and intitle: make queries more precise, but Google has no truncation and ranks results opaquely, so database strings must be rewritten as short queries. Google Scholar is a useful supplement with a fixed stopping rule. Programmable search engines and stopping rules make the work more systematic and easier to report.

Reflection

Google supports these operators: quotation marks for an exact phrase; OR, written in capitals, for alternatives; a minus sign attached directly to a word to exclude it; site: to restrict results to a domain or a top-level domain such as .ca; filetype: to restrict results to a file format such as PDF; and intitle: to require a word in the page title. Google has no truncation, and operators must be typed with no space after the colon. (1) Write a Google query that finds PDF documents on the British Columbia government domain, gov.bc.ca, that contain the phrase social isolation and the word evaluation. (2) Write a query that finds pages on any .ca website with loneliness in the page title that mention either seniors or older adults, excluding job postings. (3) A colleague wrote site: canada.ca "social isolation" seniors or elderly isolat* +program. Identify every error and write a corrected version. (4) A stopping rule fixes in advance how many ranked results will be screened. State the stopping rule you would apply to these queries and explain why it should be set before searching.

Model answer

(1) "social isolation" evaluation site:gov.bc.ca filetype:pdf

(2) intitle:loneliness seniors OR "older adults" site:ca -jobs. Adding -careers would remove more postings.

(3) The query has four errors. The space after site: makes Google read "site" and "canada.ca" as ordinary words. The lower-case or is treated as a word, so the alternatives are not combined. The asterisk in isolat* does not truncate in Google, since the asterisk only stands for a whole word inside a phrase. The plus sign has been retired, so +program does not force the word to appear. A corrected version is "social isolation" seniors OR elderly program site:canada.ca; if the team wants pages about isolated people as well, it can run a second query with isolated in place of the phrase.

(4) I would screen the first 100 results of each query, or all results when fewer are shown, and stop early if two consecutive pages contain nothing relevant. Setting the rule before searching keeps the effort consistent across queries, prevents the searcher from stopping when the results happen to look favourable or tiring, and allows the report to state exactly how much of each ranking was examined.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: Which query restricts Google results to PDF files on the British Columbia government domain that contain the exact phrase social isolation?

The phrase is in quotation marks, and both operators are typed with no space after the colon. Option a has spaces after the colons, so Google reads the operators as ordinary words. Option b omits the quotation marks and uses type:, which is not the file-type operator. Option d uses database-style syntax that Google does not recognize.

Question 2: The intern types isolat* seniors site:canada.ca, expecting to retrieve pages about isolated and isolation. Why will this not work as intended?

Google does not support truncation. The asterisk is a placeholder for a whole word inside a quoted phrase. The intern should write the forms needed, such as isolated OR isolation, or rely on Google's automatic matching of some word variants. The other options describe rules that do not exist.

Question 3: In the case study by Godin and colleagues (2015) on Canadian school breakfast program guidelines, which strategy identified the most included publications?

Targeted website searching found 14 of the 15 included publications. Content experts identified 9, and the grey-literature databases found only 1. Four publications were found by only one strategy, which is why the authors recommended combining strategies.

Question 4: Why do Haddaway and colleagues recommend Google Scholar only as a supplementary source for systematic review searches?

Haddaway and colleagues (2015) found that Google Scholar missed important studies in five of six case studies, and Gusenbauer and Haddaway (2020) concluded that its queries cannot be controlled or reproduced precisely enough for it to be the principal search system. Google Scholar indexes journal articles as well as grey literature, accepts quotation marks and OR, and is free.
Section 3 of 5

Backward and Forward Citation Chasing

⏱ Estimated reading time: 35 minutes
Section 3 of 5

Backward and Forward Citation Chasing

Finding studies through their links to studies you already have.

Terminology

The TARCiS terms

Backward

It retrieves and screens the references that seed references cite.

Forward

It retrieves and screens the references that cite the seeds.

Seed references

They are normally all records included after full-text screening.

Iterative

It repeats the search with newly found eligible records as seeds.

When it helps

Hard-to-search topics benefit most

The same kind of program may be called social prescribing, community referral, befriending, friendly visiting or a telephone check-in service.

Greenhalgh and Peacock (2005) found about half the sources in a review of complex evidence through snowballing. TARCiS treats citation searching as a supplement to database searching.

Backward searching

Looking back from the seeds

  • The team assembles the seed set and records it.
  • Cited references are retrieved from a citation index, or from reference lists by hand.
  • Duplicates and records already screened are removed.
  • Remaining records are screened like any other search result.
Forward searching

Looking forward from the seeds

Where

Web of Science, Scopus, Lens.org and Google Scholar record citing documents, with different coverage.

Watch for

Citation counts grow, so record the date. Recent seeds have had little time to be cited.

Citation-network tools

Complete lists and selected maps

citationchaser

It retrieves complete cited and citing lists from Lens.org and exports them for screening.

Map-building tools

Connected Papers, Litmaps and similar tools show selected similar papers, which is useful for exploration.

Co-citation (Small, 1973) and bibliographic coupling (Kessler, 1963) underlie the maps.

Cedar Valley

Citation chasing in the completed review

2,448records retrieved
1,160new records screened
23full texts assessed
4studies added (38 to 42)
Carry forward

From searching to documenting

  • All included records serve as seeds for backward and forward searching.
  • Results are de-duplicated against everything already screened.
  • The decision to iterate is made deliberately and reported.
  • Section 4 shows how to plan, log and report every search.

Learning Objectives for this section

  • Define citation chasing and distinguish backward from forward citation searching using the terminology of the TARCiS statement.
  • Explain when citation chasing adds most to a search and summarize the evidence on its yield.
  • Carry out backward and forward citation searching from a set of seed references, with de-duplication against records already screened.
  • Describe how co-citation and bibliographic coupling underlie citation-network tools, and judge what those tools can and cannot contribute to a review.
  • Record citation searching so that another team could repeat it.

3.1 What Citation Chasing Is

Every research article cites earlier work and is later cited by newer work. Citation chasing uses these links to find studies that keyword searches miss. If a study included in the Cedar Valley review cites an evaluation of a befriending program, that evaluation is likely to be relevant too, whatever words its title and abstract happen to use. Citation chasing therefore finds studies by their relationships to known relevant studies, which makes it a useful complement to searches that depend on matching words.

The method has gone by many names, including snowballing, pearl growing, reference list checking, citation tracking and citation searching. To bring order to these terms, Hirt and colleagues (2024) ran a Delphi consensus study with international methods experts and published the TARCiS statement (Terminology, Application and Reporting of Citation Searching), which contains ten recommendations. This lesson uses "citation chasing" as the everyday name and the TARCiS terms for the specific methods. The cards define them.

Citation searchingClick to explore
Seed referencesClick to explore
Backward citation searchingClick to explore
Forward citation searchingClick to explore
Direct and indirect searchingClick to explore
Iterative citation searchingClick to explore

3.2 When Citation Chasing Helps

Citation chasing adds most when a topic is hard to search with words, as loneliness interventions are. The same kind of program may be called social prescribing, community referral, a link worker service, a community navigator program, befriending, friendly visiting or a telephone check-in service, and many evaluations describe their outcomes as well-being or social participation, with loneliness mentioned only in the full text. A search string cannot anticipate every label, but studies on the same topic tend to cite one another.

Greenhalgh and Peacock (2005) audited the sources of a systematic review of complex evidence on the diffusion of innovations in health service organizations. They found that fewer than a third of the primary sources had come from the database and hand searches set out in the protocol, while about half had been found by "snowballing", that is, by following references of references. That review covered a scattered literature from many disciplines, so its proportions are unusually high, but it showed how much a word-based search can miss. A Cochrane methodology review by Horsley, Dingwall and Sampson (2011) found limited but supportive evidence that checking reference lists identifies additional studies for systematic reviews. A scoping review by Hirt and colleagues (2023) mapped how citation tracking is used and reported, and it informed the TARCiS statement.

TARCiS turns this evidence into practical advice. For topics that are difficult to search, the team should seriously consider backward and forward citation searching as supplementary methods. For topics with a clear vocabulary and a highly sensitive search, they are not explicitly recommended, although checking the reference lists of included records can still test whether the database search was sensitive enough. Citation searching should not replace extensive database searching in a review that aims to find all relevant studies, because it can only find documents that are linked to the seeds.

Two limits to keep in mind

Citation chasing can reinforce the biases of the seed set: if the included studies come mostly from one country or research group, their citations tend to lead back to the same networks. Forward chasing also favours older seeds, because recent studies have had little time to be cited. Both limits are reasons to use citation chasing alongside the other searches.

3.3 Backward Citation Searching

Backward searching starts from the seed references and retrieves everything they cite. TARCiS recommends, where possible, retrieving the cited references from a citation index so that titles and abstracts can be screened, since a reference list read by hand usually gives only titles. The steps are as follows.

  1. Assemble the seed set, normally all studies included after full-text screening, and record the list.
  2. Retrieve the cited references from one or more citation indexes, such as Web of Science or Lens.org, and check by hand the reference lists of any seeds that the indexes do not cover.
  3. Export the retrieved records to the reference manager, remove duplicates within the set, and remove records that were already screened from the database and grey-literature searches.
  4. Screen the remaining records against the eligibility criteria, using the same screening process as for the main search (Lesson 7).

Backward searching is also one of the best ways to find grey literature. Authors of included studies often cite program reports, government documents and theses, and these citations point to documents that no index covers. The Cedar Valley team therefore reads the reference lists of the included grey-literature documents by hand as well, and counts any new documents found this way with its grey-literature results.

3.4 Forward Citation Searching

Forward searching retrieves the documents that cite each seed. It is the only systematic way to move forward in time from a known study, and it often finds recent evaluations and follow-up studies. Web of Science shows a "Times Cited" count and list, Scopus and Google Scholar show "Cited by" links, and Lens.org records citing works. These indexes cover different sets of documents, so the same seed can have quite different citing lists in each. TARCiS suggests considering two citation indexes to widen coverage, especially when some seeds are missing from one index. Google Scholar's citing lists include more theses and reports than the subscription indexes, but they cannot be exported in bulk, so they are most practical for a small number of seeds.

Citation counts grow over time, so the date on which forward searching was done must be recorded. Some indexes also let a user set an alert that sends an email whenever a seed is newly cited, which is useful for keeping a review current, as Lesson 12 discusses.

Seed Co-cited Co-citing Backward: what the seed cites Forward: what cites the seed Older Newer
Arrows point from a citing document to the document it cites. Backward searching follows the seed's own arrows to older work, forward searching finds the newer documents whose arrows point to the seed, and the dashed arrows show the indirect relationships of co-citation and bibliographic coupling.

3.5 Citation-Network Tools

Two older ideas from information science underlie most citation-network tools. Co-citation, described by Small (1973), links two documents that are cited together by a later document: the more often they are cited together, the more closely related they are assumed to be. Bibliographic coupling, described by Kessler (1963), links two documents that cite the same earlier work: the more references they share, the more closely related they are assumed to be. Both build on the citation index, which Eugene Garfield proposed in 1955 and which became the Science Citation Index, the ancestor of Web of Science. In TARCiS terms, retrieving co-cited documents and co-citing documents are the indirect methods of citation searching.

Citation-network tools retrieve these relationships automatically and display them as lists or maps. The tabs describe examples available in 2026. Products change quickly, so the point is to understand what kind of relationship each tool uses and what data it draws on.

citationchaser is a free, open-source R package and web application developed by Haddaway, Grainger and Gray (2022) for transparent forward and backward citation chasing. The user enters the identifiers (for example DOIs) of the seed references, and the tool retrieves their cited and citing references from Lens.org and exports them in a standard format for de-duplication and screening, which fits the TARCiS recommendations well.

Connected Papers builds a graph of documents similar to one seed paper, using co-citation and bibliographic coupling to judge similarity, so that related papers appear near each other even when they do not cite one another. It is useful for exploring a topic and finding terms to add to a search, but it shows a selection chosen by its own algorithm and cannot serve as a complete backward or forward search.

Tools such as Litmaps, ResearchRabbit and Inciteful build maps from one or more seed papers, showing citation links over time and suggesting further papers, and some can report new papers that cite a map. Their suggestions depend on data sources and ranking methods that may change between versions.

Web of Science and Scopus are subscription citation indexes with curated journal coverage. Lens.org is a free platform that aggregates records from several open sources, and OpenAlex is an open index launched in 2022 by the non-profit OurResearch, on which many newer tools draw. Google Scholar has the broadest coverage of grey material, but its citing lists cannot be exported in bulk. Mapping software such as VOSviewer (van Eck and Waltman, 2010) can draw co-citation and coupling maps from exported records.

Whichever tool is used, the team records its name, the date of use, the data source it drew on, the seeds entered and any settings, and it exports the full list of retrieved records so that they can be de-duplicated and screened like any other search result. Tools that show only a selection of related papers are best used during scoping and search development. Lesson 6 extends this discussion to artificial intelligence tools that summarize the context of citations and suggest related papers, and it explains how to verify what they report.

3.6 Worked Example: Citation Chasing for Cedar Valley

A check of the database search

Before screening began, the librarian used citation relationships to check the sensitivity of the database search. The team had found three earlier reviews of interventions for loneliness and social isolation in older people during scoping: Dickens and colleagues (2011), Gardiner, Geldenhuys and Gott (2018), and Fakoya, McCorry and Donnelly (2020). The intern listed the studies cited in these reviews that appeared to meet the Cedar Valley criteria and checked whether each one was among the 2,480 database records. Suppose that two such studies were missing and that both described their programs as "friendly visiting". The librarian would add that phrase and related terms to the search strategy and rerun it, as Lesson 4 described, and the team would report the check and the change.

Direct citation searching after full-text screening

Once screening is complete (Lesson 7), the team will use the 38 studies included from the database searches as seed references and run backward and forward searching on all of them, retrieving records through citationchaser (Lens.org) and Web of Science and checking by hand the reference lists of two seeds that neither index covered fully. The table shows how the records are reduced to a set for screening. These numbers come from the completed review, which later lessons describe.

StepBackward searchingForward searching
Seed references38 included studies38 included studies
Records retrieved from both indexes1,296 cited references1,152 citing records
After removing duplicates within each direction884903
Removed because already screened from the database and grey-literature searches263318
New records in each direction621585
Records found in both directions: 46. New records screened by title and abstract: 621 + 585 − 46 = 1,160. Full texts assessed: 23. Studies included: 4.

The 4 studies found by citation chasing bring the review from 38 to 42 included studies. TARCiS suggests considering a further iteration with the newly included studies as seeds. The Cedar Valley team decided against a second iteration because of its twelve-week deadline and because the first iteration had added few studies relative to the number screened, and it reported this decision and its reason. A team with more time, or one whose first iteration had found many new studies, might reasonably have continued.

Reporting citation searching

TARCiS recommends reporting the seed references (and a justification if they differ from the included records), the direction of searching, the dates, the number of iterations and the reason for stopping, every index and tool used, how de-duplication was done, how records were screened, and the numbers in the PRISMA 2020 flow diagram. The PRISMA-S item on citation searching asks for similar information. A methods paragraph for Cedar Valley might read: "We used the 38 studies included from the database searches as seed references for one iteration of backward and forward citation searching on [date], retrieving records from Lens.org with citationchaser and from Web of Science, and checking two reference lists by hand. The included grey-literature documents are not covered by citation indexes, so we checked their reference lists by hand and counted new documents with the grey-literature results. Records were de-duplicated against each other and against all records already screened, and the remaining 1,160 were screened by two reviewers. We did not run a second iteration because of the review's timeline."

Try it: plan a citation search

A team has 12 included studies on peer-support programs for family caregivers. Three of the studies were published in the past year, and two are evaluation reports that no citation index covers. Write four sentences that state which documents will serve as seeds, which methods and indexes will be used for each group of seeds, how the results will be de-duplicated, and what the team will do if the search finds new eligible studies.

Suggested answerv

All 12 included documents will serve as seed references. Backward and forward searching for the ten indexed studies will be run through citationchaser (Lens.org) and one subscription index, while the reference lists of the two evaluation reports will be checked by hand, and Google Scholar's "Cited by" lists will be checked for them and for the three recent studies, whose citations the other indexes may not yet capture. All retrieved records will be de-duplicated against each other and against every record already screened before two reviewers screen them. If new eligible studies are found, the team will consider a second iteration using them as seeds and will report the number of iterations and why it stopped.

Summary of Section 3

Citation chasing finds studies through their citation links to known relevant studies: backward searching retrieves what the seeds cite, forward searching retrieves what cites them, and indirect methods use co-citation and bibliographic coupling. TARCiS supports it as a supplement for topics that are hard to search with words. Seeds are normally all included studies, results are de-duplicated against everything already screened, and every index, tool, date and iteration is reported.

Reflection

The TARCiS statement defines backward citation searching as retrieving and screening the references that seed references cite, and forward citation searching as retrieving and screening the references that cite the seeds. Seed references are normally all records included after full-text screening, and iterative citation searching repeats the process with newly found eligible records as seeds. A review team has included 15 documents on peer-support programs for older adults who live alone. Two of the included documents are evaluation reports posted on organizational websites, which citation indexes do not cover, and four of the included studies were published in the past twelve months. The team has access to Web of Science, to Lens.org through the free tool citationchaser, and to Google Scholar, whose citing lists cannot be exported in bulk. Write a citation-searching plan of 150 to 250 words that states (1) the seed references, (2) how backward and forward searching will be done for each group of seeds and with which indexes or tools, (3) how duplicates will be handled, (4) how the team will decide whether to run a second iteration, and (5) what will be reported. Then (6) explain in two sentences why a tool that draws a map of papers similar to one seed could not replace this plan.

Model answer

All 15 included documents will serve as seed references. For the 13 indexed studies, the team will run backward and forward searching in citationchaser, which retrieves cited and citing records from Lens.org, and in Web of Science, because the two indexes cover different documents. For the two evaluation reports, a reviewer will check the reference lists by hand and look up each report in Google Scholar's "Cited by" list. The four recent studies will also be checked in Google Scholar, since the other indexes may not yet record citations to them.

All retrieved records will be imported into the reference manager and de-duplicated against each other and against every record already screened from the database and grey-literature searches. Two reviewers will screen the remaining records with the review's screening form.

If the first iteration finds new eligible studies, the team will consider a second iteration with those studies as seeds, weighing the number of new studies found against the time needed, and it will record the decision and its reason.

The report will state the seeds, the directions searched, the dates, every index and tool used, the de-duplication method, the screening method, the number of iterations and the reason for stopping, and the counts in the PRISMA 2020 flow diagram.

A mapping tool shows a selection of related papers chosen by its own algorithm, so it does not return the complete lists of cited and citing references. Its results also depend on data and ranking methods that may change, which makes the search hard to report and repeat.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: In the terminology of the TARCiS statement, what is forward citation searching?

Forward citation searching moves forward in time to the newer documents that cite a seed. Option b describes backward citation searching, option c describes co-citing searching (bibliographic coupling), and option d describes iterative citation searching.

Question 2: Two evaluations published in 2021 both cite the same five earlier studies, although neither cites the other. Which relationship links the two evaluations?

Bibliographic coupling, described by Kessler (1963), links documents that share references. Co-citation, described by Small (1973), would apply if a later document cited both evaluations together. Forward and backward citation describe direct links between a document and the documents that cite it or that it cites, and the two evaluations have no direct link.

Question 3: According to TARCiS, which documents should normally serve as the seed references for citation searching in a systematic or scoping review?

TARCiS recommends using all records included after full-text screening, with any deviation justified in the report. Records retained after title and abstract screening include many that will later be excluded, and choosing seeds by citation count or by a tool's ranking would introduce selection that is hard to justify.

Question 4: After backward and forward searching on its 38 included studies, the Cedar Valley team has retrieved 2,448 records. What should it do before screening them?

TARCiS recommends de-duplicating the results of backward and forward searching against each other and, where possible, against the records already screened from the primary search. In the Cedar Valley example this reduced 2,448 retrieved records to 1,160 new records. The other options would discard eligible studies on grounds unrelated to the eligibility criteria.
Section 4 of 5

Documenting Web and Grey-Literature Searches

⏱ Estimated reading time: 40 minutes
Section 4 of 5

Documenting Web and Grey-Literature Searches

Search plans, web search logs, preservation and reporting.

The problem

Web searches are usually reported poorly

300systematic reviews examined
108reported web searching
6reached the highest reporting standard

The figures come from Briscoe (2015), a study of UK Health Technology Assessment programme reviews.

Two standards

Transparent and reproducible

Transparent

A reader can see exactly what was done, from sources and queries to stopping rules and results.

Reproducible

Repeating the steps gives the same results, which web searches rarely achieve.

The goal for web searching is full transparency and as much reproducibility as the sources allow.

The search plan

Decide before you search

  • The plan lists sources with a reason for each one.
  • The plan states the method and queries for each source.
  • The plan sets a stopping rule for every ranked source.
  • The plan names who searches and checks each source, and when.
The web search log

One row per search, written at the time

W04 · 2026-02-18 · Google (google.ca, private window, signed out) · intitle:loneliness seniors program evaluation site:ca filetype:pdf -jobs · first 100 screened · 11 saved (W04-01 to W04-11)

Searches that find nothing are logged too, so readers know the source was covered.

Preserving

Protecting the evidence from link rot

Identify

Name each item with its search ID, such as W04-03.

Record

Save the address and access date in the reference manager.

Archive

Save an archived copy of each web page.

Capture

Screenshot the first results page of each query.

Reporting

PRISMA-S and the flow diagram

Cedar Valley counts

Grey-literature sources gave 771 records and items, 557 remained after de-duplication, and 26 documents were included in the completed review.

PRISMA 2020 flow diagram

Registers sit with databases. Websites, organizations and citation searching sit under other methods.

Next steps

Reflection, final assessment and Lesson 6

Reflection and final page

Rewrite a vague methods sentence, then plan searches for a new question and complete the fifteen-question knowledge check.

Next: Lesson 6

Lesson 6 adds AI-assisted searching and its search and verification log.

Learning Objectives for this section

  • Explain why poor documentation is the main weakness of grey-literature and web searching, using evidence from studies of how reviews report these searches.
  • Distinguish a transparent search from a reproducible one.
  • Write a grey-literature search plan that states sources, methods, stopping rules, responsibilities and timing.
  • Keep a web search log and preserve the documents found so that the evidence base can be checked later.
  • Report grey-literature, web and citation searches with PRISMA-S and the PRISMA 2020 flow diagram.

4.1 The Documentation Problem

A database search is reported by pasting the full strategy for each database, a practice Lesson 4 taught with PRISMA-S. Grey-literature and web searches are much harder to report, and studies of published reviews show that they are often reported poorly. Briscoe (2015) examined systematic reviews published in the United Kingdom's Health Technology Assessment programme between 2004 and 2013. Of 300 systematic reviews, 108 reported searching the web. Most of these gave only the names of the websites (54 reviews) or the search engines (33 reviews) used, and only 6 reached the highest standard of reporting, which included the search terms, dates and results. Briscoe concluded that this level of reporting did not allow readers to judge or repeat the searches, and that full reproducibility is difficult in any case because websites and search engines change.

Mahood, Van Eerd and Irvin (2014) found that grey-literature searching was hard to document because each source had its own interface and few allowed results to be exported. Godin and colleagues (2015) noted that web documents and their addresses are transient, that Google results could not be exported into their record-management spreadsheet, and that personalization of search results may have affected what they saw. For these reasons, documentation is often the weakest part of a grey-literature search. The remedy is a plan written before searching and a log kept during it.

4.2 Transparent and Reproducible Searches

Two related standards apply. A search is transparent when a reader can see exactly what was done: which sources were searched, with which queries and settings, on which dates, how far down each list of results the searcher went, and what was kept. A search is reproducible when another person could repeat those steps and obtain the same results. Lesson 3 explained why web searches cannot be fully reproducible: search engines personalize and re-rank results, their indexes change daily, and web pages move or disappear. A database search run on the same date with the same strategy is close to reproducible, while a Google search rarely is.

The practical goal is therefore full transparency and as much reproducibility as the sources allow. The team records enough detail for someone else to rerun each search, and it keeps copies of everything it saved so that the evidence can be checked after the original pages change.

PRISMA-S itemWhat to record for grey-literature, web and citation searches
Study registriesEach registry searched, the search terms and fields used, and the date.
Online resources and browsingEach website or online resource searched or browsed, with its address, and how it was searched (browsing, site search or a search-engine query).
Citation searchingWhether cited and citing references were examined, and the methods and tools used to retrieve them.
ContactsWhether authors, organizations or experts were contacted to find additional studies or documents, and how.
Other methodsAny other method, such as hand searching conference abstract supplements.
Full search strategiesThe exact query for every search, copied as it was run, including operators.
Limits and restrictionsAny date, language or file-type limits, and the stopping rule for ranked results.
Dates of searchesThe date each source was searched.
Total records and deduplicationThe number of records or items found in each source and how duplicates were removed.

4.3 The Grey-Literature Search Plan

A grey-literature search plan sets out, before searching begins, which sources will be searched, how, how far, by whom and when. It belongs in the protocol (Lesson 2), and any later change is recorded with its reason. Writing the plan first forces the team to justify each source from the review question, fixes stopping rules before anyone sees the results and divides the work among team members. The plan has the following elements.

  • The plan lists the source groups and the specific sources in each, with a reason for including each one.
  • The plan states the method for each source, such as exporting all records, browsing publication pages, using site search or running site-restricted queries.
  • The plan gives the queries or search terms, any date, language or file-type limits, and a stopping rule for every ranked source.
  • The plan names who will search and check each source and when, and how found items will be captured, logged and passed to screening.

4.4 The Web Search Log

A web search log is a table with one row for every search of a website, search engine or other online source, completed at the time the search is run. It is the record from which the methods section and the search appendix are written. The template shows the fields.

FieldWhat to enterWhy it matters
Search IDA short code, such as W04Links the row to saved items and screenshots
Date and searcherThe date (and time, if several searches are run in a day) and the person searchingWeb content and rankings change over time
Source and addressThe search engine, database or website, with its addressLets a reader find the same source
SettingsThe Google domain (for example google.ca), browser mode, whether signed in, apparent location, language and any filtersSettings change what a search engine returns
Query or pathThe exact query as typed, or the pages browsed and site-search terms usedParaphrased queries cannot be rerun
Results screenedThe number of results examined and the stopping rule applied, for example "first 100" or "all 58 shown"Shows how much of each ranking was examined
Items savedThe number saved and their item IDs, for example W04-01 to W04-11Connects each document to the search that found it
NotesProblems and changes to the query, such as a site search that failedExplains deviations from the plan

Results counts reported by Google are estimates and can be noted, but they are not a measure of how many results were examined. Google now often shows an AI-generated overview above the results. The log can note that one appeared, but the overview is not a source, and Lesson 6 explains how to treat AI-generated answers.

4.5 Capturing and Preserving What You Find

Health authority and government websites are reorganized regularly, and documents are removed when programs end. Addresses that stop working, a problem known as link rot, can make a review's evidence base impossible to check. The team therefore saves a copy of every item at the time it is found and records where it came from.

  • Each saved item receives an ID that combines the search ID with a sequence number, such as W04-03, and the file is saved in the team's shared folder under that ID.
  • Each item is added to the reference manager with its address and access date; Zotero (Lesson 7) records the access date automatically.
  • For web pages, as opposed to PDF files, the team also saves an archived copy, for example with the Internet Archive's Wayback Machine "Save Page Now" function, and records the archived address.
  • For each search-engine query, the team saves a screenshot or PDF of the first page of results, which documents the ranking it saw.
  • Any persistent identifier, such as a DOI or a repository handle, is recorded when one exists.
A documented grey-literature search 1. Plan Sources, methods and stopping rules in protocol 2. Search Run each query and screen to the stopping rule 3. Capture Save file, address, access date and archived copy 4. Log One row per search, completed at the time 5. Screen De-duplicate and screen with all records (Lesson 7) 6. Report PRISMA-S items and PRISMA 2020 flow diagram
The plan is written before searching, each search is captured and logged as it is run, and the log supplies the details for the report.

4.6 Worked Example: The Cedar Valley Plan and Log

This worked example shows the grey-literature search plan and the web search log for the fictional Cedar Valley review. The numbers are illustrative.

The search plan

Source groupSources and reasonMethod and stopping ruleWho and when
ThesesProQuest Dissertations & Theses Global and Theses Canada, because graduate evaluations of local programs are common and often unpublishedSimplified keyword versions of the Lesson 4 concepts; export all ProQuest records; screen Theses Canada results in place and save relevant onesLibrarian, week 5
Trial registriesClinicalTrials.gov and the World Health Organization ICTRP search portal, to find completed but unpublished trialsKeyword search of condition and intervention fields; export all recordsLibrarian, week 5
PreprintsmedRxiv and PsyArXiv, for recent studiesKeyword search; export or record all resultsIntern, week 5
Conference abstractsAbstract supplements of the Gerontological Society of America meetings in Innovation in Aging, 2022 to 2025, since Embase already covers many other meetingsHand search of abstract titles using a list of key termsIntern, week 6
Search enginesGoogle, for reports on government, health authority and organizational sites; Google Scholar, for theses and reports on academic sitesFive Google queries with site:, filetype: and intitle: (first 100 results each); one Google Scholar query (first 200 results)Intern, week 5, with the first page of each query checked by the evidence officer
Targeted websitesSix organizations that report on older adults, social isolation or social prescribing and whose documents a site-restricted query might missBrowse all items on publication pages, or use site search with simple termsIntern, week 5
ContactsOrganizations running community programs, which are approached for the environmental scan (Lesson 11)Ask each for unpublished evaluation reports and record the request, reminder and replyEvidence officer, weeks 7 to 9
Citation searchingIncluded studies and included grey-literature documents (Section 3)Backward and forward searching, one iterationLibrarian and intern, week 9

The web search log

All Google and Google Scholar searches were run on google.ca in a private browsing window, signed out of all Google accounts, with English as the language, Verbatim off and no date filter; the location Google reported was Burnaby, British Columbia. Google Scholar searches excluded citations and patents. A screenshot of the first results page was saved for each query.

IDDateSourceQuery or path, exactly as runScreenedSaved
W012026-02-17Google"social prescribing" OR "community connector" loneliness seniors site:canada.ca OR site:gc.caFirst 1006
W022026-02-17Googleloneliness OR "social isolation" seniors evaluation site:gov.bc.ca filetype:pdfAll 58 shown5
W032026-02-17Googleloneliness seniors program site:fraserhealth.ca OR site:interiorhealth.ca OR site:islandhealth.ca OR site:northernhealth.ca OR site:vch.ca OR site:fnha.caFirst 1008
W042026-02-18Googleintitle:loneliness seniors program evaluation site:ca filetype:pdf -jobsFirst 10011
W052026-02-18Google"social isolation" "older people" intervention evaluation site:who.intAll 47 shown2
W062026-02-18Google Scholar"social prescribing" OR "community connector" OR "link worker" loneliness "older adults"First 20014
W072026-02-19Office of the Seniors Advocate (seniorsadvocatebc.ca)Browsed the reports page; every item listed41 listed3
W082026-02-19National Seniors Council (canada.ca)Browsed the publications and reports page; every item listed23 listed2
W092026-02-19United Way British ColumbiaSite search: isolation; then evaluation36 results4
W102026-02-20Canadian Institute for Social PrescribingBrowsed the resources page; every item listed29 listed5
W112026-02-20National Institute on AgeingBrowsed the reports page; every item listed34 listed3
W122026-02-20Statistics Canada (statcan.gc.ca)Site search: loneliness seniorsAll 52 shown0

The log records W12, which found nothing relevant, because readers need to know the source was searched. Twelve searches saved 63 items: 32 from the five Google queries, 14 from Google Scholar and 17 from the six websites. The table combines these with the other sources in the plan.

Source groupRecords or items carried forward
Web searching (W01 to W12)63
Theses (164 records exported from ProQuest Dissertations & Theses Global; 6 saved from Theses Canada)170
Trial registries (231 from ClinicalTrials.gov; 198 from the ICTRP search portal)429
Preprints (medRxiv and PsyArXiv)88
Conference abstracts (Innovation in Aging supplements)21
Total771
Duplicates removed, within these sources and against the database records214
Records and items passed to screening557

The PRISMA 2020 flow diagram, which Lesson 7 builds, places study registers beside databases in its left-hand column and counts websites, organizations and citation searching in a separate column for other methods. The Cedar Valley team ran its registry, thesis, preprint and conference searches in week 5 as part of the grey-literature search, after the database records had been de-duplicated and screening had begun, and it screened the 557 records and items in the table above as one set. It therefore reported all of these sources in the column for other methods and explained the adaptation in its methods. A team that searches registers together with its databases and screens their records as one set counts them in the left-hand column, as the template intends. Of the 557 records and items, 61 were sought as full reports, and none of the registry records led to an additional study. In the completed review, grey-literature and website searching contributed 26 documents, which the team charted separately from the research studies, and citation searching added 4 studies to the 38 found through the databases, for a total of 42 included studies.

Reporting the web searches

From the log, the team writes a short methods paragraph and places the full log in an appendix. "Between 17 and 20 February 2026, one reviewer searched Google (google.ca) with five queries and Google Scholar with one query, in a private browsing window while signed out; Appendix 3 gives each query exactly as run. We screened the first 100 Google results and the first 200 Google Scholar results for each query, or all results when fewer were shown, and a second reviewer checked the first page of results for each query. We searched six organizational websites by browsing their publication pages or using their site search. All saved documents were archived with their addresses and access dates." Together with the appendix and the plan, it addresses the PRISMA-S items in Section 4.2.

Gap: naming only the search enginev

"We searched Google" cannot be checked or repeated. Give every query as run, with its date, settings and stopping rule.

Gap: paraphrased queriesv

"We searched Google for reports on social prescribing for seniors" hides the operators and terms. Copy each query from the log exactly, including quotation marks and operators.

Gap: no stopping rulev

Without a stated stopping rule, readers cannot tell whether five or five hundred results were examined. State the rule and report the number screened for each query.

Gap: saved items not linked to searchesv

If the files are not named with search IDs, the team cannot say which search found which document, and the flow diagram counts cannot be checked. Use item IDs such as W04-03 from the start.

Gap: unrecorded contactsv

Requests to organizations for unpublished reports are a search method under PRISMA-S. Record whom the team contacted, when, whether a reminder was sent and what was received.

Summary of Section 4

Web and grey-literature searches are usually reported in too little detail to be checked. A web search cannot be fully reproducible, but it can be fully transparent: a plan written before searching fixes sources, methods and stopping rules, a log completed at the time records every search, and archived copies protect the evidence base from link rot. PRISMA-S sets out what to report.

Three records from this lesson

For a rapid scoping review and environmental scan, the methods of this lesson produce three records. The first is a grey-literature search plan that covers at least four source groups, for example targeted websites, Google, Google Scholar, and theses or trial registries, and gives for each source the reason for including it, the method, the stopping rule, who will search it and when; for a British Columbia question it usually lists at least five organizational websites. The second is a web search log in which each Google query, written with site:, filetype: and intitle: where they help, is recorded with every field from the template in Section 4.4, and each item kept is saved under its log ID. The third is a record of a backward and forward citation search from two or more seed articles, giving the citation index, the date and the number of records found. Lesson 6 adds the AI-assisted search and verification log.

Reflection

A draft review contains this sentence in its methods: "We also searched Google and relevant websites for grey literature." The team's notes show the following. On 3 March 2026, one reviewer searched google.ca in a private browsing window while signed out, using the query "falls prevention" seniors program evaluation site:gov.bc.ca filetype:pdf; the reviewer screened all 64 results shown and saved 5 documents. On the same day, the reviewer ran intitle:falls seniors "community program" site:ca -jobs, screened the first 100 results and saved 9 documents. On 4 March 2026, the reviewer browsed every item on the reports page of the Office of the Seniors Advocate of British Columbia website (41 items) and saved 2. No archived copies of web pages were made, and the saved files were named by their titles. (1) Rewrite the sentence as a methods paragraph that would let a reader repeat the searches. (2) Set out the three searches as rows of a web search log with columns for ID, date, source, settings, query or path, results screened and items saved. (3) Name two weaknesses in how the documents were preserved and say how to fix each.

Model answer

(1) "On 3 and 4 March 2026, one reviewer searched the web for grey literature. Google (google.ca) was searched in a private browsing window while signed out, with two queries given in full in Appendix 2; for each query, the reviewer screened the first 100 results, or all results when fewer were shown. The reviewer also browsed every item listed on the reports page of the Office of the Seniors Advocate of British Columbia. In total, 164 search results and 41 listed reports were screened, and 16 documents were saved for eligibility screening."

(2) W01 | 2026-03-03 | Google | google.ca, private window, signed out | "falls prevention" seniors program evaluation site:gov.bc.ca filetype:pdf | all 64 shown | 5 (W01-01 to W01-05).
W02 | 2026-03-03 | Google | same settings | intitle:falls seniors "community program" site:ca -jobs | first 100 | 9 (W02-01 to W02-09).
W03 | 2026-03-04 | Office of the Seniors Advocate website | not applicable | browsed reports page, every item | 41 listed | 2 (W03-01 to W03-02).

(3) First, the files are named by title, so no one can tell which search found each document. They should be renamed with item IDs such as W02-04, and the IDs should be entered in the log. Second, no archived copies were made, so the documents may become unavailable if the websites change. The team should record each address and access date and save an archived copy, for example with the Wayback Machine's Save Page Now function.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: In Briscoe's (2015) study of systematic reviews in the United Kingdom's Health Technology Assessment programme, what did most reviews that searched the web report?

Of 108 reviews that reported web searching, most gave only the names of the websites (54) or search engines (33), and only 6 reached the highest standard of reporting. This is the documentation problem that the web search log is designed to solve.

Question 2: Which web search log entry best allows another team to repeat the search?

Only option d gives the date, the Google domain, the browser setting, the exact query, the stopping rule and item IDs that link the saved documents to the search. Option a lacks a date, query and counts, option b has no stopping rule and no query, and option c omits the date, the query and the settings.

Question 3: Which statement describes the difference between a transparent and a reproducible web search?

Transparency means a reader can see the sources, queries, settings, dates, stopping rules and results. Reproducibility means that repeating the steps gives the same results, which web searches rarely achieve because rankings and pages change. The goal for web searching is full transparency and as much reproducibility as the sources allow.

Question 4: Where does the PRISMA 2020 flow diagram place documents found through websites, organizations and citation searching?

The PRISMA 2020 flow diagram has one column for studies identified via databases and registers and another for studies identified via other methods, such as websites, organizations and citation searching. Documents from these sources are screened and counted like any other records.
Section 5 of 5

Final Assessment

⏱ Estimated time: 25 minutes

Bringing It All Together

This lesson extended the Cedar Valley search beyond bibliographic databases. Grey literature is defined by who controls its publication, and it matters to a review because publication bias and time-lag bias distort a journal-only evidence base, because community programs are often evaluated in reports that never reach a journal, and because reports describe who runs and funds programs. Its quality varies, so the AACODS checklist offers a consistent first judgement of each document.

Finding grey literature combines grey-literature databases, trial registries, preprint servers, targeted websites, search engines, contacts and citation searching, since each finds documents the others miss. Google and Google Scholar operators make web searching more precise, but ranked results must be sampled with a stopping rule. Backward and forward citation searching, described in the terminology of the TARCiS statement, find studies through their links to included studies, and citation-network tools can support exploration without replacing complete citation searches.

Documentation holds these methods together. A plan written before searching, a web search log completed during it and archived copies of every item make the search transparent, and PRISMA-S and the PRISMA 2020 flow diagram set out how to report it. The final reflection asks you to plan these searches for a new question, and the knowledge check integrates all four sections.

Key Takeaways from this lesson

  • Grey literature is material whose publication is not controlled by commercial publishers, and its quality must be judged separately from its publication status.
  • Searching grey literature reduces publication bias and time-lag bias and finds practice-based evaluations of community programs that journals rarely publish.
  • Trial registries reveal completed trials whose results were never published, and they are counted with databases in the PRISMA 2020 flow diagram.
  • Canadian grey-literature sources include federal agencies, provincial ministries and health authorities, community organizations, national centres and thesis repositories.
  • Godin and colleagues showed that combining grey-literature databases, targeted websites, search engines and contacts finds documents that any single strategy would miss.
  • Google operators such as quotation marks, OR, the minus sign, site:, filetype: and intitle: are typed without a space after the colon, and Google has no truncation.
  • Google and Google Scholar rank results opaquely, so each query needs a stopping rule set before searching, and Google Scholar serves as a supplement to database searching.
  • Backward citation searching retrieves what seed references cite, forward citation searching retrieves what cites them, and TARCiS recommends using all included records as seeds.
  • Citation-network tools use co-citation and bibliographic coupling to suggest related papers, but they do not provide the complete lists that citation searching requires.
  • A grey-literature search plan, a web search log and archived copies of saved items make web searching transparent and allow it to be reported with PRISMA-S.

Core Concepts Reviewed

Section 1: grey literature and its definition, publication bias, time-lag bias, government and agency reports, program evaluations, theses, conference abstracts, preprints, trial registries, Canadian and British Columbian sources, and the AACODS checklist.

Section 2: the four complementary strategies of Godin and colleagues, targeted website searching, the Grey Matters checklist, Google operators (quotation marks, OR, the minus sign, site:, filetype:, intitle:), Google Scholar's strengths and limits, programmable search engines and stopping rules.

Section 3: citation chasing, the TARCiS terminology of backward, forward, direct, indirect and iterative citation searching, seed references, the evidence on yield, co-citation, bibliographic coupling and citation-network tools.

Section 4: the documentation problem, transparent and reproducible searches, PRISMA-S items for grey-literature and web searching, the grey-literature search plan, the web search log, preserving documents against link rot, and the PRISMA 2020 flow diagram.

The final reflection asks you to plan grey-literature, web and citation searches for a new review question about support for family caregivers of people living with dementia.

Reflection

A health authority in British Columbia has asked for a rapid scoping review and environmental scan of community-based programs that support family caregivers of people living with dementia. The database search is complete, and the team now needs to search grey literature, the web and citations. Use these definitions. Grey literature is material whose publication is not controlled by commercial publishers, such as government and agency reports, program evaluations, theses, conference abstracts, preprints and trial registry records. Google operators include quotation marks (exact phrase), OR in capitals (alternatives), the minus sign (exclude a word), site: (restrict to a domain), filetype: (restrict to a file format) and intitle: (require a word in the title), typed with no space after the colon. Backward citation searching retrieves the references that included studies cite, and forward citation searching retrieves the documents that cite them. A stopping rule fixes in advance how many ranked results will be screened. A web search log records the date, source, settings, exact query, results screened and items saved for every search. Write a plan of 200 to 300 words that (1) names four kinds of grey literature you would search and one specific Canadian or British Columbian source for each, (2) gives two correct Google queries and your stopping rule, (3) describes your citation-searching approach, including the seed references and how duplicates will be handled, and (4) explains how you will log and preserve what you find so that the search is transparent.

Model answer

The plan covers four kinds of grey literature. For government and agency reports, the team will search the Public Health Agency of Canada's publications on Canada.ca, including A Dementia Strategy for Canada: Together We Aspire (2019), and the BC Ministry of Health. For program evaluations, it will browse the publication pages of the Alzheimer Society of B.C. and Family Caregivers of British Columbia. For theses, it will search Theses Canada and the Summit and cIRcle repositories. For trial registry records, it will search ClinicalTrials.gov for completed trials of caregiver support programs.

Two Google queries are dementia caregiver support program evaluation site:gov.bc.ca filetype:pdf and intitle:caregivers dementia "support group" OR "respite" site:ca -jobs. The stopping rule is the first 100 results of each query, or all results when fewer are shown, set before searching so that effort is consistent and reportable.

After full-text screening, all included studies will be the seed references. Backward and forward searching will use citationchaser (Lens.org) and Web of Science, and the reference lists of included reports will be checked by hand. Retrieved records will be de-duplicated against each other and against all records already screened, and the team will consider a second iteration only if the first finds new eligible studies.

Every search will be entered in the web search log when it is run, with google.ca, a private window and a signed-out browser noted as settings. Each saved document will receive an ID such as W02-03, be stored under that ID with its address and access date, and be archived with the Wayback Machine. The report will follow PRISMA-S, with the full log in an appendix.

Minimum 30 characters required.

✓ Reflection saved

Final Knowledge Assessment

Final Assessment, this lesson: Grey Literature, Web Searching and Citation Chasing (15 Questions)

Question 1: A Cedar Valley query must cover both of the main Government of Canada web domains. Which restriction does this?

Government of Canada content is split between the canada.ca domain and older gc.ca domains, such as the one used by Statistics Canada. Option c includes the British Columbia government domain in place of canada.ca, and the domains in options a and d are not Government of Canada domains.

Question 2: Which source is most likely to reveal a completed but unpublished trial of a befriending program for older adults?

Trial registries record trials whether or not they are published, and the ICTRP search portal brings together records from many registries. Citing lists and reference lists contain only documents that have been cited, which an unpublished trial rarely is, and a statistical report does not evaluate interventions.

Question 3: Which statement about grey literature and quality is accurate?

The definition of grey literature concerns who controls publication. A grey document may be rigorous or weak, so its credibility is judged with a tool such as AACODS and its design with the risk-of-bias tools in Lesson 8. Government reports are not necessarily peer reviewed, and scoping reviews regularly include grey literature.

Question 4: Why does the Cedar Valley plan search the websites of organizations that fund and deliver community programs for older adults?

Evaluations of community programs are often written for funders and posted on organizational websites, and they are rarely indexed in bibliographic databases. No guidance sets a fixed number of websites, and contacting organizations remains a separate method, because many evaluations are never posted at all.

Question 5: What does the query "community connector" OR "social prescribing" seniors site:ca -jobs ask Google to return?

OR joins the two phrases as alternatives, seniors is required through Google's implied AND, site:ca restricts results to the .ca top-level domain, which includes non-government sites, and -jobs excludes pages containing the word jobs.

Question 6: In Google Scholar, which query finds records with loneliness in the title and the phrase older adults anywhere in the record?

Google Scholar supports intitle: and quotation marks. Option c uses PubMed's field tag syntax, option d uses a database-style field code and truncation, which Google Scholar does not support, and option b uses syntax that Google Scholar does not recognize.

Question 7: TARCiS notes that checking reference lists by hand usually gives less information than retrieving the cited references from a citation index. What does the hand check usually lack?

A reference list gives titles, authors and years, but no abstracts, so screening has to rely on titles alone. Retrieving cited references from a citation index such as Lens.org or Web of Science provides titles and abstracts for screening.

Question 8: Citation searching on the Cedar Valley seeds found 4 new eligible studies. What does TARCiS suggest the team consider next?

TARCiS suggests considering another iteration when citation searching finds new eligible records, weighing the time needed against the likely benefit. The Cedar Valley team decided against a second iteration because of its deadline and reported that decision. Studies found by citation searching are included like any others.

Question 9: Why can Connected Papers and similar map-building tools not replace complete backward and forward citation searching?

Map-building tools display papers that their algorithms judge to be similar, often using co-citation and bibliographic coupling, so they do not return the complete lists of cited and citing references that direct citation searching requires. They work with journal articles, recent papers and existing data sources.

Question 10: Lesson 3 explained that web searches cannot be fully reproduced. What is the practical goal for documenting them?

Because rankings and pages change, the goal is full transparency, with the date, source, settings, exact query, stopping rule and results recorded for each search, and as much reproducibility as the sources allow, supported by saved and archived copies of each item. Abandoning web searching would miss practice-based evidence.

Question 11: The intern saves a PDF evaluation report from a health authority website. Which details should be recorded with it?

The item ID links the document to the log row for the search that found it, and the address, access date and archived copy allow the evidence to be checked after the website changes. Google's result estimate is unstable, and the program details belong in the charting table.

Question 12: Two evaluations published in 2021 both cite the same five earlier studies, although neither cites the other. Which relationship links them?

Documents that share references are bibliographically coupled (Kessler, 1963). Co-citation (Small, 1973) would apply if a later document cited both evaluations. In TARCiS terms, retrieving documents that share references with a seed is co-citing citation searching.

Question 13: Hopewell and colleagues (2007) compared trials published in journals with trials found only in the grey literature. Which finding and implication is correct?

Trials published in journals tended to show larger intervention effects than grey trials, a sign of publication bias, so a review limited to journals can overestimate effects. Searching grey literature reduces this bias at the search stage, and HSCI 230 Lesson 2 teaches statistical checks for it.

Question 14: In the PRISMA 2020 flow diagram, where are records from trial registries such as ClinicalTrials.gov counted?

PRISMA 2020 groups registers with databases in the column for studies identified via databases and registers. Websites, organizations and citation searching are counted in the column for other methods. Registry records are screened and counted like other records.

Question 15: Which item belongs in a grey-literature search plan written before searching begins?

A plan fixes sources, methods, stopping rules, responsibilities and timing before anyone sees the results, so that effort is consistent and reportable. The number of included documents and the results of searches are not known in advance, and expected conclusions have no place in a search plan.
✦ Complete the final reflection above before submitting

Congratulations!

You have successfully completed this lesson: Grey Literature, Web Searching and Citation Chasing.

You can now explain why evidence reviews search for grey literature, identify the main kinds of grey literature and their Canadian sources, judge a grey document's credibility, write correct Google and Google Scholar queries, plan backward and forward citation searching, and keep a search plan and web search log that make your searches transparent and reportable.

Lesson 6 turns to artificial intelligence in evidence retrieval and synthesis. It explains how large language models and retrieval-augmented tools produce answers, why they sometimes fabricate citations, and how to verify their output and disclose their use, and it shows how to keep an AI-assisted search and verification log.

Continue to Lesson 6 →
Reference

Glossary: Key Terms, People & Frameworks

📚 Reference page, available throughout the lesson

Terms, tools and people introduced in this lesson on grey literature, web searching and citation chasing.

Core Concepts
Grey literature Material produced by government, academia, business and industry whose publication is not controlled by commercial publishers, such as reports, theses, conference abstracts, preprints and trial registry records.
Publication bias Distortion of the available evidence that occurs when a study's results influence whether, where or how quickly it is published.
Time-lag bias The tendency for studies with favourable results to be published sooner than studies with null results, so that recent evidence is incomplete.
Preprint A complete manuscript posted on a public server, such as medRxiv, before peer review.
Trial registry A public database of records describing trials, their planned outcomes, their status and sometimes their results, such as ClinicalTrials.gov.
Targeted website searching Searching the websites of organizations chosen in advance as likely publishers of relevant documents, by browsing publication pages, using site search and running site-restricted queries.
Search operator A word or symbol that tells a search engine how to treat part of a query, such as site: or quotation marks.
Stopping rule A decision made before searching about how many ranked results will be screened for each query, such as the first 100 Google results.
Citation chasing Finding relevant documents through their citation links to known relevant documents; TARCiS calls this citation searching.
Backward citation searching Retrieving and screening the references that seed references cite.
Forward citation searching Retrieving and screening the references that cite the seed references, using a citation index.
Reference list checking Backward citation searching done by reading the reference lists of seed references by hand.
Seed references The known relevant documents from which a citation search starts, normally all records included after full-text screening.
Iterative citation searching One or more repetitions of a citation search, usually with newly found eligible records as new seeds.
Co-citation The relationship between two documents that are cited together by a later document (Small, 1973).
Bibliographic coupling The relationship between two documents that cite the same earlier work (Kessler, 1963).
Citation index A database that records which documents cite which, allowing forward and backward searching, such as Web of Science, Scopus or Lens.org.
Grey-literature search plan A plan, written before searching and placed in the protocol, that states sources, methods, queries, stopping rules, responsibilities and timing.
Web search log A table with one row per web search that records the date, source, settings, exact query, results screened and items saved.
Transparent and reproducible searches A transparent search shows exactly what was done; a reproducible search gives the same results when its steps are repeated, which web searches rarely achieve.
Link rot The loss of access to documents when web addresses stop working, which makes archived copies necessary.
Frameworks & Tools
AACODS checklist A checklist by Tyndall (2010) for judging grey literature on authority, accuracy, coverage, objectivity, date and significance.
Grey Matters A checklist of health-related grey-literature sources organized by type and jurisdiction, developed by CADTH and now published by Canada's Drug Agency.
Google advanced operators Operators such as quotation marks, OR, the minus sign, site:, filetype: and intitle: that make Google queries more precise.
Google Scholar A free search engine for scholarly material, including theses and reports, that ranks results opaquely, displays at most 1,000 results and is best used as a supplementary source.
Programmable Search Engine A Google service, formerly called Custom Search Engine, that builds a search box restricted to a chosen list of websites.
ClinicalTrials.gov and the ICTRP ClinicalTrials.gov is a large trial registry run by the United States National Library of Medicine; the World Health Organization's International Clinical Trials Registry Platform search portal combines records from many registries.
Theses Canada A collaborative program of Library and Archives Canada and Canadian universities, begun in 1965, that provides a search of Canadian theses; recent coverage varies by university.
citationchaser A free, open-source R package and web application (Haddaway, Grainger and Gray, 2022) that retrieves cited and citing references from Lens.org for transparent citation chasing.
Wayback Machine The Internet Archive's service for archiving web pages; its Save Page Now function records a copy of a page at a given time.
PRISMA-S An extension of the PRISMA statement for reporting literature searches (Rethlefsen et al., 2021), with items on registries, online resources, citation searching, contacts and dates.
TARCiS statement Guidance on the terminology, application and reporting of citation searching, with ten recommendations agreed in a Delphi study (Hirt et al., 2024).
Key People
Eugene Garfield Information scientist who proposed citation indexing in 1955 and founded the Institute for Scientific Information, which produced the Science Citation Index, the forerunner of Web of Science.
Henry Small Information scientist who introduced co-citation as a measure of the relationship between documents in 1973.
M. M. Kessler Researcher at the Massachusetts Institute of Technology who introduced bibliographic coupling in 1963.
Trisha Greenhalgh British primary care researcher who, with Richard Peacock (2005), showed that snowballing found about half the sources in a review of complex evidence.
Neal Haddaway Evidence synthesis methodologist who evaluated Google Scholar for grey-literature searching (2015) and co-developed citationchaser (2022).
Melissa Rethlefsen Health sciences librarian who led the development of PRISMA-S for reporting literature searches (2021).
Katelyn Godin Canadian researcher who led a 2015 case study applying systematic search methods to grey literature with four complementary strategies.
Simon Briscoe Information specialist whose studies of systematic reviews showed that web searching and citation searching are often reported in too little detail.
No matching entries. Try a different search term.