Data Extraction, Risk of Bias and Certainty of Evidence
Finding & Synthesizing Health Evidence
Learning objectives for this lesson:
- Design a data extraction form whose domains follow the review question, using the TIDieR checklist for intervention fields and a data dictionary for every field.
- Pilot an extraction form with two extractors, calculate percent agreement, and revise the form in response to the pattern of disagreement.
- Distinguish risk of bias from reporting quality, imprecision and applicability, and explain why domain-based judgements have replaced summary quality scores.
- Describe the structure and judgement scales of RoB 2, ROBINS-I, ROBINS-E, the Newcastle-Ottawa Scale, the JBI checklists and the Mixed Methods Appraisal Tool, and select a tool for each design in a review.
- Apply a risk-of-bias tool with a second reviewer through calibration, written decision rules, recorded support and consensus, and calculate Cohen's kappa for the judgements.
- Present risk-of-bias judgements in a traffic-light plot and explain how a review uses them in synthesis.
- Rate the certainty of a body of evidence with GRADE, from its starting level through the five domains for rating down and the three for rating up, and write a matching plain-language statement.
- Plan the charting form, the appraisal approach and, where the question concerns effects, a GRADE evidence profile for a rapid scoping review.
This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University, drawing on the Cochrane Handbook for Systematic Reviews of Interventions and the JBI Manual for Evidence Synthesis.
Designing and Piloting a Data Extraction Form
Learning Objectives for this section
- Explain what data extraction is, why a review uses a standard form, and how extraction relates to the charting step of a scoping review.
- List the main domains of an extraction form and use the TIDieR checklist to plan the fields that describe an intervention.
- Write a data dictionary with definitions, allowed values and instructions so that two people fill in each field the same way.
- Plan and run a pilot of an extraction form, calculate percent agreement for the pilot, and revise the form in response.
- Choose between dual independent extraction and single extraction with verification, and handle multiple reports of one study.
Introduction
By the end of Lesson 7, a review team has a list of included studies and a PRISMA flow diagram showing how it got there. The next task is to take the information the review needs out of each study and record it in a form that can be compared, checked and synthesized. This task is called data extraction: the systematic recording of the same set of details from every included study onto a standard form. Much of the work is clerical, but the form decides what the review can say later. A characteristic that nobody extracted cannot be used to group studies, explain differences between them or describe whom the evidence applies to.
The lesson follows the order of the work. Section 1 designs and pilots the extraction form. Section 2 introduces the tools that assess the risk of bias in each study, matched to its design. Section 3 shows how two reviewers apply a tool consistently. Section 4 rates the certainty of the whole body of evidence for an outcome with GRADE. These tools rest on a few concepts of bias, chiefly selection bias, information bias and confounding. Section 2 summarizes them, and HSCI 230 Lessons 7 to 11 teach them in depth for students who take that course.
The fictional Cedar Valley Health Authority in British Columbia serves about 210,000 residents, of whom about 46,000 are aged 65 and older, across Cedar City, several smaller communities and 24 primary care clinics. Before launching a community connector (social prescribing) program for older adults, its planning team asked a small evidence team for a rapid scoping review and environmental scan within twelve weeks. The evidence team is a health authority evidence officer, a university librarian and a student intern. The review asks which community-based interventions have been evaluated for reducing loneliness or social isolation among adults aged 65 and older, and with what outcomes.
The database searches retrieved 2,480 records. After 610 duplicates were removed, the team screened 1,870 titles and abstracts, assessed 142 full texts and included 38 studies. Citation chasing added 4 studies, for 42 included studies in total, and grey literature and website searching added 26 documents. For this lesson, the 42 studies comprise 14 randomized trials (4 of them cluster-randomized), 10 non-randomized controlled studies, 9 uncontrolled before-and-after studies, 6 qualitative studies and 3 mixed-methods studies. These design counts, like the other numbers in this lesson beyond the shared figures above, are illustrative.
In a scoping review, the JBI method calls this step data charting (Peters et al., 2020). Charting tends to be more descriptive and more iterative than extraction for a systematic review of effects, and Lesson 10 returns to it. The design principles in this section apply to both. The Cedar Valley team uses one form for its 42 studies, because the planning team asked about outcomes as well as about the range of programs studied, and a shorter descriptive form for the 26 grey-literature documents, most of which are short program evaluation reports whose methods are too briefly described to classify by design, and which also feed the environmental scan in Lesson 11.
1.1 What an Extraction Form Is For
An extraction form does four jobs. It describes the included studies, which become the characteristics table that readers use to judge whom the evidence covers (Lesson 12). It collects the information the synthesis needs, such as intervention categories and outcome results (Lesson 9). It records the details on which risk-of-bias assessment depends, such as how participants were allocated and how many were lost to follow-up (Sections 2 and 3). And it leaves an audit trail, so that a reader or a later update can see where every number came from.
The unit of extraction is the study. A single study often produces several reports, such as a registry entry, a protocol, a main results paper and a later follow-up paper. These are called companion reports, and the Cochrane Handbook advises collating all reports of a study so that each study is counted once and its information is complete (Li, Higgins and Deeks, 2019). The reverse problem also occurs, when two papers that look like separate studies report overlapping samples. The Cedar Valley form therefore begins with a study identifier and a field that lists every report linked to it.
1.2 What to Collect: The Domains of a Form
Most extraction forms are organized into domains that follow the elements of the review question. The table shows the domains of the Cedar Valley form.
| Domain | Example fields | Cedar Valley example (illustrative) |
|---|---|---|
| Source | Study identifier, linked reports, extractor, date, contact with authors | CV-017; main paper and registry entry |
| Eligibility check | Confirmation that population, concept and context criteria are met | Mean age 74 years; community setting; loneliness measured |
| Methods | Design, country, setting, recruitment, dates, length of follow-up | Cluster-randomized trial in 12 seniors' centres |
| Participants | Number enrolled and analyzed, age, gender, living arrangement, baseline loneliness | 240 randomized, 210 analyzed; 61 percent living alone |
| Intervention and comparator | TIDieR fields: what, who, how, where, how much, tailoring | Weekly volunteer-led group walks for 16 weeks, compared with usual activities |
| Outcomes | Outcome, instrument, scoring, time points | Six-item De Jong Gierveld Loneliness Scale at 16 weeks and 12 months |
| Results | Numbers analyzed and statistics exactly as reported | Means, standard deviations and group sizes at each time point |
| Other | Funding, conflicts of interest, notes for appraisal, queries | Municipal grant; no conflicts declared |
The intervention fields need the most planning, because community programs for loneliness vary in ways that matter for synthesis. Hoffmann and colleagues (2014) published the Template for Intervention Description and Replication, known as TIDieR, a 12-item checklist for reporting interventions. Reviewers can use its items to structure extraction, which also shows when a study left out something important. The accordion shows how the Cedar Valley team turned TIDieR items into fields.
The form records the program's stated rationale, any materials used and the activities participants take part in. A free-text summary is paired with a closed field for the intervention category so that studies can be grouped later.
The form records the type of provider (paid staff, trained volunteer, peer or health professional), the mode of delivery (in person, telephone or video), and whether the program was delivered to individuals or groups. For connector programs, it records whether the connector was based in a clinic or a community organization.
The form records the setting, the number, length and frequency of sessions, and the total duration, so that a six-week telephone program can be compared with a year-long group program on the same terms.
The form records whether the program was adapted to individuals, whether it changed during the study, and whether the authors measured how well it was delivered. A program delivered as planned to few participants can explain a null result, so these fields matter for appraisal too.
1.3 Writing Fields That Two People Fill In the Same Way
A field that two people interpret differently produces data that cannot be trusted. The main safeguard is a data dictionary, sometimes called a codebook, which defines every field: its name, a definition, the allowed values, instructions for unusual cases and an example. Several principles make fields reliable. Each field should hold one fact, so "age and gender of participants" becomes two fields. Closed fields with allowed values suit anything that will be used to group or count studies, and free text suits descriptions that readers will need in the authors' own terms. Every field should allow "not reported" and "not applicable" as answers. Extractors should record the page or table where each number was found, and results should be copied exactly as reported, with the type of statistic, the time point and the group sizes; any conversion is done later and documented.
| Field | Definition and allowed values | Instruction for unusual cases |
|---|---|---|
| Setting | Where the program was delivered: community-dwelling; residential care; mixed; not reported. | Use "mixed" only when the study includes both settings and does not separate the results. |
| Intervention category | The main component: group activity; one-to-one befriending or telephone contact; community connector or social prescribing; intergenerational; technology-based; other. | Assign the category of the component that takes the most contact time, and record any second component in the secondary category field. |
| Loneliness instrument | The scale used, with its version and score range. | Record the number of items, for example the three-item scale of Hughes et al. (2004) or the six-item scale of De Jong Gierveld and van Tilburg (2006). |
| Primary time point | The measurement closest to the end of the intervention. | Extract every reported time point, flag the one closest to the end of the program as primary, and flag the longest follow-up up to twelve months as secondary. |
Three answers that mean different things
"Not reported" means the study should have said something and did not, such as the number lost to follow-up. "Not applicable" means the field does not apply, such as the number of clusters in an individually randomized trial. "Zero" is a reported value. If all three are left as blank cells, nobody can tell them apart later, and missing information about attrition is itself a reason for concern in appraisal.
1.4 Piloting the Form
Piloting means that two or more people use the draft form on a small, varied set of included studies, compare what they recorded, and revise the form and the data dictionary where they differed. The figure shows where piloting fits in the extraction workflow.
The Cedar Valley team piloted its form with the student intern and the evidence officer. Each extracted the same five studies independently: an individually randomized trial, a cluster-randomized trial, a non-randomized controlled study, a qualitative study and a mixed-methods study. The draft form had 24 fields that could be compared directly, giving 5 × 24 = 120 paired entries, and the extractors disagreed on 17 of them.
Percent agreement in the pilot
Percent agreement = (number of entries on which both extractors recorded the same value ÷ total number of paired entries) × 100
First pilot: 103 ÷ 120 × 100 = 85.8 percent. Second pilot, after revision, on three new studies: 67 ÷ 72 × 100 = 93.1 percent.
The pattern of disagreement mattered more than the total. Six of the 17 disagreements concerned setting, because two studies recruited people living at home and in assisted living without separating their results. Five concerned the time point, because studies reported several follow-ups and the extractors chose different ones. Four concerned the intervention category, because one program combined group activities with one-to-one support, and two were copying errors. The team added the "mixed" setting value, the time-point rule and the secondary category field shown in the data dictionary, then piloted the revised form on three new studies (3 × 24 = 72 paired entries), with 5 disagreements. There is no universal threshold for acceptable agreement, so a team should decide in advance what it will accept and should look at which fields disagree, since a high overall percentage can hide one field that is often wrong.
A study describes a program in which a volunteer visits each older adult at home once a week for an hour, and once a month brings the person to a two-hour group lunch. Using the Cedar Valley rule, which category and secondary category would you record? Over a four-week month, the visits provide about four hours of contact and the lunch about two, so the primary category is one-to-one befriending and the secondary category is group activity. Write one sentence for the notes field that would let a later reader check your decision.
1.5 Managing Extraction: Who, How and With What
Errors in extraction are common. Buscemi and colleagues (2006) found that single extraction produced more errors than double extraction, although it took less time. Cochrane guidance therefore recommends that two people extract data independently, particularly outcome data, and resolve differences by discussion. A common compromise is for one person to extract and a second to check every entry against the source, and the Cochrane Rapid Reviews Methods Group recommends this approach for rapid reviews (Garritty et al., 2021). Lesson 10 discusses the trade-offs that rapid-review shortcuts make.
The Cedar Valley protocol, as amended in week 7 (Lesson 10), uses both approaches. For the 24 comparative studies (the 14 randomized trials and the 10 non-randomized controlled studies), two people extract the results fields independently, because errors in effect estimates would change what the brief says about effectiveness. One person extracts the descriptive fields for those studies, and all fields for the other 18 studies, and the other person checks them in full. A discrepancy log records each disagreement, its cause and its resolution, and the team emails authors when key information is missing, recording the request and any reply on the form.
The software matters less than the form. Spreadsheets work for small reviews if the columns are restricted to the allowed values. Screening platforms such as Covidence (Lesson 7) include extraction modules that support dual extraction, and other tools available in 2026 include SRDR+ (the Systematic Review Data Repository Plus, hosted by the United States Agency for Healthcare Research and Quality) and REDCap, a data capture platform that many universities provide. The cards describe errors that the Cedar Valley team watched for.
What Carries Forward
The form now holds, for each of the 42 studies, its design, its methods and the details that bear on bias, such as allocation, blinding, attrition and the outcome measure. The design field decides which appraisal tool applies to each study, and Section 2 introduces those tools.
Reflection
A student team drafted an extraction form for a review of community programs for loneliness among adults aged 65 and older. The draft includes these four fields: (1) Participants (age, gender, number), completed as free text; (2) Setting, completed as free text; (3) Outcome results, completed as free text; and (4) Was the program effective? (yes or no). Two extractors piloted the form on five varied studies. The form had 24 comparable fields, giving 120 paired entries, and the extractors disagreed on 17 of them. Identify the problem with each of the four fields and rewrite each as one or more fields with a definition, allowed values or an instruction. Then calculate the percent agreement in the pilot and state what the team should do before beginning full extraction.
Field 1 combines three facts, so extractors will record them in different orders and formats. It should become separate fields: mean or median age (with the statistic named), percentage of women, number randomized or enrolled, and number analyzed, each allowing "not reported". Field 2 as free text will produce many spellings of the same setting; it should be a closed field with the values community-dwelling, residential care, mixed and not reported, with the rule that "mixed" is used only when results are not separated by setting. Field 3 needs several fields: outcome instrument and version, time point, group sizes, type of statistic, the values for each group exactly as reported, and the table or page where they appear, with any conversion done later and documented. Field 4 records an interpretation; it should be replaced by a field quoting the authors' conclusion, separate from the numerical results, so that the review reaches its own judgement.
Agreement was (120 − 17) ÷ 120 = 103 ÷ 120 = 85.8 percent. Before full extraction, the team should list which fields produced the 17 disagreements, revise those definitions and rules in the data dictionary, and pilot the revised form on a few new studies against an agreement level it set in advance.
Minimum 20 characters required.
Question 1: In the Cedar Valley pilot, two extractors compared 120 paired entries and disagreed on 17 of them. What was their percent agreement?
Question 2: Why does the Cedar Valley extraction form begin with a study identifier and a list of linked reports?
Question 3: A study gives no information about how many participants were lost to follow-up. What should the extractor record in the attrition field?
Question 4: Buscemi and colleagues (2006) compared single and double data extraction. Which statement reflects their finding and its implication for review teams?
Risk-of-Bias Tools Matched to Study Design
Learning Objectives for this section
- Distinguish risk of bias from reporting quality, imprecision and applicability, and explain why domain-based tools replaced summary quality scores.
- Name the five domains of RoB 2 and explain how signalling questions lead to judgements of low risk, some concerns or high risk for a specific result.
- Describe the seven domains of ROBINS-I, the idea of a target trial, and the judgement levels from low to critical.
- Describe the purpose and structure of ROBINS-E, the Newcastle-Ottawa Scale, the JBI checklists and the Mixed Methods Appraisal Tool (MMAT).
- Select an appropriate tool for each design in a review and justify the choice in a protocol.
Introduction
A study can report its methods clearly and enrol thousands of people and still give a misleading answer, because its design or conduct pushed the result away from the truth in a consistent direction. That systematic error is bias, and the risk of bias of a result is the likelihood that features of the study's design or conduct have made it systematically too large or too small. HSCI 230 Lessons 7 to 11 teach the sources of bias: measurement (Lesson 7), selection (Lesson 8), information bias (Lesson 9), design-specific and temporal biases (Lesson 10) and confounding (Lesson 11). The tools in this section turn those concepts into structured questions that a reviewer answers for each included study.
Background: three sources of bias
Confounding is bias caused by a confounder: a common cause of the exposure (here, receiving a program) and the outcome, or a proxy for one, that is not on the causal pathway between them. In a Cedar Valley study that compares adults referred to a connector program with adults in clinics without one, baseline loneliness can influence both who is referred and later loneliness, which is why the team listed it as a confounding domain. ROBINS-I addresses this in its domain on bias due to confounding; in a trial, randomization addresses it, and RoB 2 checks the randomization process in domain 1.
Selection bias arises when entry into or retention in a study depends jointly on the exposure and the outcome, or on their causes. One Cedar Valley trial lost 35 percent of its participants, with more losses in the comparison group; if the loneliest participants were the least likely to return the twelve-week questionnaire, retention would depend on both group and outcome. RoB 2 addresses this in its domain on missing outcome data, and ROBINS-I in its domains on selection of participants and missing data. A sample that is unlike older adults in general raises a different question, external validity.
Information (measurement) bias arises from error in measuring the exposure or the outcome. Measurement error has two separate dimensions: it can be random or systematic, and it can be differential (different in the groups compared) or non-differential. Participants who know they received a befriending program may report their loneliness differently on a self-completed scale, which RoB 2 addresses in its domain on measurement of the outcome.
HSCI 230 Lessons 7 to 11 are optional fuller reading, in particular Lesson 8 on selection, Lesson 9 on information bias and Lesson 11 on confounding.
Risk of bias should be kept separate from three other properties of a study. Reporting quality describes how completely a paper reports its methods, and a poorly reported trial may have been conducted well. Imprecision is random error, shown by a wide confidence interval, and a small unbiased trial is imprecise without being biased. Applicability asks whether the study's population, intervention and setting match the review question. Risk-of-bias tools address bias alone, and GRADE in Section 4 considers imprecision and applicability (which it calls indirectness) across the body of evidence.
Earlier instruments often added points into a quality score. Jüni and colleagues (1999) applied 25 quality scales to the same set of trials and found that conclusions about whether high-quality trials showed different effects depended on which scale was used. Current tools are domain-based: the reviewer judges each type of bias separately, explains each judgement and combines them with explicit rules. The broader term critical appraisal, used particularly by JBI, covers assessments of conduct and relevance as well as bias.
Should a scoping review appraise studies at all?
Scoping reviews usually do not assess risk of bias. PRISMA-ScR lists critical appraisal as an optional item to be reported, with its rationale, if it is done (Tricco et al., 2018), and JBI guidance for scoping reviews generally does not require it (Peters et al., 2020). The Cedar Valley team decided in its protocol to appraise its studies, because the planning team needs to know how far to trust the outcome findings before funding a program, and to report appraisal descriptively without excluding any study because of it.
2.1 Choosing a Tool by Design
The ways in which a randomized trial can go wrong differ from the ways a cohort study or a qualitative study can go wrong, so each design has its own tools. The figure maps common designs to the tools in this section.
Authors' design labels are sometimes wrong, for example when a "pilot trial" allocated participants by alternation, so the extraction form records the design as the reviewers judge it.
2.2 RoB 2 for Randomized Trials
The revised Cochrane risk-of-bias tool for randomized trials, RoB 2 (Sterne et al., 2019), replaced the original Cochrane tool (Higgins et al., 2011), which rated items such as sequence generation and blinding as low, high or unclear risk. RoB 2 assesses a specific result, such as loneliness at twelve weeks, so one trial can receive different judgements for different outcomes. The reviewer states the effect of interest first: the effect of assignment to the intervention (the intention-to-treat effect, which most reviews of public health programs estimate) or the per-protocol effect of adhering to it. Judgements are reached through signalling questions, factual questions answered "yes", "probably yes", "probably no", "no" or "no information". An algorithm maps the answers in each domain to a proposed judgement of low risk of bias, some concerns or high risk of bias, which the reviewer can override with a written reason. The tabs describe the five domains.
Bias arising from the randomization process. This domain asks whether the allocation sequence was random, whether it was concealed until participants were enrolled and assigned, and whether baseline differences suggest a problem. Allocation concealment means that the people enrolling participants could not know or predict the next assignment, so they could not steer particular people into a group. Example: "Was the allocation sequence concealed until participants were enrolled and assigned to interventions?" In one Cedar Valley trial, a coordinator assigned participants from a list she could see in advance, and the groups differed at baseline in the proportion living alone.
Bias due to deviations from intended interventions. For the effect of assignment, this domain asks whether participants and program staff knew the assignments, whether deviations from the intended intervention arose because of the trial context, and whether the analysis kept participants in the groups to which they were randomized. Example: "Was an appropriate analysis used to estimate the effect of assignment to intervention?"
Bias due to missing outcome data. This domain asks whether outcome data were available for all or nearly all randomized participants and, if not, whether missingness could depend on the true value of the outcome. Example: "Could missingness in the outcome depend on its true value?" The loneliest participants may be the least likely to return a follow-up questionnaire. One Cedar Valley trial lost 35 percent of its participants, with more losses in the comparison group.
Bias in measurement of the outcome. This domain asks whether the measurement method was appropriate, whether it could have differed between groups, and whether knowledge of the assigned intervention could have influenced the assessment. Example: "Were outcome assessors aware of the intervention received by study participants?" For self-reported loneliness, the participant is the outcome assessor and usually knows which group he or she is in. Section 3 shows the decision rule the Cedar Valley reviewers wrote for this domain.
Bias in selection of the reported result. This domain asks whether the result was analyzed according to a plan finalized before unblinded outcome data were available, and whether it may have been selected from several scales, time points or analyses. A registry entry or protocol is the usual evidence for this domain.
The overall judgement follows explicit rules. A result is at low risk of bias overall when all five domains are at low risk. It has some concerns when at least one domain raises some concerns and none is at high risk. It is at high risk when at least one domain is at high risk, or when several domains raise some concerns in a way that substantially lowers confidence in the result. A variant for cluster-randomized trials adds a domain on bias arising from the identification or recruitment of participants into clusters, and a variant for crossover trials addresses carry-over and period effects. The Cedar Valley team used the cluster variant for its 4 cluster-randomized trials. Because the developers revise these tools from time to time, a protocol should name the version used.
2.3 ROBINS-I for Non-randomized Studies of Interventions
The Risk Of Bias In Non-randomised Studies of Interventions tool, ROBINS-I (Sterne et al., 2016), assesses studies that compare people who received a program with people who did not, using as a reference a hypothetical randomized trial that would answer the same question without bias. This target trial is a reference point, and it may be hypothetical even when such a trial would be impractical or unethical (Hernán and Robins, 2016). Describing it makes the reviewer specify the population, intervention, comparator and outcome that the study is trying to estimate. The seven domains are arranged by the stage at which bias arises.
Background: the target trial protocol
A target trial is described with the same elements as the protocol of a real randomized trial (Hernán and Robins, 2016): the eligibility criteria; the treatment strategies compared, such as referral to a connector program or usual care; the assignment procedure, which in the target trial is random; time zero, when eligibility is met, a strategy is assigned and follow-up starts; the follow-up period; the outcome, such as loneliness at twelve weeks; the causal contrast, either the effect of assignment or the effect of adhering to a strategy; and the analysis plan. An observational study designed to mimic each element is said to emulate the target trial.
ROBINS-I asks the reviewer to write the target trial for the review question first and then judges each study by how far it departs from it. Without random assignment, the groups may differ in causes of the outcome, so the confounding domain asks how well the study dealt with them. When eligibility, assignment and the start of follow-up do not coincide, the domain on selection of participants applies. When the strategies are defined with information collected after time zero, the domain on classification of interventions applies. The causal contrast decides how deviations from intended interventions are judged, and the last three domains compare the study's follow-up, outcome measurement and reported analysis with the protocol. HSCI 230 Lesson 3, Section 1 (Introduction to Observational Studies), applies the same reasoning to the design of observational studies and is optional reading.
Bias due to confounding arises when factors that predict the outcome also influence who receives the intervention. Reviewers list the important confounding domains in the protocol before reading the studies, then judge whether each study measured and controlled for them. The Cedar Valley team listed baseline loneliness, depressive symptoms, living alone, mobility or functional limitation, age, and prior social participation. Bias in selection of participants into the study arises when entry into the study or the analysis is related to both intervention and outcome, for example when follow-up begins weeks after people start a program and early dropouts are never counted.
Bias in classification of interventions arises when it is unclear or misrecorded who received the program, particularly when classification is decided with knowledge of the outcome, such as defining "participants" after the results are known as people who attended at least four sessions.
The last four domains parallel RoB 2. Bias due to deviations from intended interventions includes imbalances in co-interventions, which are other services that may affect the outcome, such as home support. Bias due to missing data, bias in measurement of outcomes and bias in selection of the reported result follow the same logic as in RoB 2.
Each domain and the overall result are judged at low risk (comparable to a well-performed randomized trial), moderate risk (sound for a non-randomized study but not comparable to a well-performed randomized trial), serious risk (some important problems) or critical risk (too problematic to provide useful evidence on the effect), with a further option of no information. The overall judgement is generally the most severe domain judgement. Because confounding can rarely be ruled out without randomization, moderate risk is a good result for a non-randomized study.
2.4 ROBINS-E for Studies of Exposures
Some questions concern exposures that nobody assigns, such as air pollution, shift work or living alone. The Risk Of Bias In Non-randomized Studies of Exposures tool, ROBINS-E (Higgins et al., 2024), was developed for follow-up (cohort) studies of such exposures. It adapts the ROBINS-I domains to bias due to confounding, bias arising from measurement of the exposure, bias in selection of participants into the study or analysis, bias due to post-exposure interventions, bias due to missing data, bias arising from measurement of the outcome, and bias in selection of the reported result. Judgements range from low risk through some concerns and high risk to very high risk, with an additional category of low risk except for concerns about uncontrolled confounding, which recognizes that residual confounding can almost never be excluded in exposure studies. The Cedar Valley question concerns interventions, so the team does not use ROBINS-E.
2.5 The Newcastle-Ottawa Scale
The Newcastle-Ottawa Scale (NOS), developed by Wells and colleagues at the Universities of Newcastle (Australia) and Ottawa, appraises cohort and case-control studies and appears in many published reviews. The cohort version has three categories: selection, with four items (representativeness of the exposed cohort, selection of the non-exposed cohort, ascertainment of exposure, and absence of the outcome at the start); comparability, with one item worth up to two stars for control of the most important factor and any additional factor; and outcome, with three items (assessment of the outcome, adequate length of follow-up and adequacy of follow-up). The case-control version replaces outcome with an exposure category. The maximum is nine stars.
Stang (2010) criticized the scale for unclear items and cut-offs with no stated basis, and Hartling and colleagues (2013) found low agreement between individual reviewers using it. Many reviews convert stars into "good", "fair" or "poor" ratings with thresholds that vary between reviews, which reproduces the problem with summary scores that Jüni and colleagues (1999) described. A review that uses the NOS should report the stars awarded on each item and explain what earns each star for its question. For non-randomized studies of interventions, ROBINS-I is the more structured choice.
2.6 The JBI Checklists
JBI (formerly the Joanna Briggs Institute, at the University of Adelaide) publishes a family of critical appraisal checklists for many designs, including randomized trials, quasi-experimental studies, cohort, case-control and analytical cross-sectional studies, prevalence studies, case series, case reports, qualitative research, economic evaluations, and text and opinion papers. Questions are answered "yes", "no", "unclear" or "not applicable". The analytical cross-sectional checklist has eight questions, including questions on valid measurement and on identifying and dealing with confounders. The qualitative checklist has ten, most of which ask about congruity between the study's philosophical perspective, methodology, methods, analysis and interpretation, along with the researcher's influence and the representation of participants' voices. JBI has revised its checklists for randomized and quasi-experimental studies so that their questions are grouped by the type of bias they address (Barker et al., 2023).
The family covers designs, such as prevalence studies and case series, that the Cochrane tools do not. JBI guidance asks reviewers to decide in advance how appraisal results will be used, for example which answers would lead to exclusion or to reporting with caution (Aromataris and Munn, 2020), and reporting the answer to each question is more informative than a count of "yes" answers. The Cedar Valley team uses the JBI quasi-experimental checklist for its 9 uncontrolled before-and-after studies.
The CASP checklists, from the Critical Appraisal Skills Programme, are a widely used alternative family, with checklists for randomized trials, cohort, case-control and qualitative studies, among others; HSCI 230 and HSCI 841 Lesson 2 use them. Whichever family a review uses, it names the checklist it applies to each design in its protocol.
2.7 The Mixed Methods Appraisal Tool
A review that includes qualitative, quantitative and mixed-methods studies may prefer one tool for all of them. The Mixed Methods Appraisal Tool (MMAT), first developed by Pluye and colleagues (2009) and revised as version 2018 by Hong and colleagues (2018), begins with two screening questions: whether the study has clear research questions, and whether the collected data allow those questions to be addressed. It then offers five design categories, each with five criteria answered "yes", "no" or "can't tell": qualitative studies, quantitative randomized controlled trials, quantitative non-randomized studies, quantitative descriptive studies and mixed-methods studies. A mixed-methods study is appraised against the mixed-methods criteria, which ask about the rationale for combining methods and the integration of components, and against the criteria for its qualitative and quantitative components. The developers discourage an overall score and recommend reporting the rating for each criterion. The Cedar Valley team uses the MMAT for its 6 qualitative and 3 mixed-methods studies. The table summarizes the allocation of all 42 studies.
| Design (number of studies) | Tool | Assessment process |
|---|---|---|
| Randomized trials (14, including 4 cluster-randomized) | RoB 2, with the cluster variant where needed, for loneliness at the primary time point | Two reviewers independently, then consensus |
| Non-randomized controlled studies (10) | ROBINS-I, with confounding domains listed in the protocol | Two reviewers independently, then consensus |
| Uncontrolled before-and-after studies (9) | JBI checklist for quasi-experimental studies | One reviewer, checked in full by a second |
| Qualitative studies (6) | MMAT, qualitative category | One reviewer, checked in full by a second |
| Mixed-methods studies (3) | MMAT, mixed-methods and component categories | One reviewer, checked in full by a second |
Name a tool for each study. (a) Twenty seniors' centres are randomly allocated to a peer-led walking group or usual activities. (b) A health authority compares loneliness among adults referred to a connector program with similar adults in clinics without it, adjusting for baseline loneliness. (c) A cohort follows adults aged 65 and older for eight years to see whether hearing loss predicts loneliness. (d) Researchers interview 18 participants about a befriending program. Suggested answers: (a) RoB 2 with the cluster variant; (b) ROBINS-I; (c) ROBINS-E; (d) the JBI qualitative checklist, or the MMAT qualitative category in a review that mixes designs.
What Carries Forward
Each of these tools asks for judgement, and two careful reviewers can answer the same question differently. Section 3 describes the procedures that make appraisal consistent.
Reflection
A review of programs for loneliness among older adults includes the following four studies. (a) An individually randomized trial of a weekly telephone befriending program compared with usual care, with loneliness self-reported by participants at twelve weeks. (b) A study comparing 300 adults referred to a community connector program with 300 adults from clinics without the program, adjusting for age and sex only. (c) An eight-year cohort study of whether living alone predicts later loneliness, which the review includes as background on risk factors. (d) A study that combines a participant survey with interviews about a group exercise program, in a review that includes many designs and wants one tool for qualitative, quantitative and mixed-methods studies. For each study, name an appropriate appraisal tool, identify one domain or criterion most likely to raise concern and explain why, and state the judgement scale or response options the tool uses.
(a) RoB 2 fits a randomized trial. Domain 4, measurement of the outcome, is the likely concern, because participants report their own loneliness and know whether they received calls. The domain and overall judgements are low risk, some concerns or high risk, reached through signalling questions answered yes, probably yes, probably no, no or no information.
(b) ROBINS-I fits a non-randomized study of an intervention. Confounding is the main concern, because clinicians refer people for reasons related to loneliness, and adjusting only for age and sex leaves baseline loneliness, depressive symptoms and living alone uncontrolled; the study is likely at serious risk. ROBINS-I uses low, moderate, serious and critical risk, with a no-information option.
(c) ROBINS-E fits a follow-up study of an exposure that nobody assigns. Confounding by health and income, and measurement of the exposure over eight years, are likely concerns. Its judgements run from low risk through some concerns and high to very high risk, with a category of low risk except for concerns about uncontrolled confounding. The Newcastle-Ottawa Scale is an alternative if the review must match earlier reviews, reported item by item.
(d) The MMAT fits, using the mixed-methods criteria together with the qualitative and quantitative descriptive categories. Integration of the survey and interview components is a likely concern. Each criterion is answered yes, no or can't tell, and no overall score is calculated.
Minimum 20 characters required.
Question 1: Which list gives the five domains of RoB 2?
Question 2: A reviewer is appraising a study that compares older adults referred to a connector program with similar adults who were not referred. Which tool and reference point fit this study best?
Question 3: Which statement about the Newcastle-Ottawa Scale is accurate?
Question 4: How does the 2018 version of the Mixed Methods Appraisal Tool appraise a mixed-methods study?
Applying a Tool Consistently Between Two Reviewers
Learning Objectives for this section
- Explain why two trained reviewers can reach different risk-of-bias judgements for the same study, and why reviews use two independent assessors.
- Prepare for appraisal by reading the tool's guidance, specifying the effect of interest, gathering all reports of each study and running a calibration exercise.
- Write decision rules that make recurring judgements consistent, and record the support for each answer with a quotation, its location and a reason.
- Compare two reviewers' judgements, resolve disagreements by consensus or arbitration, and calculate percent agreement and Cohen's kappa for the judgements.
- Present risk-of-bias results in a traffic-light plot and describe how a review uses them in synthesis without summary scores.
Introduction
The tools in Section 2 structure judgement without removing it. A signalling question such as "Is it likely that assessment of the outcome was influenced by knowledge of intervention received?" asks a reviewer to weigh the outcome, the comparator, the setting and what participants were told. Studies of the tools show that reviewers often disagree. Hartling and colleagues (2013) reported low agreement between individual reviewers using the Newcastle-Ottawa Scale, and Minozzi and colleagues (2020) reported low agreement between reviewers applying RoB 2 and described the difficulties they met in applying it. Much of this disagreement comes from reviewers reading different parts of a report or making different assumptions that can be written down and agreed. The procedures in this section are designed to remove those avoidable sources of disagreement and to make the remaining judgements transparent.
The standard approach is for two reviewers to assess each study independently, compare their answers and reach consensus, with a third person available to arbitrate. The Cedar Valley protocol uses this approach for the 24 comparative studies, which are the studies whose risk of bias most affects what the evidence brief will say about effectiveness, and uses one assessor with full checking by a second for the other 18 studies (Section 2).
3.1 Preparing to Appraise
Preparation prevents most disagreements before they happen. Both reviewers should read the full guidance document for each tool, since the short signalling questions depend on definitions given in the guidance. For RoB 2, the team agrees on the effect of interest before starting, and the Cedar Valley team chose the effect of assignment to the intervention because the planning team wants to know what happens when a program is offered to older adults, including those who attend rarely. For ROBINS-I, the team lists the important confounding domains and co-interventions in the protocol, as Section 2 described. For every study, the reviewers gather all available sources: the main paper, its supplementary files, the trial registry entry, any published protocol or statistical analysis plan, and companion reports identified during extraction. A judgement about selective reporting made without the registry entry is likely to differ from one made with it.
The team then runs a calibration exercise, in which both reviewers independently appraise the same small set of studies, usually two or three per tool, and discuss every difference before appraisal begins. Calibration serves the same purpose as piloting the extraction form in Section 1, and it usually produces the first decision rules. The figure shows the full process.
3.2 Writing Decision Rules
A decision rule is a written agreement about how the team will answer a signalling question in a situation that recurs across studies. Decision rules leave the tool itself unchanged and record how the team interprets the tool's guidance for its own topic, so that the same situation receives the same answer in every study and so that readers can see the reasoning. The rules belong in an appendix to the review or in a protocol amendment. The table shows the rules the Cedar Valley team wrote during calibration and the first round of comparison.
| Tool and domain | Recurring situation | Cedar Valley decision rule |
|---|---|---|
| RoB 2, domain 4 (outcome measurement) | Loneliness is self-reported, and participants know which group they are in. | Answer 4.3 "yes" and 4.4 "probably yes" for every self-reported loneliness outcome. Answer 4.5 "probably no" (which leads to some concerns) unless the report gives a specific reason to expect that knowledge of assignment shaped responses, such as a comparison group told it was on a waiting list or recruitment materials promising less loneliness, in which case answer "probably yes" (which leads to high risk). |
| RoB 2, domain 3 (missing outcome data) | Some participants did not complete the follow-up questionnaire. | Treat data as available for nearly all participants when at least 95 percent of randomized participants are analyzed. The RoB 2 guidance explains that what counts as nearly all depends on the outcome and context, so the team states its threshold. |
| RoB 2, domain 5 (selection of the reported result) | No registry entry, protocol or analysis plan can be found. | Answer 5.1 "no information". If there is no sign of selection among several scales, time points or analyses, the domain is judged as some concerns. |
| RoB 2 cluster variant, domain 1b | Participants in cluster trials are recruited after centres are randomized. | Record who recruited participants and whether they knew the centre's allocation. Recruitment after randomization by people who knew the allocation leads to at least some concerns. |
| ROBINS-I, confounding | A study adjusts for some confounding domains listed in the protocol. | A study that does not control for baseline loneliness, depressive symptoms and living alone is at serious risk of confounding unless it shows that the groups were similar on the omitted factors. |
The first rule came from the most common disagreement in calibration. In two trials whose comparison groups attended a different social activity of similar length, the student intern had answered 4.4 "probably no", reasoning that both groups expected some benefit, and so had judged domain 4 at low risk. The evidence officer had answered "probably yes", because a participant who knows that she is in the "new program" group may still report her loneliness differently. The team agreed that the active comparator is better considered at question 4.5, where it reduces the likelihood of influence without removing the possibility, and wrote the rule accordingly.
3.3 Recording the Support for Each Judgement
Every answer should be supported by a record that another reader could check: a short quotation or summary from the source, its location, and the reviewer's reason when the answer involves judgement. RoB 2 and ROBINS-I provide a free-text box for this purpose beside each signalling question. The support record turns an opinion into an argument, and it makes consensus meetings faster because reviewers can see where their sources differed. The example shows the support record for domain 3 of one Cedar Valley trial, labelled Trial G.
| Signalling question | Answer | Support (source, location and reason) |
|---|---|---|
| 3.1 Were data for this outcome available for all, or nearly all, randomized participants? | No | Of 354 participants recruited in the 12 randomized centres, 230 (65 percent) completed the twelve-week questionnaire: 133 of 177 in the intervention centres and 97 of 177 in the comparison centres (Table 2, page 6). |
| 3.2 Is there evidence that the result was not biased by missing outcome data? | No | Only a complete-case analysis is reported, with no sensitivity analysis for missing data (page 7). |
| 3.3 Could missingness in the outcome depend on its true value? | Probably yes | Reasons for dropout are not reported, and lonelier participants may be less willing to return a questionnaire. |
| 3.4 Is it likely that missingness in the outcome depended on its true value? | Probably yes | Losses were almost twice as high in the comparison group (80 compared with 44), consistent with disengagement among people offered no program. |
| Domain judgement | High risk | The answers follow the path in the RoB 2 algorithm that leads to high risk of bias for this domain. |
3.4 Comparing Judgements and Reaching Consensus
After independent assessment, the two reviewers compare their answers question by question and their judgements domain by domain. It helps to classify each disagreement. A disagreement of fact occurs when one reviewer found information that the other missed, such as a registry entry or a table in a supplementary file, and it is resolved by looking at the source together. A disagreement of interpretation occurs when both reviewers saw the same information and answered differently, and it is resolved by discussion with reference to the guidance. If the same interpretive disagreement recurs, the team writes a decision rule and rechecks the studies already assessed. If the reviewers cannot agree, a third person with methodological experience arbitrates, and the record notes that arbitration was used. When the report lacks information that would change a judgement, the team can write to the authors, using the same log as in Section 1.
Many reviews report agreement between the reviewers' independent judgements, before consensus, as an indicator of how difficult the appraisal was. Lesson 7 introduced percent agreement and Cohen's kappa for screening decisions, and the same statistics apply here. In the Cedar Valley review, the two reviewers made 5 domain judgements for each of 14 trials, giving 70 judgements. They disagreed on 13, so they agreed on 57 of 70 (81.4 percent), and 7 of the 13 disagreements were in domain 4. Their independent overall judgements for the 14 trials are cross-tabulated below.
| Intern (rows) by evidence officer (columns) | Low risk | Some concerns | High risk | Total |
|---|---|---|---|---|
| Low risk | 0 | 2 | 0 | 2 |
| Some concerns | 0 | 6 | 2 | 8 |
| High risk | 0 | 1 | 3 | 4 |
| Total | 0 | 9 | 5 | 14 |
Agreement on the overall judgements
Observed agreement, po = (0 + 6 + 3) ÷ 14 = 9 ÷ 14 = 0.643, or 64.3 percent.
Agreement expected by chance, pe = [(2 × 0) + (8 × 9) + (4 × 5)] ÷ 142 = 92 ÷ 196 = 0.469.
Cohen's kappa, κ = (po − pe) ÷ (1 − pe) = (0.643 − 0.469) ÷ (1 − 0.469) = 0.174 ÷ 0.531 = 0.33.
On the benchmarks of Landis and Koch (1977), a kappa between 0.21 and 0.40 indicates fair agreement. The two trials that the intern rated at low risk overall were the attention-control trials described in Section 3.2, and the decision rule resolved both.
A kappa of 0.33 is lower than the 81.4 percent agreement on domain judgements might suggest, for two reasons. The overall judgement takes the most severe domain judgement into account, so a single disagreement in any of five domains can change it. And because most trials fell into one category, chance agreement is high, which lowers kappa for a given level of observed agreement. After consensus, the 14 trials were judged at some concerns (9 trials) or high risk of bias (5 trials) for loneliness, and none reached low risk overall. This pattern is common in reviews of psychosocial programs, because participants report their own loneliness and know which program they received, so domain 4 rarely reaches low risk.
3.5 Common Errors in Applying the Tools
Several errors recur in published appraisals, and the Cedar Valley team reviewed them during calibration.
3.6 Presenting Risk of Bias and Using It in Synthesis
Risk-of-bias results are usually presented in two figures. A traffic-light plot shows each study as a row and each domain as a column, with a coloured symbol for each judgement. A summary plot shows, for each domain, the proportion of studies at each level of risk. Reviewers can draw these figures by hand, in a spreadsheet, or with robvis, a free tool released as an R package and a web application (McGuinness and Higgins, 2021). The plot below shows the consensus judgements for the 8 Cedar Valley trials of group-based programs, labelled Trial A to Trial H.
The review then puts the assessment to use. It should present each study's judgement beside its results, so that readers can see whether the studies at higher risk show different effects. It can stratify or order studies by risk of bias in tables and figures (Lesson 9 shows how), and a review with a meta-analysis can run a sensitivity analysis restricted to studies at lower risk (HSCI 230 Lesson 2). Excluding studies because of their risk of bias is defensible only when the protocol specified the rule in advance, since an exclusion decided after seeing the results can be used to steer the conclusions. Finally, the judgements feed the risk-of-bias domain of GRADE, which Section 4 describes.
A Cedar Valley trial randomizes 140 older adults to a weekly telephone befriending call or to a waiting list, and participants complete the three-item loneliness scale themselves at twelve weeks. The recruitment flyer said that the program "helps people feel less lonely". Using the team's decision rule from Section 3.2, answer signalling questions 4.3, 4.4 and 4.5 and state the domain judgement. Suggested answer: 4.3 is "yes" because the participants are the assessors and know their group; 4.4 is "probably yes" because loneliness is subjective; and 4.5 is "probably yes" because the comparison group knew it was waiting and the flyer set an expectation of benefit, so the domain is at high risk of bias. Write the support record for 4.5 in one or two sentences.
What Carries Forward
The team now has a consensus judgement, with its support, for every comparative study. Those judgements describe individual studies. Section 4 asks a different question: taking all the studies of one intervention and one outcome together, how confident can the team be in what they show? GRADE answers that question for each outcome.
Reflection
Two reviewers independently applied RoB 2 to ten randomized trials of group programs whose outcome was loneliness, self-reported by participants who knew their group. For two trials whose comparison groups attended a different social activity of similar length, Reviewer 1 answered signalling question 4.4 ("Could assessment of the outcome have been influenced by knowledge of intervention received?") as "probably no" and judged domain 4 at low risk; Reviewer 2 answered "probably yes". In RoB 2, a "probably yes" to 4.4 leads to question 4.5 ("Is it likely that assessment of the outcome was influenced by knowledge of intervention received?"), where "probably no" gives some concerns and "probably yes" gives high risk. Their independent overall judgements were: Reviewer 1 rated 2 trials low risk, 6 some concerns and 2 high risk; Reviewer 2 rated 0 low risk, 7 some concerns and 3 high risk. They agreed on 5 trials rated some concerns and 2 rated high risk, and on none rated low risk. (a) Write a decision rule for domain 4 that would resolve the disagreement and apply to future trials. (b) Describe what each reviewer should record as support for question 4.5. (c) Calculate observed agreement and Cohen's kappa, using kappa = (observed agreement − chance agreement) ÷ (1 − chance agreement), where chance agreement is the sum across categories of (Reviewer 1 proportion × Reviewer 2 proportion). (d) Interpret the result.
(a) Rule: for self-reported loneliness, answer 4.3 "yes" and 4.4 "probably yes", since participants are the assessors and the outcome is subjective. Consider the comparator at 4.5: answer "probably no" (some concerns) when the comparison group received an active program of similar contact, and "probably yes" (high risk) when it was on a waiting list or the trial promoted an expectation of benefit. Under this rule both disputed trials have some concerns in domain 4.
(b) Each reviewer records the comparator as described, a quotation about what participants were told, the page or table, and a sentence explaining why influence is or is not likely.
(c) Observed agreement = (0 + 5 + 2) ÷ 10 = 0.70. Chance agreement = (0.2 × 0.0) + (0.6 × 0.7) + (0.2 × 0.3) = 0 + 0.42 + 0.06 = 0.48. Kappa = (0.70 − 0.48) ÷ (1 − 0.48) = 0.22 ÷ 0.52 = 0.42.
(d) A kappa of 0.42 indicates moderate agreement on the Landis and Koch benchmarks. It is lower than 70 percent agreement suggests because most trials fall in one category, which raises chance agreement. The disagreements came from one interpretive question, which the rule should remove, and the team should recheck trials already assessed against it.
Minimum 20 characters required.
Question 1: One reviewer finds a trial registry entry that the other reviewer missed, and their answers to a domain 5 question differ as a result. How should the team classify and resolve this disagreement?
Question 2: Two reviewers' independent overall judgements for 14 trials agree on 9, and chance agreement calculated from their marginal totals is 0.469. What is Cohen's kappa, to two decimal places?
Question 3: A trial paper does not describe how allocation was concealed, and no protocol or registry entry can be found. How should the reviewer answer the RoB 2 signalling question about allocation concealment?
Question 4: Why did the Cedar Valley team write a decision rule for domain 4 of RoB 2?
Rating the Certainty of a Body of Evidence with GRADE
Learning Objectives for this section
- Explain what the certainty of evidence means in GRADE, and why it is rated for each outcome across a body of studies.
- State the starting level of certainty for randomized and non-randomized evidence, and describe the alternative start used with ROBINS-I.
- Describe the five reasons for rating certainty down (risk of bias, inconsistency, indirectness, imprecision and publication bias) and the three reasons for rating it up (a large effect, a dose-response gradient, and plausible confounding that would reduce the effect).
- Build a GRADE evidence profile for a body of evidence without a pooled estimate, and write an informative statement that matches each level of certainty.
- Explain how certainty ratings inform an evidence brief, and plan the extraction, appraisal and certainty steps of a rapid scoping review.
Introduction
Risk-of-bias tools describe single studies. A decision-maker asks a different question: taking all the relevant studies together, how confident can we be about the effect of this intervention on this outcome? The Grading of Recommendations Assessment, Development and Evaluation approach, known as GRADE, answers that question with a rating of the certainty of evidence for each outcome. GRADE was developed by an international working group (Guyatt et al., 2008) and is used by Cochrane, the World Health Organization and many guideline developers. Students who take HSCI 230 Lesson 2, before or after this course, meet GRADE there in a short call-out; this section shows how a review team applies it.
Three features of GRADE shape its use. Certainty is rated for a body of evidence, which means all the studies that address one comparison and one outcome, so the same review can have high certainty for one outcome and very low certainty for another. Certainty is rated for outcomes that matter to the people who will use the review, which for the Cedar Valley planning team means loneliness first. And GRADE separates the certainty of evidence from the strength of any recommendation that follows from it. Guideline panels and decision-makers make recommendations after weighing certainty together with benefits, harms, costs, values and feasibility, so a recommendation can be strong even when certainty is low.
4.1 Four Levels of Certainty
GRADE uses four levels. The descriptions below paraphrase those given by Balshem and colleagues (2011), and the symbols are those used in GRADE tables.
| Level | Symbol | Meaning |
|---|---|---|
| High | ⊕⊕⊕⊕ | The review team is very confident that the true effect lies close to the estimate. |
| Moderate | ⊕⊕⊕◯ | The team is moderately confident: the true effect is likely to be close to the estimate, but it could be substantially different. |
| Low | ⊕⊕◯◯ | The team's confidence is limited: the true effect may be substantially different from the estimate. |
| Very low | ⊕◯◯◯ | The team has very little confidence: the true effect is likely to be substantially different from the estimate. |
GRADE originally called these levels the quality of evidence. The current term, certainty, makes clear that the rating describes confidence in an estimate of effect, which can be low even when every study was carefully conducted.
4.2 Where the Rating Starts
The rating begins from the design of the studies. A body of evidence from randomized trials starts at high certainty, because randomization protects against confounding. A body of evidence from non-randomized studies, including cohort studies, controlled before-and-after studies and other observational designs, starts at low certainty, because confounding and selection can rarely be excluded. The team then considers reasons to rate the evidence down and, less often, reasons to rate it up, as the figure shows.
A second starting point exists for teams that use ROBINS-I. Because ROBINS-I compares each study with a target trial, the GRADE Working Group has described an approach in which non-randomized evidence assessed with ROBINS-I starts at high certainty and is then rated down for risk of bias by as many levels as the ROBINS-I judgements warrant (Schünemann et al., 2019). Studies at serious risk of bias usually bring the evidence to low certainty or below under this approach, so the two starting points tend to reach similar results. The Cedar Valley protocol states that it uses the conventional low starting point, which is simpler to explain to the planning team.
Two kinds of evidence in the Cedar Valley review sit outside this profile. The 9 uncontrolled before-and-after studies have no comparison group, so they cannot separate the program's effect from change over time, and the team reports them as supporting information. The 6 qualitative studies and the qualitative parts of the 3 mixed-methods studies describe participants' experiences and are assessed with GRADE-CERQual, the companion approach for qualitative evidence that Lesson 10 introduces.
4.3 Five Reasons to Rate Down
For each domain, the team judges whether there is no serious concern, a serious concern (rate down one level) or a very serious concern (rate down two levels). Certainty cannot fall below very low, and the same problem should be counted in only one domain. The accordion describes each domain.
This domain draws on the study-level judgements from Sections 2 and 3, considered across the body of evidence. The question is whether the limitations of the studies, weighted by how much each study contributes, lower confidence in the overall result. If most of the information comes from studies with some concerns or high risk of bias, the team rates down. It helps to check whether studies at high risk show different results from the others, since an effect that appears only in studies at high risk warrants more concern.
Inconsistency is unexplained variation in results across studies. Signs include point estimates that differ widely, confidence intervals that barely overlap, and, in a meta-analysis, statistical measures of heterogeneity such as I-squared, which HSCI 230 Lesson 2 teaches. When variation can be explained by a pre-specified characteristic, such as program length or setting, the team can rate the subgroups separately and need not rate down. When studies differ in the size of a benefit but agree on its direction, the concern is smaller than when some show benefit and others show harm.
Indirectness arises when the evidence differs from the review question in its population, intervention, comparator or outcome. Examples include trials in residential care when the question concerns older adults living at home, or a measure of social contact used in place of loneliness.
Imprecision is random error. The team asks whether the confidence interval, or the range of results across studies, is narrow enough to support a decision. If it includes both an important benefit and no effect, the evidence is imprecise. GRADE also uses the idea of an optimal information size: the number of participants that a single adequately powered trial would need. When the total number of participants in the body of evidence falls short of it, the team usually rates down even if the result is statistically significant.
Publication bias arises when the studies that reach the review differ systematically from those conducted, usually because studies with favourable results are published more often or more quickly. Suspicion rises when the evidence consists of a few small studies with positive results, when registered trials have not reported, or when a funnel plot is asymmetric. Funnel plots and tests for asymmetry, taught in HSCI 230 Lesson 2, are generally not used with fewer than ten studies. Searching trial registries and grey literature (Lessons 4 and 5) reduces the problem. GRADE describes publication bias as undetected or strongly suspected, and a team usually rates down by one level when it is strongly suspected.
4.4 Three Reasons to Rate Up
GRADE allows the certainty of non-randomized evidence to be rated up in three circumstances (Guyatt et al., 2011). Rating up is usually considered only when no domain has been rated down, since a large effect in studies at serious risk of bias may itself be a product of bias.
4.5 Rating Certainty Without a Pooled Estimate
GRADE is often described in terms of a meta-analysis, with a pooled estimate and its confidence interval. Many reviews, including the Cedar Valley review, synthesize results without pooling them, using the methods and the SWiM reporting guideline that Lesson 9 teaches. Murad and colleagues (2017) described how to rate certainty in this situation. The same five domains apply, but the team judges them from the pattern of results across studies: the spread of effects and their direction for inconsistency, the precision of individual studies and the total number of participants for imprecision, and so on. The team explains each judgement in a footnote, since a reader cannot check it against a single number. Rating certainty in a scoping review is unusual, because scoping reviews are designed to map evidence and seldom estimate effects. The Cedar Valley protocol planned no certainty ratings, and a dated amendment in week 8 added them, limited to loneliness for three intervention categories, because the planning team asked directly how confident it could be; the brief labels the ratings as provisional.
4.6 Worked Example: The Cedar Valley Evidence Profile
The team rated certainty for one outcome, loneliness measured with a validated self-report scale at the end of the program, and for three comparisons: group-based programs compared with usual activities or no program (8 randomized trials, 1,236 participants), one-to-one befriending or telephone programs compared with usual care or a waiting list (6 randomized trials, 742 participants), and community connector (social prescribing) programs compared with usual care (5 non-randomized controlled studies, 1,480 participants). The other 5 non-randomized controlled studies evaluated intergenerational and technology-based programs and were too few in each category to rate. All results described here are illustrative and belong to the fictional case.
An evidence profile sets out the judgement for each domain, the resulting certainty and a footnote explaining each decision.
| Comparison (studies; participants) | Risk of bias | Inconsistency | Indirectness | Imprecision | Publication bias | Certainty |
|---|---|---|---|---|---|---|
| Group-based programs (8 randomized trials; 1,236) | Seriousa | Seriousb | Not serious | Not serious | Undetectedc | ⊕⊕◯◯ Low |
| One-to-one befriending or telephone (6 randomized trials; 742) | Seriousd | Not serious | Not serious | Seriouse | Undetectedc | ⊕⊕◯◯ Low |
| Community connector programs (5 non-randomized studies; 1,480) | Seriousf | Not serious | Not serious | Not serious | Undetectedc | ⊕◯◯◯ Very low |
a All 8 trials have some concerns (6) or high risk of bias (2) for loneliness, mainly because participants who know their group report their own loneliness (Section 3). b Five trials reported small reductions in loneliness, two reported little or no difference, and one reported a larger reduction; program length and setting did not explain the variation. c Fewer than ten studies per comparison, so funnel plots were not used; registry searches found no completed but unreported trials. d Three of the 6 trials are at high risk of bias and three have some concerns; their results were similar, so the team rated down one level. e The total sample is modest, and most trials' confidence intervals include both no effect and a meaningful reduction. f Three of the 5 studies are at serious risk of confounding by ROBINS-I, because clinicians referred people they judged likely to benefit; the other two are at moderate risk.
Group-based programs start at high certainty as randomized evidence. The team rated down one level for risk of bias and one for inconsistency, giving low certainty. It judged indirectness not serious because all trials enrolled community-dwelling adults aged 65 and older, and imprecision not serious because the trials together enrolled more than a thousand participants. One-to-one programs also start at high, and the team rated down one level for risk of bias and one for imprecision, again giving low certainty. Community connector programs start at low certainty as non-randomized evidence and were rated down one level for risk of bias, giving very low certainty. Two of these studies showed larger reductions in loneliness among people with more connector contacts, and the team considered rating up for a dose-response gradient. It did not rate up, because the risk of bias was serious and people whose loneliness improved early may have chosen to keep meeting their connector, which would produce the same gradient.
4.7 Communicating Certainty
Certainty ratings help decision-makers when the review states them in consistent, plain terms. Santesso and colleagues (2020) proposed standard wording for the statements in reviews and summaries, linking each level of certainty to a verb phrase. The Cedar Valley evidence brief uses this wording, as the table shows.
| Certainty | Wording pattern | Cedar Valley statement |
|---|---|---|
| High | The intervention reduces the outcome. | No Cedar Valley comparison reached high certainty. |
| Moderate | The intervention probably reduces the outcome. | No Cedar Valley comparison reached moderate certainty. |
| Low | The intervention may reduce the outcome. | Group-based programs may reduce loneliness slightly among older adults living in the community. One-to-one befriending or telephone programs may reduce loneliness. |
| Very low | The evidence is very uncertain about the effect of the intervention on the outcome. | The evidence is very uncertain about the effect of community connector programs on loneliness. |
In a full review, these statements appear in a summary of findings table, which presents, for each important outcome, the number of studies and participants, the effect and the certainty. Software such as GRADEpro GDT, available in 2026, helps teams build these tables, and Lesson 12 shows how to present them to decision-makers.
The Cedar Valley ratings have a direct implication for the planning team. The community connector model it intends to launch rests on very uncertain evidence about loneliness, so the true effect could be substantially larger or smaller than existing studies suggest. That rating leaves open whether the program works, and it strengthens the case for launching the program with a built-in evaluation, for example by staggering its introduction across clinics, which the evidence brief in Lesson 12 recommends.
Reflection
A review team rates the certainty of evidence for one comparison: intergenerational programs compared with usual activities, outcome loneliness among older adults living in the community. The evidence is 4 randomized trials with 510 participants. Risk of bias for loneliness: 1 trial at low risk, 2 with some concerns and 1 at high risk, with most participants in the trials that have some concerns or high risk. Results: two trials found moderate reductions in loneliness and two found slight increases, and no planned subgroup explains the difference. Two of the four trials were conducted in residential care homes. The confidence interval of every trial includes no effect. There are fewer than ten trials, and a registry search found no completed but unreported trials. In GRADE, randomized evidence starts at high certainty, each serious concern rates it down one level and each very serious concern two levels, and certainty cannot fall below very low. Rate each of the five domains with a reason, state the final certainty, and write a plain-language statement using the GRADE wording for that level. Then explain what the rating does and does not tell a planning team.
Risk of bias: serious, because most of the information comes from trials with some concerns or high risk, mainly from self-reported loneliness in unblinded trials (rate down one level). Inconsistency: serious, because two trials suggest benefit and two suggest slight harm, with no explanation (rate down one). Indirectness: serious, because half the trials enrolled residents of care homes, while the question concerns older adults living in the community (rate down one). Imprecision: serious, because the total sample is modest and every confidence interval includes no effect. Publication bias: undetected, because funnel plots are not informative with four trials and the registry search found no unreported trials.
Starting at high, the first three domains already bring the rating to very low, and imprecision cannot lower it further, so the certainty is very low. Statement: "The evidence is very uncertain about the effect of intergenerational programs on loneliness among older adults." A team could defensibly judge risk of bias not serious, given one low-risk trial, and the rating would still be very low.
The rating tells the planning team that the true effect could be substantially different from what these trials suggest, in either direction. The rating leaves open whether intergenerational programs work, so a team that adopted one would need to evaluate it.
Minimum 20 characters required.
Question 1: In the conventional GRADE approach, at what level does certainty start for a body of evidence from randomized trials and from non-randomized studies?
Question 2: Which list contains only reasons for rating certainty down in GRADE?
Question 3: A body of evidence from eight randomized trials is rated down one level for risk of bias and one level for inconsistency, with no other concerns. What is its certainty, and which statement fits it?
Question 4: Why did the Cedar Valley team decline to rate up the community connector evidence for a dose-response gradient?
Final Assessment
Bringing It All Together
This lesson followed the Cedar Valley evidence team from a list of 42 included studies to a statement about how confident the planning team can be in what those studies show. The extraction form, built on the TIDieR items and a data dictionary and piloted until two people recorded the same information, determined what the review could later compare and appraise. The design recorded on that form decided which appraisal tool applied: RoB 2 for the randomized trials, ROBINS-I for the non-randomized controlled studies, the JBI quasi-experimental checklist for the uncontrolled before-and-after studies, and the MMAT for the qualitative and mixed-methods studies, with ROBINS-E and the Newcastle-Ottawa Scale described for the exposure and observational questions where students will meet them.
The tools structure judgement without removing it, which is why calibration, decision rules, recorded support, independent assessment and consensus matter. In the Cedar Valley review, no trial reached low risk of bias overall for loneliness, because participants report their own loneliness and know which program they received. GRADE then moved the question from single studies to bodies of evidence: group-based and one-to-one programs reached low certainty, and community connector programs, the model the health authority intends to launch, reached very low certainty. That rating leaves the program's effect open, which supports launching it with an evaluation built in.
Key Takeaways from this lesson
- Data extraction records the same details from every included study on a standard form, and the form determines what the review can later compare, appraise and report.
- A data dictionary with definitions, allowed values and rules for unusual cases, together with a pilot by two extractors, makes extraction consistent and shows which fields need revision.
- Risk of bias concerns systematic error from a study's design or conduct, and it is assessed separately from reporting quality, imprecision and applicability.
- RoB 2 assesses a specific result from a randomized trial across five domains, using signalling questions and an algorithm to reach judgements of low risk, some concerns or high risk.
- ROBINS-I compares a non-randomized study of an intervention with a target trial across seven domains, and ROBINS-E adapts that approach to follow-up studies of exposures.
- The Newcastle-Ottawa Scale is widely used for cohort and case-control studies, but its summary stars have known problems, and item-level reporting is preferable.
- The JBI checklists cover many designs that the Cochrane tools do not, and the MMAT offers one tool for reviews that mix qualitative, quantitative and mixed-methods studies.
- Consistent appraisal depends on calibration, written decision rules, recorded support for each answer, independent assessment by two reviewers and a consensus process.
- GRADE rates certainty for each outcome across a body of evidence, starting at high for randomized and low for non-randomized evidence, rating down for five reasons and up for three.
- Certainty ratings should be communicated with standard wording, so that decision-makers can tell what the evidence shows and how much weight it can bear.
Core Concepts Reviewed
Section 1: data extraction and charting, companion reports, the domains of an extraction form, TIDieR, the data dictionary, piloting and percent agreement, and dual extraction compared with single extraction and verification.
Section 2: risk of bias compared with reporting quality and imprecision, the problems with summary scores, RoB 2 and its five domains, ROBINS-I and the target trial, ROBINS-E, the Newcastle-Ottawa Scale, the JBI checklists and the MMAT.
Section 3: preparation and calibration, decision rules, support records, disagreements of fact and of interpretation, consensus and arbitration, Cohen's kappa for appraisal judgements, traffic-light plots and the use of risk of bias in synthesis.
Section 4: certainty of evidence, the four GRADE levels, starting levels, the five domains for rating down and three for rating up, rating without a pooled estimate, evidence profiles and informative statements.
The final reflection asks you to write the extraction, appraisal and certainty methods for a new rapid scoping review, drawing on all four sections.
Reflection
A health authority asks a two-person evidence team for a twelve-week rapid scoping review of walking-group programs for loneliness among adults aged 65 and older. Preliminary screening suggests the included studies will be randomized trials (some cluster-randomized), non-randomized controlled studies, uncontrolled before-and-after studies and qualitative interview studies. The authority wants to know what has been studied and how confident it can be that walking groups reduce loneliness. Write the methods text (about 200 words) for the extraction, appraisal and certainty steps of the protocol. It should state how data will be extracted and checked, whether studies will be appraised and why, which tool will be used for each design, how the two reviewers will make their judgements consistent, how the appraisal will be used in the synthesis, and whether and how GRADE will be applied, including the starting levels. End with the sentence the team would use in the brief if the randomized evidence were rated low certainty for a small reduction in loneliness.
Data will be charted on a form with a data dictionary, structured by the TIDieR items for intervention fields and piloted by both reviewers on five varied studies before full charting. One reviewer will chart descriptive fields and the second will check them in full; outcome results from comparative studies will be extracted independently by both. Although appraisal is optional in scoping reviews, studies will be appraised because the authority must judge how far to trust reported outcomes. Randomized trials will be assessed with RoB 2 (with the cluster variant where needed), non-randomized controlled studies with ROBINS-I against the confounding domains listed in this protocol, uncontrolled before-and-after studies with the JBI quasi-experimental checklist, and qualitative studies with the JBI qualitative checklist or the MMAT. Both reviewers will calibrate on two studies per tool, record decision rules and supporting quotations, assess comparative studies independently and resolve disagreements by consensus, with a third person arbitrating. Judgements will be reported by domain in a traffic-light plot, without summary scores, and no study will be excluded on that basis. Certainty for loneliness will be rated with GRADE, starting at high for randomized and low for non-randomized evidence, using guidance for evidence without a pooled estimate. Brief wording: "Walking-group programs may reduce loneliness slightly among older adults."
Minimum 30 characters required.
Final Knowledge Assessment
Question 1: A review team wants to describe each included study's intervention in enough detail to compare programs. Which resource gives the most suitable structure for those extraction fields?
Question 2: A Cedar Valley trial reports self-reported loneliness and emergency department visits taken from health records. Why might RoB 2 give these two results different judgements?
Question 3: Which pairing of appraisal tool and study design is appropriate?
Question 4: How do risk of bias and imprecision differ?
Question 5: A scoping review team asks whether it must appraise its included studies. Which answer reflects PRISMA-ScR and JBI guidance for scoping reviews?
Question 6: In ROBINS-I, which description matches a study judged at moderate risk of bias overall?
Question 7: Which finding would justify rating down the Cedar Valley evidence for indirectness?
Question 8: After seeing the results, a team decides to exclude all studies at high risk of bias from its synthesis. Why is this a problem?
Question 9: Which feature distinguishes ROBINS-E from ROBINS-I?
Question 10: In an extraction pilot, overall agreement is 93 percent, but the primary time-point field disagrees in most studies. What should the team do?
Question 11: Which statement about funnel plots in GRADE's publication bias domain matches the lesson?
Question 12: Why do RoB 2 assessments of community programs for loneliness rarely reach low risk in domain 4?
Question 13: An evidence brief states: "The evidence is very uncertain about the effect of community connector programs on loneliness." What does this wording tell the planning team?
Question 14: Why should a review avoid adding Newcastle-Ottawa stars or checklist answers into a single score?
Question 15: Starting from low certainty, which situation would most plausibly justify rating non-randomized evidence up by one level?
Glossary: Key Terms, People & Frameworks
📚 Reference page, available throughout the lesson
These terms, tools and people appear in this lesson on data extraction, risk of bias and certainty of evidence.