# Lesson 8: Data Extraction, Risk of Bias and Certainty of Evidence

*Companion-podcast transcript, Sarah and Kiffer*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This is the episode for Lesson eight of Health Sciences two forty-one. Once a review team has its included studies, it has to take the same information out of every study, judge how far each study can be trusted, and then ask how confident it can be in the evidence as a whole. That is data extraction, risk of bias and the certainty of evidence.

**Sarah:** Before we start, remind listeners where the Cedar Valley team is.

**Kiffer:** The Cedar Valley evidence review is our fictional running case. A health authority in British Columbia is planning a community connector program, a form of social prescribing, for adults aged sixty-five and older. It asked a small evidence team for a rapid scoping review and an environmental scan within twelve weeks, on which community-based interventions have been evaluated for loneliness or social isolation in that age group, and with what outcomes.

**Sarah:** And by now they have finished screening.

**Kiffer:** The database searches found two thousand four hundred and eighty records. After removing six hundred and ten duplicates, they screened one thousand eight hundred and seventy titles and abstracts, read one hundred and forty-two full texts, and included thirty-eight studies. Citation chasing added four more, so there are forty-two included studies. Grey literature searching also added twenty-six documents, which are mostly short program evaluation reports.

**Sarah:** What kinds of studies are the forty-two?

**Kiffer:** For this lesson, fourteen are randomized trials, four of them randomized by cluster, such as by seniors' centre. Ten are non-randomized controlled studies, nine are before-and-after studies without a comparison group, six are qualitative studies and three are mixed-methods studies. That mix is the reason we need several appraisal tools later on.

**Sarah:** Let's start with section one, then. What is data extraction?

**Kiffer:** Data extraction is the systematic recording of the same set of details from every included study onto a standard form. In a scoping review, JBI, the organization formerly called the Joanna Briggs Institute, calls this step data charting, and Lesson ten comes back to that. The design principles are the same.

**Sarah:** It sounds like clerical work. What does the form actually do?

**Kiffer:** Much of it is clerical, but the form decides what the review can say later. It does four jobs. It describes the included studies, which becomes the characteristics table in the final report. It collects what the synthesis needs, such as the intervention category and the results. It records the details that the risk-of-bias tools will ask about, such as how participants were allocated. And it leaves an audit trail, so a reader can see where every number came from.

**Sarah:** The lesson makes a point about the unit of extraction.

**Kiffer:** The unit is the study, and one study can produce several reports. There might be a registry entry, a protocol, a main results paper and a later paper on long-term follow-up. We call those companion reports. The Cochrane Handbook advises collating them so that each study is counted once and its information is complete. If you treat a main paper and its follow-up as two studies, you count the same participants twice.

**Sarah:** What goes on the form?

**Kiffer:** Most forms follow the elements of the review question. There is a source section, a check that the study really meets the eligibility criteria, then methods, participants, the intervention and comparator, outcomes, results, and a few other fields such as funding and conflicts of interest.

**Sarah:** Which part needs the most thought?

**Kiffer:** The intervention fields, because community programs for loneliness vary a great deal. The Cedar Valley team used the Template for Intervention Description and Replication, known as TIDieR, a twelve-item checklist that Tammy Hoffmann and colleagues published in twenty fourteen. Its items, such as who delivered the program, how, where, how often and whether it was tailored, became fields on the form.

**Sarah:** And then there is the data dictionary.

**Kiffer:** A data dictionary sits beside the form and defines every field: what it means, the allowed values, rules for unusual cases, and an example. Without one, two people will fill in the same field differently. A field like participants, age, gender and number, written as free text, is a good example of what goes wrong. It holds three facts, and every extractor will record them in a different order.

**Sarah:** The lesson has a box about three answers that mean different things.

**Kiffer:** Not reported, not applicable and zero. Not reported means the study should have told you something and did not, such as how many people dropped out. Not applicable means the field does not apply, such as the number of clusters in an individually randomized trial. Zero is a reported value. Left as blank cells, the three cannot be told apart later.

**Sarah:** How does a team know the form works?

**Kiffer:** It pilots the form. Two people use the draft on a small and varied set of included studies, compare what they recorded, and revise the form where they differed. The Cedar Valley intern and the evidence officer each extracted the same five studies: an individually randomized trial, a cluster trial, a non-randomized study, a qualitative study and a mixed-methods study.

**Sarah:** And how did it go?

**Kiffer:** The form had twenty-four fields that could be compared directly, so five studies gave one hundred and twenty paired entries. They disagreed on seventeen, so they agreed on one hundred and three, which is about eighty-six percent.

**Sarah:** Is that good enough?

**Kiffer:** There is no universal threshold, and the total matters less than the pattern. Six of the disagreements were about setting, because two studies mixed people living at home with people in assisted living. Five were about the time point, because studies reported several follow-ups. Four were about the intervention category, because one program combined group activities with one-to-one visits. Two were plain copying errors. Each of the first three had a cause the team could fix with a rule.

**Sarah:** So they wrote rules.

**Kiffer:** They added a mixed setting value with a rule for its use, a time-point rule that flags the measurement closest to the end of the program as primary, and a secondary category field for programs with two components. Then they piloted again on three new studies and agreed on about ninety-three percent of entries.

**Sarah:** Who does the extraction after that? Is it always two people?

**Kiffer:** Buscemi and colleagues compared single and double extraction in two thousand six and found that single extraction produced more errors, although it was faster. Cochrane guidance therefore recommends two people extracting independently, particularly for outcome data. Rapid reviews often use one extractor with a second person checking every entry, which is what the Cochrane Rapid Reviews Methods Group recommends.

**Sarah:** What did Cedar Valley do?

**Kiffer:** Both. For the twenty-four comparative studies, meaning the fourteen trials and the ten non-randomized controlled studies, two people extracted the results independently, because a mistake in an effect estimate would change what the brief says. Everything else was extracted by one person and checked in full by the other. They kept a discrepancy log and wrote to authors when key information was missing.

**Sarah:** Let's move to section two. What exactly is risk of bias?

**Kiffer:** Bias is systematic error, a push away from the truth in a consistent direction because of how a study was designed or conducted. Health Sciences two thirty, in Lessons seven to eleven, teaches where bias comes from: measurement, selection, information bias, design-specific biases and confounding. This lesson is about the tools that turn those ideas into structured questions.

**Sarah:** The lesson is careful to separate risk of bias from some other things.

**Kiffer:** Three other things. Reporting quality is how completely a paper describes what was done, and a poorly reported trial may have been run perfectly well. Imprecision is random error, the kind you see in a wide confidence interval, and a small trial can be imprecise without being biased. Applicability is whether the study matches your question. The risk-of-bias tools deal with bias alone, and GRADE brings the others in later.

**Sarah:** Why not just give each study a quality score out of ten?

**Kiffer:** That used to be common. Peter Jüni and colleagues applied twenty-five quality scales to the same set of trials in nineteen ninety-nine and found that the conclusions depended on which scale they used. A score gives arbitrary weights to unrelated items, and a study with one serious flaw can still score well. Current tools judge each kind of bias separately.

**Sarah:** Before we get to the tools, a question that I suspect students will ask. The Cedar Valley project is a scoping review. Do scoping reviews even appraise studies?

**Kiffer:** Usually they do not. The reporting guideline for scoping reviews treats critical appraisal as optional, to be reported with a rationale if it is done, and the JBI guidance generally does not require it. The Cedar Valley team decided in its protocol to appraise, because the planning team will use the review to decide whether to fund a program and needs to know how far to trust the outcome findings. The appraisal is reported descriptively, and no study is excluded because of it.

**Sarah:** All right. Which tool for which design?

**Kiffer:** For randomized trials, the main tool is the second version of the Cochrane risk-of-bias tool, called RoB two, published by Jonathan Sterne and colleagues in twenty nineteen. It replaced the original Cochrane tool from twenty eleven.

**Sarah:** What changed?

**Kiffer:** Three things matter most. RoB two assesses a specific result, such as loneliness at twelve weeks, so one trial can get different judgements for different outcomes. The reviewer first states the effect of interest, usually the effect of assignment to the program. And judgements come from signalling questions, answered yes, probably yes, probably no, no, or no information, which an algorithm turns into low risk, some concerns or high risk.

**Sarah:** And there are five domains.

**Kiffer:** The randomization process, deviations from the intended interventions, missing outcome data, measurement of the outcome, and selection of the reported result. A result is at low risk overall only if all five are at low risk.

**Sarah:** Which domain gives the Cedar Valley team the most trouble?

**Kiffer:** Measurement of the outcome, domain four. Loneliness is self-reported, so the participant is the person assessing the outcome, and in a community program the participant almost always knows which group she is in. That knowledge could shape how she answers a loneliness questionnaire. We will come back to how the team handled it.

**Sarah:** What about studies that were not randomized?

**Kiffer:** For non-randomized studies of interventions, there is the Risk Of Bias In Non-randomized Studies of Interventions tool, known as ROBINS-I, also led by Sterne, published in twenty sixteen. Its central idea is the target trial, the randomized trial that would answer the same question without bias. It can be hypothetical, even when a real trial would be impractical or unethical, and describing it makes the reviewer clear about the population, intervention, comparison and outcome.

**Sarah:** What are its domains?

**Kiffer:** There are seven, arranged by when the bias arises. Before the intervention, there is confounding and the selection of participants into the study. At the intervention, there is classification of who received it. After it starts, there are deviations from the intended intervention, missing data, measurement of outcomes and selection of the reported result. The judgements are low, moderate, serious or critical risk. Low means comparable to a well-performed randomized trial, so moderate is a good result for a non-randomized study.

**Sarah:** And the team lists confounders in advance?

**Kiffer:** In the protocol, before reading the studies. For loneliness, the Cedar Valley list was baseline loneliness, depressive symptoms, living alone, mobility or functional limitation, age, and prior social participation. A connector study that adjusts only for age and sex leaves most of that list uncontrolled.

**Sarah:** The lesson also mentions ROBINS-E.

**Kiffer:** That is the version for exposures, published by Julian Higgins and colleagues in twenty twenty-four, for follow-up studies of things nobody assigns, such as living alone. It adapts the seven domains and adds a judgement of low risk except for concerns about uncontrolled confounding, because residual confounding almost never goes away in exposure studies. Cedar Valley's question is about interventions, so the team does not use it.

**Sarah:** Students will also see the Newcastle-Ottawa Scale in a lot of papers.

**Kiffer:** They will. George Wells and colleagues developed it for cohort and case-control studies. It awards stars across three categories, selection, comparability, and outcome or exposure, up to a maximum of nine.

**Sarah:** But the lesson is cautious about it.

**Kiffer:** Andreas Stang criticized it in twenty ten for unclear items and cut-offs with no stated basis, and Hartling and colleagues found low agreement between reviewers using it. Many reviews turn the stars into good, fair or poor ratings, which brings back the summary score problem. If you use it, report the stars item by item.

**Sarah:** And the JBI checklists?

**Kiffer:** JBI publishes a checklist for each of many designs, from trials and cohort studies to prevalence studies, case series and qualitative research, answered yes, no, unclear or not applicable. The Cedar Valley team uses the JBI quasi-experimental checklist for its nine before-and-after studies, a design the Cochrane tools do not cover.

**Sarah:** Then there is the MMAT.

**Kiffer:** The Mixed Methods Appraisal Tool, first developed by Pierre Pluye and colleagues and revised in twenty eighteen by Quan Nha Hong and colleagues. It suits a review that wants one tool for qualitative, quantitative and mixed-methods studies. After two screening questions, it offers five design categories with five criteria each, answered yes, no or can't tell, and the developers discourage an overall score.

**Sarah:** So, putting it together for Cedar Valley.

**Kiffer:** RoB two, with its cluster variant where needed, for the fourteen trials. ROBINS-I for the ten non-randomized controlled studies. The JBI quasi-experimental checklist for the nine before-and-after studies. The MMAT for the six qualitative and three mixed-methods studies. The trials and the non-randomized controlled studies are assessed by two reviewers independently, and the rest by one reviewer with full checking by the other.

**Sarah:** Section three is about those two reviewers. Why do they disagree in the first place?

**Kiffer:** Because the tools structure judgement without removing it. Hartling and colleagues reported low agreement between individual reviewers using the Newcastle-Ottawa Scale, and Minozzi and colleagues reported low agreement for RoB two and described how hard it was to apply. Much of that disagreement is avoidable. One reviewer reads the supplementary file and the other does not, or the two make different assumptions that nobody wrote down.

**Sarah:** How do you prevent it?

**Kiffer:** Preparation first. Both reviewers read the full guidance for each tool, and for each study they gather the paper, the supplements, the registry entry and any protocol. Then both appraise the same two or three studies and discuss every difference. That is called a calibration exercise.

**Sarah:** And that is where decision rules come from.

**Kiffer:** Usually, yes. A decision rule is a written agreement about how the team will answer a question in a situation that keeps recurring. It leaves the tool unchanged and records how the team interprets the guidance for its own topic, so the same situation gets the same answer every time and readers can see the reasoning.

**Sarah:** Give me the Cedar Valley example.

**Kiffer:** It is domain four. In two trials, the comparison group attended a different social activity of similar length. The intern reasoned that both groups expected some benefit, answered that knowledge of the assigned program probably could not have influenced the loneliness reports, and judged the domain at low risk. The evidence officer answered that it probably could, because someone who knows she is in the new program may still report differently.

**Sarah:** Who was right?

**Kiffer:** The team agreed that the active comparison matters, but at a later question. For self-reported loneliness, the reviewers now answer that participants knew their group and that this could have influenced their reports. At the final question, which asks whether influence is likely, they answer probably no, giving some concerns, unless the report gives a specific reason, such as a waiting list or a promise of less loneliness, which gives high risk.

**Sarah:** So the intern's low-risk judgements became some concerns.

**Kiffer:** Yes, and under that rule no Cedar Valley trial reaches low risk overall for loneliness. That is common in reviews of psychosocial programs, and it is an honest result.

**Sarah:** The lesson also stresses recording support for each answer.

**Kiffer:** Every answer needs something another reader could check: a quotation or summary, where it was found, and a reason when judgement is involved. For Trial G, the support for the missing data domain shows that two hundred and thirty of three hundred and fifty-four participants completed follow-up, that losses were almost twice as high in the comparison group, and that only a complete-case analysis was reported. That leads to high risk for the domain.

**Sarah:** And when the two reviewers compare?

**Kiffer:** They go question by question and sort each disagreement. A disagreement of fact means one person missed something, and they look at the source together. A disagreement of interpretation is settled by discussion, and if it recurs, by a new rule, after which they recheck the studies already done. If they still cannot agree, a third person arbitrates.

**Sarah:** Do reviews report how often the reviewers agreed?

**Kiffer:** Many do, using the same statistics students met for screening in Lesson seven, percent agreement and Cohen's kappa. The Cedar Valley reviewers made five domain judgements for each of fourteen trials, seventy in all, and agreed on fifty-seven, about eighty-one percent. On the overall judgement for each trial, they agreed on only nine of fourteen, and Cohen's kappa was about zero point three three.

**Sarah:** Why so much lower?

**Kiffer:** The overall judgement depends on the most severe domain judgement, so one disagreement in any of five domains can change it. And most trials fell into one category, which makes chance agreement high and pulls kappa down. On the usual benchmarks, that is fair agreement. After consensus, nine trials had some concerns and five were at high risk for loneliness.

**Sarah:** How are the results shown?

**Kiffer:** Usually in a traffic-light plot, with studies as rows, domains as columns and a coloured symbol in each cell. For the eight trials of group-based programs, every trial has some concerns in domain four, Trial G is at high risk because of missing data, and Trial H is at high risk because a coordinator could see the allocation list in advance.

**Sarah:** And then the judgements are used how?

**Kiffer:** The review shows each judgement beside each study's results, and it can order or stratify studies by risk of bias, which Lesson nine shows. What it should not do is decide to exclude high-risk studies after seeing the results, because that can steer the conclusions. If exclusion is planned, it belongs in the protocol.

**Sarah:** That brings us to section four and GRADE.

**Kiffer:** GRADE stands for Grading of Recommendations Assessment, Development and Evaluation. It was developed by an international working group, with Gordon Guyatt among its founders, and it is used by Cochrane, the World Health Organization and many guideline developers. It rates the certainty of evidence, meaning how confident we can be that an estimate of effect is close to the true effect.

**Sarah:** And it rates each outcome separately.

**Kiffer:** Each outcome, across the body of evidence, which means all the studies that address one comparison and one outcome. The same review can have high certainty for one outcome and very low certainty for another. GRADE also separates certainty from the strength of a recommendation. A guideline panel weighs certainty together with benefits, harms, costs, values and feasibility.

**Sarah:** What are the levels?

**Kiffer:** High, moderate, low and very low. High means we are very confident the true effect is close to the estimate. Moderate means it is probably close, but could be substantially different. Low means our confidence is limited, and the true effect may be substantially different. Very low means the true effect is likely to be substantially different.

**Sarah:** Where does the rating start?

**Kiffer:** With the design. A body of randomized trials starts at high certainty, and a body of non-randomized studies starts at low. Holger Schünemann and colleagues described an alternative in twenty nineteen, in which non-randomized evidence assessed with ROBINS-I starts at high and is rated down for risk of bias. The two approaches tend to agree, and the Cedar Valley team uses the conventional low start because it is easier to explain.

**Sarah:** And then there are five reasons to rate down.

**Kiffer:** Risk of bias, inconsistency, indirectness, imprecision and publication bias. A serious concern lowers the rating by one level and a very serious concern by two. Inconsistency is unexplained variation across studies. Indirectness is a mismatch with the question, such as trials in care homes when the question concerns people at home. Imprecision is random error, judged partly against the optimal information size. Publication bias is suspected when a few small positive studies dominate, and funnel plots are generally not used with fewer than ten studies.

**Sarah:** And three reasons to rate up.

**Kiffer:** A large effect, a dose-response gradient, and a situation where all plausible residual confounding would reduce the effect that was observed. As a guide, GRADE suggests that a relative risk above two or below one half can justify rating up one level, when the evidence is consistent and free of serious problems. Rating up is usually considered only for non-randomized evidence that has not been rated down.

**Sarah:** The Cedar Valley review did not pool its results. Does GRADE still work?

**Kiffer:** It does. Murad and colleagues described in twenty seventeen how to rate certainty without a single pooled estimate. You judge the same five domains from the pattern of results across studies and explain each judgement in a footnote. Rating certainty in a scoping review is unusual, so the team limited it to loneliness for three kinds of program, because the planning team asked directly how confident it could be, and it labels the ratings as provisional.

**Sarah:** Walk me through the evidence profile.

**Kiffer:** Remember that these results belong to the fictional case. Group-based programs had eight randomized trials with one thousand two hundred and thirty-six participants. They start at high. The team rated down one level for risk of bias, because every trial had some concerns or high risk, and one level for inconsistency, because five trials found small reductions, two found little or no difference and one found a larger reduction, with no explanation from setting or program length. That gives low certainty.

**Sarah:** And the one-to-one programs?

**Kiffer:** Six trials, seven hundred and forty-two participants, of befriending or telephone programs. They also start at high. Three trials were at high risk and three had some concerns, but their results were similar, so the team rated down one level for risk of bias. It rated down another level for imprecision, because the total sample is modest and most trials' confidence intervals include both no effect and a meaningful reduction. That is also low certainty.

**Sarah:** And the community connector programs, which is what Cedar Valley actually wants to launch.

**Kiffer:** Five non-randomized studies with one thousand four hundred and eighty participants. They start at low. Three were at serious risk of confounding, because clinicians referred people they judged likely to benefit, so the team rated down one level, to very low.

**Sarah:** Did anything justify rating up?

**Kiffer:** Two of the connector studies showed larger reductions in loneliness among people who had more contacts with their connector, which looks like a dose-response gradient. The team did not rate up, because the risk of bias was serious and because people whose loneliness improved early may have chosen to keep meeting their connector. That would produce the same pattern without more contact causing more benefit.

**Sarah:** How should the team say all this to the planning team?

**Kiffer:** Santesso and colleagues proposed standard wording in twenty twenty. High certainty uses plain verbs, the program reduces loneliness. Moderate uses probably. Low uses may. Very low says the evidence is very uncertain. So the brief says that group-based programs may reduce loneliness slightly among older adults, that one-to-one programs may reduce loneliness, and that the evidence is very uncertain about the effect of community connector programs on loneliness.

**Sarah:** That last sentence could alarm a planning team that is about to launch one.

**Kiffer:** It could, which is why the explanation matters. Very low certainty means the true effect could be substantially larger or smaller than the existing studies suggest. It leaves open whether the program works. For a health authority, the sensible response is to launch with an evaluation built in, for example by introducing the program in clinics in stages so outcomes can be compared. Lesson twelve shows how the brief makes that recommendation.

**Sarah:** Let's finish by putting it together. How would a small team plan these steps for a new rapid scoping review?

**Kiffer:** First, build a charting form with a data dictionary, including the TIDieR fields that suit the topic, have two people pilot it on three included studies, and record their percent agreement and what they changed. Second, write a short protocol amendment that says whether the team will appraise its studies, why, and which tool it will use for each design. Appraisal is optional in a scoping review, so the reason matters.

**Sarah:** And then?

**Kiffer:** If the team is appraising, two reviewers assess the same studies independently, compare their answers, and keep their support records and any decision rules they agree. And fourth, if the question is about the effects of an intervention, the team drafts a GRADE evidence profile for its main outcome, with a footnote for each judgement. If the question is not about effects, the protocol explains why a certainty rating does not apply.

**Sarah:** Any last advice?

**Kiffer:** Write things down. Most of the trouble in this lesson comes from decisions that lived in one person's head, such as which time point to extract or how to treat a self-reported outcome. Once a decision is written as a rule, a partner can apply it and a reader can check it.

**Sarah:** Next time, Lesson nine, synthesizing quantitative findings.

**Kiffer:** We will take the extracted results and the appraisal and put them together, without a meta-analysis, and we will see why counting studies by whether they were statistically significant misleads.

**Sarah:** Thanks, Kiffer. And thanks to everyone for listening to Office Hours.
