# Lesson 5: Indicators, Process Evaluation and Mixed Methods

*Companion-podcast transcript, Sarah and Kiffer*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This week we are working on Lesson five, which covers indicators, process evaluation and mixed methods.

**Sarah:** Where does this lesson sit in the course? Last week was about interest holders and evaluation questions.

**Kiffer:** That's right. In Lesson four, the Cedar Valley team worked with the people who will use the evaluation to agree on a short list of evaluation questions. We also saw the structure of the evaluation matrix, with columns for indicators, data sources, methods, timing and the person responsible. This lesson fills in those columns.

**Sarah:** And we are still using the Cedar Valley Connector program.

**Kiffer:** We are. Cedar Valley is a fictional community connector program, a form of social prescribing, run by a fictional health authority in British Columbia. Primary care clinicians in twelve first-wave clinics refer adults aged sixty-five and older who score six or higher on the three-item University of California, Los Angeles Loneliness Scale, which I'll call the UCLA scale. A connector meets each person up to six times over twelve weeks, builds a plan with them, and links them to groups, volunteer roles, transportation help and services.

**Sarah:** And we have some numbers from the first six months.

**Kiffer:** We do. Three hundred and twelve older adults were referred, two hundred and forty-one attended a first meeting, one hundred and eighty-eight had both a baseline and a follow-up loneliness score, and among those people the mean score fell from seven point one to six point three. Spending was four hundred and twenty thousand dollars.

**Sarah:** Let's start with section one. What exactly is an indicator?

**Kiffer:** An indicator is a specific, observable and measurable characteristic that shows whether a program component is in place or whether an expected change has occurred. The evaluation question states what we want to know, and the indicator states what we will look at in order to answer it.

**Sarah:** Students often use measure and indicator interchangeably. Is there a difference?

**Kiffer:** There is, and it helps to keep them apart. A measure is the instrument that produces a value, such as the UCLA scale. An indicator is the defined quantity you calculate from one or more measures, such as the mean change in that score from baseline to twelve weeks. The baseline is the value at the start, and the target is the value the program intends to reach by a stated date.

**Sarah:** The section sorts indicators into types. Can you walk through them?

**Kiffer:** The easiest way is to read them off the logic model from Lesson three. Input indicators describe resources, such as how many connector positions are filled. Process indicators describe how activities are carried out, such as the number of days from referral to first contact. Output indicators count the direct products of activities, such as the number of older adults linked to a volunteer role. Outcome indicators describe change in participants or systems, such as loneliness scores or emergency department visits.

**Sarah:** Why does the distinction matter in practice?

**Kiffer:** Because each type answers a different question. Process indicators arrive early and help managers fix delivery problems, but they cannot show whether anyone benefited. Outcome indicators show whether change happened, but on their own they cannot show how much of the change the program caused. That needs a design, which is the work of Lessons six to eight.

**Sarah:** The section spends a lot of time on specifying an indicator. Why is that so important?

**Kiffer:** Because a name like reach or engagement is too vague to calculate. If you give two analysts the same name, they will produce two different numbers. A full specification, often written on an indicator reference sheet, gives the numerator, the denominator, the data source, how often it is calculated, how it is broken down by subgroup, the baseline, the target and who is responsible.

**Sarah:** Give me the Cedar Valley example.

**Kiffer:** Take the proportion of referred older adults who attend a first meeting. The numerator is the number of unique older adults referred in the period whose first meeting is recorded in the program database. The denominator is the number of unique older adults referred. In the first six months that is two hundred and forty-one out of three hundred and twelve, or seventy-seven point two percent.

**Sarah:** That sounds like a decent result. Does it tell us the program is reaching lonely older adults?

**Kiffer:** It tells us how well the program converts referrals into participation. It tells us nothing about the people who were never referred. So we need a second indicator with a different denominator. Lesson two estimated that the twelve first-wave clinics have about twenty-three thousand patients aged sixty-five and older. If the regional survey estimate of twenty-four point five percent scoring six or higher applies to them, roughly five thousand six hundred and thirty-five patients meet the referral rule. The three hundred and twelve referrals are then about five and a half percent of the estimated eligible population.

**Sarah:** So the same program looks very different depending on the denominator.

**Kiffer:** Exactly, and that is where most indicator disputes begin. The second indicator also rests on an assumption, that the regional prevalence applies to these clinic panels, and the reference sheet should say so. An evaluation of reach usually needs both indicators.

**Sarah:** How should a program set targets?

**Kiffer:** There are several sources. You can project improvement from the baseline, benchmark against similar programs, use a clinical or policy standard, or negotiate with the primary intended users. Whichever you use, the target should be agreed before the data arrive, and the plan should record who agreed and why. For Cedar Valley, the steering committee set a target of eighty percent for the first-meeting proportion by the end of year two, and a plan can also set a target for the gap between rural and urban participants.

**Sarah:** How do we judge whether an indicator is any good?

**Kiffer:** One practical checklist comes from Kusek and Rist, who wrote a World Bank guide to results-based monitoring in two thousand and four. They used the acronym CREAM. A good indicator is clear, relevant, economic, adequate and monitorable. Clear means two people would calculate the same value, and monitorable means someone else can validate it because the source and calculation are documented.

**Sarah:** And for outcome measures there are the usual measurement properties.

**Kiffer:** Validity, reliability and sensitivity to change. The UCLA scale is a good example of the trade-offs. Hughes and colleagues developed it for large surveys, and its three items are quick to administer. But a total that runs from three to nine is coarse. Scores move in whole points, and someone who already scores three cannot improve on the scale.

**Sarah:** You also raise a problem with interpreting the fall from seven point one to six point three.

**Kiffer:** Yes, because of how people enter the program. The referral rule selects people who score six or higher on one occasion. Part of a high score on any given day is temporary distress, so some of those people will score lower next time even without a connector. That is regression to the mean, which Lesson seven covers as a threat to validity. The fall is a valid description of change among the one hundred and eighty-eight people, but on its own it cannot show how much of the change the program caused.

**Sarah:** The last part of section one contrasts performance measurement with evaluation. Aren't they the same activity?

**Kiffer:** They are related, and they do different jobs. Performance measurement is the ongoing, routine tracking of indicators against targets, mostly for managers. Harry Hatry wrote a well-known book on it. Evaluation is periodic and asks why results occurred, for whom, and how much of the change the program caused. Performance measurement raises questions, and evaluation investigates them.

**Sarah:** The section ends with a warning about indicators changing behaviour.

**Kiffer:** It does. Donald Campbell observed in nineteen seventy-nine that the more a quantitative indicator is used for social decisions, the more it is subject to corruption pressures and the more it distorts the processes it is meant to monitor. People now call that Campbell's law. It shows up as gaming, goal displacement and creaming. If Cedar Valley paid clinics per referral or ranked them monthly, it would invite all three, so balanced indicators reported privately to each clinic are a safer choice.

**Sarah:** Let's move to section two, process evaluation. Why do we need it?

**Kiffer:** Because an outcome evaluation on its own treats the program as a black box. If loneliness does not fall, we cannot tell whether the program's theory was wrong or whether the program was never really delivered. Lesson one introduced the distinction between implementation failure and theory failure. Process evaluation is the study of how a program is implemented and received, and it provides the evidence to tell those two explanations apart.

**Sarah:** There's a term in the reading I hadn't seen before, a Type three error.

**Kiffer:** Dobson and Cook used that term in nineteen eighty for the mistake of evaluating a program that was not actually implemented and then concluding that the program does not work. Imagine a health authority deciding that connectors do not reduce loneliness, when in fact most participants met a connector only once. That conclusion would be a Type three error.

**Sarah:** So what does a process evaluation actually measure?

**Kiffer:** A widely used list comes from Steckler and Linnan's edited volume from two thousand and two. Their components are context, reach, dose delivered, dose received, fidelity, implementation and recruitment. Later work, including the Medical Research Council guidance we'll come to, adds adaptation.

**Sarah:** Dose delivered and dose received sound almost the same.

**Kiffer:** They are easy to confuse, so here is the simplest way to separate them. Dose delivered is what the provider offers. Dose received is what the participant takes up. If a connector offers six meetings and the person attends three, dose delivered is six and dose received is three. If the connector's caseload only allows three meetings, both are three, and the shortfall lies with the program. Recording meetings offered as well as meetings attended tells you which situation you are in.

**Sarah:** What did the Cedar Valley data show on dose?

**Kiffer:** By the six-month review, two hundred and five participants had reached the end of their twelve weeks. Together they attended eight hundred and thirty-six meetings, which is a mean of about four point one. The steering committee had set four meetings as the minimum dose needed to complete a plan and receive follow-up after a linkage. One hundred and thirty-five of the two hundred and five reached it, which is about sixty-six percent.

**Sarah:** And there was a rural and urban difference.

**Kiffer:** A large one. At the four rural clinics, participants attended a mean of about three point four meetings, and forty-eight point four percent reached four or more. At the eight urban clinics, the mean was about four point four, and seventy-three point four percent reached four or more. That is a gap of twenty-five percentage points, and the numbers alone cannot tell us why.

**Sarah:** Let's talk about fidelity. How did Cedar Valley measure it?

**Kiffer:** The program team built a checklist of five core components into the program database: a co-developed plan, a linkage, a transportation assessment, a follow-up contact after the first linkage, and a closing summary to the referring clinician.

**Sarah:** And the results?

**Kiffer:** The plan was recorded for eighty-three point nine percent and the linkage for eighty-eight point three percent. Transportation assessment was at seventy-two point seven percent, follow-up at seventy-three point five percent, and the closing summary at seventy-seven point one percent. The steering committee's standard for acceptable delivery was eighty percent for each component, so three components fell short.

**Sarah:** The follow-up figure uses a different denominator, doesn't it?

**Kiffer:** It does, and that is a detail students often miss. Only the one hundred and eighty-one participants who were linked could receive follow-up after a linkage, so they form the denominator. Counting the twenty-four people who were never linked would treat them as failures of a step they could not receive.

**Sarah:** Can we trust a checklist that connectors fill in about their own work?

**Kiffer:** That's the right question. A self-completed checklist measures documented delivery, which may overstate actual delivery. It also records whether something happened, and it cannot judge quality. A plan can be ticked as co-developed even if the participant contributed very little. The usual remedy is to audit a random sample of case notes against the checklist.

**Sarah:** Fidelity can sound rigid. Programs always change when they spread to new places.

**Kiffer:** They do, and the reading uses an idea from Hawe, Shiell and Riley from two thousand and four to handle that. They suggested standardizing the function of each component, meaning the step in the change process it is meant to bring about, and allowing its form to vary with context. For Cedar Valley, one function is reducing barriers to attending, and its form might be a transport voucher in town or a volunteer driver in a rural area. Fidelity is judged against the function, and adaptations are logged with their reasons.

**Sarah:** How does an evaluator plan all of this?

**Kiffer:** Saunders, Evans and Joshi published a practical guide in two thousand and five with six steps. The second step, defining complete and acceptable delivery, is the heart of it. For Cedar Valley, complete delivery means every participant receives all five components and is offered six meetings. Acceptable delivery means eighty percent of participants receive each component and everyone is offered at least four meetings.

**Sarah:** And the Medical Research Council guidance?

**Kiffer:** The United Kingdom's Medical Research Council published guidance on process evaluation of complex interventions in two thousand and fifteen, with Graham Moore as lead author. It organizes process evaluation around three components. Implementation covers how delivery is achieved and what is delivered, including fidelity, dose, adaptation and reach. Mechanisms of impact covers how participants respond, the mediating processes through which change happens, and unexpected consequences. Context covers the outside factors that affect all of this.

**Sarah:** Is there any advice in it that surprised you?

**Kiffer:** One recommendation is worth remembering. The guidance suggests that, where possible, process data should be analyzed before the outcome results are known. If you already know the program failed, it is easy to go looking for delivery problems that confirm your explanation. Analyzing process data first protects the credibility of the interpretation.

**Sarah:** The reading ends section two with a clinic that sent very few referrals.

**Kiffer:** Yes. One rural clinic referred six older adults in six months, while the other rural clinics referred between twenty and thirty-five. The quarterly report showed the problem but did not explain it. A short process evaluation found that the screening template had never been installed in that clinic's electronic medical record and that most visits were with locum physicians who had not been oriented. Without that work, the low count might have been read as low need in that community.

**Sarah:** Section three is mixed methods. Start with a definition.

**Kiffer:** Mixed methods research is the collection and analysis of both quantitative and qualitative data within one study, with the two deliberately integrated so that the combined result says more than either part. That definition follows Creswell and Plano Clark. The key word is integrated.

**Sarah:** Why do evaluations combine methods?

**Kiffer:** Greene, Caracelli and Graham identified five purposes in a review of mixed methods evaluations published in nineteen eighty-nine. Triangulation seeks agreement between methods that study the same thing. Complementarity uses one method to elaborate the results of another. Development uses one method to build the other, for example using interviews to write survey items. Initiation looks for contradictions that reframe the question. Expansion uses different methods for different parts of the evaluation.

**Sarah:** Then the three core designs.

**Kiffer:** In a convergent design, you collect quantitative and qualitative data in the same period, analyze them separately, and then merge the results. For Cedar Valley, that might be the twelve-week survey alongside interviews with about twenty participants in the same weeks. In an explanatory sequential design, the quantitative results come first and qualitative work explains them. The rural attendance gap is the obvious case. The evaluator interviewed fourteen rural participants, eight with three or fewer meetings and six with five or six, together with the three rural connectors.

**Sarah:** And the third design?

**Kiffer:** In an exploratory sequential design, qualitative work comes first and is used to build something that is then measured, such as a set of indicators. Cedar Valley is co-designing a land-based connection pathway with one First Nation, and standard loneliness items may not capture what connection means in community terms. A qualitative phase led by the Nation, using talking circles and conversations with Elders and participants, could identify how connection is recognized. Whether to quantify it at all is a decision for the Nation under its own governance and the principles of ownership, control, access and possession, known as OCAP.

**Sarah:** You keep saying integration is the defining feature. What does integration look like?

**Kiffer:** Fetters, Curry and Creswell described integration at three levels. At the design level, it is built into the choice of design. At the methods level, the data sets can be connected, as when quantitative results choose the interview sample, or built, merged or embedded. At the level of interpretation, it happens through narrative, data transformation or joint displays.

**Sarah:** And they describe how the results fit together.

**Kiffer:** They describe three kinds of fit. Confirmation means both strands lead to the same conclusion. Expansion means each extends what the other shows. Discordance means they conflict. Discordance should be reported and investigated, because it often points to something important.

**Sarah:** Give us a joint display example from Cedar Valley.

**Kiffer:** A joint display is a table that puts the two strands side by side and states a meta-inference, which is the conclusion that draws on both. In one row, the quantitative result is the rural attendance gap. The qualitative finding is that rural participants and connectors described volunteer rides cancelled in bad weather and long drives to meetings in the regional centre, and several people said they had not been offered telephone meetings and would have accepted them. The fit is expansion. The meta-inference is that transportation explains much of the gap, and that offering telephone or home meetings routinely to rural participants is an adaptation worth testing.

**Sarah:** And an example of discordance?

**Kiffer:** At twelve weeks, one hundred and seventy-one of the one hundred and eighty-eight participants, or ninety-one percent, rated the program good or excellent. Yet some interviewed participants said they felt obliged to try groups that did not suit them and did not want to disappoint a connector they liked. Those results conflict. The meta-inference is that the high ratings may reflect appreciation of the connector more than of the linkages, so the evaluation adds a survey question about choice and reviews how plans are co-developed.

**Sarah:** The section ends with two methods for outcomes. Start with most significant change.

**Kiffer:** The most significant change technique, described in a guide by Rick Davies and Jess Dart in two thousand and five, collects stories of change from participants and staff. Panels at different levels of the program then select the stories they consider most significant and record why. Its main weakness is a tendency toward positive stories, which users address by adding a domain for negative or unexpected change and by analyzing all the stories, including the ones not selected.

**Sarah:** And outcome harvesting?

**Kiffer:** Outcome harvesting, described by Ricardo Wilson-Grau and Heather Britt in two thousand and twelve, works backward. You collect evidence of observed changes in the behaviour, relationships, actions, policies or practices of people and organizations, and then work out whether and how the program contributed. For Cedar Valley, it suits community partners, such as a seniors' centre that added a weekday drop-in after several referrals, or a volunteer driver society that changed its booking rules.

**Sarah:** Section four is about data systems. Where does an evaluator start?

**Kiffer:** With the data collection plan, which turns the matrix into operational instructions. For each indicator, the plan says which instrument or record is used, who collects it, from whom, when, how it is entered and stored, and who checks it. Alongside it sits a data dictionary that defines every variable, so that something like first meeting date means the same thing in every clinic.

**Sarah:** What does that look like for one Cedar Valley participant?

**Kiffer:** Connectors collect everything from referral to twelve weeks as part of their routine work, including the baseline survey, the meeting log and the fidelity checklist. At six months, evaluation staff who do not deliver the program conduct a telephone survey, which reduces the pressure to report improvement to your own connector. At twelve months, the analyst obtains linked emergency department and primary care data, using the Personal Health Number to link records.

**Sarah:** Talk about administrative and medical record data. What are their strengths?

**Kiffer:** Administrative data are records created to run or pay for services, such as physician billing claims, and medical record data are created by clinicians during care. Both add no burden for participants, cover whole populations, and reach back in time for baselines and comparison groups.

**Sarah:** And their weaknesses?

**Kiffer:** They were created for other purposes. Coding varies between clinics and between record systems, some information sits in free text, and access can be slow. Population Data BC facilitates research access to linked provincial holdings such as Medical Services Plan billing and hospital discharge abstracts, but approvals from each data steward can take many months. And no billing record measures loneliness, so administrative data supplement participant reports and cannot replace them.

**Sarah:** The section uses a framework for data quality.

**Kiffer:** Weiskopf and Weng reviewed how researchers assessed the quality of electronic health record data and identified five dimensions: completeness, correctness, concordance, plausibility and currency. In plain terms: is the value there, is it true, do sources agree, does it make sense, and is it up to date.

**Sarah:** And the Cedar Valley analyst actually ran those checks.

**Kiffer:** Before the six-month report, yes. The raw extract held three hundred and nineteen referral records, and seven were duplicates, which left the three hundred and twelve referrals we have been using. Three baseline loneliness totals were recorded as ten or higher, which is impossible on a scale that runs from three to nine, and the item responses showed data entry errors. Five follow-up dates came before the baseline date because day and month had been swapped.

**Sarah:** And there was a mismatch with the clinic records.

**Kiffer:** The clinic records listed three hundred and twenty-seven referral codes for the same period, so fifteen referrals made in clinics had no program record. Some were probably lost in transmission. That is a data quality finding and also a process finding about the referral pathway, which is a nice example of how the sections of this lesson connect.

**Sarah:** What about privacy and data sharing?

**Kiffer:** Evaluation data about individuals are personal health information. A health authority in British Columbia is a public body under the Freedom of Information and Protection of Privacy Act, which requires a privacy impact assessment for new initiatives involving personal information. When data move between organizations, for example from the health authority to a university partner, a data sharing agreement sets the purpose, the minimal list of variables, security, permitted uses, publication rules such as suppressing small cells, retention and breach reporting.

**Sarah:** And for First Nations participants?

**Kiffer:** One connector position is hosted by a First Nations health centre, and some participants are First Nations. Data about them follow the OCAP principles and the specific governance of the partner Nation and the host health centre. They decide how those data are used and reported, and findings about First Nations participants are reviewed with them before release.

**Sarah:** The section also introduces dashboards.

**Kiffer:** A dashboard shows a small number of indicators so that a manager can monitor them at a glance. Each indicator appears with its target and its trend, broken down where equity matters. The Cedar Valley dashboard shows the three hundred and twelve referrals, the first-meeting proportion of seventy-seven point two percent against a target of eighty, about sixty-six percent reaching four meetings against a target of seventy, and twelve-week follow-up at about ninety-two percent. It does not rank clinics, and it suppresses small counts.

**Sarah:** Then everything comes together in the evaluation matrix.

**Kiffer:** It does. The Cedar Valley matrix has five evaluation questions, covering reach and its equity, implementation, outcomes for participants, emergency department and primary care visits, and experience and mechanisms. The fifth question draws on interviews, most significant change stories, an outcome harvest with partners, and indicators for the land-based pathway defined with the First Nation. The economic question is added in Lesson ten.

**Sarah:** What makes it a good matrix, beyond being complete?

**Kiffer:** Every indicator traces to a question and to the logic model, every question about how or why has a qualitative source, and linked data are requested early because access takes time. And the whole thing was checked against capacity. With one half-time analyst, the quantitative indicators stay with the analyst, while the interviews, the outcome harvest and the six-month telephone survey go to an external evaluator.

**Sarah:** So when an evaluator writes up a matrix like this, what should it contain?

**Kiffer:** It covers every evaluation question the users prioritized. Each question gets one to four clearly defined indicators with baselines and targets where they make sense, and each indicator gets a data source, method, timing and responsible person. A short paragraph alongside it covers mixed methods, data quality checks, and the agreements or approvals the plan needs.

**Sarah:** Any advice for someone writing one?

**Kiffer:** Start from your questions, and resist the temptation to list every indicator you can think of. Make sure each indicator has a denominator someone else could reproduce. Name the subgroups you will report on, because equity has to be planned. And check your plan against the people and time you actually have, because a matrix that cannot be delivered will fail in the same way as an under-delivered program.

**Sarah:** And next week?

**Kiffer:** Lesson six turns to randomized designs for health services interventions, including cluster and stepped-wedge trials. The outcome questions in the matrix ask whether the program caused the changes you observe, and that is the question those designs are built to answer.

**Sarah:** Thanks, Kiffer. That's all for this week's Office Hours.

**Kiffer:** Thanks, Sarah, and thanks to everyone for listening.
