Indicators, Process Evaluation and Mixed Methods
Program Planning & Evaluation
Learning objectives for this lesson:
- Distinguish input, process, output and outcome indicators and place each on a program's logic model.
- Specify an indicator in full, with a definition, numerator, denominator, data source, frequency, disaggregation, baseline and target, and judge its quality against the CREAM criteria and measurement properties.
- Contrast performance measurement with evaluation and explain how indicators tied to rewards or rankings can distort the work they measure.
- Plan a process evaluation that measures reach, dose delivered, dose received, fidelity, adaptation, recruitment and context, drawing on Steckler and Linnan, Saunders and colleagues, and the Medical Research Council guidance.
- Select a convergent, explanatory sequential or exploratory sequential mixed methods design for an evaluation question and integrate the results in a joint display.
- Describe how most significant change and outcome harvesting produce qualitative evidence about program outcomes.
- Write a data collection plan that uses program, electronic medical record and administrative data, with data quality checks and the agreements needed to share data lawfully.
- Design a simple dashboard that reports a small set of indicators with targets to program managers.
- Complete an evaluation matrix with indicators, data sources and methods for each evaluation question, following the Cedar Valley worked example.
This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University, drawing on Rossi, P. H., Lipsey, M. W., & Henry, G. T. (2019). Evaluation: A Systematic Approach (8th ed.). SAGE; and Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and Quasi-Experimental Designs for Generalized Causal Inference. Houghton Mifflin.
Indicators and Performance Measurement
Learning Objectives for this section
- Define an indicator and explain how indicators connect evaluation questions to a program's logic model.
- Distinguish input, process, output and outcome indicators and classify indicators for a health program.
- Specify an indicator in full, with a definition, numerator, denominator, data source, frequency, disaggregation, baseline and target.
- Judge the quality of an indicator against the CREAM criteria and against measurement properties such as validity, reliability and sensitivity to change.
- Contrast performance measurement with evaluation and describe the risks that arise when indicators are tied to rewards or rankings.
Introduction
Lesson 4 ended with a set of prioritized evaluation questions and the structure of the evaluation matrix, which has a row for each question and columns for the indicator, data source, method, timing and person responsible. This lesson fills in those columns. Section 1 covers indicators, Section 2 process evaluation, Section 3 mixed methods, and Section 4 data systems, ending with a complete evaluation matrix for the running case.
The running case is the Cedar Valley Connector program, a fictional community connector (social prescribing) program run by the fictional Cedar Valley Health Authority in British Columbia. Primary care clinicians in 12 first-wave clinics refer adults aged 65 and older who score 6 or higher on the three-item UCLA Loneliness Scale, or who are judged by the clinician to be isolated, to a community connector. The connector meets each person up to six times over twelve weeks, co-develops a plan, and links the person to community groups, volunteer roles, transportation help and services. In its first six months the program received 312 referrals, 241 older adults attended a first meeting, 188 had both a baseline and a follow-up loneliness score, the mean score among those 188 fell from 7.1 to 6.3, and spending was $420,000. These figures, first reported in Lesson 1, recur throughout this lesson.
1.1 What an Indicator Is
An indicator is a specific, observable and measurable characteristic that shows whether a program component is in place or whether an expected change has occurred. An evaluation question states what the evaluation wants to know, and an indicator states what will be observed in order to answer it. A question about reach cannot be answered until the evaluator decides what will count as evidence of reach, such as the proportion of referred older adults who attend a first meeting.
Three terms are often confused. A measure or instrument is the tool that produces a value, such as the three-item UCLA Loneliness Scale. An indicator is the defined quantity calculated from one or more measures, such as the mean change in UCLA score between baseline and twelve weeks among participants with both measurements. A target is the value of the indicator that the program intends to reach by a stated date, and a baseline is the value at the start, against which change is judged. In the terms of Lesson 1's evaluative reasoning, indicators make criteria observable, and targets are one kind of standard.
Indicators can be quantitative or qualitative. Quantitative indicators are counts, proportions, rates, means and ratios. Qualitative indicators describe the presence, quality or nature of something, such as whether a clinic has a documented referral pathway, or the themes in participants' accounts of what changed for them. Evaluation matrices commonly include qualitative indicators for questions about experience, context and mechanism, and Section 3 returns to them.
1.2 Process, Output and Outcome Indicators
Indicators are easiest to organize when they are attached to the components of a logic model, which Lesson 3 taught. Each column of a logic model suggests a different type of indicator, and each type answers a different kind of question.
Input indicators describe the resources a program has, such as the number of connector positions filled or the proportion of the budget spent. Process indicators describe whether and how the program's activities are carried out, such as the number of meetings held per participant, the time from referral to first contact, or the proportion of plans co-developed within two meetings. Output indicators count the direct products of the activities, such as the number of older adults who completed a plan or the number of linkages made to community groups. Outcome indicators describe changes in the people, organizations or systems the program aims to affect, such as a fall in loneliness scores or in emergency department visits. Some frameworks separate outcomes from impacts, using impact for long-term population-level change, and some combine process and output indicators into a single category. The labels matter less than consistency within an evaluation plan.
Process indicators show whether the program is being delivered as intended. For Cedar Valley they include the median number of days from referral to first contact, the number of meetings each participant was offered, the proportion of participants whose transportation needs were assessed, and the number of adaptations logged by connectors. They are available early and help managers correct delivery problems quickly. They cannot show whether participants benefited, because a program delivered exactly as planned can still fail if its theory is wrong.
Output indicators count what the activities produced. For Cedar Valley they include the number of older adults who attended a first meeting (241 in the first six months), the number who completed a plan, and the number of linkages made to community groups, volunteer roles and transportation help. Outputs are often required in reports to funders, although a high count of linkages says nothing about whether older adults attended the groups or whether attending reduced their loneliness.
Outcome indicators describe change in participants or systems. For Cedar Valley the short-term outcome indicators include the mean change in UCLA Loneliness Scale score from baseline to twelve weeks and the proportion of participants who scored 6 or higher at baseline and scored 5 or lower at follow-up. Longer-term indicators include social participation, self-rated health, and emergency department and primary care visits. Outcome indicators show whether change occurred, and a design from Lessons 6 to 8 is needed to judge how much of it the program caused.
A leading indicator changes early and predicts later results, such as attendance at a first meeting, and a lagging indicator, such as emergency department visits, may take a year or more to change. A proxy indicator stands in for something that cannot be measured directly or affordably; the number of community groups accepting referrals is a proxy for the community's capacity to absorb new members.
Classify each of the following Cedar Valley indicators as an input, process, output or outcome indicator. (1) The proportion of the transport fund spent by month six. (2) The median number of days from referral to first connector contact. (3) The number of older adults linked to at least one volunteer role. (4) The proportion of participants reporting good, very good or excellent self-rated health at six months. (5) The proportion of connector meetings held in the participant's home. (6) The number of community partners that received a partner grant. (7) The rate of emergency department visits per 1,000 person-years among participants. (8) The number of connectors who completed the training module on Indigenous cultural safety. Suggested answers follow.
Items 1 and 8 are input indicators, although item 8 could be treated as a process indicator if training is counted as a program activity. Items 2 and 5 are process indicators, items 3 and 6 are output indicators, and items 4 and 7 are outcome indicators (intermediate and longer-term). The evaluation plan should state the convention it uses for borderline items.
1.3 Specifying an Indicator
An indicator name such as “reach” is too vague to measure, because two analysts given the same name will calculate different numbers. A full specification, often recorded on an indicator reference sheet, removes that ambiguity. Many agencies use a reference sheet of this kind, and federal performance information profiles ask for similar information. The table shows a completed reference sheet for one Cedar Valley indicator.
| Element | Specification for the Cedar Valley first-meeting indicator |
|---|---|
| Name | Proportion of referred older adults who attend a first connector meeting |
| Evaluation question | EQ1: To what extent does the program reach the older adults it was designed to serve? |
| Type | Output indicator (a leading indicator of reach) |
| Numerator | Number of unique older adults referred in the reporting period whose first meeting with a connector is recorded in the program database within 60 days of referral |
| Denominator | Number of unique older adults referred in the reporting period, after removal of duplicate referrals |
| Unit and calculation | Percentage, calculated quarterly and cumulatively |
| Data source | Program case-management database (referral and meeting records) |
| Disaggregation | Clinic, rural or urban location, age group, gender, preferred language and self-identified Indigenous identity, with the last reported only as agreed with First Nations partners |
| Baseline | First six months of wave one: 241 of 312 referred older adults (77.2 percent) |
| Target | 80 percent by the end of year two, agreed by the steering committee |
| Responsibility | The program analyst calculates it, and the program coordinator reviews it with connectors |
| Limitations | Referrals lost before reaching the program are not counted, and older adults who decline the referral before it is sent are invisible to the indicator |
The formula and its six-month value are as follows.
Worked calculation: first-meeting proportion
First-meeting proportion = (number of referred older adults who attended a first meeting) ÷ (number of unique older adults referred) × 100
= 241 ÷ 312 × 100 = 77.2 percent
The choice of denominator is where most indicator disputes begin. The first-meeting proportion uses referrals as its denominator, so it shows how well the program converts referrals into participation. It says nothing about the older adults who were never referred. A second reach indicator uses the estimated eligible population as its denominator. Lesson 2 estimated that the 12 first-wave clinics have about 23,000 patients aged 65 and older on their panels. If the regional survey estimate that 24.5 percent of older adults score 6 or higher applies to these clinics, about 5,635 patients would meet the referral rule, and the 312 referrals in six months represent about 5.5 percent of them (312 ÷ 5,635 × 100). This second indicator rests on an assumption that the regional prevalence applies to these clinic panels, and the reference sheet should state that assumption. The two indicators answer different questions, and an evaluation of reach usually needs both.
Baselines and targets
A target set without a baseline has little empirical basis, and a baseline without a target leaves the evaluator unable to say whether performance is adequate. Baselines come from the program's early data, the pilot, routine data from before the program began, or a comparison population. Targets can be set by projecting improvement from the baseline, by benchmarking against similar programs, by reference to a clinical or policy standard, or by negotiation with the primary intended users about what performance would justify continued investment. Targets should be set before the data are seen, with a record of who agreed to them and why. Equity targets deserve particular attention. If rural older adults attend first meetings less often than urban older adults, a single overall target can be met while the gap widens, so the plan may set a target for the gap itself.
Lesson 10 shows how indicators, targets and qualitative evidence are combined in an evaluative rubric to reach a judgement of merit or worth.
1.4 Judging Indicator Quality
Kusek and Rist (2004), in a World Bank guide to results-based monitoring, summarized the qualities of a good performance indicator with the acronym CREAM, drawing on Schiavo-Campo and Tommasi (1999). The criteria are useful as a checklist when a planning team has more candidate indicators than it can collect.
Outcome indicators built from measurement instruments must also meet the standards of measurement that students met in their epidemiology courses. Validity is the extent to which the instrument measures the intended construct. Reliability is the consistency of its results across occasions, raters or items. Sensitivity to change, also called responsiveness, is the instrument's ability to detect change when change has occurred. The three-item UCLA Loneliness Scale was developed for large surveys by Hughes and colleagues (2004). Each item asks how often the respondent feels a lack of companionship, left out, or isolated from others, with responses scored from 1 (hardly ever) to 3 (often), giving a total from 3 to 9. Its brevity makes it practical for clinicians and connectors, although its seven possible values make it coarse: scores move only in whole points, and a participant who scores 3 cannot improve on the scale. A team wanting finer-grained evidence might add a longer instrument for a subsample, at the cost of additional burden.
Interpretation matters as well. The referral rule selects older adults who score 6 or higher on one occasion, and some will score lower when measured again even without the program, because a single high score includes temporary distress. This regression to the mean is a threat to validity taught in Lesson 7. The fall from 7.1 to 6.3 among the 188 participants with both scores describes change, and it cannot by itself show how much of the change the program caused.
Background: the minimal important difference
A change in a score can be detectable and still too small to matter to the people who experience it. The minimal important difference (MID) is the smallest change in a score that the people measured perceive as important. It is estimated in two main ways. Anchor-based methods compare score changes with an external judgement of change, such as a question asking whether the respondent feels a little better, about the same or worse, and take the typical change among those who report a small but noticeable improvement. Distribution-based methods express change relative to the spread of scores, for example as half a standard deviation of baseline scores, or as the standard error of measurement, which is calculated from the standard deviation and the reliability of the instrument. Distribution-based values describe the size of a change relative to variability or measurement error, and only anchor-based values link a change to what people notice, so anchor-based estimates are preferred where they exist.
Is a mean fall of 0.8 points on the 3 to 9 scale likely to be meaningful? (Answer: The lesson gives neither an anchor nor the spread of scores, so the question cannot be settled from these figures. Because individual scores move in whole points, a mean fall of 0.8 combines participants who improved by one point or more with others who did not change. The evaluation would judge it against a published MID for the three-item scale, or against an anchor question added to the twelve-week survey, and would remember that part of the fall reflects regression to the mean.)
Optional reading: HSCI 410 Lesson 7 Section 3 (Validity) treats interpretability and the minimal important difference in more depth.
1.5 Performance Measurement and Evaluation
Performance measurement is the ongoing, routine collection and reporting of indicators to track a program's activities, outputs and outcomes against targets. Hatry (2006) describes it as regular measurement of the results and efficiency of services. Evaluation, as Lesson 1 defined it, is the systematic determination of a program's merit, worth or significance, usually through a study designed to answer specific questions, including why results occurred and how much of the change the program caused. Performance measurement supplies routine data and raises questions, and evaluation explains the patterns it detects.
| Feature | Performance measurement | Evaluation |
|---|---|---|
| Main question | What is happening, and are targets being met? | Why is it happening, for whom, and did the program cause it? |
| Frequency | Continuous, with monthly or quarterly reports | Periodic, timed to decisions |
| Data | Routine program and administrative data | Routine data plus data collected for the study, including qualitative data |
| Comparison | Against targets and previous periods | Against a counterfactual, a standard or a rubric |
| Main users | Program managers and funders | Decision-makers, program staff, participants and the public |
| Cedar Valley example | A quarterly report of referrals, first meetings and follow-up scores by clinic | A study of whether the program reduced loneliness and service use, and why results differed between rural and urban clinics |
Under the Treasury Board Policy on Results (2016), federal departments maintain performance information for their programs and also plan evaluations that examine relevance and effectiveness. Lesson 1 described that structure, with routine indicators tracked continuously and evaluations scheduled periodically to explain them, as a sound model for a health authority. Cedar Valley follows it: the analyst produces a quarterly performance report, and the evaluation plan specifies the studies that will inform the second-wave decision.
When indicators change behaviour
Indicators influence the behaviour of the people whose work they measure. Campbell (1979) observed that the more a quantitative social indicator is used for social decision-making, the more it becomes subject to corruption pressures and the more it distorts the processes it is intended to monitor. This observation is now known as Campbell's law. Its consequences are familiar in health care and appear in several forms.
Staff change how data are recorded without changing what happens. If Cedar Valley rewarded clinics for referrals, a clinic might record every older adult who accepted a pamphlet as a referral. Clear definitions, audit of a sample of records, and indicators that are hard to inflate (such as first meetings attended) reduce the risk.
Staff concentrate on what is measured and neglect what is not. If the quarterly report tracks only the number of linkages, connectors may make many quick linkages and spend less time on the follow-up contact that helps older adults keep attending. A balanced set of process and outcome indicators makes this harder.
Programs select participants who are easiest to serve or most likely to succeed. A target for the proportion of participants who complete six meetings could lead connectors to prioritize mobile, urban, English-speaking referrals. Equity indicators, such as the gap between rural and urban participation, make creaming visible.
League tables of clinics treat random variation as performance. A clinic with 15 referrals in a quarter can move from top to bottom of a ranking because of two or three people. Lesson 8 introduces statistical process control charts, which distinguish common-cause variation from special-cause variation and are a better tool for comparing clinics over time.
Before the second wave, a member of the Cedar Valley executive proposes that every clinic should reach 90 percent attendance at first meetings and that a monthly ranking of clinics should be circulated to clinic managers. The evaluator would note that the current value is 77.2 percent and that the agreed target is 80 percent by the end of year two. A 90 percent target with rankings would invite distortion, such as delaying the recording of referrals until a first meeting is booked. The evaluator might propose a balanced set of indicators reported privately to each clinic, run charts in place of rankings, an equity indicator for rural clinics, and a commitment that missed targets will prompt a review of causes supported by the process evaluation described in Section 2.
Summary
Indicators translate evaluation questions into observable quantities and descriptions. Process, output and outcome indicators correspond to the columns of the logic model. A full specification allows an indicator to be calculated consistently and checked by others, and the CREAM criteria and measurement properties help a team choose among candidates. Performance measurement tracks indicators routinely, and evaluation explains them. Section 2 turns to process evaluation, which asks why a program is or is not working as intended.
Reflection
The Cedar Valley Connector program is a fictional community connector (social prescribing) program run by a fictional health authority in British Columbia. Primary care clinicians refer adults aged 65 and older who score 6 or higher on the three-item UCLA Loneliness Scale (scored 3 to 9, higher scores meaning greater loneliness) to a connector, who meets them up to six times over twelve weeks, co-develops a plan, and links them to community groups, volunteer roles, transportation help and services. In its first six months it received 312 referrals, and connectors record every linkage in a program database. Connectors survey participants at twelve weeks, and evaluation staff survey them by telephone at six months. The steering committee has added an evaluation question: To what extent do participants remain involved in the community activities they were linked to after their twelve weeks with the connector end? Write a full specification for one indicator that answers this question, giving its name, type, numerator, denominator, data source, frequency, disaggregation, how the baseline will be obtained, a proposed target with its justification, and the person responsible. Then assess the indicator against the CREAM criteria (clear, relevant, economic, adequate and monitorable) and name one limitation.
The indicator is the proportion of linked participants still attending a linked activity at six months, an intermediate outcome indicator. The numerator is the number of participants who report in the six-month telephone survey that they attended at least one activity they were linked to during the previous four weeks. The denominator is the number of participants with at least one linkage recorded in the program database who completed the six-month survey. Data come from the program database and one added survey item, calculated quarterly and disaggregated by rural or urban clinic, type of linkage, gender, language and Indigenous identity, with the last reported as agreed with First Nations partners. The baseline is the first cohort's value at six months. A target, perhaps 60 percent, would be agreed with the steering committee once the baseline is known. The analyst is responsible.
The indicator is clear because the four-week window and the denominator are defined. It is relevant because sustained participation is the step in the program theory between linkage and reduced loneliness, and it is economic because it adds one item to an existing survey. It is monitorable because linkages and survey responses are stored and can be audited. It is only partly adequate, since attendance says nothing about the quality of participation, so interviews should accompany it. Its main limitation is that people who stopped attending may also be less likely to answer the survey, inflating the value, so follow-up completeness should be reported with it.
Minimum 20 characters required.
Question 1: Which of the following is an output indicator for the Cedar Valley Connector program?
Question 2: In six months Cedar Valley received 312 unique referrals, 241 older adults attended a first meeting, and about 5,635 patients in the first-wave clinics are estimated to meet the referral rule. Which statement is correct?
Question 3: A proposed indicator would be calculated by each connector from private notes, with no documented calculation and no records that anyone else could check. Which CREAM criterion does it fail most directly?
Question 4: Which statement best describes the relationship between performance measurement and evaluation?
Process Evaluation
Learning Objectives for this section
- Explain why outcome findings cannot be interpreted without evidence about implementation, using the ideas of implementation failure, theory failure and Type III error.
- Define reach, dose delivered, dose received, fidelity, adaptation, recruitment and context, following Steckler and Linnan, and write a process question and indicator for each.
- Calculate and interpret dose and fidelity indicators from program records.
- Apply the six steps of Saunders and colleagues to plan a process evaluation, including a definition of complete and acceptable delivery.
- Describe the Medical Research Council framework for process evaluation of complex interventions and its three components of implementation, mechanisms of impact and context.
2.1 Why Process Evaluation Matters
An outcome evaluation that reports only whether outcomes changed treats the program as a black box. If loneliness does not fall, the evaluation cannot say whether the program's theory was wrong or whether the program was never delivered as designed. If loneliness does fall, the evaluation cannot say which parts of the program mattered or whether the result would hold in other clinics. Lesson 1 introduced Weiss's distinction between implementation failure, in which the program was not delivered as planned, and theory failure, in which the program was delivered but the expected causal process did not occur, and Lesson 3 located the two failures in Chen's action and change models. Process evaluation is the study of how a program is implemented and received, and it supplies the evidence needed to tell these explanations apart.
Dobson and Cook (1980) used the term Type III error for the mistake of evaluating a program that was not actually implemented, and then attributing the absence of effects to the program itself. A health authority that concluded from a null outcome evaluation that community connectors do not work, when in fact most participants met a connector only once, would commit this error. The table shows how process and outcome evidence combine.
| Implementation | Outcomes improved | Outcomes did not improve |
|---|---|---|
| Delivered as intended | The finding is consistent with the program theory. A causal claim still requires a suitable design (Lessons 6 to 8). | The finding suggests theory failure, although weak measurement or a short follow-up may also explain it. |
| Poorly or partly delivered | Something other than the program as designed may explain the change, such as an adaptation that worked or a change in the community. The evaluation should investigate. | The finding suggests implementation failure. The program theory has not been tested, and concluding that the program does not work would be a Type III error. |
Process evaluation has three common uses. Formative use improves delivery while the program is running, for example by identifying clinics where referrals are lost. Interpretive use explains outcome findings, as in the table above. Transfer use informs decisions about scale and adaptation, which is directly relevant to Cedar Valley because the evidence from the first wave will shape how the program is delivered in the second-wave clinics.
2.2 The Components of Process Evaluation
Steckler and Linnan's edited volume Process Evaluation for Public Health Interventions and Research (Steckler & Linnan, 2002) set out a widely used list of process evaluation components. Linnan and Steckler's overview chapter defined seven: context, reach, dose delivered, dose received, fidelity, implementation and recruitment. Later work, including the Medical Research Council guidance described in Section 2.5, added adaptation as a component in its own right. Each card defines one component and applies it to Cedar Valley.
Each component becomes a process question with one or more indicators. The table gives examples for Cedar Valley, which Section 4 carries into the evaluation matrix.
| Component | Process question | Example indicator | Data source |
|---|---|---|---|
| Recruitment | How do clinics identify and refer older adults? | Referrals per 1,000 patients aged 65 and older, by clinic and quarter | Program database, clinic panel counts |
| Reach | Who attends, and who is missing? | Proportion of referrals attending a first meeting, by subgroup | Program database |
| Dose delivered | How much of the program do connectors provide? | Mean meetings offered per participant | Connector meeting log |
| Dose received | How much do participants take up? | Proportion attending four or more meetings | Connector meeting log |
| Fidelity | Are the core components delivered? | Proportion of participants receiving each of five core components | Fidelity checklist in the program database |
| Adaptation | How and why is the program changed? | Number and type of adaptations, with reasons | Adaptation log, connector interviews |
| Context | What helps or hinders delivery? | Barriers and facilitators reported by staff and partners | Interviews, steering committee minutes |
2.3 Measuring Dose and Fidelity
Dose
Dose delivered and dose received are often confused because both can be expressed as a number of meetings. The distinction is the source of the shortfall. If a connector offers six meetings and the participant attends three, dose delivered is six and dose received is three. If the connector's caseload allows only three meetings, both are three, and the shortfall lies with the program. Recording meetings offered as well as meetings attended lets the evaluator tell these situations apart.
By the six-month review, 205 of the 241 older adults who attended a first meeting had reached the end of their twelve-week period. Their connector logs give the following distribution of meetings attended. The four rural clinics and eight urban clinics are shown separately because the steering committee asked whether transportation limits participation.
| Meetings attended | 1 | 2 | 3 | 4 | 5 | 6 | Total | Mean |
|---|---|---|---|---|---|---|---|---|
| Rural clinics | 12 | 9 | 11 | 12 | 9 | 9 | 62 | 3.39 |
| Urban clinics | 12 | 10 | 16 | 26 | 32 | 47 | 143 | 4.38 |
| All first-wave clinics | 24 | 19 | 27 | 38 | 41 | 56 | 205 | 4.08 |
Worked calculation: dose received
Mean meetings attended = total meetings ÷ participants = (24 × 1 + 19 × 2 + 27 × 3 + 38 × 4 + 41 × 5 + 56 × 6) ÷ 205 = 836 ÷ 205 = 4.08
Proportion attending four or more meetings = (38 + 41 + 56) ÷ 205 × 100 = 135 ÷ 205 × 100 = 65.9 percent
Rural clinics: 30 ÷ 62 × 100 = 48.4 percent. Urban clinics: 105 ÷ 143 × 100 = 73.4 percent.
The steering committee had set a threshold of four meetings as the minimum dose at which a participant can complete a plan and receive follow-up after a linkage. About two thirds of participants reached it, with a gap of 25 percentage points between rural and urban clinics. The numbers describe the gap. They do not explain it, since rural participants may live farther away, may have fewer groups to be linked to, or may have been referred with more complex needs. Section 3 shows how qualitative follow-up can explain the gap.
Fidelity
Dane and Schneider (1998) reviewed how prevention programs measured implementation and identified five dimensions of fidelity: adherence to the program's components, exposure (the amount delivered), quality of delivery, participant responsiveness and program differentiation, which is the extent to which the program's distinctive features are present. Carroll and colleagues (2007) proposed a conceptual framework in which adherence, measured as content, frequency, duration and coverage, is the core of fidelity, and in which intervention complexity, facilitation strategies, quality of delivery and participant responsiveness moderate it. Both frameworks show that fidelity has more than one dimension, and that counting meetings captures exposure without capturing whether the essential work of the meetings occurred.
Cedar Valley's program team identified five core components and built a checklist into the program database, which connectors complete for each participant. The table reports the first six months for the 205 participants who had reached twelve weeks.
| Core component | Criterion | Delivered | Percent |
|---|---|---|---|
| Co-developed plan | Plan written with the participant by the second meeting | 172 of 205 | 83.9 |
| Linkage | At least one linkage matched to a goal in the plan | 181 of 205 | 88.3 |
| Transportation assessment | Transportation needs recorded at the first or second meeting | 149 of 205 | 72.7 |
| Follow-up after linkage | Contact within two weeks of the first linkage, among participants linked | 133 of 181 | 73.5 |
| Closing summary | Summary sent to the referring clinician at the end of twelve weeks | 158 of 205 | 77.1 |
Three cautions apply to checklist data of this kind. Connectors record their own performance, so the checklist measures documented delivery, which may overstate actual delivery; a periodic audit of a random sample of case notes against the checklist would test this. A checklist records whether a component occurred and cannot judge its quality; a plan can be written by the connector with little input from the participant and still be ticked. The denominator must match the component, as the follow-up row shows: only the 181 participants who were linked could receive follow-up after a linkage.
Fidelity and adaptation
Programs delivered in many settings are always adapted, and a narrow view of fidelity would treat every adaptation as a failure. Hawe, Shiell and Riley (2004) proposed a resolution for complex interventions: standardize the function of each component, meaning the step in the change process it is meant to bring about, and allow its form to vary with local context. Under this view, fidelity is judged against functions. The table shows how this applies to Cedar Valley.
| Function to standardize | Forms that may vary by clinic or participant |
|---|---|
| A person-centred conversation identifies what matters to the older adult and what stands in the way of connection. | Meetings at the clinic, at home, by telephone or at a community venue |
| A co-developed, written connection plan records the participant's own priorities. | A paper or electronic plan, in the participant's preferred language |
| The participant is actively linked to at least one community resource that matches the plan. | Groups, volunteer roles, faith communities, land-based activities co-designed with the First Nation partner |
| Follow-up identifies and reduces barriers to attending and revises the plan. | Transport vouchers, volunteer drivers, accompaniment to a first meeting, telephone check-ins |
An adaptation log records each change, who made it, why, and whether it preserves the component's function. Lesson 9 introduces the FRAME approach, which gives a structured vocabulary for this record.
2.4 Planning a Process Evaluation
Saunders, Evans and Joshi (2005) published a practical guide to developing a process evaluation plan for health promotion programs. Its central idea is that the evaluator should first define what complete and acceptable delivery of the program would look like, and then design questions and methods to assess it. Their six steps are summarized below, with the Cedar Valley application for each. The authors note that steps three to five are iterative, because questions are revised as methods and resources are considered.
The evaluator records the program's purpose, theory, objectives, strategies and expected outcomes. For Cedar Valley, this is the program description from Lesson 1 and the logic model and theory of change from Lesson 3.
The team specifies, for each core component, what full delivery would be and what level would be acceptable. For Cedar Valley, complete delivery means that every participant receives all five core components and is offered six meetings; acceptable delivery was set at 80 percent of participants receiving each component and at least four meetings offered to every participant.
Questions are drafted for each process component, such as “What proportion of participants receive a follow-up contact within two weeks of a linkage?” and “How do connectors adapt the program for participants without transportation?”.
For each question, the team chooses the data source, instrument, timing and analysis, using program records, checklists, logs, observation, surveys and interviews.
The plan is checked against the resources available. With one half-time analyst, Cedar Valley chose to build the checklist into the program database and to sample case notes for audit, instead of observing meetings.
The team agrees on a final set of questions, indicators and methods, ranked by priority, and adds them to the evaluation matrix.
Against the acceptable-delivery standard, the first six months show two components below 80 percent (transportation assessment at 72.7 percent and follow-up after linkage at 73.5 percent) and one close to it (closing summary at 77.1 percent). These are formative findings, and the program coordinator can act on them before the second wave.
The Cedar Valley steering committee plans to add accompaniment, in which a volunteer goes with a participant to the first session of a linked group. Using Step 2 of Saunders and colleagues, write a definition of complete delivery and of acceptable delivery for this component, name one fidelity indicator with its numerator and denominator, and state the data source. A suggested answer is in the accordion below.
Complete delivery would be that every participant who is linked to a group, and who reports anxiety about attending alone, is offered accompaniment by a trained volunteer to the first session. Acceptable delivery might be that 80 percent of such participants are offered accompaniment and that 60 percent of those offered it are accompanied within four weeks of the linkage. A fidelity indicator is the proportion of eligible participants offered accompaniment, with a numerator of participants offered accompaniment and a denominator of linked participants who reported anxiety about attending alone at the linkage meeting. The data source is a new field in the program database completed by the connector, verified against the volunteer coordinator's schedule.
2.5 The Medical Research Council Guidance on Process Evaluation
In 2015 the United Kingdom Medical Research Council published guidance on process evaluation of complex interventions (Moore et al., 2015). Complex interventions have several interacting components, require behaviours from the people who deliver and receive them, act at several levels, and are often adapted to local settings. A community connector program meets every part of that description. The guidance applies to process evaluation the broader MRC framework for developing and evaluating complex interventions, introduced in Lesson 2 Section 3.5. It organizes process evaluation around three components.
Implementation covers how delivery is achieved, such as training, resources and support for connectors, and what is delivered, described by fidelity, dose, adaptations and reach. Mechanisms of impact covers how participants respond to and interact with the program, the mediating processes through which it produces change, and any unexpected pathways and consequences. For Cedar Valley, a program theory of the kind developed in Lesson 3 might propose that a plan built on the participant's own priorities increases motivation, that linkage with accompaniment and transportation help reduces practical barriers, and that repeated contact with a group builds relationships that reduce loneliness. A process evaluation might measure whether participants who attend a linked group three or more times report new relationships, which is a mediator on that pathway. Context covers the factors external to the program that affect implementation, mechanisms and outcomes, and which the program may in turn affect.
The guidance makes several recommendations about how a process evaluation should be conducted. The process evaluation should be planned alongside the outcome evaluation and should be based on a clear description of the program and its causal assumptions. Quantitative and qualitative methods should be combined, because quantitative data show how much was delivered and qualitative data show how and why. The guidance also suggests that, where possible, process data should be analyzed before the outcome results are known, so that interpretations of implementation are not shaped by knowledge of whether the program worked. It discusses the relationship between process evaluators and the people who designed and deliver the program, which needs to be close enough for the evaluator to understand the program and independent enough for the findings to be credible. The 2021 framework for developing and evaluating complex interventions (Skivington et al., 2021) extended this approach by treating program theory, context and engagement with interest holders as core elements at every phase of evaluation.
The MRC framework overlaps with realist evaluation (Lessons 1 and 3), which states program theory as context-mechanism-outcome configurations and tests them. Both lead the evaluator to specify, before data collection begins, which mechanisms and contextual factors the evaluation will examine.
At the six-month review, one rural Cedar Valley clinic had referred 6 older adults while the other rural clinics had referred between 20 and 35. The quarterly performance report showed the problem without explaining it. A short process evaluation, consisting of interviews with the clinic manager, two physicians and the connector and a review of the clinic's electronic medical record configuration, found that the screening template had never been installed in that clinic's record system and that most visits were with locum physicians who had not received the orientation. The finding concerns recruitment and context, and it suggested that second-wave clinics should test the template before launch and include locum physicians in orientation. Without it, the low referral count might have been read as low need.
Summary
Process evaluation explains how a program was delivered and received, which outcome findings need for their interpretation. Steckler and Linnan's components, with adaptation added by later work, structure the process questions and indicators. Dose and fidelity can be calculated from program records if denominators are chosen carefully and records are audited. Saunders and colleagues organize planning around complete and acceptable delivery, and the Medical Research Council guidance places implementation, mechanisms of impact and context in one framework. Questions such as why rural participants attend fewer meetings need both numbers and accounts from the people involved, and Section 3 turns to the designs that combine them.
Reflection
The Cedar Valley Connector program is a fictional community connector (social prescribing) program run by a fictional health authority in British Columbia. Primary care clinicians refer adults aged 65 and older who score 6 or higher on the three-item UCLA Loneliness Scale (scored 3 to 9, higher scores meaning greater loneliness) to a connector, who meets them up to six times over twelve weeks, co-develops a plan, and links them to community groups, volunteer roles, transportation help and services. By the six-month review, 205 participants had reached the end of their twelve weeks. Program records show that a plan was co-developed by the second meeting for 172 of 205; at least one linkage was made for 181 of 205; transportation needs were recorded at the first or second meeting for 149 of 205; a follow-up contact within two weeks of the first linkage was made for 133 of the 181 linked participants; and a closing summary was sent to the referring clinician for 158 of 205. The steering committee defined acceptable delivery as 80 percent of participants receiving each component. Participants at four rural clinics attended a mean of 3.39 meetings (48.4 percent attended four or more), compared with 4.38 meetings (73.4 percent) at eight urban clinics. (a) Calculate the percentage for each component and identify those below the standard. (b) Explain why the follow-up component uses a different denominator. (c) Propose two process evaluation questions, each tied to a named component (context, reach, dose delivered, dose received, fidelity, recruitment or adaptation), with a method for answering each. (d) Explain how this evidence helps the steering committee avoid a Type III error, which is concluding that a program does not work when it was not delivered as intended.
(a) The percentages are 83.9 for the plan, 88.3 for linkage, 72.7 for transportation assessment, 73.5 for follow-up after linkage and 77.1 for the closing summary. Three components fall below the 80 percent standard: transportation assessment, follow-up and the closing summary.
(b) Only participants who were linked can receive a follow-up contact after a linkage, so the denominator is the 181 linked participants. Using 205 would count the 24 unlinked participants as failures of a component they could not receive.
(c) The first question concerns dose delivered and dose received: are rural participants offered fewer meetings, or offered as many but attending fewer? I would add meetings offered to the meeting log, compare offered and attended meetings by location, and interview the rural connectors and a purposive sample of rural participants with low and high attendance. The second question concerns fidelity: is transportation assessment omitted, or done and not recorded? I would audit a random sample of about 10 percent of case notes against the checklist and ask connectors in a group discussion how the assessment fits into first meetings.
(d) If later outcome analyses show little change among rural participants, these data show that many received less than the four-meeting minimum and that a core component aimed at their main barrier was often missing. The steering committee could then attribute a weak result partly to implementation, fix delivery, and avoid concluding that the program theory has failed.
Minimum 20 characters required.
Question 1: A health authority concludes that community connectors do not reduce loneliness after an evaluation in which most participants met a connector only once. Which error does this illustrate?
Question 2: A connector offers a participant six meetings, and the participant attends three. How should dose be recorded?
Question 3: Under the approach proposed by Hawe, Shiell and Riley (2004), what should be standardized across Cedar Valley clinics?
Question 4: Which three components organize process evaluation in the Medical Research Council guidance (Moore et al., 2015)?
Mixed Methods in Evaluation
Learning Objectives for this section
- Define mixed methods evaluation and explain the purposes for combining methods identified by Greene, Caracelli and Graham.
- Apply convergent, explanatory sequential and exploratory sequential designs to evaluation questions for a health program.
- Describe how integration occurs at the design, methods and interpretation levels, and classify the fit of integrated results as confirmation, expansion or discordance.
- Construct a joint display that states a meta-inference for each domain of an evaluation.
- Describe the most significant change technique and outcome harvesting, and judge when each is suited to an evaluation.
3.1 Why Evaluations Combine Methods
Section 2 ended with a gap that numbers alone could not explain: participants at rural Cedar Valley clinics attended a mean of 3.39 connector meetings, compared with 4.38 at urban clinics. The program records show the size of the gap, and interviews with participants and connectors can show what produces it. Evaluation has used mixed methods for decades, because evaluation questions so often ask both whether something happened and how or why it happened. The two Background boxes below restate the mixed methods and qualitative ideas this section builds on.
Background: mixed methods, purposes, notation and designs
Mixed methods research collects and analyzes both quantitative and qualitative data within one study and deliberately integrates them, so that the combined result says more than either part (Creswell & Plano Clark, 2018). Greene, Caracelli and Graham (1989) identified five purposes for combining methods in evaluation: triangulation seeks convergence of results; complementarity uses one method to elaborate or clarify another; development uses one method’s results to build the other; initiation seeks contradictions that reframe questions; and expansion uses different methods for different components. Each purpose implies a different design.
Morse’s (1991) notation writes QUAN and QUAL in capitals for the method given priority and in lower case for a supplementary method, with a plus sign for concurrent collection and an arrow for sequence. Creswell and Plano Clark describe three core designs: convergent (QUAN + QUAL), which merges data collected in the same period; explanatory sequential (QUAN → qual), which explains quantitative results; and exploratory sequential (QUAL → quan), which builds a measure or feature that is then tested. Fetters, Curry and Creswell (2013) place integration at the design level, the methods level (connecting, building, merging or embedding) and the interpretation level (narrative, data transformation or joint displays), and they describe the fit of integrated results as confirmation, expansion or discordance.
Optional reading: HSCI 207 Lesson 6 Section 2 (Choosing an Approach and Managing a Research Project) teaches these ideas in depth.
Background: a short refresher on qualitative methods
Students whose training is mainly quantitative may find a brief summary useful. Qualitative evaluation collects data as words, observations and documents through individual interviews, focus groups, observation of program activities and review of program documents. It samples purposively, choosing participants for the information they can provide, such as participants with high and low attendance or connectors at rural and urban clinics. Malterud, Siersma and Guassora (2016) proposed that sample size be guided by information power, which is higher when the study aim is narrow, participants are highly specific to the aim, the study is supported by established theory, the dialogue is strong, and the analysis examines a few cases in depth. Thematic analysis (Braun & Clarke, 2006) is a common analytic approach that moves through six phases, which Braun and Clarke (2022) name familiarization with the data, coding, generating initial themes, developing and reviewing themes, refining, defining and naming themes, and writing up. Evaluation timelines are often short, and rapid qualitative techniques replace full transcription and line-by-line coding with structured summaries and matrices (Vindrola-Padros & Johnson, 2020). In a common form, the analyst completes a summary template organized by the topics of the interview guide soon after each interview, transfers the summaries into a matrix with one row per participant and one column per topic, and compares the rows across groups, such as rural and urban participants. This is the method the Section 4.6 evaluation matrix names for rapid qualitative analysis.
Optional reading: HSCI 841 Lesson 5 Section 1.4 teaches thematic analysis, HSCI 841 Lesson 7 Section 1.7 teaches rapid qualitative analysis as a matrix method, and HSCI 207 Lesson 11 Section 4 gives an introductory overview of analytic approaches.
3.2 Three Core Designs
The three core designs recalled in the Section 3.1 Background box can be combined into more complex designs, such as a mixed methods evaluation spanning several phases. The figure shows the three designs, and the tabs apply each one to an evaluation question for Cedar Valley.
In a convergent design (QUAN + QUAL), the evaluator collects quantitative and qualitative data in the same period, analyzes them separately, and then merges the results to compare or relate them. The purpose is usually triangulation or complementarity. For Cedar Valley, the twelve-week follow-up survey measures loneliness, social participation and satisfaction, while semi-structured interviews with a purposive sample of 20 participants, conducted in the same weeks, ask how the program affected their social lives. The results are merged in a joint display by domain. The design is efficient when time is short, but it requires the evaluator to plan in advance how the two sets of results will be compared, and it can leave discrepancies that the evaluation has no time to resolve.
In an explanatory sequential design (QUAN → qual), quantitative results come first, and a qualitative phase follows to explain them. The integration point is connecting: the quantitative results determine the questions and the sample for the qualitative phase. For Cedar Valley, the six-month review found that 48.4 percent of rural participants and 73.4 percent of urban participants attended four or more meetings. The evaluator then interviewed 14 rural participants, chosen so that 8 had attended three or fewer meetings and 6 had attended five or six, together with the three connectors who serve the rural clinics. The design suits questions that arise from routine data, which makes it a natural partner for performance measurement. It takes longer than a convergent design, because the second phase cannot begin until the first is analyzed.
In an exploratory sequential design (QUAL → quan), qualitative work comes first and is used to build something that is then tested or measured quantitatively, such as an instrument, a set of indicators or a program feature. For the land-based connection pathway being co-designed with one First Nation, standard loneliness items may not capture what connection means in community terms. A qualitative phase, led by the Nation and using talking circles and conversations with Elders and participants, could identify what connection means and how it might be recognized, and the results could be built into a short set of indicators to be used in the pathway. Whether to quantify at all, and how the data are held and reported, are decisions for the Nation under its own data governance and the principles of OCAP® discussed in Lesson 4. The design is the most time-consuming of the three, and it is the right choice when existing measures do not fit the population or program.
The three designs can be combined. A multiphase evaluation of Cedar Valley would use an exploratory sequential component for the land-based pathway, an explanatory sequential component for the rural attendance gap, and a convergent component at each annual follow-up. Creswell and Plano Clark (2018) describe such combinations as complex designs, and they note that an evaluation spanning several years often takes this form.
Integration and fit
Mixed methods are defined by integration, which the Section 3.1 Background box placed at the design, methods and interpretation levels (Fetters, Curry & Creswell, 2013), and integrated results fit as confirmation, expansion or discordance, where discordance is itself a finding to be reported and investigated. HSCI 207 Lesson 6 Section 2.4 (Integration and Joint Displays) is optional reading on both ideas.
3.3 Joint Displays
A joint display is a table or figure that brings quantitative and qualitative results together so that they can be compared and interpreted as one (Guetterman, Fetters & Creswell, 2015). Its most important column is the meta-inference, the conclusion that draws on both strands. The joint display below integrates findings from the first six months of Cedar Valley. The interview results are illustrative summaries of the kind of themes an evaluation might report.
| Domain | Quantitative result | Qualitative finding | Fit | Meta-inference |
|---|---|---|---|---|
| Dose in rural clinics | Rural participants attended a mean of 3.39 meetings (48.4 percent attended four or more), compared with 4.38 (73.4 percent) at urban clinics. | Rural participants and connectors described volunteer rides cancelled in poor weather and long drives to meetings in the regional centre. Several participants said they had not been offered telephone meetings and would have accepted them. | Expansion | Transportation explains much of the gap. Offering telephone or home meetings routinely to rural participants is an adaptation the program could test, while preserving the plan-building function of the meetings. |
| Change in loneliness | Among 188 participants with both scores, the mean three-item UCLA score fell from 7.1 to 6.3. | Participants who joined groups described new acquaintances and regular reasons to leave home. Several who did not join a group said the connector's visits themselves made them feel less alone. | Expansion | The relationship with the connector may be a mechanism in its own right, separate from group attendance. The evaluation should examine change separately for participants with and without a sustained linkage. |
| Referrals at one rural clinic | One rural clinic referred 6 older adults in six months, compared with 20 to 35 at the other rural clinics. | Staff interviews found that the screening template had not been installed and that locum physicians had not been oriented. | Expansion | Low referral reflects a recruitment failure at the clinic, and the count should not be read as low need. |
| Satisfaction | At twelve weeks, 171 of 188 participants (91.0 percent) rated the program good or excellent. | Some interviewed participants described feeling obliged to try groups that did not suit them and being reluctant to say so to a connector they liked. | Discordance | High ratings may reflect appreciation of the connector more than of the linkages. The evaluation should add a survey item on choice of activities and review how plans are co-developed. |
The calculation behind the satisfaction figure is 171 ÷ 188 × 100 = 91.0 percent. Guetterman and colleagues (2015) suggest that a good joint display organizes rows by a shared dimension, such as a domain, a theme or a participant subgroup, places the meta-inference in its own column, and states the fit explicitly. A joint display can also be organized by case, with each row a participant or clinic, or by statistical result, with each row a quantitative finding and the qualitative data that explain it. For an explanatory sequential design, a useful layout groups participants by the quantitative result used to select them, such as low and high attendance, and shows the themes in each group.
The six-month survey shows that 64 percent of participants linked to a volunteer role were still volunteering at six months, compared with 41 percent of participants linked to a social group who were still attending it. Interviews suggest that volunteer roles gave participants a defined responsibility and a reason to return each week, while some social groups were described as welcoming at first and harder to join over time. Write the domain, the fit and a meta-inference for this row, and state one implication for the program and one for the evaluation.
The domain could be named “Persistence in linked activities.” The fit is expansion, because the interviews extend the quantitative difference by suggesting a mechanism: a defined role and a weekly obligation sustain attendance, while open social groups depend on whether members make room for newcomers. A meta-inference would state that the type of linkage appears to affect persistence, possibly through role-based belonging. For the program, connectors could favour role-based linkages where they fit a participant's goals, or partner groups could appoint a member to welcome new arrivals. For the evaluation, the comparison should be checked for confounding, because participants who choose volunteer roles may differ in health and mobility from those who choose social groups, so the evaluator would compare baseline characteristics before drawing conclusions.
3.4 Most Significant Change and Outcome Harvesting
Some program outcomes cannot be specified in advance. Participants may value changes that the logic model did not anticipate, and community partners may change their own practices in response to the program. Two qualitative methods were designed to collect evidence of such outcomes systematically.
Most significant change
The most significant change technique, described in a guide by Davies and Dart (2005), collects stories of change from participants and staff and then involves people at different levels of a program in selecting, from those stories, the ones they judge most significant, with their reasons recorded. The technique is participatory, it does not rely on predefined indicators, and the discussion of why one story is chosen over another makes the values of the program's interest holders visible.
The program agrees on a few broad domains, such as changes in participants' social connections, changes in their health and well-being, and any other changes, and on how often stories will be collected. Cedar Valley chose these three domains and a six-month cycle.
Participants, connectors or volunteers are asked, for each domain, what they consider the most significant change over the reporting period and why it matters to them. Stories are recorded in the teller's words with consent, and identifying details are removed before stories are shared.
Stories pass through a hierarchy of selection panels, for example connectors at each clinic, then the program team, then the steering committee. Each panel selects one story per domain and records its reasons for the choice.
The selections and reasons are fed back to the people who told the stories and to staff. Selected stories may be verified with the teller or others. A secondary analysis examines all stories, including those not selected, for themes and for the types of change that are absent.
The technique's main limitation is a bias toward positive stories, because tellers and selectors tend to favour success. Users of the technique reduce this bias by adding a domain for negative or unexpected changes and by analyzing all stories, including those not selected (Davies & Dart, 2005). Most significant change complements quantitative indicators and is rarely sufficient on its own to judge whether a program met its objectives.
Outcome harvesting
Outcome harvesting, developed by Ricardo Wilson-Grau and colleagues and described by Wilson-Grau and Britt (2012), works backward from observed change. The evaluator collects evidence of what has changed and then determines whether and how the program contributed. An outcome is defined as an observable change in the behaviour, relationships, actions, policies or practices of an individual, group, community, organization or institution. The method has six steps: design the harvest with its primary users, review documentation and draft outcome descriptions, engage with the people who know about each outcome to refine the descriptions, substantiate a selection of outcomes with independent informants, analyze and interpret the outcomes in relation to the evaluation questions, and support the use of the findings.
Outcome harvesting suits outcomes among organizations and systems, where change is hard to predict. For Cedar Valley, a harvest among community partners might document that a seniors' centre added a weekday morning drop-in after several referrals, that a volunteer driver society changed its booking rules to accept requests from connectors, and that a faith community began a visiting program. Each outcome would be described with what changed, who changed, when and where, how the program contributed, and why the change matters, and a sample would be confirmed with someone outside the program team. Lesson 3's contribution analysis gives the logic for judging the program's contribution to such outcomes.
| Feature | Most significant change | Outcome harvesting |
|---|---|---|
| What is collected | Stories of change told by participants and staff | Descriptions of observed changes in the behaviour or practice of social actors |
| How evidence is judged | Selection panels choose and justify the most significant stories | Selected outcomes are substantiated by independent informants |
| Main strength | Reveals what participants value and makes interest holders' values explicit | Captures unplanned changes in organizations and systems and traces contribution |
| Main limitation | Tends to favour positive stories | Depends on available documentation and informants, and substantiation takes time |
| Cedar Valley use | Stories from participants and connectors every six months | A harvest among community partners at the end of year one |
3.5 Quality and Reporting
A mixed methods evaluation should be judged on the quality of each strand and on the quality of its integration. O'Cathain, Murphy and Nicholl (2008) proposed the Good Reporting of A Mixed Methods Study (GRAMMS) guideline for health services research. It asks authors to justify the use of mixed methods, to describe the design in terms of purpose, priority and sequence, to describe each method's sampling, data collection and analysis, to describe where and how integration occurred and who took part in it, to describe any limitation of one method associated with the presence of the other, and to describe the insights gained from mixing. HSCI 207 Lesson 12 Section 2.3 (Reporting Guidelines) places GRAMMS among the reporting guidelines for other designs and is optional reading. An evaluation plan can use the same items as a checklist before data collection begins. For Cedar Valley, the plan states that integration will occur in joint displays reviewed by the steering committee, including its older adult members, and in a review of findings with the First Nations partners before results about First Nations participants are reported.
Summary
Mixed methods evaluation integrates quantitative and qualitative evidence so that an evaluation can say both what happened and why. Greene and colleagues' five purposes guide the choice among convergent, explanatory sequential and exploratory sequential designs. Integration can be designed at three levels, and the fit of integrated results, whether confirmation, expansion or discordance, is reported in a joint display with explicit meta-inferences. Most significant change and outcome harvesting add systematic qualitative evidence about outcomes that indicators did not anticipate. All of these methods depend on data that are collected reliably, stored securely and shared lawfully, which is the subject of Section 4.
Reflection
The Cedar Valley Connector program is a fictional community connector (social prescribing) program run by a fictional health authority in British Columbia. Primary care clinicians refer adults aged 65 and older who score 6 or higher on the three-item UCLA Loneliness Scale (scored 3 to 9, higher scores meaning greater loneliness) to a connector, who meets them up to six times over twelve weeks, co-develops a plan, and links them to community groups, volunteer roles, transportation help and services. After twelve weeks, the mean three-item UCLA score among 188 participants fell from 7.1 to 6.3. A clinic-level breakdown shows that the mean fall ranged from about 0.2 points at some of the 12 clinics to about 1.4 points at others. The steering committee asks why change differs across clinics. The evaluation has four months to report and an external qualitative evaluator available for about 25 days. Three mixed methods designs are available: convergent (quantitative and qualitative data collected in the same period and merged), explanatory sequential (quantitative results first, followed by qualitative work to explain them) and exploratory sequential (qualitative work first, used to build a measure or tool that is then tested quantitatively). Choose a design and justify the choice, describe the qualitative sample and the point of integration, and write one illustrative joint display row with a domain, a quantitative result, a qualitative finding, the fit (confirmation, expansion or discordance) and a meta-inference.
An explanatory sequential design fits best, because the quantitative result already exists and the question is why it occurred. A convergent design would collect qualitative data without knowing which clinics to compare, and an exploratory design would take too long and answers a different question. Before sampling, the analyst should check whether the clinic differences exceed what chance would produce with small clinic samples, because some apparent variation will be random. I would then select two clinics with the largest falls and two with the smallest, matched where possible on rurality, and interview the connector, the clinic manager, a physician and three or four participants at each, about 20 interviews in total, which suits 25 days with rapid qualitative analysis. Integration occurs first by connecting, because the quantitative results determine the sample, and then in a joint display organized by clinic group.
An illustrative row has the domain community capacity. The quantitative result is that clinics with larger falls recorded more linkages to volunteer roles and weekly groups per participant. The qualitative finding is that connectors at the clinics with smaller falls described few groups with space for new members and long waits for volunteer drivers. The fit is expansion. The meta-inference is that local community capacity may moderate the program's effect, which suggests directing partner grants toward communities with few activities and adding a capacity indicator to the matrix.
Minimum 20 characters required.
Question 1: After finding that rural participants attend fewer meetings, the Cedar Valley evaluator interviews rural participants selected for low and high attendance. Which design and form of integration does this represent?
Question 2: At twelve weeks, 91.0 percent of participants rated the program good or excellent, while interviewed participants described feeling obliged to try groups that did not suit them. How should the fit of these results be classified?
Question 3: Which method is best suited to documenting unplanned changes in the practices of Cedar Valley's community partners, such as a seniors' centre adding a weekday drop-in?
Question 4: Using what is learned in talking circles led by a First Nation to draft indicators of connection for later measurement illustrates which purpose for mixing methods described by Greene, Caracelli and Graham (1989)?
Data Systems for Evaluation
Learning Objectives for this section
- Write a data collection plan that states what will be collected, from whom, by whom, when and with what instrument.
- Describe the strengths and limitations of program records, electronic medical record data and administrative health data in British Columbia.
- Apply data quality checks for completeness, correctness, concordance, plausibility and currency to program data.
- Identify the contents of a data sharing agreement and the privacy and governance requirements that apply to evaluation data.
- Design a simple dashboard and complete an evaluation matrix with indicators, data sources and methods for each evaluation question.
4.1 The Data Collection Plan
An indicator is only as good as the data behind it. The data collection plan turns the indicator and data source columns of the evaluation matrix into operational instructions: for each indicator, which instrument or record is used, who collects the data, from whom, at what point, in what mode, how the data are entered and stored, and who checks them. The plan is usually accompanied by a data dictionary, which defines every variable, its permitted values and its source, so that a variable such as “first meeting date” means the same thing in every clinic. The figure shows the collection points for each Cedar Valley participant.
Two principles shape the plan. Data collection should be built into routine work wherever possible, because data that connectors record as part of their job are more complete than data collected separately. Burden on participants and staff should be proportionate, because every additional item lowers response and competes with program delivery. Cedar Valley's baseline survey therefore takes under ten minutes: the three UCLA items, a social participation item set, a self-rated health item and an intake form.
4.2 Administrative and Electronic Medical Record Data
Background: administrative data, EMR data and record linkage
Administrative data are records created to run or pay for services, such as physician billing claims and hospital discharge abstracts. Electronic medical record (EMR) data are records created by clinicians in the course of care. Record linkage joins records for the same person across sources. Deterministic linkage matches records on an exact identifier, such as the Personal Health Number in British Columbia, and probabilistic linkage weighs agreement on several fields, such as name, birth date and postal code, when identifiers are missing or inconsistent. Linkage errors are not random, because records are missed more often for people who move frequently or whose names are recorded inconsistently.
Optional reading: HSCI 207 Lesson 8 Section 3 (administrative records) and HSCI 207 Lesson 9 Section 3 (record linkage) treat these sources and methods in depth.
For an evaluation, both kinds of record are produced without additional burden on participants, cover whole populations, and extend backward in time, which allows baselines and comparison groups to be constructed. Both were created for purposes other than evaluation, so their variables, coding and completeness reflect those purposes. The cards describe the sources available to the Cedar Valley evaluation.
Because linkage errors fall unevenly across groups, the evaluator should report linkage rates by subgroup, and moving records between organizations for linkage requires the agreements described in Section 4.4. Administrative data also miss important outcomes: no billing record measures loneliness, so administrative data supplement participant-reported outcomes and cannot replace them.
4.3 Data Quality Checks
Weiskopf and Weng (2013) reviewed how studies assessed the quality of EHR data for research reuse and identified five dimensions, which describe the quality of EHR data and apply equally to program data. Surveillance systems are judged on a related set of attributes (Centers for Disease Control and Prevention, 2001), in which data quality is defined by completeness and validity and timeliness corresponds to currency. The table applies Weiskopf and Weng’s dimensions to Cedar Valley.
| Dimension | Question | Cedar Valley check |
|---|---|---|
| Completeness | Is the value present? | Proportion of participants missing each baseline item, by clinic |
| Correctness | Is the value true? | Audit of a random sample of fidelity checklists against case notes |
| Concordance | Do sources agree? | EMR referral codes matched to program referral records |
| Plausibility | Does the value make sense? | UCLA totals within 3 to 9; follow-up dates after baseline dates |
| Currency | Is the value up to date? | Days between a meeting and its entry in the database |
Before the six-month report, the Cedar Valley analyst ran these checks on the raw extract. The extract held 319 referral records, of which 7 were duplicates (the same Personal Health Number and referral date), leaving the 312 unique referrals reported in Section 1. Three baseline UCLA totals were recorded as 10 or higher, which is impossible on a 3 to 9 scale; the item-level responses showed data entry errors, which were corrected. Five follow-up dates preceded the baseline date because day and month had been transposed. The concordance check found 327 referral codes in the clinic EMR extracts for the same period, so 15 referrals recorded in clinics had no program record. Some of these referrals were probably lost in transmission, which is a process finding about the referral pathway as well as a data quality finding. Each correction should be logged, the original extract retained, and the checks scripted so that they run identically every quarter.
4.4 Data Sharing Agreements and Privacy
Evaluation data about individuals are personal health information, and their collection, use and disclosure are governed by law and policy. In British Columbia, health authorities are public bodies under the Freedom of Information and Protection of Privacy Act (FIPPA), which requires a privacy impact assessment for new programs and initiatives that involve personal information. Private physician practices are generally governed by the Personal Information Protection Act. Where the evaluation is research, TCPS 2 applies, and Lesson 1 described how Article 2.5 distinguishes program evaluation and quality improvement from research requiring ethics review. Data that move between organizations, such as from clinics to the health authority or from the health authority to an academic partner, require a data sharing agreement.
The agreement names the parties, states the purpose for which data are shared, and cites the legal authority for the disclosure and collection.
It lists the variables to be shared and limits them to those needed for the stated purpose. Direct identifiers are removed or replaced with study codes as early as possible.
It states who may access the data, where they are stored, how they are protected, and that they may not be linked to other data without approval.
It sets rules for publication, including suppression of small cells, the retention period, the method of destruction, and the procedure for reporting a privacy breach.
For data about First Nations participants, the agreement reflects the principles of ownership, control, access and possession (OCAP®) set out by the First Nations Information Governance Centre, and the specific governance of the partner Nation and the health centre that hosts a connector. Lesson 4 discussed these principles.
4.5 Simple Dashboards
A dashboard displays a small number of indicators so that a manager can monitor them at a glance (Few, 2006). A useful program dashboard shows each indicator with its target, shows change over time, disaggregates where equity matters, and is updated on a known schedule. It avoids ranking clinics, for the reasons given in Section 1, and suppresses counts small enough to identify individuals; many organizations suppress cells below a threshold such as five. The figure sketches a quarterly dashboard for Cedar Valley built from the indicators in this lesson. Lesson 8 adds run charts for tracking indicators over many periods, and Lesson 10 treats data visualization for decision-makers.
The dashboard values come from earlier sections: 241 of 312 referrals attended a first meeting (77.2 percent), 135 of 205 participants attended four or more meetings (65.9 percent), and 188 of those 205 had a twelve-week score (91.7 percent). By quarter, 98 of 131 referrals (74.8 percent) and 143 of 181 referrals (79.0 percent) attended a first meeting.
4.6 Worked Example: The Cedar Valley Evaluation Matrix
Lesson 4 set out the structure of the evaluation matrix. The matrix below completes it for the five evaluation questions that Cedar Valley's primary intended users might prioritize to inform the second-wave decision. Each question has indicators with definitions and targets, data sources, methods, timing and a responsible person. The economic question, what the program costs per participant and per unit of outcome, is added in Lesson 10.
| Indicator, definition and target | Data source | Method | Timing | Responsibility |
|---|---|---|---|---|
| EQ1. Reach: To what extent does the program reach eligible older adults, and does reach differ by rurality, language, gender and Indigenous identity? | ||||
| 1a. Referral proportion: unique referrals ÷ estimated eligible patients aged 65 and older. Six months: 5.5 percent. Target: 15 percent in year one. | Program database; clinic panel counts | Proportion, cumulative and by clinic | Quarterly | Analyst |
| 1b. First-meeting proportion: referrals attending a first meeting ÷ unique referrals. Baseline 77.2 percent. Target: 80 percent by end of year two. | Program database | Proportion with 95 percent confidence interval | Quarterly | Analyst; coordinator reviews with connectors |
| 1c. Equity of reach: indicators 1a and 1b by rurality, language, gender and Indigenous identity. Target: rural and urban first-meeting proportions within 5 percentage points. | Intake form (self-identified); clinic panel profiles | Disaggregated proportions; First Nations data reported as agreed with partners | Every six months | Analyst; First Nations partners |
| 1d. Unmatched referrals: EMR referral codes with no program record. Six months: 15 of 327. | Clinic EMR queries; program database | Record matching by Personal Health Number | Quarterly | Analyst; clinic managers |
| EQ2. Implementation: Is the program delivered as intended, how is it adapted, and what helps or hinders delivery? | ||||
| 2a. Dose delivered: participants offered four or more meetings ÷ participants reaching twelve weeks. Target: 100 percent. | Connector meeting log | Proportion by clinic | Quarterly | Coordinator |
| 2b. Dose received: participants attending four or more meetings ÷ participants reaching twelve weeks. Six months: 65.9 percent. Target: 70 percent. | Connector meeting log | Proportion and mean, rural and urban | Quarterly | Analyst |
| 2c. Fidelity: proportion receiving each of five core components. Target: 80 percent for each. | Fidelity checklist; audit of a 10 percent random sample of case notes | Proportions; agreement between checklist and audit | Quarterly; audit every six months | Coordinator; analyst |
| 2d. Adaptations: number, type and reason, and whether the core function is preserved. | Adaptation log; monthly connector meetings | Content analysis using FRAME categories (Lesson 9) | Monthly log; six-month review | Coordinator |
| 2e. Context: barriers and facilitators to delivery. | Interviews with clinic staff, connectors and partners; steering committee minutes | Rapid qualitative analysis; explanatory follow-up of gaps in 2b | Months 6 and 12 | External qualitative evaluator |
| EQ3. Outcomes: How much do loneliness, social participation and self-rated health change among participants, and for whom? | ||||
| 3a. Mean change in three-item UCLA Loneliness Scale score from baseline to twelve weeks and six months. Twelve weeks: 7.1 to 6.3 (188 participants). Standard set in the rubric (Lesson 10). | Baseline and twelve-week surveys by connectors; six-month telephone survey by evaluation staff | Paired change; mixed models; comparison with second-wave clinics (Lessons 6 to 8) | Baseline, 12 weeks, 6 months | Analyst |
| 3b. Participants scoring 6 or higher at baseline who score 5 or lower at six months ÷ participants scoring 6 or higher at baseline with a six-month score. | As for 3a | Proportion with 95 percent confidence interval, by subgroup | 6 months | Analyst |
| 3c. Social participation: community activities attended in the past four weeks. | Survey item set | Mean change | Baseline, 12 weeks, 6 months | Analyst |
| 3d. Self-rated health: proportion reporting good, very good or excellent. | Single survey item | Change in proportion | Baseline, 6 months | Analyst |
| 3e. Follow-up completeness: participants with a follow-up score ÷ participants reaching that time point, at twelve weeks and six months. Twelve weeks: 91.7 percent. Target: 80 percent at each point. | Program database | Proportion; baseline comparison of completers and non-completers | Quarterly | Analyst |
| EQ4. Health service use: Does the program change emergency department and primary care visits in the twelve months after referral, compared with older adults in second-wave clinics? | ||||
| 4a. Emergency department visits per 1,000 person-years, twelve months before and after referral. | Health authority emergency department records, linked by Personal Health Number | Difference-in-differences with second-wave clinics (Lesson 7) | Annually; request submitted in month 3 | Analyst with academic partner |
| 4b. Primary care visits per person-year, twelve months before and after referral. | Medical Services Plan billing via Population Data BC, or clinic EMRs | As for 4a | Annually | Analyst; privacy office for the data sharing agreement |
| EQ5. Experience and mechanisms: How do participants, connectors, clinicians and partners experience the program, which changes do participants value most, and through what mechanisms do changes occur? | ||||
| 5a. Themes on mechanisms, such as the connector relationship, role-based belonging and transportation. | Interviews with 20 participants at twelve weeks and 14 rural participants after the six-month review | Thematic analysis; joint display with EQ2 and EQ3 results | Months 6 and 12 | External qualitative evaluator |
| 5b. Most significant change stories in three domains, with reasons for selection. | Stories from participants and connectors | Selection panels; secondary analysis of all stories | Every six months | Coordinator; steering committee |
| 5c. Changes in the practices of community partners. | Partner documents and interviews | Outcome harvesting with substantiation | End of year one | External qualitative evaluator |
| 5d. Land-based pathway: indicators of connection defined with the First Nation. | Talking circles and conversations led by the Nation | Exploratory sequential design; data governed by the Nation under OCAP® | Set with the Nation | First Nation partner; connector at the health centre |
Several features of the matrix are deliberate. Every indicator traces to a column of the logic model and to one evaluation question. Each question that asks how or why has at least one qualitative source. Targets are stated where the steering committee has agreed them, and outcome standards are left to the rubric of Lesson 10. Indicators that require linked data are scheduled early because access takes months. The matrix was checked against capacity: the half-time analyst carries the quantitative indicators, while interviews, the outcome harvest and the six-month telephone survey are assigned to an external evaluator funded from the evaluation budget. The First Nation partner holds responsibility for the land-based pathway's indicators.
Checking an evaluation matrix
An evaluation matrix turns prioritized evaluation questions into a plan for gathering credible evidence on each of them. A draft can be checked against five points, which the Cedar Valley matrix was built to meet.
- Every indicator is clearly defined and traces to an evaluation question and a logic model component.
- Process and outcome questions both have suitable indicators, and questions about how or why have a qualitative source.
- Data sources are realistic for a British Columbia program, with privacy and governance requirements identified.
- Timing and responsibilities are feasible for the evaluation’s resources.
- Equity is addressed through named disaggregations.
Lessons 6 to 8 add the designs that answer the outcome questions, Lesson 9 adds implementation outcomes, and Lesson 10 adds costing and the rubric that turns the indicators into a judgement.
Reflection
The Cedar Valley Connector program is a fictional community connector (social prescribing) program run by a fictional health authority in British Columbia. Primary care clinicians refer adults aged 65 and older who score 6 or higher on the three-item UCLA Loneliness Scale (scored 3 to 9, higher scores meaning greater loneliness) to a connector, who meets them up to six times over twelve weeks, co-develops a plan, and links them to community groups, volunteer roles, transportation help and services. The program is run by a health authority, which is a public body under British Columbia's Freedom of Information and Protection of Privacy Act. The health authority plans to link its program database (referrals, meeting dates, UCLA scores and Personal Health Numbers) to its emergency department records, and to send an extract to a university partner for a difference-in-differences analysis. One connector position is hosted by a First Nations health centre, and some participants are First Nations. Before linkage, the analyst finds that a raw extract holds 319 referral records, of which 7 are duplicates; that 3 UCLA totals are recorded above 9; that 5 follow-up dates are earlier than the baseline dates; and that clinic electronic medical records list 327 referrals for the same period. (a) For each problem, name the data quality dimension involved (completeness, correctness, concordance, plausibility or currency) and the action you would take. (b) List the contents you would require in the data sharing agreement with the university and the privacy step required before the linkage. (c) Describe how data about First Nations participants should be governed.
(a) The duplicates are a correctness problem, since they inflate counts; I would remove records sharing a Personal Health Number and referral date, leaving 312, and log the rule. The totals above 9 are a plausibility problem; I would recalculate them from the item responses and add a validation rule restricting entries to 3 to 9. The date reversals are also a plausibility problem; I would check them against meeting logs, correct transposed days and months, and add a rule that follow-up must follow baseline. The 15 referrals in the clinic records without a program record are a concordance problem; I would trace them with the clinics, since they may be lost referrals, which is also a process finding.
(b) A privacy impact assessment is required before the linkage. The data sharing agreement should name the parties, purpose and legal authority; list a minimal set of variables; specify that the health authority performs the linkage and replaces Personal Health Numbers with study codes before release; state who may access the data, where they are stored and how they are secured; prohibit further linkage without approval; set rules for small-cell suppression and publication review; and set retention, destruction and breach-reporting procedures.
(c) Data about First Nations participants should be governed by an agreement with the partner Nation and the host health centre that reflects OCAP®. They decide whether these data go to the university, how results are reported, and they review findings before release.
Minimum 20 characters required.
Question 1: Three baseline UCLA totals in the Cedar Valley extract are recorded as 10 or higher. Which data quality dimension do these values fail?
Question 2: Clinic electronic medical records show 327 referral codes for the first six months, while the program database holds 312 unique referrals. What is the best interpretation?
Question 3: Under British Columbia's Freedom of Information and Protection of Privacy Act, what must a health authority complete for a new program or initiative that involves personal information?
Question 4: Which design choice for a Cedar Valley quarterly dashboard best follows the principles in Section 4?
Final Assessment
Bringing It All Together
This lesson filled in the columns of the evaluation matrix that Lesson 4 set out. Indicators turn evaluation questions into observable quantities and descriptions, and their specification, with numerators, denominators, data sources, disaggregations, baselines and targets, lets them be calculated consistently and checked by others. Performance measurement tracks those indicators routinely, while evaluation explains the patterns they reveal and judges what the program achieved.
Process evaluation supplies the evidence needed to interpret outcomes. Reach, dose delivered, dose received, fidelity, adaptation, recruitment and context describe how the program was delivered and received, and the Medical Research Council guidance adds mechanisms of impact. Mixed methods designs integrate this evidence with qualitative accounts, so that the Cedar Valley evaluation can explain, for example, why rural participants attended fewer meetings. Data systems, from program databases and electronic medical records to linked administrative data, make all of this possible when they are planned, checked and governed carefully.
The final reflection asks you to write matrix rows for a new Cedar Valley question, and the knowledge check below integrates ideas from all four sections.
Key Takeaways from this lesson
- An indicator is a specific, observable and measurable characteristic that shows whether a program component is in place or whether an expected change has occurred.
- Input, process, output and outcome indicators correspond to the columns of a logic model, and each answers a different kind of question.
- A full indicator specification states the numerator, denominator, data source, frequency, disaggregation, baseline, target and responsible person, and the choice of denominator determines what question the indicator answers.
- The CREAM criteria and the measurement properties of validity, reliability and sensitivity to change help a team choose among candidate indicators.
- Performance measurement tracks indicators continuously against targets, while evaluation explains patterns and judges effects, and indicators tied to rewards or rankings are prone to the distortions described by Campbell.
- Process evaluation measures how a program was delivered and received, and without it an evaluation risks a Type III error.
- Fidelity is best judged against the functions of a program's core components, which allows their forms to be adapted to local context.
- Convergent, explanatory sequential and exploratory sequential designs integrate quantitative and qualitative data at different points, and joint displays state the fit and meta-inference for each domain.
- Program, electronic medical record and administrative data each have strengths and limitations, and they require systematic quality checks, privacy impact assessment, data sharing agreements and respect for Indigenous data governance.
- A complete evaluation matrix links every indicator to an evaluation question and states its data source, method, timing and responsibility within the resources available.
Core Concepts Reviewed
Section 1: indicators; input, process, output and outcome indicators; indicator reference sheets; baselines and targets; the CREAM criteria; performance measurement compared with evaluation; and Campbell's law.
Section 2: process evaluation; Type III error; reach, dose delivered, dose received, fidelity, implementation, recruitment, adaptation and context; complete and acceptable delivery; and the Medical Research Council framework of implementation, mechanisms of impact and context.
Section 3: the purposes of mixed methods; convergent, explanatory sequential and exploratory sequential designs; integration and fit; joint displays and meta-inferences; most significant change; and outcome harvesting.
Section 4: data collection plans and data dictionaries; program, electronic medical record and administrative data; data quality dimensions; data sharing agreements, privacy impact assessments and OCAP®; dashboards; and the evaluation matrix.
The final reflection asks you to apply the whole lesson by writing evaluation matrix rows for a new question about the Cedar Valley Connector program.
Reflection
The Cedar Valley Connector program is a fictional community connector (social prescribing) program run by a fictional health authority in British Columbia. Primary care clinicians refer adults aged 65 and older who score 6 or higher on the three-item UCLA Loneliness Scale (scored 3 to 9, higher scores meaning greater loneliness) to a connector, who meets them up to six times over twelve weeks, co-develops a plan, and links them to community groups, volunteer roles, transportation help and services. The program operates in four rural and eight urban first-wave clinics. In the first six months, rural participants attended a mean of 3.39 connector meetings (48.4 percent attended four or more), compared with 4.38 meetings (73.4 percent) at urban clinics, and connectors recorded transportation needs for 149 of 205 participants. The transport fund is $40,000 a year. Connectors record meetings, linkages and a fidelity checklist in a program database, evaluation staff conduct a six-month telephone survey, the program has one half-time analyst, and an external qualitative evaluator is available. The steering committee adds a new evaluation question: Does the transportation help component reduce barriers to attendance for participants at rural clinics? Write the evaluation matrix rows for this question. Include at least one process indicator and one outcome indicator with definitions and targets; the data source, method, timing and person responsible for each; a mixed methods component with its point of integration; one data quality check; and one equity consideration.
Process indicator. Transportation assessment: participants with transportation needs recorded at the first or second meeting ÷ all participants, by rural or urban clinic. The baseline is 72.7 percent (149 of 205) and the target is 90 percent. Source: fidelity checklist. Method: proportion by clinic. Timing: quarterly. Responsibility: coordinator.
Output indicator. Transport support received: rural participants with an identified need who received a voucher or volunteer ride ÷ rural participants with an identified need. Target: 85 percent. Source: transport fund ledger linked to the program database. Timing: quarterly. Responsibility: analyst.
Outcome indicator. Rural dose received: rural participants attending four or more meetings ÷ rural participants reaching twelve weeks. Baseline 48.4 percent, target 65 percent, and a target rural-urban gap of 10 percentage points or less. Source: meeting log. Method: proportions with confidence intervals, compared before and after any change to transport support. Responsibility: analyst.
Mixed methods. An explanatory sequential component: after the next quarterly report, the external evaluator interviews about 12 rural participants, selected by whether they received transport support and by attendance, together with the rural connectors. Integration occurs by connecting (the sample) and in a joint display. Timing: months 6 to 9.
Data quality and equity. The analyst checks concordance between ledger entries and meeting dates each quarter. Results are disaggregated by gender and language, and the First Nations host clinic's results are reviewed with partners before reporting.
Minimum 30 characters required.
Final Knowledge Assessment
Question 1: Which of the following is the best-specified indicator?
Question 2: The mean three-item UCLA score among 188 Cedar Valley participants fell from 7.1 to 6.3. Why can this change not be attributed to the program on its own?
Question 3: Which indicator would make creaming, the selection of participants who are easiest to serve, most visible in Cedar Valley's reporting?
Question 4: At which step of Saunders, Evans and Joshi's (2005) process for planning a process evaluation does the team define complete and acceptable delivery?
Question 5: Why does the Medical Research Council guidance suggest that, where possible, process data be analyzed before outcome results are known?
Question 6: Which design is best suited to developing indicators of connection for the land-based pathway that reflect what connection means in the partner Nation's own terms?
Question 7: In a joint display, which column states the conclusion that draws on both the quantitative and the qualitative results?
Question 8: Which data source is appropriate for indicator 4a in the Cedar Valley matrix, emergency department visits per 1,000 person-years in the twelve months before and after referral?
Question 9: Connectors complete the Cedar Valley fidelity checklist themselves, so it may overstate delivery. What is the most suitable response?
Question 10: A member of the executive proposes paying clinics a bonus for each referral. What does Campbell's law predict?
Question 11: What is the main limitation of the most significant change technique, and how do its users address it?
Question 12: Of 205 Cedar Valley participants who reached twelve weeks, 135 attended four or more meetings. The target is 70 percent. Which statement is correct?
Question 13: Which question belongs to evaluation more than to performance measurement?
Question 14: How should a data sharing agreement treat data about First Nations participants in the Cedar Valley evaluation?
Question 15: Which feature of the Cedar Valley evaluation matrix shows that it was checked for feasibility?
Glossary: Key Terms, People & Frameworks
📚 Reference page, available throughout the lesson
These terms, frameworks and people appear in Lesson 5, and the definitions follow the way the lesson uses them.