Foundations of Program Planning and Evaluation
Program Planning & Evaluation
Learning objectives for this lesson:
- Define evaluation as the systematic determination of the merit, worth or significance of a program, and distinguish evaluation from research by purpose and audience.
- Use criteria, standards and a rubric to move from evidence to an evaluative judgement about a program.
- Describe needs assessment, formative, process, outcome, impact and economic evaluation, and place each in the program life cycle.
- Compare utilization-focused, developmental, participatory, empowerment, theory-driven, realist, culturally responsive and Indigenous approaches using Alkin and Christie’s evaluation theory tree.
- Describe the Treasury Board Policy on Results, the Canadian Evaluation Society competencies and Credentialed Evaluator designation, and the Program Evaluation Standards.
- Apply Article 2.5 of the Tri-Council Policy Statement and the ARECCI screening tool to decide what ethical review an evaluation activity needs.
- Describe the steps, standards and cross-cutting actions of the CDC evaluation framework in its 1999 and 2024 versions, and apply them to the fictional Cedar Valley Connector program.
- Write a program description covering a program’s purpose, population, activities, resources, setting and history, using the Cedar Valley Connector program as a worked example.
Quantitative background for later lessons:
Lessons 6 to 8 use linear and logistic regression, product terms and clustered standard errors to estimate program effects. Each of those lessons restates the method in a short Background box where it is first used, and HSCI 410 Lessons 3 to 5 (linear and logistic regression, generalized linear models, and modelling dependent data) offer fuller optional reading.
This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University, drawing on Rossi, P. H., Lipsey, M. W., & Henry, G. T. (2019). Evaluation: A Systematic Approach (8th ed.). SAGE; and Kidder, D. P., et al. (2024). CDC Program Evaluation Framework, 2024. MMWR Recommendations and Reports, 73(6).
What Evaluation Is: Merit, Worth and the Program Life Cycle
Learning Objectives for this section
- Define evaluation as the systematic determination of the merit, worth or significance of a program, and distinguish the three terms with examples.
- Distinguish evaluation from research by purpose, audience and the use made of findings, and recognize where the two overlap.
- Explain evaluative reasoning as a movement from criteria to standards to evidence to a synthesized judgement, and apply a simple rubric.
- Describe needs assessment, formative, process, outcome, impact and economic evaluation, and place each in the program life cycle.
- Describe the fictional Cedar Valley Connector program that serves as the running case for the course.
Introduction
A program is an organized set of activities, delivered by people with resources, intended to change something for a defined population. A school vaccination clinic, a provincial quit-smoking line and a community kitchen are all programs, and each rests on claims that someone has to check: that the program is needed, that it is delivered as intended, that it reaches the people it was designed for, that it changes their health or circumstances, and that its benefits justify its costs. Program evaluation checks such claims systematically and turns the evidence into judgements that people can act on.
This section defines evaluation, separates it from research, introduces the reasoning that turns evidence into a judgement of value, and sets out the main types of evaluation across a program’s life. Sections 2 to 4 then survey approaches to evaluation, the Canadian context, and the evaluation cycle that organizes the course.
1.1 The Running Case: The Cedar Valley Connector Program
Each lesson uses one invented program as its main worked example, so that concepts from different weeks can be seen working together on the same case.
The Cedar Valley Connector program is a fictional community connector (social prescribing) program run by the fictional Cedar Valley Health Authority in British Columbia. Primary care clinicians refer adults aged 65 and older who screen as socially isolated or lonely to a community connector. The connector meets each person up to six times over twelve weeks, works with them on a plan, and links them to community groups, volunteer roles, transportation help and services. The program uses the three-item UCLA Loneliness Scale (Hughes et al., 2004), scored from 3 to 9, at intake and at twelve weeks, and refers people who score 6 or higher or whom a clinician judges to be socially isolated.
The health authority launched the program in 12 of the region’s 24 primary care clinics in a first wave. The other 12 clinics are scheduled to join in a second wave a year later. The first wave employs seven connectors, one of whom is hosted by a First Nations health centre under an Indigenous health partnership with local First Nations, and its annual budget is $840,000. A steering committee that includes four older adults with lived experience of loneliness oversees the program. The outcomes of interest are loneliness, social participation, self-rated health, emergency department visits and primary care visits. All figures about Cedar Valley in this course are illustrative.
The case raises most of the questions this course addresses. Managers want to know whether referred adults come to a first meeting. The steering committee wants to know whether people feel less lonely, and for whom. The First Nations partners want to know whether the program is culturally safe. The executive must decide whether the second wave proceeds, and the staggered launch may let an evaluator compare clinics with and without the program, a possibility that Lessons 6 and 7 examine.
1.2 Defining Evaluation
The most widely used definition of evaluation comes from the philosopher and evaluation theorist Michael Scriven, for whom evaluation is the systematic determination of the merit, worth or significance of something (Scriven, 1991; Davidson, 2005). The definition places judgement at the centre of the work. An evaluation that describes what a program did, without reaching any conclusion about how good, valuable or important it was, has stopped short of evaluation in Scriven’s sense. The three terms in the definition name different kinds of value.
Merit is the intrinsic quality of a program, judged against what a program of its kind should do, without reference to cost or to a particular setting. A connector program has merit if its connectors are well trained, its plans are tailored to each person, its links to community groups are appropriate, and participants become less lonely. A program can have high merit and still be a poor investment for a given health authority.
Worth is the value of a program to a particular organization or population in a particular context, which brings in costs, alternatives and local needs. The Cedar Valley program might have high merit and still have limited worth to the health authority if a cheaper volunteer visiting program produced similar benefits, or if the clinics it serves already have strong links to community services.
Significance is the importance of a program or of its effects, including its importance beyond the immediate setting. A modest reduction in loneliness among older adults might be significant if loneliness is common in the region and if few other programs address it, or if the program shows how primary care and community services can work together in ways that other regions could adopt.
Other definitions emphasize different features of the same activity. Carol Weiss (1998) defined evaluation as the systematic assessment of the operation or outcomes of a program, compared with explicit or implicit standards, as a means of contributing to its improvement. Rossi, Lipsey and Henry (2019) define it as the use of social research methods to investigate the effectiveness of programs in ways adapted to their political and organizational settings. The definitions share three elements: evaluation is systematic, it compares findings with a standard of what ought to be, and it is meant to inform decisions about a particular program.
Formative and summative purposes
Scriven (1967) also introduced the distinction between formative and summative evaluation. Formative evaluation is conducted to improve a program while it is being developed or delivered, and its main users are the people who run the program. Summative evaluation is conducted to reach an overall judgement about a program, usually to inform a decision about whether to continue, expand, change or end it, and its main users are funders and decision-makers. Robert Stake’s analogy, which Scriven (1991) reports, compares formative evaluation to the cook tasting the soup and summative evaluation to the guests tasting it. The same method can serve either purpose: a survey of Cedar Valley participants is formative if connectors use it to adjust their practice, and summative if the executive uses it to decide whether the second wave proceeds.
1.3 Evaluation and Research
Evaluation and research use the same methods: surveys, interviews, administrative data, experiments and statistical models. Students who come to evaluation from research training often treat the two as the same activity. Mathison (2008) argues that the two share methods and differ most clearly in purpose, and the difference in purpose gives each a different primary audience. Research aims to produce knowledge that holds beyond the setting in which it was produced, and its primary audience is other researchers. Evaluation aims to inform judgements and decisions about a particular program, and its primary audience is the people who make or are affected by those decisions.
| Feature | Evaluation | Research |
|---|---|---|
| Primary purpose | To judge the merit, worth or significance of a particular program and inform decisions about it | To produce knowledge that generalizes beyond the setting of the study |
| Primary audience | Program managers, funders, participants, communities and policy-makers | Researchers and the scientific literature |
| Who sets the questions | Usually negotiated with the people who will use the findings | Usually the investigator, guided by gaps in the literature |
| Role of values | Values are explicit, since criteria and standards define what counts as good | Values shape the choice of topic, and findings are usually reported without a judgement of worth |
| Timing | Set by decision cycles, budgets and program schedules | Set mainly by the logic of the study and funding cycles |
| Test of success | Whether findings are credible and used to improve or decide about the program | Whether findings are valid, replicable and add to theory |
The distinction is a matter of emphasis, and many studies sit between the two. An impact evaluation that compares first-wave and second-wave Cedar Valley clinics would inform the health authority’s decision and could also add to the international evidence on social prescribing. The difference still matters in practice: it affects who frames the questions, how quickly findings are needed, and, as Section 3 explains, whether research ethics board review is required.
Two meanings of “standard”
Evaluators use the word standard in two senses. A performance standard is a level of performance on a criterion that defines what counts as good enough, such as “at least 80 percent of referred adults attend a first meeting.” An evaluation standard is a principle for judging the quality of an evaluation itself, such as the Program Evaluation Standards introduced in Section 3. This section uses the first sense.
1.4 Evaluative Reasoning: Criteria, Standards and Rubrics
If evaluation ends in a judgement, an evaluator needs a defensible way of getting from evidence to that judgement. Scriven described a general logic of evaluation, which Fournier (1995) set out as four steps. The evaluator establishes the criteria of merit, which are the dimensions on which the program will be judged. The evaluator then constructs standards, which specify how well the program must perform on each criterion to be judged excellent, adequate or poor. The evaluator measures performance and compares it with the standards. Finally, the evaluator synthesizes the results across criteria into an overall judgement of merit or worth.
Where criteria come from
A good evaluation states where its criteria came from. The program’s objectives are the obvious source, although Davidson (2005) notes that objectives can omit important effects or be set at levels that are easy to reach, and Scriven’s proposal of goal-free evaluation asked evaluators to judge a program’s actual effects, intended or unintended, against the needs of the people it serves. Other sources include assessed needs, ethical and legal requirements, professional standards of practice, what comparable programs achieve, and the values of interest holders, including participants and communities. For Cedar Valley, a criterion of cultural safety comes from the Indigenous health partnership, a criterion of reach from the concern that referrals be acted on, and a criterion of cost from the executive’s need to compare the program with other uses of the money.
Rubrics
An evaluative rubric sets out the criteria in rows and describes, for each criterion, what performance at each level looks like. Rubrics make the evaluator’s reasoning visible, and they allow interest holders to agree on standards before the evidence arrives, which protects the evaluation from the temptation to move the standards once results are known (Davidson, 2005; King et al., 2013). The rubric below is the one the Cedar Valley steering committee agreed for a six-month review of the first wave. Lesson 10 returns to rubrics and to the synthesis of mixed evidence in more depth.
| Criterion | Excellent | Good | Adequate | Poor |
|---|---|---|---|---|
| Reach: percentage of referred adults who attend a first meeting | 80 percent or more | 65 to 79 percent | 50 to 64 percent | Below 50 percent |
| Change in mean loneliness score (3 to 9) from intake to twelve weeks | Fall of 1.0 point or more | Fall of 0.5 to 0.9 points | Fall of 0.2 to 0.4 points | Fall below 0.2 points, or a rise |
| Cultural safety, as judged by Indigenous participants and partners | Described as safe and respectful, with no unresolved concerns | Generally safe, with concerns resolved promptly | Mixed reports, with concerns acknowledged and a plan in place | Reports of disrespect or harm without a response |
| Cost per adult who attends a first meeting | $1,200 or less | $1,201 to $1,600 | $1,601 to $2,000 | More than $2,000 |
The committee also agreed a minimum requirement for synthesis. The program could not be rated better than adequate overall if cultural safety was rated poor, whatever its performance on the other criteria. Such a requirement is sometimes called a bar or a hurdle, and it records the committee’s view that a program which harms some participants cannot be rescued by good results for others.
In the first six months, clinicians referred 312 older adults, and 241 attended a first meeting. Reach was 241 ÷ 312 = 77.2 percent, which the rubric rates as good. Among the 188 participants with loneliness scores at intake and at twelve weeks, the mean score fell from 7.1 to 6.3, a fall of 0.8 points, which the rubric also rates as good. Program spending over the six months was $420,000, so the cost per adult who attended a first meeting was $420,000 ÷ 241 = $1,743, which the rubric rates as adequate. A talking circle with Indigenous participants and a review meeting with the First Nations partners described the program as generally safe, and one concern about the wording of an intake question had been resolved, which the committee rated as good.
The committee synthesized these ratings as good overall, with two qualifications recorded in the report. The cost per participant was expected to fall as connectors reached full caseloads, since start-up costs were spread over a small number of early participants. The fall in loneliness scores could not be attributed to the program, because people are referred when their scores are high, and high scores tend to fall on remeasurement even without intervention. This pattern, regression to the mean, is one of the threats to validity that Lesson 7 teaches.
Suppose that in a later six-month period clinicians referred 400 older adults and 268 attended a first meeting, spending was $380,000, and the mean loneliness score among participants with both measurements fell from 7.0 to 6.6. Use the rubric to rate reach, change in loneliness and cost. (Reach is 268 ÷ 400 = 67.0 percent, which is good. The fall in loneliness is 0.4 points, which is adequate. The cost per adult attending is $380,000 ÷ 268 = $1,418, which is good.) Consider what the committee should conclude if, in the same period, the First Nations partners reported a pattern of disrespectful treatment at one clinic that had not been addressed.
1.5 Types of Evaluation Across the Program Life Cycle
Programs pass through a life cycle. A need is identified, a program is designed, it is implemented and adjusted, it settles into stable delivery, and eventually it is reviewed, scaled up, changed or ended. Different evaluation questions make sense at different stages, and the most common types of evaluation can be placed on that cycle.
Rossi, Lipsey and Henry (2019) arrange these types as a hierarchy in which each level depends on those beneath it: an assessment of need, then of program design and theory, then of implementation, then of outcomes and impact, and finally of cost and efficiency. The ordering has a practical lesson. An impact evaluation of a program that was never delivered as designed tells decision-makers little, because a null result cannot distinguish a poor idea from a poorly implemented one. Section 2 returns to this point through Weiss’s distinction between implementation failure and theory failure.
A note on “outcome” and “impact”
These two words are used inconsistently. In parts of the health promotion literature, impact evaluation refers to a program’s immediate effects and outcome evaluation to its longer-term effects. In development economics and much of the causal inference literature, impact evaluation refers specifically to estimating a causal effect against a counterfactual. This course uses the second meaning. When you read an evaluation report, check how its authors define the terms before comparing its findings with another report.
Daniel Stufflebeam’s CIPP model divides the cycle in a similar way, organizing evaluation around a program’s context, input, process and product. Several neighbouring activities are often confused with evaluation.
Performance measurement is the routine tracking of a small set of indicators, such as the number of referrals per month or the percentage of participants who complete twelve weeks. It tells managers what is happening and alerts them to change. Evaluation is a periodic, more intensive inquiry that asks why results look as they do and whether the program is good or worthwhile. The two depend on each other, since good performance data make evaluation cheaper and evaluation tells managers which indicators matter. Lesson 5 develops the distinction.
Quality improvement uses rapid, small-scale cycles of change and measurement, such as Plan-Do-Study-Act cycles, to improve a process within an organization. Its methods overlap with formative and process evaluation. Lesson 8 compares quality improvement designs with the evaluation designs taught in Lessons 6 to 8, and Section 3 of this lesson explains how Canadian research ethics policy treats both activities.
The Cedar Valley program reports monthly referrals, first-meeting attendance and completions to the health authority, which is monitoring. The six-month review above became evaluation when the steering committee applied criteria and standards to those data, added evidence on cultural safety and cost, and reached a judgement.
An evaluation plan usually combines several of these types. A plan for Cedar Valley might combine a process evaluation of the first wave, an impact evaluation comparing first-wave and second-wave clinics, and an economic evaluation that draws on both. Which questions take priority depends partly on the evaluator’s approach, the subject of Section 2.
Reflection
A steering committee agreed the following rubric before reviewing a community connector program for older adults. Reach (percentage of referred adults who attend a first meeting): excellent 80 percent or more, good 65 to 79 percent, adequate 50 to 64 percent, poor below 50 percent. Change in mean loneliness score on a scale from 3 to 9, from intake to twelve weeks: excellent a fall of 1.0 point or more, good a fall of 0.5 to 0.9, adequate a fall of 0.2 to 0.4, poor a fall below 0.2 or a rise. Cost per adult who attends a first meeting: excellent $1,200 or less, good $1,201 to $1,600, adequate $1,601 to $2,000, poor more than $2,000. In one quarter, 180 adults were referred and 99 attended a first meeting, program spending was $150,000, and among participants with both measurements the mean loneliness score fell from 7.4 to 6.2. Adults are referred only if they score 6 or higher. Rate each criterion and show your calculations, state an overall judgement, and explain what the loneliness result can and cannot tell the committee about the program’s effect.
Reach is 99 ÷ 180 = 55.0 percent, which the rubric rates as adequate. Cost per adult attending is $150,000 ÷ 99 = $1,515, which is good. The mean loneliness score fell by 7.4 − 6.2 = 1.2 points, which meets the descriptive standard for excellent.
An overall judgement of good, with a clear concern about reach, is defensible. Almost half of referred adults never reached a first meeting, which limits whatever benefit the program offers, so the committee should look at why referrals are not converted, for example by asking clinicians and connectors where contact fails and whether transportation or timing is the barrier.
The loneliness result describes change among participants who had both measurements. It cannot show that the program caused the change. Because people are referred only when they score 6 or higher, their scores would be expected to fall on remeasurement even without the program, which is regression to the mean. The result also excludes people who dropped out, who may have fared worse. A claim about the program’s effect needs a comparison with a credible counterfactual, such as clinics that have not yet started the program. A strong answer might also note that a rubric agreed in advance protects the committee from adjusting its standards after seeing the results.
Minimum 20 characters required.
Question 1: In Scriven’s definition of evaluation, how does the worth of a program differ from its merit?
Question 2: According to Mathison (2008), evaluation and research differ most clearly in which respect?
Question 3: A rubric rates reach as good from 65 to 79 percent. Of 250 referred adults, 170 attend a first meeting. How is reach rated, and why is it useful to agree the rubric before the data arrive?
Question 4: Among Cedar Valley participants with both measurements, the mean loneliness score fell from 7.1 to 6.3. Which type of evaluation would be needed to claim that the program caused the fall?
Approaches to Evaluation and the Evaluation Theory Tree
Learning Objectives for this section
- Describe Alkin and Christie’s evaluation theory tree and explain what its methods, use and valuing branches represent.
- Explain utilization-focused, developmental, participatory and empowerment evaluation, and the role each gives to intended users and communities.
- Explain theory-driven and realist evaluation, including the distinction between implementation failure and theory failure and the context-mechanism-outcome configuration.
- Describe culturally responsive evaluation and Indigenous evaluation, and explain why they treat culture, relationships and self-determination as matters of validity and ethics.
- Select and combine approaches for a given evaluation and justify the choice.
Introduction
Two competent evaluators given the same program can design very different evaluations. One may begin by asking the health authority’s executive what decision the evaluation must inform. Another may begin by mapping how the program is supposed to produce change. A third may begin by asking the communities served what a good program would mean to them. Each is following an evaluation approach, a set of prescriptions about whose questions an evaluation should answer, what counts as credible evidence, how judgements of value should be made, and what role the evaluator should play. Evaluation theorists use the word theory for these prescriptive models (Alkin & Christie, 2004). This section surveys the main approaches and organizes them on the evaluation theory tree.
2.1 The Evaluation Theory Tree
Marvin Alkin and Christina Christie (2004) drew the history of evaluation theory as a tree. Its roots are the two traditions from which evaluation grew: a demand for accountability, meaning that those who spend public or charitable money should show what it achieved, and systematic social inquiry, which supplied the methods for doing so. The trunk divides into three branches according to the main concern of each theorist. The methods branch, which grew from Donald Campbell’s work on experimental and quasi-experimental designs, is concerned above all with producing credible knowledge about program effects. The use branch is concerned with ensuring that evaluations are used by the people who make decisions. The valuing branch, which grows from Michael Scriven’s insistence that evaluation must reach judgements of value, is concerned with how those judgements are made and whose values count.
The tree has been revised several times. Later editions of Evaluation Roots moved theorists along and between branches, and the third edition (Alkin & Christie, 2023) categorizes approaches in place of individual theorists, placing culturally responsive, transformative and Indigenous approaches on the valuing branch. Mertens and Wilson (2012) proposed a fourth branch for approaches centred on social justice. The placements describe each approach’s primary emphasis. Every serious evaluation attends to methods, use and values to some degree, and the tree is most useful as a reminder of which of the three an approach is prepared to trade away when they conflict.
2.2 The Use Branch
The approaches on the use branch begin from the observation that many evaluations are commissioned, completed and then ignored. They respond by building the evaluation around the people expected to act on it.
Michael Quinn Patton’s utilization-focused evaluation holds that an evaluation should be judged by its usefulness to its intended users, and that it should be designed from the start for intended use by intended users (Patton, 2008). Patton’s early research on federal health evaluations identified what he called the personal factor: evaluations were most often used when an identifiable person or group cared about the findings and was involved in shaping the study. The evaluator therefore begins by identifying the primary intended users, negotiates with them the purpose and questions of the evaluation, and involves them in decisions about methods and interpretation. For Cedar Valley, the primary intended users are the steering committee and the director who will decide how the second wave is rolled out.
Developmental evaluation supports the development of innovations in complex and changing conditions, where the program’s form is still emerging and may never settle into a fixed model (Patton, 2011). The evaluator works as part of the innovation team, brings data and evaluative questions into decisions as they arise, and documents how and why the innovation changes. Patton distinguishes it from formative evaluation, which improves a model that is expected to stabilize and later be judged. A Canadian primer by Gamble (2008), published by the J.W. McConnell Family Foundation, helped introduce the approach to community organizations. At Cedar Valley, a land-based connection pathway being co-designed with one First Nation is a candidate for developmental evaluation, since neither partner yet knows what its final form will be.
Participatory evaluation shares the work of evaluation with people who are not professional evaluators. Cousins and Whitmore (1998) distinguished practical participatory evaluation, which involves program staff and managers in order to increase the use of findings, from transformative participatory evaluation, which involves community members and participants in order to redistribute power and support social change. They described participatory approaches along three dimensions: who controls technical decisions, how diverse the participating groups are, and how deeply they take part. At Cedar Valley, the older adults on the steering committee could help design the participant survey, review its wording for clarity, and interpret the results.
Empowerment evaluation, developed by David Fetterman (1994), helps program staff and communities to evaluate their own work, with the evaluator acting as coach and critical friend. Fetterman and Wandersman (2005) set out ten principles: improvement, community ownership, inclusion, democratic participation, social justice, community knowledge, evidence-based strategies, capacity building, organizational learning and accountability. Critics, including Stufflebeam (1994), argued that the approach weakens the independence on which credible judgements depend. Its supporters reply that self-evaluation builds lasting capacity and that independence can be preserved through external review.
Use-branch approaches draw attention to the different ways in which evaluations are used. Instrumental use occurs when findings directly inform a decision, such as changing the connector training. Conceptual use occurs when findings change how people think about a problem, for example by showing managers that transport is the main barrier to participation. Process use, a term Patton introduced, refers to the learning that occurs among people who take part in an evaluation, regardless of its findings. Symbolic use occurs when an evaluation is commissioned or cited to legitimize a decision already made. Lesson 10 examines the conditions under which evaluations are used.
2.3 The Methods Branch
The methods branch descends from Campbell’s work on experimental and quasi-experimental designs and from the view that an evaluation’s first duty is to produce credible evidence about whether a program causes the changes it claims. Lessons 6 to 8 of this course belong to this tradition. Two approaches on the branch focus on explaining how programs work.
Theory-driven evaluation
Huey-Tsyh Chen (1990, 2015) argued that evaluations which test only whether outcomes changed treat the program as a black box. Theory-driven evaluation makes the program’s theory explicit and tests its links: what the program does, what it is expected to change first, and how those early changes are expected to lead to the final outcomes. Chen separates the change model, which states the causal process the program relies on, from the action model, which states the arrangements needed to deliver it. Lesson 3 teaches both.
Program theory helps an evaluator interpret disappointing results. Weiss (1998) distinguished implementation failure, in which a program is not delivered as planned, from theory failure, in which the program is delivered as planned but the causal process it relies on does not occur. Suppose loneliness scores at Cedar Valley do not fall. If connectors held few meetings and made few linkages, the problem is implementation. If meetings and linkages happened as planned, but people who joined groups felt no less lonely, the problem lies in the theory, and the program itself needs rethinking.
Realist evaluation
Ray Pawson and Nick Tilley (1997) argued that programs do not work the same way for everyone, so the useful question is what works, for whom, in what circumstances, and why. Realist evaluation holds that programs produce outcomes by triggering mechanisms, which are the ways participants respond to the resources a program offers, and that mechanisms operate only in certain contexts. Its analytic unit is the context-mechanism-outcome configuration. Reporting standards for realist evaluations, known as RAMESES II, were published by Wong and colleagues (2016). This lesson places realist evaluation on the methods branch beside theory-driven evaluation because its main contribution is an account of causal explanation.
For older adults who have recently lost a spouse and have no existing link to community groups (context), a connector who goes with them to the first meeting of a walking group (resource) may ease their anxiety about entering a room of strangers (reasoning), so that they keep attending and feel less lonely (outcome). For older adults in rural areas who have stopped driving (context), the program’s transportation help (resource) makes attendance feasible and predictable (reasoning), with the same outcome. The two groups reach a similar result through different mechanisms, and an evaluation that reported only the average change would miss both. Lesson 3 teaches how to build and test such configurations.
2.4 The Valuing Branch
The valuing branch starts from Scriven’s argument, introduced in Section 1, that evaluation must reach judgements of value and that the evaluator is responsible for making them defensible. Later theorists on the branch asked whose values should define merit. Robert Stake’s responsive evaluation organized evaluations around the concerns of the people involved in a program, and Ernest House and Jennifer Greene argued that evaluations should give fair representation to the interests of less powerful groups. Two families of approaches on this branch are central to health program evaluation in Canada.
Culturally responsive evaluation
Culturally responsive evaluation places the culture and context of the community served at the centre of every stage of an evaluation, from framing the questions to sharing the findings (Hood, Hopson & Kirkhart, 2015). Hood and his colleagues trace its roots to African American evaluators and scholars of education in the United States, whose work showed that evaluations designed without attention to culture could misread programs serving their communities. It asks evaluators to examine their own cultural position, to choose methods and measures that make sense to participants, and to involve community members in interpreting results. Kirkhart (1995) argued that culture is a matter of validity, since conclusions drawn through measures or relationships that participants do not recognize can be wrong. The American Evaluation Association’s Public Statement on Cultural Competence in Evaluation (2011) took a similar position. Mertens’s transformative approach extends this reasoning by orienting the whole evaluation toward human rights and social justice.
Indigenous evaluation
Indigenous evaluation is evaluation grounded in Indigenous knowledge systems, values and governance, and led by or conducted in partnership with Indigenous communities. Its principles include relationships built on respect and trust, accountability to the community, self-determination over questions, data and interpretation, and the recognition of Indigenous ways of knowing as legitimate evidence. Kirkness and Barnhardt’s (1991) four Rs of respect, relevance, reciprocity and responsibility, first written about First Nations students in higher education, are often cited as a starting point. LaFrance and Nichols (2009) developed an Indigenous Evaluation Framework with tribal colleges in the United States, and Chouinard and Cousins (2007) reviewed evaluation practice with Indigenous communities in Canada and elsewhere. Cram and Chouinard (2023) describe culturally responsive Indigenous evaluation as an approach in its own right on the valuing branch.
Indigenous evaluation in this course
Lesson 4 teaches the principles of relational accountability, Two-Eyed Seeing and OCAP® (ownership, control, access and possession), and Section 3 of this lesson covers the research ethics policy for work with First Nations, Inuit and Métis Peoples. In the Cedar Valley case, Indigenous evaluation means that the First Nations partners help decide which outcomes matter, such as connection to family, culture and land, how data about their members are held and shared, and how findings are interpreted and reported. Frameworks from one Nation should not be assumed to apply to another, and evaluators learn the relevant protocols from the communities involved.
2.5 Choosing and Combining Approaches
Approaches are tools, and a single evaluation often combines several. The choice depends on the purpose of the evaluation, the stage of the program, the decisions to be made, the relationships among those involved, and the resources available. The table compares the approaches in this section.
| Approach | Central question | Role of the evaluator | Use at Cedar Valley |
|---|---|---|---|
| Utilization-focused | What do the primary intended users need to know to act? | Facilitator and negotiator with intended users | Overall frame for the evaluation of the first wave |
| Developmental | How is the innovation developing, and what should change next? | Embedded member of the innovation team | The land-based pathway co-designed with one First Nation |
| Participatory | How can those affected share in the evaluation? | Partner and trainer | Older adults on the steering committee shape the survey and interpretation |
| Empowerment | How can the program evaluate itself? | Coach and critical friend | Connectors build routine self-assessment into practice |
| Theory-driven | Do the links in the program theory hold? | Analyst of causal processes | Structure of the outcome evaluation |
| Realist | What works, for whom, in what circumstances, and why? | Theory builder and tester | Explaining variation across clinics and groups |
| Culturally responsive and Indigenous | Whose values and knowledge define merit, and is the evaluation valid for this community? | Partner accountable to the community | Criteria, data governance and interpretation with First Nations partners |
Approaches on the use branch bring evaluators close to programs, while many methods-branch evaluators prize distance from them. Close involvement improves relevance and use, and it can also make it harder to report unwelcome findings. Evaluators manage the tension by being explicit about their role, documenting decisions, and arranging external review where independence matters to decision-makers.
Designs that produce strong causal evidence often require fixed protocols, while developmental and participatory approaches expect the evaluation to change as it proceeds. A combined evaluation can protect a fixed design for its impact questions, such as the comparison of first-wave and second-wave clinics, while allowing other strands to adapt.
Different interest holders may hold different criteria of merit. The executive may weigh cost and emergency department use, older adults may weigh friendship and purpose, and First Nations partners may weigh cultural continuity. Valuing-branch approaches ask evaluators to make these differences visible in the criteria and rubric, and Lesson 4 teaches how to work with them when setting evaluation questions.
Three requests reach the Cedar Valley evaluation team. The executive asks whether the program should be extended to the second-wave clinics without changes. A First Nations partner asks for help learning, month by month, how the new land-based pathway is working while it is still being designed. The steering committee asks why some clinics have much higher first-meeting attendance than others. For each request, name the approach that fits best and give one reason. (A strong answer pairs the first with a utilization-focused evaluation built around the executive’s decision, using a design from the methods branch for its impact question; the second with developmental evaluation conducted in partnership with, and accountable to, the Nation; and the third with a realist or theory-driven analysis of how clinic context shapes referral and attendance.)
The approaches surveyed here have grown up in many countries, but evaluation in Canada also takes shape within particular policies, professional bodies and ethical rules. Section 3 turns to that context.
Reflection
Consider five evaluation approaches. Utilization-focused evaluation designs the evaluation around the intended use of named primary intended users. Developmental evaluation embeds the evaluator in an innovation team to support a program whose form is still emerging in complex, changing conditions. Realist evaluation asks what works, for whom, in what circumstances and why, using context-mechanism-outcome configurations. Empowerment evaluation helps program staff and communities evaluate their own work, with the evaluator acting as coach. Indigenous evaluation is grounded in Indigenous knowledge systems and governance and is led by, or conducted in partnership with, Indigenous communities. Now consider three situations. (a) A provincial ministry must decide in nine months whether to renew funding for a falls-prevention exercise program. (b) A Métis community organization is designing a new peer support initiative for family caregivers and expects its form to change as it learns. (c) A school mental health literacy program shows good average results, but teachers report that it works poorly in some schools. For each situation, choose the approach or combination that fits best, justify your choice, and name one risk of the approach you chose.
(a) A utilization-focused evaluation fits, because the ministry has a defined decision and deadline. The evaluator would identify the officials who will make the renewal decision, agree with them the questions that bear on it, such as reach, falls outcomes and cost, and plan reporting before the nine-month deadline. One risk is that the evaluation adopts only the ministry’s criteria and neglects what older participants value.
(b) Developmental evaluation conducted as Indigenous evaluation fits best. The initiative is still being designed, so an embedded evaluator can feed data into decisions as the form develops, and because the organization is Métis-led, the evaluation should be governed by the organization, with its knowledge and protocols shaping questions, data stewardship and interpretation. One risk is that close involvement makes it harder to report unwelcome findings, which the partners can manage through agreed reporting rules.
(c) A realist evaluation fits, because the question is why the same program works in some schools and not others. The evaluator would develop and test context-mechanism-outcome configurations, for example about teacher preparation or class size. One risk is that the approach needs rich data from enough schools to compare contexts, which takes time and money.
Minimum 20 characters required.
Question 1: On Alkin and Christie’s evaluation theory tree, which branch is chiefly concerned with how evaluation findings are used by decision-makers?
Question 2: What did Patton call the personal factor?
Question 3: Cedar Valley loneliness scores stay level, while records show that connectors held meetings and made linkages as planned. In Weiss’s terms, this pattern points to which kind of failure?
Question 4: Which question is a realist evaluator of the Cedar Valley program most likely to ask?
Evaluation in Canada: Policy, Profession, Standards and Ethics
Learning Objectives for this section
- Describe the Treasury Board Policy on Results (2016) and the main requirements it places on federal departments.
- Describe the Canadian Evaluation Society, the five domains of its competencies for evaluation practice, and the Credentialed Evaluator designation.
- Name the five attributes of the Program Evaluation Standards and apply them to judge the quality of an evaluation.
- Apply Article 2.5 of the Tri-Council Policy Statement to decide whether an evaluation activity requires research ethics board review, and describe how the ARECCI ethics screening tool supports that decision.
- Identify ethical issues that arise in evaluation practice and the ways evaluators address them.
Introduction
Evaluation in Canada has a particular history. The federal government has required departments to evaluate their programs since the 1970s, Canada has a national evaluation society with its own competencies and a professional designation, and Canadian research ethics policy draws an explicit line between program evaluation and research. A graduate student planning an evaluation of a British Columbia health program will meet all three. This section covers federal evaluation policy, the evaluation profession, the standards by which evaluations are judged, and the ethics of evaluation.
3.1 Federal Evaluation Policy and the Policy on Results
The Treasury Board issued its first formal program evaluation policy in 1977, and an Office of the Comptroller General created in 1978 housed a program evaluation branch. Successive policies followed, including the 2009 Policy on Evaluation. The current instrument is the Policy on Results, which took effect on July 1, 2016 and replaced the Policy on Evaluation and its related instruments. The policy is supported by a Directive on Results, which sets out the detailed requirements.
The Policy on Results joins performance measurement and evaluation in one framework. Its aim is to improve the achievement of results across government and to improve understanding of the results government seeks, the results it achieves and the resources it uses to achieve them. Each department organizes its work in a Departmental Results Framework, which sets out its core responsibilities, the departmental results it intends to achieve and the indicators that will track them, and in a Program Inventory, which lists its programs. Each program has a Performance Information Profile that states its expected results, indicators and data sources. Deputy heads designate a Head of Performance Measurement and a Head of Evaluation and establish a Performance Measurement and Evaluation Committee. Each year they approve a five-year departmental evaluation plan.
The policy gives departments more discretion than earlier policies over which programs to evaluate, so that evaluation can follow risk, need and priority. One requirement remains fixed. Under section 42.1 of the Financial Administration Act, ongoing programs of grants and contributions must be reviewed for relevance and effectiveness every five years, and the policy requires departments to evaluate such programs with five-year average actual expenditures of $5 million or more per year at least once in that period. Federal evaluations typically address questions of relevance, effectiveness and efficiency, and departments publish their evaluation reports with management responses. The evaluation reports of the Public Health Agency of Canada and Health Canada are useful models of the format.
British Columbia’s health authorities are provincial bodies, so the Policy on Results does not govern the Cedar Valley Connector program. The policy becomes relevant in two ways. If a federal department funded part of the program through a contribution agreement, the program would report results that feed the department’s Performance Information Profile and could be included in the department’s evaluation of its contribution program. More generally, the structure of the policy, in which routine indicators are tracked continuously and evaluations are scheduled periodically to explain them, is a sound model for a health authority, and Lesson 5 uses it when distinguishing performance measurement from evaluation.
3.2 The Evaluation Profession in Canada
The Canadian Evaluation Society (CES) was formed in 1981. It is a bilingual national association with regional chapters, and it publishes the Canadian Journal of Program Evaluation. Evaluation is not a regulated profession in Canada, so anyone may call themselves an evaluator. The CES has addressed this through a statement of competencies and a professional designation.
Competencies for Canadian evaluation practice
The CES competencies were developed as part of its credentialing program and revised in 2018. They are grouped into five domains, which together describe what a competent evaluator knows and does.
The Credentialed Evaluator designation
The CES introduced its Professional Designations Program in 2009, and through it awards the Credentialed Evaluator (CE) designation. Applicants must hold a graduate degree, or a graduate certificate or diploma in program evaluation, must have the equivalent of two years of full-time evaluation-related work experience within the previous ten years, and must document, with examples from their work, how they demonstrate the competencies in each domain. Credentialed evaluators maintain the designation through continuing professional learning. The designation is voluntary. Some employers and funders ask for it, and it gives graduate students a clear map of the competencies they need to build.
Ethical guidance for evaluators
The CES publishes ethical guidelines for its members, and many Canadian evaluators also use the American Evaluation Association’s Guiding Principles for Evaluators, revised in 2018, which set out five principles: systematic inquiry, competence, integrity, respect for people, and common good and equity. These guidelines apply whether or not a project falls under research ethics review, a point that becomes important in Section 3.4.
3.3 The Program Evaluation Standards
The Program Evaluation Standards are the most widely used statement of what makes an evaluation good. They are developed by the Joint Committee on Standards for Educational Evaluation, a body of professional associations from the United States and Canada, and the current third edition was written by Yarbrough, Shulha, Hopson and Caruthers (2011). Lyn Shulha, one of its authors, worked at Queen’s University. The edition contains 30 standards organized under five attributes. The first four attributes appeared in earlier editions, and the 1999 CDC framework presented in Section 4 adopted them. The third edition added evaluation accountability.
| Attribute | Standards | What it asks | Example standards | Cedar Valley illustration |
|---|---|---|---|---|
| Utility | 8 (U1 to U8) | Does the evaluation serve the information needs of its intended users? | U2 Attention to Interest Holders (title paraphrased with the newer term); U4 Explicit Values; U7 Timely and Appropriate Communicating and Reporting | Findings on first-wave reach arrive before the second-wave decision is made. |
| Feasibility | 4 (F1 to F4) | Is the evaluation realistic, practical and efficient in its setting? | F2 Practical Procedures; F3 Contextual Viability; F4 Resource Use | Data collection fits into connector meetings without crowding out the work itself. |
| Propriety | 7 (P1 to P7) | Is the evaluation proper, fair, legal and just? | P2 Formal Agreements; P3 Human Rights and Respect; P5 Transparency and Disclosure; P6 Conflicts of Interests | A written agreement with the First Nations partners sets out data ownership and review of reports. |
| Accuracy | 8 (A1 to A8) | Are the information and conclusions dependable and truthful? | A2 Valid Information; A6 Sound Designs and Analyses; A7 Explicit Evaluation Reasoning | The rubric and its standards are published with the findings, so readers can follow the reasoning. |
| Evaluation accountability | 3 (E1 to E3) | Is the evaluation documented and itself evaluated? | E1 Evaluation Documentation; E2 Internal Metaevaluation; E3 External Metaevaluation | An evaluator from another health authority reviews the evaluation plan before data collection begins. |
The standards guide planning as well as review. During planning, an evaluator can use them as a checklist of risks: a design that answers the executive’s question too late fails on utility, a survey that doubles the length of each connector meeting fails on feasibility, and an agreement that leaves data ownership unstated fails on propriety. After an evaluation, the standards support metaevaluation, the evaluation of an evaluation. The attributes can conflict. A larger sample improves accuracy and reduces feasibility, and full disclosure of findings for small groups improves transparency while risking the identification of participants. The standards leave such conflicts to the evaluator’s judgement, and the evaluator’s task is to make the trade-off explicit.
3.4 The Ethics of Evaluation
When evaluation requires research ethics review
The Tri-Council Policy Statement: Ethical Conduct for Research Involving Humans (TCPS 2), whose current edition dates from 2022, is the joint research ethics policy of the Canadian Institutes of Health Research, the Natural Sciences and Engineering Research Council and the Social Sciences and Humanities Research Council. Institutions eligible for funding from these agencies, which include universities and many hospitals and health authorities, must apply it to research conducted under their auspices. Article 2.1 defines research as an undertaking intended to extend knowledge through a disciplined inquiry or systematic investigation. Article 2.5 states that quality assurance and quality improvement studies, program evaluation activities, and performance reviews do not constitute research under the policy, and do not fall within the scope of research ethics board (REB) review, when they are “used exclusively for assessment, management or improvement purposes.”
Three further points in the policy shape how Article 2.5 is applied. If data collected for evaluation are later proposed for use in research, that secondary use may require REB review. The application of Article 2.1 states that the choice of methodology and the intent or ability to publish findings are not factors that determine whether an activity is research requiring review, so neither a randomized design nor a plan to publish settles the question by itself. Finally, the policy notes that activities outside REB review can still raise ethical issues that would benefit from careful consideration by an individual or body able to give independent guidance. An evaluation exempt from REB review is still subject to ethical scrutiny.
When research involves First Nations, Inuit or Métis Peoples, Chapter 9 of TCPS 2 applies, and researchers must seek engagement with the relevant community where the research is likely to affect its welfare. Evaluations that fall under Article 2.5 are outside the formal scope of that chapter, but its principles of community engagement, respect for Indigenous governance and collaboration remain the appropriate standard for any evaluation involving Indigenous communities.
Screening evaluation and quality improvement projects
Because evaluation projects outside REB review still need ethical scrutiny, several Canadian organizations use structured screening tools. The best known is the ARECCI (A pRoject Ethics Community Consensus Initiative) set of decision support tools, developed in Alberta and offered by Alberta Innovates. The ARECCI Ethics Screening Tool takes a project lead through four steps: preliminary questions, the project’s primary purpose, a set of risk filters, and the screening results. It produces a summary score that indicates the level of risk to participants and suggests the appropriate type of review, which may be a research ethics board, an ARECCI second opinion review, or an internal organizational review. A companion ARECCI Ethics Guideline Tool asks project teams to work through six ethics considerations for their project’s objectives, methods and outcomes. In British Columbia, Interior Health’s project ethics policy asks quality improvement and evaluation teams to complete the ARECCI guidelines while developing a project and to use the screening tool to determine its primary purpose and level of risk, with the level of review scaled to that risk.
The program’s analyst reviews routine referral and attendance data to rebalance connector caseloads. This activity is used exclusively for management and improvement, so it falls under Article 2.5. The program surveys participants about their satisfaction with connector meetings to improve practice, which also falls under Article 2.5, although the health authority should screen it for risk because it asks older adults to comment on staff they depend on. A university team proposes to compare first-wave and second-wave clinics in order to publish evidence on social prescribing for older adults. That activity is intended to extend knowledge beyond the program, so it is research and requires REB review, and Chapter 9 engagement requirements apply to any part involving the First Nations partners. A graduate student later asks to use the analyst’s routine data for a thesis, which is a secondary use of evaluation data for research and may require REB review.
Ethical issues in evaluation practice
Evaluations often use data collected for service delivery, such as referral records and intake scores, which participants did not provide for evaluation. Evaluators should use only the data the question requires, remove identifiers as early as possible, follow the organization’s privacy policies and data sharing agreements, and tell participants how program data may be used for evaluation.
Small programs and small communities make participants easy to identify. A table showing outcomes for Indigenous participants at a single clinic might describe five people. Evaluators should set minimum cell sizes for reporting, combine categories where necessary, and agree reporting rules with community partners in advance.
Participants who depend on a service may hesitate to criticize it, and may fear that criticism will affect their care. Evaluators should collect feedback through someone other than the person who delivers the service where possible, explain that participation in the evaluation will not affect services, and offer anonymous channels.
Surveys of evaluators, beginning with Morris and Cohn (1993), have found that pressure from interest holders to alter or soften findings is among the most frequently reported ethical problems. Formal agreements made at the start, covering ownership of the report and the right to publish findings, and the propriety standards on transparency and disclosure and on conflicts of interests, give evaluators a basis for resisting such pressure.
Designs that compare people who receive a program with people who do not raise questions of fairness, especially when the program is believed to help. A staggered rollout such as Cedar Valley’s, in which every clinic eventually receives the program, can ease this concern. Lesson 6 examines the ethics of randomizing access to a program.
The Cedar Valley steering committee proposes to interview twenty participants about why they stopped attending connector meetings, in order to redesign the follow-up process, and the evaluation lead would like to present the findings at a national conference. Decide whether TCPS 2 Article 2.5 applies, what screening you would recommend, and what you would put in writing before the interviews begin. (A strong answer notes that the stated purpose is improvement of the program, so Article 2.5 can apply, and that the plan to present findings does not by itself make the activity research. It recommends ARECCI or equivalent screening because the interviews ask dependent participants about a service they may need again, and an agreement covering consent, confidentiality and reporting. If the team wants the findings to answer questions beyond the program, it should seek REB review before collecting data.)
The policies, competencies, standards and ethical rules in this section describe the conditions under which an evaluation is carried out. Section 4 turns to the sequence of tasks that make up an evaluation and to the evaluation plan that sets them out.
Reflection
Under the Tri-Council Policy Statement (TCPS 2, 2022), research is an undertaking intended to extend knowledge through a disciplined inquiry or systematic investigation (Article 2.1). Article 2.5 states that program evaluation and quality improvement activities do not require research ethics board (REB) review when they are used exclusively for assessment, management or improvement purposes. Data collected for evaluation and later proposed for research may require REB review as a secondary use, and the policy states that the intent to publish does not by itself determine whether an activity is research. Chapter 9 requires community engagement for research likely to affect the welfare of First Nations, Inuit or Métis communities. The Program Evaluation Standards group quality into five attributes: utility, feasibility, propriety, accuracy and evaluation accountability. A health authority’s evaluation team plans three activities for a connector program for older adults: (1) a monthly analysis of program records to adjust connector caseloads; (2) interviews with fifteen participants who left the program early, to redesign follow-up, with a plan to present findings at a national conference; (3) a partnership with a university to compare clinics with and without the program and publish evidence for other provinces, using data that include members of a partner First Nation. For each activity, state whether REB review is required and what other review you would recommend. Then name the attribute of the Program Evaluation Standards most at risk in activity 2 and explain how you would protect it.
Activity 1 is used exclusively for management, so it falls under Article 2.5 and needs no REB review. It should still follow the health authority’s privacy rules, use the minimum data required, and pass any internal screening the organization requires.
Activity 2 has an improvement purpose, so Article 2.5 can apply, and the conference presentation does not by itself make it research. Because interviewees depend on the service and may fear consequences for criticizing it, I would recommend an ethics screening such as the ARECCI tool and an organizational review scaled to the risk it identifies.
Activity 3 is intended to extend knowledge beyond the program, so it is research and needs REB review before data collection. Because it uses data about members of a First Nation, Chapter 9 applies, and the team should engage the Nation and agree how its data will be governed and how findings will be reviewed before publication.
In activity 2, propriety is most at risk, particularly the standard on human rights and respect. I would have someone other than the participant’s connector conduct the interviews, state in writing that taking part will not affect services, remove identifying details from quotations, and avoid reporting results for groups so small that individuals could be recognized.
Minimum 20 characters required.
Question 1: Which requirement is part of the Treasury Board Policy on Results, which took effect on July 1, 2016?
Question 2: Which attribute was added to the Program Evaluation Standards in the third edition (Yarbrough et al., 2011)?
Question 3: A health authority team interviews participants solely to redesign program follow-up and plans to present the findings at a conference. How does TCPS 2 treat this activity?
Question 4: What is the ARECCI Ethics Screening Tool designed to do?
The Evaluation Cycle and the Program Description
Learning Objectives for this section
- Describe the six steps and four standards of the 1999 CDC Framework for Program Evaluation.
- Describe how the 2024 CDC Program Evaluation Framework revised the steps, added three cross-cutting actions, and adopted five evaluation standards.
- Apply the evaluation cycle to the Cedar Valley Connector program.
- Describe how the parts of an evaluation plan follow the steps of the cycle.
- Write a program description that covers purpose, population, activities, resources, setting and history.
Introduction
Evaluators need a map of the tasks an evaluation involves and the order in which they are usually done. This course uses the framework published by the United States Centers for Disease Control and Prevention (CDC) for that purpose. The framework is a United States federal document, but its steps are general enough to organize evaluation work in Canadian settings, and it is widely taught in public health. This section presents the original 1999 framework and its 2024 update, applies the cycle to the Cedar Valley Connector program, and shows how to write the program description on which an evaluation plan rests.
4.1 The CDC Framework for Program Evaluation (1999)
The Framework for Program Evaluation in Public Health was published in 1999 in the CDC’s Morbidity and Mortality Weekly Report series of recommendations and reports (Centers for Disease Control and Prevention, 1999). A CDC working group wrote it to summarize the essential elements of program evaluation in a form that public health practitioners could use. It describes six steps arranged in a cycle and four groups of standards for judging the quality of an evaluation.
| Step (1999 label) | Elements the framework lists |
|---|---|
| 1. Engage interest holders | Those involved in program operations, those served or affected by the program, and the primary users of the evaluation |
| 2. Describe the program | Need, expected effects, activities, resources, stage of development, context and logic model |
| 3. Focus the evaluation design | Purpose, users, uses, questions, methods and agreements |
| 4. Gather credible evidence | Indicators, sources, quality, quantity and logistics |
| 5. Justify conclusions | Standards, analysis and synthesis, interpretation, judgement and recommendations |
| 6. Ensure use and share lessons learned | Design, preparation, feedback, follow-up and dissemination |
The four standards, utility, feasibility, propriety and accuracy, were taken from the second edition of the Joint Committee’s Program Evaluation Standards, introduced in Section 3. The steps are presented as a cycle because each depends on the ones before it and because evaluation of a program usually continues through several rounds. A program description written in Step 2 shapes the questions chosen in Step 3, and the findings shared in Step 6 feed the next round of engagement and description. In practice, evaluators move back and forth among the steps.
4.2 The 2024 Update
In 2024 the CDC published the CDC Program Evaluation Framework, 2024 (Kidder et al., 2024), also in its recommendations and reports series. The update keeps the six-step structure and the practical, nonprescriptive character of the original, and it makes four main changes.
First, the framework begins with a new step, Assess context, which asks evaluators to consider four factors before planning: readiness for evaluation, the interest holders involved, the place in which the program operates, and the evaluation capacity available. Second, several steps were renamed and revised. Step 2 remains Describe the program, covering need, inputs, activities, outcomes, contextual factors and stage of development, often summarized in a logic model. Step 3 becomes Focus the evaluation questions and design, and its products are a purpose statement, a statement of the type of evaluation, a list of intended users and uses, the evaluation questions and a description of the overall design. Step 4 remains Gather credible evidence. Step 5 becomes Generate and support conclusions. Step 6 becomes Act on findings, whose key elements are planning, preparing findings for use, and facilitating the move from insights to action.
Third, the update adds three cross-cutting actions that apply at every step: engage collaboratively, advance equity, and learn from and use insights. The authors describe these actions as the largest change from the 1999 framework. Engagement, which was the first step in 1999, now runs through the whole cycle. Fourth, the four standards are replaced by the five federal evaluation standards adopted in the United States under the Foundations for Evidence-Based Policymaking Act of 2018 and set out by the Office of Management and Budget: relevance and utility, rigour, independence and objectivity, transparency, and ethics. The update also replaces an older term, built on the word stake, with interest holder, explaining that the older term can imply a power differential and has a violent connotation for some American Indian and Alaska Native tribes and tribal members. This course uses the term interest holder for similar reasons, and it uses the newer term when describing the 1999 framework and other sources that use the older one.
| Feature | 1999 framework | 2024 framework |
|---|---|---|
| First step | Engage interest holders | Assess context (engagement becomes a cross-cutting action) |
| Steps 3, 5 and 6 | Focus the evaluation design; justify conclusions; ensure use and share lessons learned | Focus the evaluation questions and design; generate and support conclusions; act on findings |
| Cross-cutting actions | None | Engage collaboratively; advance equity; learn from and use insights |
| Standards | Utility, feasibility, propriety, accuracy | Relevance and utility, rigour, independence and objectivity, transparency, ethics |
| Term for affected groups | An older term built on the word stake | Interest holders |
For Canadian evaluators, the CDC framework is a useful organizing cycle, while the Program Evaluation Standards, the CES competencies and TCPS 2 from Section 3 remain the main Canadian references for quality and ethics. The two sets of standards are compatible, and an evaluation plan for a British Columbia program can cite both.
4.3 Applying the Cycle to the Cedar Valley Connector Program
Each step of the cycle corresponds to work that the Cedar Valley evaluation team must do, and to one or more lessons of this course. The tabs use the 2024 step labels.
The team assesses the program’s readiness for evaluation (routine data exist for the first wave, but no comparison data have been collected), identifies interest holders (the executive, the steering committee, connectors, clinicians, community partners, participants and the First Nations partners), considers place (a small city and rural communities with uneven transportation), and takes stock of capacity (one half-time analyst). Lesson 4 teaches interest holder mapping.
The team writes an agreed description of the program’s need, inputs, activities, outcomes, context and stage of development. The first wave is in early implementation, and the land-based pathway is still in co-design. Section 4.5 gives this description in full, and Lessons 2 and 3 add the needs statement, objectives and logic model.
With the primary intended users, the team sets the purpose of the evaluation (to inform the second-wave decision and improve delivery), the types of evaluation (process, outcome, impact and economic), the evaluation questions, and the overall design, which may exploit the staggered rollout. Lessons 4 and 6 to 9 teach this step.
The team specifies indicators, data sources and collection methods for each question, such as loneliness scores at intake and twelve weeks, attendance records, health authority data on emergency department visits, and talking circles with Indigenous participants. Lesson 5 teaches evaluation matrices, indicators and data systems.
The team analyzes the data, applies the rubric agreed in advance, synthesizes qualitative and quantitative findings, and states how confident it is in each conclusion. Section 1 of this lesson introduced the reasoning, and Lesson 10 develops synthesis and judgement.
The team plans how findings will reach each audience: a briefing for the executive before the second-wave decision, a plain-language summary for older adults, and a report reviewed with the First Nations partners before release. Lesson 10 covers reporting, knowledge translation and use.
The cross-cutting actions apply to every tab. Engaging collaboratively means that the steering committee and the First Nations partners shape each step. Advancing equity means asking at each step whether the program and the evaluation serve older adults who face the greatest barriers, such as those without transportation or those whose first language is neither English nor French. Learning from and using insights means that the six-month review feeds back into program design.
4.4 From the Cycle to an Evaluation Plan
An evaluation plan sets out, before any data are collected, how an evaluation will be carried out. Its parts follow the cycle in the order an evaluator works, and a plan prepared for a health authority usually ends in a proposal with costing and an executive summary written for decision-makers. The table shows where this course teaches each part, using the Cedar Valley Connector program as the worked example throughout.
| Lesson | Part of the plan | Main step in the cycle |
|---|---|---|
| 1 | Program description | Describe the program |
| 2 | Needs statement supported by local data, and SMART objectives | Describe the program |
| 3 | Logic model and theory of change narrative with stated assumptions | Describe the program |
| 4 | Interest holder map and prioritized evaluation questions | Assess context; focus the questions |
| 5 | Evaluation matrix with indicators, data sources and methods | Gather credible evidence |
| 6 | Assessment of whether a randomized design is feasible | Focus the design |
| 7 | Comparison groups and the threats to validity they leave open | Focus the design |
| 8 | Choice and justification of a design | Focus the design |
| 9 | Implementation framework and implementation outcomes | Focus the design; gather evidence |
| 10 | Costing, synthesis, reporting and an executive summary | Generate conclusions; act on findings |
A program is easiest to describe and evaluate when it is specific, such as a single service or initiative, and when enough information about it is available from program documents, public reports, websites or funding announcements. It should also be at a stage where evaluation decisions remain open. Teaching cases such as Cedar Valley are realistic composites based on real models, and a description of a composite says so clearly.
An internal evaluator who works for the organization that runs the program uses only information they are permitted to share, and the description includes no identifiable information about clients or staff. When it is unclear what may be shared, the evaluator checks with the program manager before the description circulates.
A description of a program led by or serving First Nations, Inuit or Métis communities relies on information the community or organization has made public, describes the program in the terms it uses for itself, and notes where community governance of data and interpretation would apply. Lesson 4 sets out the principles involved.
4.5 Writing a Program Description
A program description is the shared account of what a program is, on which every later part of an evaluation depends. If interest holders disagree about what the program does or is meant to achieve, the evaluation questions will be contested and the findings disputed. A good description is specific, using numbers, names of activities and eligibility rules. It is written in a neutral descriptive voice and separates what the program intends from what it has been shown to achieve. It covers six elements: purpose, population, activities, resources, setting and history. These map closely onto the elements of the CDC’s second step.
Purpose. The Cedar Valley Connector program aims to reduce loneliness and social isolation among adults aged 65 and older by connecting them with community groups, volunteer roles, transportation help and services. Its longer-term aims are to increase social participation and improve self-rated health. The program’s planners also expect fewer emergency department visits and fewer primary care visits made mainly for social reasons, although these expectations have not yet been tested. The program describes its practice as person-directed, community-based and culturally safe: participants decide what they want to work on, connectors draw on resources that already exist in the community, and the program is accountable to its First Nations partners for how it serves their members.
Population. The program serves adults aged 65 and older who attend a participating primary care clinic and who score 6 or higher on the three-item UCLA Loneliness Scale (range 3 to 9), or whom a clinician judges to be socially isolated. People who need urgent mental health care are referred to mental health services, and residents of long-term care homes are not eligible because their homes provide social programming. Connectors can work with participants through interpreters and can meet at home with people who have limited mobility. In the first six months, clinicians referred 312 older adults, and 241 attended a first meeting.
Activities. Clinicians screen older adults during routine visits and refer through the electronic medical record. A connector contacts each person within ten business days and meets with them up to six times over twelve weeks, in person, by telephone or at home. Together they identify what matters to the person and write a connection plan. The connector links the person to groups, volunteer roles and services, accompanies them to a first activity where helpful, and arranges transportation help. Connectors record each contact, repeat the loneliness scale at twelve weeks, and send a summary to the referring clinician. The coordinator holds a monthly case review with connectors, and community partners receive small grants to make their groups easier for newcomers to join, for example by offering a buddy at the first session. Participants may leave the program at any time and may be referred again later.
Resources. The first wave has an annual budget of $840,000. This covers seven connector positions ($595,000, including benefits), a program coordinator ($105,000), a half-time data and evaluation analyst ($55,000), a transportation fund ($40,000), small grants to community partners ($30,000), and training, travel and data systems ($15,000). One connector position is hosted by a First Nations health centre under the program’s Indigenous health partnership. Connectors complete training in person-centred conversations, community resource mapping and cultural safety before taking referrals. Community partners include seniors’ centres, a volunteer centre, faith communities and recreation programs.
Setting. The program is run by the Cedar Valley Health Authority in a region of British Columbia that includes a small city and several rural communities, where public transit is limited outside the city. The region has 24 primary care clinics, ranging from solo practices to a multidisciplinary team clinic in the city, and three of them serve communities close to the partner First Nations. Twelve joined the program in the first wave, and the remaining twelve are scheduled to join a year later. A twelve-member steering committee oversees the program, including four older adults with lived experience of loneliness, two representatives of the First Nations partners, clinic and community partner representatives, the coordinator and the health authority’s director of primary care.
History. The program grew from a regional survey in which about one in four older adult respondents scored 6 or higher on the loneliness scale, and from clinicians’ reports that they had no clear referral route for lonely patients. An eighteen-month pilot in two clinics tested the referral pathway and connector role. The Indigenous health partnership agreement was signed before the first wave launched, and the partners are co-designing a land-based connection pathway. The program reports monthly referral and attendance figures to the health authority, and its steering committee completed a six-month review of the first wave. No comparison data have yet been collected from clinics outside the program, and the health authority has asked for an evaluation to inform the second-wave rollout.
What makes the example work
The description gives numbers that later steps of an evaluation can use, such as the budget lines, the referral threshold and the clinic counts. It states the planners’ expectations about emergency department and primary care visits as expectations, without claiming them as results. It names the program’s stage of development, which tells a reader that a full impact evaluation of the land-based pathway would be premature. A real description would also cite its sources, such as program documents and public reports.
Reflection
The 2024 CDC Program Evaluation Framework has six steps: (1) assess context, (2) describe the program, (3) focus the evaluation questions and design, (4) gather credible evidence, (5) generate and support conclusions, and (6) act on findings. It also has three cross-cutting actions that apply at every step: engage collaboratively, advance equity, and learn from and use insights. Assign each of the following tasks from an evaluation of a community connector program to one step: (a) agreeing with the executive that the evaluation will inform the decision on a second wave of clinics, and listing the evaluation questions; (b) checking whether routine attendance data are complete and whether the half-time analyst has time for the evaluation; (c) writing an agreed account of the program’s need, inputs, activities, outcomes and stage of development; (d) extracting emergency department visit counts from health authority records; (e) applying an agreed rubric to six-month data and stating confidence in each conclusion; (f) preparing a plain-language summary for older adults and briefing the executive before the decision. The draft plan schedules one meeting with the program’s First Nations partners, after the report is written. Identify the cross-cutting action this neglects and describe how you would change the plan.
Task (a) belongs to Step 3, focus the evaluation questions and design, because it sets the purpose, intended use and questions. Task (b) belongs to Step 1, assess context, because it concerns readiness for evaluation and evaluation capacity. Task (c) is Step 2, describe the program. Task (d) is Step 4, gather credible evidence. Task (e) is Step 5, generate and support conclusions. Task (f) is Step 6, act on findings, since it prepares findings for use by each audience.
A single meeting after the report is written neglects the cross-cutting action to engage collaboratively, and it also weakens the action to advance equity. The First Nations partners would have no say in the program description, the questions, the criteria, the handling of data about their members, or the interpretation of findings. I would invite the partners to help assess context and describe the program at the start, agree with them which outcomes matter and how their data will be held, include their criteria in the rubric, and schedule a joint review of findings before the report is finalized. These changes also reflect the Indigenous evaluation principles of relationships, accountability and self-determination.
Minimum 20 characters required.
Question 1: What is the first step of the 1999 CDC Framework for Program Evaluation?
Question 2: Which set lists the three cross-cutting actions of the 2024 CDC Program Evaluation Framework?
Question 3: Which list gives the evaluation standards used in the 2024 CDC framework?
Question 4: Which statement best describes the program description written at Step 2 of the CDC cycle?
Final Assessment
Bringing It All Together
This lesson set out the foundations on which the rest of the course builds. Evaluation is the systematic determination of a program’s merit, worth or significance, and it differs from research mainly in its purpose and audience. Its reasoning moves from criteria to standards to evidence to a synthesized judgement, and a rubric agreed in advance makes that reasoning visible. The types of evaluation, from needs assessment to economic evaluation, answer different questions at different stages of a program’s life, and an impact claim requires a credible counterfactual.
Evaluators approach this work from different positions. The evaluation theory tree organizes approaches by whether they emphasize credible methods, the use of findings, or the values that define merit, and real evaluations often combine approaches. In Canada, evaluation takes place within the Treasury Board Policy on Results, the competencies and designation of the Canadian Evaluation Society, the Program Evaluation Standards, and the ethics rules of TCPS 2, including Article 2.5 and the screening tools that support it.
The CDC framework, in its 1999 and 2024 versions, organizes the evaluator’s tasks into a cycle, and the 2024 cross-cutting actions of engaging collaboratively, advancing equity, and learning from and using insights apply throughout. The fictional Cedar Valley Connector program, introduced in this lesson, will carry these ideas through every lesson that follows, beginning with the program description in Section 4.5.
Key Takeaways from this lesson
- Evaluation is the systematic determination of the merit, worth or significance of a program, and it ends in a judgement that people can act on.
- Merit is a program’s intrinsic quality, worth is its value in a particular context including cost, and significance is its importance.
- Evaluation and research share methods, and they differ most clearly in purpose and primary audience.
- Evaluative reasoning moves from criteria to standards to evidence to synthesis, and rubrics agreed in advance make that reasoning transparent.
- Needs assessment, formative, process, outcome, impact and economic evaluation answer different questions at different stages of a program’s life, and only impact evaluation addresses attribution through a counterfactual.
- Alkin and Christie’s evaluation theory tree organizes approaches on methods, use and valuing branches, and most real evaluations combine approaches from more than one branch.
- Culturally responsive and Indigenous evaluation treat culture, relationships and self-determination as matters of validity and ethics.
- The Treasury Board Policy on Results, the CES competencies and Credentialed Evaluator designation, and the Program Evaluation Standards shape evaluation practice in Canada.
- Under TCPS 2 Article 2.5, evaluation used exclusively for assessment, management or improvement does not require REB review, although it still requires ethical scrutiny, which tools such as ARECCI support.
- The CDC framework organizes evaluation into six steps, and its 2024 update added a step to assess context, three cross-cutting actions and five standards.
Core Concepts Reviewed
Section 1: evaluation as the determination of merit, worth and significance; evaluation compared with research; criteria, standards and rubrics; and needs assessment, formative, process, outcome, impact and economic evaluation across the program life cycle.
Section 2: the evaluation theory tree; utilization-focused, developmental, participatory and empowerment evaluation; theory-driven and realist evaluation; and culturally responsive and Indigenous evaluation.
Section 3: the Treasury Board Policy on Results; the Canadian Evaluation Society, its competencies and the Credentialed Evaluator designation; the Program Evaluation Standards; and TCPS 2 Article 2.5, the ARECCI tools and ethical issues in evaluation.
Section 4: the six steps and four standards of the 1999 CDC framework; the 2024 update with its new first step, cross-cutting actions and five standards; and the parts of an evaluation plan and the program description.
The final reflection asks you to apply ideas from all four sections to a manager’s plan for evaluating the Cedar Valley Connector program.
Reflection
A health authority manager sends you this message about a fictional community connector program for adults aged 65 and older: “This is just program evaluation, so we don’t need any ethics process. We’ll compare participants’ loneliness scores before and after their twelve weeks, and if the scores go down, we’ll call the program a success and roll it out to every clinic. We don’t have time to involve the steering committee or the First Nations partners.” Relevant facts: participants are referred when they score 6 or higher on the three-item UCLA Loneliness Scale (range 3 to 9); the program runs in 12 of 24 clinics, and the other 12 are scheduled to start a year later; the program has an Indigenous health partnership with local First Nations and a steering committee that includes older adults. Under TCPS 2 Article 2.5, evaluation used exclusively for assessment, management or improvement does not require research ethics board review, although such activities can still raise ethical issues that need independent consideration. Evaluation is the systematic determination of a program’s merit, worth or significance. The 2024 CDC framework includes the cross-cutting actions engage collaboratively, advance equity, and learn from and use insights. Write a reply of 200 to 300 words that responds to each part of the manager’s plan.
Thank you for moving quickly on this. I agree that an evaluation used only to manage and improve the program falls under Article 2.5 and does not need research ethics board review. It still raises ethical questions, because participants depend on the service and some groups are small, so I suggest a quick ethics screen and agreed rules on privacy and reporting.
A before-and-after comparison will show whether scores changed, and it cannot show that the program caused the change. People enter the program only when they score 6 or higher, and high scores tend to fall on remeasurement anyway. A judgement of success also needs more than one outcome. I suggest agreeing a short rubric that covers reach, change in loneliness, cultural safety and cost, so that our conclusion reflects the program’s merit and its worth to the health authority.
The staggered start gives us a better design. Comparing the 12 clinics that have the program with the 12 that start next year offers a counterfactual and would make a rollout decision much more defensible.
Finally, involving the steering committee and the First Nations partners is part of doing this well. The partners have a say in how data about their members are used and in whether the program is culturally safe, and the older adults on the committee can tell us which outcomes matter. Engaging them early, even in two short meetings, will make the findings more credible and more likely to be used.
Minimum 30 characters required.
Final Knowledge Assessment
Question 1: A review finds that the Cedar Valley program has well-trained connectors and good outcomes, but a cheaper volunteer visiting program produces similar benefits in the same region. Which judgement is most directly affected?
Question 2: The executive uses six-month survey results to decide whether the second wave of Cedar Valley clinics will proceed. How is this use of evaluation best described?
Question 3: In the general logic of evaluation set out by Fournier (1995), which step comes immediately after establishing the criteria of merit?
Question 4: Cedar Valley refers adults who score 6 or higher on the three-item loneliness scale. Their mean score later falls from 7.1 to 6.3. Which explanation, other than a program effect, most directly accounts for a fall of this kind?
Question 5: In the hierarchy described by Rossi, Lipsey and Henry (2019), which assessment should be in place before an impact evaluation is likely to be informative?
Question 6: In the third edition of Evaluation Roots (Alkin and Christie, 2023), which approach is placed on the valuing branch?
Question 7: Cousins and Whitmore (1998) distinguished two forms of participatory evaluation. What is the main aim of practical participatory evaluation?
Question 8: The Cedar Valley program and one First Nation are co-designing a land-based connection pathway whose final form neither partner yet knows. Which approach fits best?
Question 9: Under the Policy on Results and section 42.1 of the Financial Administration Act, how often must ongoing grant and contribution programs with five-year average actual expenditures of $5 million or more per year be evaluated?
Question 10: Which domain of the Canadian Evaluation Society’s competencies covers an evaluator’s attention to a program’s history, politics, organizational culture and the cultural context of the communities served?
Question 11: A university team proposes to compare first-wave and second-wave Cedar Valley clinics in order to publish evidence on social prescribing for other provinces. The data include members of a partner First Nation. What does TCPS 2 require?
Question 12: A draft Cedar Valley report presents outcomes separately for the five Indigenous participants at one clinic. Which attribute of the Program Evaluation Standards is most directly at risk?
Question 13: In the 2024 CDC Program Evaluation Framework, what is the status of engagement with interest holders?
Question 14: In the 2024 CDC framework, which step produces a purpose statement, the type of evaluation, a list of intended users and uses, the evaluation questions and a description of the overall design?
Question 15: Which sentence is best suited to a program description written for an evaluation plan?
Glossary: Key Terms, People & Frameworks
📚 Reference page, available throughout the lesson
This glossary defines the terms, frameworks and people introduced in Lesson 1, and you can search it at any point in the lesson.