Program Theory: Logic Models and Theories of Change
Program Planning & Evaluation
Learning objectives for this lesson:
- Define program theory and distinguish Chen's change model from his action model.
- Explain, using the distinction between implementation failure and theory failure, why evaluations that ignore program theory cannot explain their own findings.
- Identify the components of a logic model and classify program statements as inputs, activities, outputs, or short-term, intermediate or long-term outcomes.
- Compare linear, nested and outcome-chain logic models and correct common errors such as activities listed as outcomes and missing assumptions.
- Construct a theory of change by backward mapping from a long-term outcome through preconditions, with stated assumptions, rationales and an accountability ceiling.
- Apply Mayne's contribution analysis to assemble and assess a contribution story where attribution is not possible.
- Write context-mechanism-outcome configurations and describe how a realist initial program theory is developed and refined.
- Produce a logic model and a one-page theory of change narrative with stated assumptions for a program, following the Cedar Valley worked example.
This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University, drawing on Rossi, P. H., Lipsey, M. W., & Henry, G. T. (2019). Evaluation: A Systematic Approach (8th ed.). SAGE; and Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and Quasi-Experimental Designs for Generalized Causal Inference. Houghton Mifflin.
Program Theory: Change Models, Action Models and Why Theory Matters
Learning Objectives for this section
- Define program theory and distinguish it from social science theory and from evaluation theory.
- Distinguish Chen's change model from his action model and map both onto a real or realistic program.
- Explain, using the distinction between implementation failure and theory failure, why an evaluation that ignores program theory cannot explain its own findings.
- Describe the sources and methods an evaluator uses to make a program's theory explicit, and the limits of theory-driven evaluation.
1.1 What Program Theory Is
Every health program carries an argument about how change happens. The fictional Cedar Valley Connector program, which runs through this course, pays seven community connectors to meet older adults who screen as lonely and to link them with groups, volunteer roles, transportation help and services. In doing so, the Cedar Valley Health Authority acts on several beliefs: that loneliness among older adults partly reflects too few opportunities for social contact, that a trusted person can help someone take up opportunities they would not take up alone, that taking part in groups produces relationships people value, and that less loneliness will in time mean better health and fewer emergency department visits. These beliefs may be well supported, partly supported or mistaken, and they are rarely written down in one place. Evaluators use the term program theory for this argument once it has been made explicit.
Leonard Bickman (1987) defined program theory as the construction of a plausible and sensible model of how a program is supposed to work. Rossi, Lipsey and Henry (2019) describe it as the set of assumptions about how a program's activities relate to the social benefits it is expected to produce, together with the strategy and tactics the program has adopted to deliver those activities. Both definitions contain two parts. The first is a causal account, which explains why the program's activities should produce the intended changes. The second is an operational account, which explains what has to be organized, by whom and for whom, so that the activities actually take place. Section 1.2 shows that Huey-Tsyh Chen built his framework around exactly this division.
Three meanings of "theory"
Students meet the word "theory" in several senses in evaluation, and confusing them causes avoidable errors. Program theory is specific to one program or one type of program: it is the account of how the Cedar Valley Connector program is meant to reduce loneliness. Social science theory is a general explanation of a phenomenon that holds across many settings, such as the work of Louise Hawkley and John Cacioppo (2010) on how loneliness shapes attention, expectations and behaviour in social situations. A good program theory often draws on social science theory, and a program theory that contradicts well-established social science theory deserves scrutiny. Evaluation theory is a theory about how to evaluate, such as the approaches placed on Alkin and Christie's evaluation theory tree in Lesson 1. This lesson is concerned with program theory, and it uses social science theory as one of its sources.
Chen (1990) originally distinguished normative theory, which describes what the program should be, from causative theory, which describes how the program is expected to work. In his later work (Chen, 2005, 2015) these became the action model and the change model, the terms used in this course.
Espoused theory and theory-in-use
Chris Argyris and Donald Schön (1974) distinguished the espoused theory that people give when asked to explain their actions from the theory-in-use that can be inferred from what they actually do. The Cedar Valley program describes its practice as person-directed, with connectors finding resources that match each participant's interests. An evaluator who observes meetings might find that some connectors steer participants toward the two or three groups they know best, because those referrals are quick and reliable. The espoused theory holds that a match with a person's interests sustains participation; the theory-in-use holds that familiar, reliable groups do. The two theories predict different patterns in the data, and an evaluation that measures only the espoused theory will misread what the program is doing.
1.2 Chen's Change Model and Action Model
Huey-Tsyh Chen's conceptual framework of program theory (Chen, 2005, 2015) is a widely used way of dividing a program's theory into parts that an evaluator can examine separately. The framework has two components. The change model states the causal process the program relies on. The action model states the arrangements the program must make so that the causal process can be set in motion.
The change model
The change model has three elements. The first is the goals and outcomes, the changes the program ultimately seeks. The second is the determinants, the factors the program attempts to change because it believes they lead to the outcomes. Determinants are sometimes called intervening variables or mediators. The third is the intervention or treatment, the set of program activities that act on the determinants. For the Cedar Valley program, the goals are lower loneliness, greater social participation and better self-rated health. The determinants are the participant's opportunities for social contact, their confidence about joining a group, their practical access to activities (in particular transportation), and the quality of the new relationships they form. The intervention is the combination of connector meetings, a co-developed connection plan, linkage and accompaniment, transportation help and small grants to community partners.
Naming the determinants is the step that most program documents omit, and it matters most to an evaluator. If the program believes that confidence is the main determinant, the evaluation should measure confidence before and after the meetings; if transportation is the main barrier, it should record whether transport was arranged and used. Without named determinants, the evaluator can measure only the final outcomes, which is the black-box situation described in Section 1.3.
The action model
The action model describes what must be organized so that the intervention reaches the intended people in the intended form. Chen identifies six elements: the implementing organization, the program implementers who deliver the intervention, the associate organizations and community partners whose cooperation the program needs, the ecological context that supports or hinders delivery, the intervention and service delivery protocols, and the target population together with the procedures for reaching and retaining it.
The table applies the action model to Cedar Valley. Each question it raises concerns implementation, which is why process evaluation (Lesson 5) is organized largely around the action model.
| Action model element | Cedar Valley Connector program | An evaluation question it raises |
|---|---|---|
| Implementing organization | The Cedar Valley Health Authority, through its primary care division and a twelve-member steering committee. | Does the health authority give the program the management attention and data support it needs? |
| Program implementers | Seven connectors trained in person-centred conversations, community resource mapping and cultural safety, supported by a coordinator. | Do connectors have manageable caseloads, and do they deliver the meetings as intended? |
| Associate organizations and community partners | Twelve primary care clinics, seniors' centres, a volunteer centre, faith communities, recreation programs, and a First Nations health centre that hosts one connector. | Do clinics screen and refer consistently, and are partner groups willing and able to welcome newcomers? |
| Ecological context | A small city and several rural communities, with limited public transit outside the city. | Are suitable activities available within reach of rural participants? |
| Intervention and service delivery protocols | Contact within ten business days, up to six meetings over twelve weeks, a connection plan, accompaniment and a summary to the referring clinician. | How many meetings do participants receive, and how often are plans and accompaniment provided? |
| Target population | Adults aged 65 and older who score 6 or higher on the three-item UCLA Loneliness Scale, or whom a clinician judges to be isolated. | Does the referral process reach the eligible population, including people who rarely visit their clinic? |
Chen's framework makes a practical point about program design. A sound change model with a weak action model produces a good idea poorly delivered, while a strong action model with a weak change model produces a well-run program that does not change the outcome. The action model is evaluated mostly with process data, such as referral counts, meeting logs and staff interviews, and the change model mostly with data on the determinants and outcomes.
Three vocabularies for the same distinction
Other authors draw the same line with different terms, and the tabs set three common vocabularies side by side.
Chen (2005, 2015) divides program theory into the change model (intervention, determinants, and goals and outcomes) and the action model (the six elements above). His framework is the most detailed of the three on implementation, which makes it useful when an evaluation must diagnose delivery problems.
Rossi, Lipsey and Henry (2019) divide program theory into program impact theory, the causal sequence from services to proximal and distal outcomes, and process theory. Process theory combines the service utilization plan, which describes how the target population comes into contact with, uses and completes the program, and the organizational plan, which describes the resources, personnel and administration the program needs.
Sue Funnell and Patricia Rogers (2011) divide program theory into a theory of change, the central processes by which change comes about, and a theory of action, the way the program is constructed to activate them. Section 3 of this lesson uses "theory of change" in a broader sense, for a full diagram and narrative of the pathway to a long-term outcome.
1.3 Why Evaluations Need Program Theory
Many outcome evaluations compare outcomes between people who did and did not receive a program. Such a design can show whether outcomes changed, but on its own it treats the program as a black box, and nothing is learned about what happened in between. Carol Weiss argued throughout her career that this kind of evaluation leaves decision-makers unable to act on its findings (Weiss, 1995, 1997, 1998). Her argument rests on a distinction that goes back to Edward Suchman (1967), between a program that fails because it was never properly delivered and a program that fails because the causal process it relies on does not work.
Implementation failure and theory failure
Implementation failure occurs when the program is not delivered as planned, so the causal process is never set in motion. Theory failure occurs when the program is delivered as planned, but the expected causal process does not occur, or it occurs without producing the outcome. The two failures call for different responses. Implementation failure calls for better delivery: more staff, better training, clearer protocols or stronger partnerships. Theory failure calls for a different program, because delivering the same activities more faithfully will not help. An evaluation that measures only final outcomes cannot tell the two apart, because both produce the same result in the outcome data.
| Pattern in the evidence | Interpretation | Cedar Valley illustration |
|---|---|---|
| The program was delivered as planned, the intermediate links occurred, and outcomes improved. | The findings are consistent with the program theory, and the evaluation can say which links carried the effect. | Participants received four or more meetings, most joined a group they valued, and loneliness scores fell. |
| The program was not delivered as planned, and outcomes did not improve. | This is implementation failure. The evaluation says little about whether the theory is sound. | Connectors in one clinic held an average of two meetings because of vacancies, and few linkages were made. |
| The program was delivered as planned, but the intermediate links did not occur or did not lead to the outcome. | This is theory failure. The program's causal assumptions need revision. | Participants joined groups, but interviews show they had little in common with other members, and loneliness did not fall. |
| Outcomes improved, but the program was poorly delivered or the intermediate links did not occur. | Something other than the program, or a mechanism outside the theory, probably produced the change. | Scores fell in a clinic where few people were linked to groups, which suggests regression to the mean or a seasonal pattern. |
Weiss made the further point that positive findings are equally hard to use without program theory. Suppose the Cedar Valley evaluation finds that loneliness fell more among participants than among comparable older adults in second-wave clinics. The health authority must then decide what to keep when the program expands, whether the transport fund is worth its $40,000, and whether the partner grants could be cut. An evaluation that tracked only the final outcome cannot say which components carried the effect. A theory-based evaluation measures the intermediate links and so can show, for example, whether people who were accompanied to a first activity were more likely to keep attending than people who received only information.
Suppose that at the twelve-week follow-up, mean loneliness scores barely changed among participants in two first-wave clinics. A black-box analysis would report the same finding for both. The program's data tell different stories. In the first clinic, a connector position was vacant for four months, participants averaged two meetings, and fewer than one in five was linked to a group, so the program was barely delivered. In the second clinic, participants averaged five meetings and most were accompanied to a first activity. Interviews show that many kept attending, but described the groups as pleasant without producing close relationships. The first clinic shows implementation failure; the second suggests a weakness in the link between participation and meaningful connection. The health authority would respond to the first by filling the vacancy and to the second by reconsidering the kind of activities connectors offer.
What a theory-based evaluation adds
Weiss (1995, 1997) set out several benefits of evaluations that trace a program's theory. They locate the point at which a chain of expected changes breaks, provide early evidence from short-term links, help interpret mixed results, and support generalization by showing what must be present for the program to work elsewhere, such as in the second-wave clinics. They also strengthen causal claims when randomization is not possible, because a program that produces each predicted intermediate change in the predicted order is a more plausible cause of the final outcome. Weiss (1995) made this argument for comprehensive community initiatives, and her paper is widely credited with bringing the term "theory of change" into evaluation. Section 3 develops theories of change, and Lessons 6 to 8 return to stronger causal designs.
Classify each of the three findings below as implementation failure, theory failure, or a pattern that cannot yet be classified, and state what additional data would settle any case you cannot classify. Suggested answers are in the accordion that follows.
(1) In a school-based physical activity program, teachers delivered only one third of the planned lessons, and children's activity levels did not change. (2) In a smoking cessation program, pharmacists delivered every counselling session as planned and participants' confidence in quitting rose, but quit rates at six months were no higher than expected. (3) In a falls prevention program, falls among participants did not decline, and the program collected no data on attendance or exercise.
Finding (1) is implementation failure, since the program was not delivered and its theory was never tested. Finding (2) is theory failure at the link between confidence and quitting: the program changed the determinant it targeted, but that determinant did not produce the outcome, which suggests that other determinants (such as nicotine dependence or the social environment) matter more. Finding (3) cannot be classified, because without attendance data the evaluator cannot tell whether participants received the exercise program; attendance records and a measure of strength or balance would settle the question.
1.4 Making Program Theory Explicit
Program theory is seldom found ready-made. The evaluator assembles it from program documents (for Cedar Valley, the first-wave business case, the report of the eighteen-month pilot and the connector training manual), from interviews and workshops with designers, managers, frontline staff and participants, from observation of practice, which reveals the theory-in-use, and from research and social science theory, which indicate whether the program's assumptions are supported elsewhere.
A deductive approach starts from research and formal theory, an inductive approach builds the theory from observing the program, and a participatory approach develops it with the people who run and use the program. Most practical work combines all three. At Cedar Valley, the evaluation team would draft a theory from the pilot report and the literature on social prescribing, test it in a workshop with connectors and the four older adults with lived experience on the steering committee, and then check the revised version against what connectors do in their meetings.
Whose theory counts is a substantive question. The First Nations partners may understand connection in relational terms that include family, community, land and culture, and the land-based pathway they are co-designing rests on that understanding. Lesson 4 discusses the Indigenous evaluation principles and governance arrangements under which partners shape how their programs are described and evaluated.
In Canada, the Treasury Board of Canada Secretariat has published guidance for federal evaluators on theory-based approaches to evaluation, and federal program evaluations under the Policy on Results (Lesson 1) commonly begin with a logic model or theory of change, the subject of Sections 2 and 3.
Limits of theory-driven evaluation
Theory-driven evaluation has well-documented limits. A systematic review by Coryn and colleagues (2011) found wide variation in how closely evaluations described as theory-driven followed the approach, with many using a theory to describe the program without using it to design data collection or analysis.
Clinicians may see the Connector program as a way of reducing visits made mainly for social reasons, while participants value it for the friendships it produces. Weiss (1997) recommended making competing theories explicit and testing the links on which they disagree.
Evaluators select the links that are most central, least supported by evidence, or most relevant to a pending decision. At Cedar Valley, the link from participation to meaningful connection is both central and uncertain.
Specifying in advance what evidence would count against each link, and examining rival explanations as contribution analysis does (Section 3), reduces the risk of looking only for confirmation.
Rogers (2008) argued that simple linear theories misrepresent programs with multiple components or sites and programs with emergent outcomes and feedback loops. Nested models, outcome chains (Section 2) and realist program theory (Section 4) are partial responses.
Summary of Section 1
Program theory is the explicit account of how a program is expected to produce its outcomes. Chen divides it into a change model and an action model. Evaluations that measure only final outcomes cannot distinguish implementation failure from theory failure or say which components carried an effect. Program theory is assembled from documents, people, observation and research.
Reflection
A regional health authority runs a twelve-week group exercise program for adults aged 70 and older who have fallen in the past year. Physiotherapists lead two strength and balance sessions a week in community centres, and participants receive a home exercise booklet. After one year, the rate of falls among participants is no lower than among similar older adults who did not take part. The evaluation recorded only falls. Using Chen's terms, (a) write the program's change model, naming the intervention, at least two determinants and the outcome; (b) name two elements of the action model and describe what could have gone wrong in each; and (c) explain, with reference to implementation failure and theory failure, what data the evaluation should have collected to interpret its finding.
(a) The intervention is twice-weekly strength and balance sessions with a home exercise booklet. The determinants are lower-limb strength, balance, and the amount of exercise participants actually do at home and in class. The outcome is a lower rate of falls. The change model holds that the sessions increase exercise, exercise improves strength and balance, and better strength and balance reduce falls.
(b) Within the program implementers, physiotherapists may have been stretched across too many sites, so sessions were cancelled or exercises were not progressed in difficulty. Within the target population, recruitment may have reached mainly fitter older adults, or attendance may have been low among frail participants because the community centres were hard to reach.
(c) The evaluation cannot tell whether the program failed because it was not delivered (implementation failure) or because the exercise did not change strength and balance enough to prevent falls (theory failure). It should have collected session records and attendance, a measure of home exercise, and a balance and strength test at baseline and twelve weeks. If attendance was low, the finding reflects implementation failure. If attendance was high and balance improved but falls did not fall, the link from balance to falls is in doubt. If attendance was high and balance did not improve, the exercise dose or content needs revision.
Minimum 20 characters required.
Question 1: In Chen's conceptual framework of program theory, which three elements make up the change model?
Question 2: An evaluation of a falls prevention program finds no change in falls and collected no data on attendance, strength or balance. Following Weiss, what can it conclude?
Question 3: In Rossi, Lipsey and Henry's terms, the part of program theory that describes how referred older adults come into contact with the Connector program, use it and complete it is the:
Question 4: The Cedar Valley program describes its practice as person-directed, but observation shows that some connectors steer participants toward the few groups they know best. This gap illustrates the difference between:
Logic Models: Components, Formats and Common Errors
Learning Objectives for this section
- Define the components of a logic model (inputs, activities, outputs, and short-term, intermediate and long-term outcomes) together with assumptions and external factors.
- Classify program statements into logic model components, separating outputs from outcomes.
- Compare linear, nested and outcome-chain formats and choose a format suited to a program's structure and audience.
- Identify and correct common errors in logic models, including activities listed as outcomes and missing assumptions.
2.1 What a Logic Model Is
A logic model is a diagram, usually fitting on a single page, that shows the sequence from the resources a program uses, through the activities it carries out and the products of those activities, to the changes it expects to produce. The W. K. Kellogg Foundation's Logic Model Development Guide (2004) and the University of Wisconsin-Extension's training materials (Taylor-Powell and Henert, 2008) established the format most evaluators now use. McLaughlin and Jordan (1999) described the logic model as a way of telling a program's performance story: a short, structured account of what the program does, for whom, and with what intended results.
A logic model is a summary of a program theory. It is strong on sequence and weak on explanation. It shows that connector meetings are expected to lead to participation and that participation is expected to lead to lower loneliness, but it does not say why, under what conditions, or for whom. The theory of change in Section 3 supplies that explanation. In practice, evaluators often draft the logic model first, because its familiar structure helps interest holders agree on what the program does, and then develop the theory of change to explain the links that matter most.
Logic models serve several purposes. In planning, they show whether the planned activities are plausibly sufficient for the intended outcomes and whether the resources match the activities. In evaluation, they show what can be measured at each stage, so that evaluation questions (Lesson 4) and indicators (Lesson 5) can be attached to specific boxes and arrows. In communication, they give funders, staff, partners and participants a shared picture of the program. In management and accountability, they underpin performance measurement frameworks in health authorities and federal departments.
2.2 The Components of a Logic Model
The standard logic model has four main components, with outcomes divided into three time frames, and two supporting elements. The University of Wisconsin-Extension model groups activities and participation together as outputs, while the Kellogg guide lists activities and outputs separately and calls long-term changes at the community or system level "impact". This course uses the six-column version shown in the table, which keeps activities and outputs apart because the distinction matters for measurement.
| Component | Definition | Question it answers | Cedar Valley example |
|---|---|---|---|
| Inputs | The resources the program uses, including funding, staff, partners, facilities, data systems and policies. | What do we invest? | A first-wave budget of $840,000, seven connector positions and a transport fund of $40,000. |
| Activities | The actions the program carries out with its inputs, written as things that staff or partners do. | What do we do? | Connectors meet participants up to six times over twelve weeks and co-develop a connection plan. |
| Outputs | The direct, countable products of activities, including the number of people reached, which are largely under the program's control. | What do we deliver, and to whom? | 312 referrals and 241 first meetings in the first six months. |
| Short-term outcomes | Early changes in participants or systems, often in knowledge, confidence, motivation or first behaviours. | What changes first? | Participants attend at least one new activity within twelve weeks. |
| Intermediate outcomes | Changes in behaviour, practice or circumstances that follow from the short-term outcomes. | What changes next? | Participants sustain participation in groups they value, and loneliness falls. |
| Long-term outcomes | Changes in health, wellbeing or conditions that the program contributes to over a longer period, often with other influences. | What ultimately changes? | Better self-rated health and fewer emergency department visits. |
| Assumptions | The beliefs about the program, the participants and the context on which the model depends. | What must be true for this to work? | Suitable groups exist within reach of rural participants. |
| External factors | Features of the environment that affect the program but lie outside its control. | What else influences the results? | Limited public transit, winter weather and other seniors' programs. |
Separating outputs from outcomes
The most important distinction in a logic model is the one between outputs and outcomes. An output describes what the program delivered; an outcome describes a change in someone or something as a result. Outputs are counted in units of service (meetings held, plans written, trips funded, people reached), and the program can increase them by working harder or spending more. Outcomes depend on how participants and partners respond, and the program can influence them without controlling them. A useful test is to ask whether the program could produce the result through its own effort alone. If it could, the result is an output. The number of older adults who attend a first meeting is an output, because it counts people reached by an activity. Whether those people later attend a community group they chose is an outcome, because it depends on their own decisions, on the group and on transport.
Outcome statements are clearest when they name who changes, what changes, in which direction and by when. "Participants report lower loneliness on the three-item UCLA Loneliness Scale at twelve weeks" is a better outcome statement than "reduced isolation", because it names the population, the measure and the time point. The time frames themselves are relative to the program. A twelve-week program can reasonably treat changes at twelve weeks as short-term, while a five-year community initiative might treat changes in the first year as short-term. The logic model should state the time frames it uses.
Reading a logic model as a chain of if-then statements
The if-then reading
If the inputs are available, then the activities can be carried out. If the activities are carried out, then the outputs will be produced. If the outputs are produced, then the short-term outcomes should follow, and if the short-term outcomes occur, then the intermediate and long-term outcomes should follow. Each "then" is a claim that can be wrong, and each claim depends on assumptions (W. K. Kellogg Foundation, 2004).
The Wisconsin model adds a further way of reading the outcome columns. Short-term outcomes are usually changes in learning (awareness, knowledge, attitudes, skills, confidence or motivation), intermediate outcomes are changes in action (behaviour, practice or decisions), and long-term outcomes are changes in conditions (health, social or economic circumstances). This sequence is a useful check: an outcome that appears early in the model but describes a change in health status, or that appears late but describes a change in knowledge, may be misplaced.
2.3 The Cedar Valley Connector Logic Model
Figure 2.1 shows a full logic model for the first wave of the Cedar Valley Connector program. It uses the facts established in Lesson 1, and it is drawn from top to bottom so that each component has room for specific statements. Most published logic models run from left to right, and either direction works as long as the components appear in order and the statements are specific.
Several features of the model reflect deliberate choices. The outputs include the actual counts from the first six months (312 referrals and 241 first meetings), which makes clear that outputs are counts of delivery and reach. The ratio between them is already informative: 241 of 312 referred older adults, or 77.2 percent, attended a first meeting, and Lesson 5 shows how such ratios become process indicators. The short-term outcomes include one change in a system as well as two changes in participants, because the partner grants are meant to change how community groups receive newcomers. Lower loneliness appears as an intermediate outcome, since the program expects it to follow from sustained participation over several months. Emergency department visits and primary care visits appear only as long-term outcomes. Lesson 1 noted that the planners' expectations about health service use have not been tested, and placing them at the end of the model, with intermediate steps in between, keeps the model honest about how far they are from the program's activities.
The model also has limits that a reader should notice. Its arrows run between rows, so it does not show which activities are expected to produce which outcomes; the transport fund, for example, is meant to act mainly on attendance, while the partner grants act on how groups receive newcomers. The assumptions are listed but not attached to particular links. The land-based connection pathway appears only as a co-design activity, because its own theory is still being developed with the First Nations partner. The theory of change in Section 3 addresses the first two limits, and a nested sub-model would address the third once the pathway is designed.
2.4 Formats: Linear, Nested and Outcome-Chain Models
Logic models come in several formats. The choice depends on the program's structure, the audience and the evaluation's purpose. Figure 2.2 sketches the three formats this course uses.
The linear format presents inputs, activities, outputs and outcomes as columns or rows in sequence, as in Figure 2.1. It is compact, familiar to funders and managers, and well suited to communication and accountability. Its weakness is that arrows between whole columns imply that every activity contributes to every outcome, which hides the specific pathways the evaluation needs to test. It also has difficulty representing feedback loops, such as a participant who joins a group and later becomes a volunteer who welcomes newcomers. A linear model is the right first draft for most programs.
A nested logic model places a high-level model of the whole program above more detailed sub-models of its components, sites or levels. Cedar Valley could have a program-level model and sub-models for the referral and connector component, the partner grants component and the land-based pathway, each with its own activities, outputs and outcomes, linked to the shared outcomes of the program. Nesting suits what Glouberman and Zimmerman (2002) and Rogers (2008) call complicated programs, which have several components, sites, or levels of organization that each require their own account. The cost is more work and a need for consistency between levels.
An outcome chain, also called an outcomes hierarchy or results chain, focuses on the causal links among outcomes. It is usually drawn from the bottom upward, with early outcomes at the base and the long-term outcome at the top, and activities are attached to the outcomes they are meant to produce (Funnell and Rogers, 2011). Because each arrow connects two specific outcomes, an outcome chain shows exactly which links the evaluation must test, and it can show several pathways converging on one outcome. It is the format closest to a theory of change, and Section 3 builds the Cedar Valley theory of change in this form. Outcome chains are harder for lay audiences to read and usually omit inputs.
The three formats are complementary. A common practice is to present a linear model in the main body of a report, a nested set of models in an appendix for program managers, and an outcome chain for the links the evaluation tests. Programs with complex features, such as outcomes that emerge from interactions among many organizations, may need models that are revised as the program develops; Rogers (2008) suggested drawing such models with explicit uncertainty and revisiting them at agreed intervals.
2.5 Common Errors in Logic Models
Logic models are easy to draw and easy to draw badly. The errors below appear often in program documents and in student work, and each has a straightforward correction.
A draft Cedar Valley model listed "connectors meet participants up to six times" in the outcome column. Meetings are something the program does, so they belong with the activities. The correction is to move the statement and ask what change the meetings are meant to produce, which yields an outcome such as "participants feel more confident about joining a group".
Statements such as "241 older adults attended a first meeting" or "120 connection plans completed" count delivery and reach. They belong with the outputs. Placing them among the outcomes allows a program to report success on measures it controls directly, which is a common way that performance reports overstate results.
A model that runs from referral to lower loneliness without stating its assumptions hides the conditions on which it depends. Cedar Valley's model depends on suitable groups existing within reach of rural participants, on participation continuing after the connector's support ends, and on loneliness being the kind that new contacts can relieve. Writing these down makes them testable.
An arrow from "connector meetings" directly to "fewer emergency department visits" skips every intermediate step. Long leaps signal that the planners have not worked out how the program is supposed to produce the outcome. The correction is to add the intermediate outcomes, which then become measurement points.
Outcomes such as "improved wellbeing" or "empowered seniors" cannot be measured as written. Each outcome should name who changes, what changes and by when, and should be specific enough that an indicator can be attached to it in Lesson 5.
Some models draw an arrow from every box to every box in the next column, which conveys no information about pathways. Arrows should represent claims the planners can explain. Where specific pathways matter, an outcome chain is the better format.
"Reduced loneliness among all older adults in the region" is a population-level change that a program serving a few hundred people a year cannot produce on its own. The model should separate outcomes for participants from population outcomes, and Section 3 introduces the accountability ceiling for this purpose.
A logic model is a working hypothesis. It should be dated, revised when the program changes or when evidence shows that a link does not hold, and kept with a record of what changed and why.
A student drafted the following logic model for a community paramedicine program in which paramedics visit older adults at home after a hospital discharge. Inputs: two paramedics; a vehicle; home visits. Activities: assess medications and home safety; refer to home care. Outputs: number of visits; fewer falls. Short-term outcomes: 400 home visits per year. Long-term outcomes: improved quality of life in the province. Identify at least four errors and propose a correction for each before reading the next paragraph.
The draft contains at least five errors. "Home visits" is an activity listed among the inputs. "Fewer falls" is an outcome listed among the outputs. "400 home visits per year" is an output listed as a short-term outcome. "Improved quality of life in the province" is a population outcome beyond the reach of two paramedics, and it is vague. The model has no intermediate outcomes, so it leaps from visits to quality of life, and it states no assumptions, such as that older adults will act on medication advice or that home care has capacity to accept referrals. A corrected model would move each statement to its proper column, add short-term outcomes such as "medication problems identified and resolved" and intermediate outcomes such as "fewer falls and fewer readmissions among visited patients within six months", restrict the long-term outcome to the program's patients, and list its assumptions.
2.6 Interactive: Build the Cedar Valley Logic Model
The builder below contains fifteen statements about the Cedar Valley program. Sort each one into the correct logic model column, then check your model. Feedback explains each misplaced statement. The statements differ from those in Figure 2.1, so the exercise tests the distinctions themselves.
2.7 From Logic Model to Evaluation
A completed logic model tells the evaluator where to look. Each output suggests a process indicator, such as the proportion of referred adults who attend a first meeting or the mean number of meetings per participant. Each outcome suggests an outcome indicator, such as the proportion of participants still attending a chosen group at six months or the mean change in the loneliness score. Each assumption suggests a question the evaluation should answer, and each external factor suggests a rival explanation the evaluation should rule out. Lesson 4 turns these into prioritized evaluation questions with the program's primary intended users, and Lesson 5 specifies the indicators, data sources and targets.
The logic model also guides the scope of the evaluation. A program in its first year, like the Cedar Valley first wave, has data on inputs, activities, outputs and some short-term outcomes, and only preliminary data on intermediate outcomes. A sensible first-year evaluation concentrates on the upper part of the model and on the earliest outcome links, while the design for estimating effects on loneliness and health service use is prepared for the second wave (Lessons 6 to 8). The logic model makes this sequencing visible to the health authority, which reduces the risk that a premature outcome evaluation is used to judge a program that is still being established.
Summary of Section 2
A logic model summarizes a program theory as a sequence from inputs and activities through outputs to short-term, intermediate and long-term outcomes, with assumptions and external factors. The output-outcome distinction is the most important one in the model: outputs count delivery and reach, while outcomes describe changes in people or systems. Linear, nested and outcome-chain formats suit different programs and audiences. Common errors, including activities listed as outcomes and missing assumptions, are easy to spot once the components are defined precisely.
Reflection
A city's community food program gives families with young children a weekly produce box, two cooking workshops a month and a referral to a dietitian, funded by a $150,000 grant. Its draft logic model reads as follows. Inputs: produce boxes; cooking workshops; a $150,000 grant. Activities: deliver boxes; 1,200 boxes delivered. Outputs: improved diet quality. Outcomes: reduced childhood obesity in the city. Identify at least four errors in this draft, then rewrite the model with correct inputs, activities, outputs, short-term, intermediate and long-term outcomes, and at least one assumption.
The draft has five errors. Cooking workshops are an activity listed as an input. The count of 1,200 boxes delivered is an output listed as an activity. Improved diet quality is an outcome listed as an output. Reduced childhood obesity in the city is a population outcome beyond the reach of one program, and the model leaps to it with no intermediate steps. No assumptions are stated.
A corrected model reads as follows. Inputs: the $150,000 grant, program staff, a produce supplier, kitchen space and a partner dietitian. Activities: pack and deliver weekly produce boxes; run two cooking workshops a month; refer families to the dietitian. Outputs: number of boxes delivered, number of families reached, workshops held and attendance, and referrals made. Short-term outcomes (three months): parents report greater confidence in cooking vegetables and families use most of the produce. Intermediate outcomes (six to twelve months): families eat more fruit and vegetables and prepare more meals at home. Long-term outcomes (two to three years): healthier diet quality and growth patterns among children in participating families. Assumptions: families have the time, equipment and storage to cook the produce, and the boxes contain foods the families are willing to eat.
Minimum 20 characters required.
Question 1: Which Cedar Valley statement belongs in the outputs column of a logic model?
Question 2: A draft logic model lists "connectors meet each participant up to six times" among the short-term outcomes. What is the error, and how should it be corrected?
Question 3: Which format is best suited to showing which specific outcomes lead to which, so that an evaluation can test individual links?
Question 4: In the University of Wisconsin-Extension reading of a logic model, short-term, intermediate and long-term outcomes usually correspond to changes in:
Theories of Change and Contribution Analysis
Learning Objectives for this section
- Distinguish a theory of change from a logic model by purpose, construction and content.
- Construct a theory of change by backward mapping from a long-term outcome through preconditions to interventions, and place an accountability ceiling.
- State the assumptions and rationales behind each link in a theory of change and judge which assumptions most need testing.
- Apply Mayne's contribution analysis to assemble and assess a contribution story where attribution is not possible.
3.1 From Logic Model to Theory of Change
The term "theory of change" entered evaluation through the work of the Aspen Institute Roundtable on Comprehensive Community Initiatives in the 1990s. Carol Weiss (1995) argued that community initiatives spanning housing, employment, health and education could not be evaluated well with designs borrowed from controlled trials, and that evaluators should ask the initiatives to spell out how and why their activities were expected to lead to their long-term goals. James Connell and Anne Kubisch (1998) developed the idea into a method and proposed that a good theory of change should be plausible, doable and testable: plausible in that the links make sense to the people involved and are consistent with evidence, doable in that the program has the resources to carry it out, and testable in that its links can be measured. Andrea Anderson's (2005) practical guide set out the backward-mapping method that most practitioners now use.
Theories of change spread to international development (Vogel, 2012) and to health research, where De Silva and colleagues (2014) showed how they support the development and evaluation of complex interventions and Breuer and colleagues (2016) reviewed their growing use in public health. The current Medical Research Council and National Institute for Health and Care Research framework for complex interventions (Skivington et al., 2021) treats an explicit program theory as a core element of intervention development and evaluation.
How a theory of change differs from a logic model
Some organizations use the two terms interchangeably. This course follows the distinction in the table, which reflects common practice (Funnell and Rogers, 2011).
| Feature | Logic model | Theory of change |
|---|---|---|
| Main purpose | Summarizes what the program does and what it expects to achieve. | Explains how and why the expected changes are supposed to happen. |
| Direction of construction | Usually built forward, from inputs and activities to outcomes. | Built backward, from the long-term outcome to the preconditions that must come first. |
| Main content | Inputs, activities, outputs and outcomes in standard columns. | Outcomes and preconditions linked in pathways, with interventions attached where they act. |
| Assumptions | Listed in a separate box, often briefly. | Attached to specific links, with the rationale and evidence for each. |
| Typical format | A one-page diagram in columns or rows. | A pathway diagram with a narrative of one or more pages. |
| Main use in evaluation | Identifies what to measure at each stage and supports performance reporting. | Identifies which links and assumptions to test and supports causal explanation. |
3.2 Backward Mapping
Backward mapping builds a theory of change by starting at the end. The planners first agree on the long-term outcome, then ask what must be in place for that outcome to occur, then ask the same question of each answer, and continue until they reach conditions that the program's interventions can produce directly. Each condition identified in this way is a precondition: an outcome that must be achieved before the outcome above it can be achieved. Anderson (2005) describes the method in the following sequence of steps, which this course adopts.
The first step is to identify the long-term outcome. It should be specific enough to measure, important to the people the program serves, and realistic for the program's scale. The second step is to map the preconditions backward. For each outcome, the planners ask what has to be true immediately before it, and they test each proposed precondition by asking whether the higher outcome could occur without it. If it could, the precondition is not necessary and may not belong on the map. When the set of preconditions under an outcome is complete, the planners ask whether those preconditions, together with the stated assumptions, would be sufficient. The third step is to identify the interventions: the program activities that produce the earliest preconditions. Gaps at this stage reveal preconditions that no activity addresses, which point to missing activities or partners. The fourth step is to state the assumptions and rationales for each link. The fifth step is to specify indicators for each precondition, naming the population, the change expected and the time frame; Lesson 5 develops this step. The sixth step is to write a narrative that explains the pathway in prose.
After the map is drawn backward, it is read forward as a test, with each link read as a sentence of the form "this happens, so that that happens". If a sentence sounds implausible, or an obvious step is missing, the map needs revision. Skipping this forward reading produces the same long leaps that Section 2 identified in logic models.
The accountability ceiling
Long-term outcomes are usually influenced by many factors outside a program's control. The accountability ceiling is a line drawn across a theory of change above which the program does not hold itself accountable for producing the outcomes, although it still expects to contribute to them. Outcomes below the ceiling are the ones against which the program's performance should be judged. Outcomes above the ceiling remain part of the theory and may still be monitored, but a failure to observe change there, especially within a short period, should not on its own be read as program failure. Placing the ceiling is a negotiation among the program, its funders and its evaluators, and it should be done before the data arrive.
Take the Cedar Valley intermediate outcome "participants keep taking part in a chosen group after the twelve weeks end". Ask what must be in place immediately before it, and list two or three preconditions. For each, check that the outcome could not occur without it. Then name the program activity that produces each precondition, or note that no current activity does. Compare your answer with the left-hand pathway in Figure 3.1.
3.3 Assumptions and Rationales
Two related terms describe the reasoning behind each link. A rationale explains why one outcome is expected to lead to the next, by citing research evidence, practice experience, pilot data or formal theory. An assumption is a condition that must hold for the link to work, but that the program does not control or has not tested. Some writers, including Anderson (2005), use "assumptions" to cover both. The practical point is that each link on the map should carry a short statement of why the planners expect it to hold and what would have to be true for it to fail.
Assumptions take several forms. Some concern the causal link itself, for example that new contacts reduce the kind of loneliness that participants experience. Some concern the context, for example that suitable groups exist within reach of rural participants. Some concern participants, for example that older adults referred by a clinician will be willing to meet a connector. Some concern implementation, for example that clinicians will screen consistently during busy appointments. Assumptions can be rated on two dimensions: how important the assumption is to the theory, and how strong the evidence for it is. Assumptions that are important and weakly supported should receive the most attention in the evaluation.
| Link and assumption | Rationale and evidence | How the evaluation tests it |
|---|---|---|
| A1. Screening to referral. Clinicians screen eligible older adults consistently during routine visits. | The eighteen-month pilot tested the referral pathway in two clinics. Whether screening holds up across twelve clinics of different sizes and staffing is untested. | Compare referrals with the number of eligible patients seen in each clinic, using electronic medical record data. |
| A2. Early preconditions to first attendance. Suitable, affordable groups and roles exist within reach of participants, including rural participants who no longer drive. | The partner list is strongest in the city. The transport fund was created because rural access was a known barrier, but the supply of rural activities is uncertain. | Connectors record requests they could not match; compare first attendance between urban and rural participants. |
| A3. First attendance to continued participation. Participation continues after the connector's support ends. | Accompaniment and welcoming practices are intended to make attendance self-sustaining. Evidence on how long participation lasts after social prescribing is limited. | Ask participants about attendance at six months; interview people who stopped attending. |
| A4. Participation and contacts to lower loneliness. The new contacts address the kind of loneliness participants feel. | A meta-analysis by Masi and colleagues (2011) found that interventions addressing maladaptive social cognition reduced loneliness more than interventions that increased opportunities for contact, so opportunity alone may help some participants and not others. | Examine change in loneliness by baseline characteristics and explore experiences in interviews; Section 4 develops this assumption into realist configurations. |
| A5. Lower loneliness to health and service use. Less loneliness leads to better self-rated health and fewer visits made mainly for social reasons. | Observational studies link loneliness and isolation with poorer health and higher mortality (Holt-Lunstad et al., 2015), but evidence that reducing loneliness improves these outcomes is limited. | Monitor self-rated health and use linked administrative data on visits; interpret above the accountability ceiling. |
3.4 The Cedar Valley Theory of Change
Figure 3.1 shows a theory of change for the first wave of the Cedar Valley Connector program, built by backward mapping. It uses the same program facts as the logic model in Figure 2.1, but it is organized as a pathway of preconditions, and it attaches assumptions A1 to A5 from the table above to specific links.
The map was built in the order backward mapping prescribes. The planners began with the long-term outcome: better self-rated health and fewer emergency department and social-reason primary care visits among older adults in the region's participating clinics. They asked what must come first and identified the intermediate outcome of lower loneliness and higher social participation among participants. They placed the accountability ceiling between these two levels, because health and service use depend on many factors beyond the program, and because the evidence that reducing loneliness improves health (assumption A5) is weaker than the evidence for the earlier links. Below the intermediate outcome, they identified two preconditions that must both be present: participants must keep taking part after the twelve weeks end, and they must form contacts they experience as meaningful. Taking part without meaningful contact, or a single meaningful conversation without continued participation, would be unlikely to reduce loneliness. Both of these depend on a middle precondition, attendance at a first activity or role that the participant chose, and that in turn depends on four early preconditions, each produced by one of the four interventions.
The forward reading of the map yields a testable narrative. Clinicians screen and refer, so that eligible older adults are identified; connectors meet participants and agree a plan, so that participants have a reason and a plan to take part; accompaniment and transport help make attendance feasible, and partner grants make groups welcoming, so that participants attend a first activity they chose; participants keep attending and form meaningful contacts, so that loneliness falls and participation rises. The map makes two points visible that the logic model did not. It shows that first attendance is a bottleneck through which every pathway passes, which makes it an early indicator of whether the theory is working. It also shows that the weakest assumptions (A3 and A4) sit on the links just below the accountability ceiling, which tells the evaluation team where to concentrate its qualitative work.
3.5 Contribution Analysis
Lessons 6 to 8 teach designs that estimate how much of an observed change a program caused, by comparing outcomes with a counterfactual: an estimate of what would have happened without the program. This is attribution. Such designs are not always feasible. A program may serve everyone eligible, leaving no comparison group; it may be too early in its development for an outcome evaluation; or its outcomes may depend on so many other actors that a single effect estimate would mean little. John Mayne, then at the Office of the Auditor General of Canada, proposed contribution analysis for these situations (Mayne, 2001). It asks a different question from attribution: whether it is reasonable to conclude that the program made an important contribution to the observed outcomes, given the evidence on its theory of change and on other influencing factors.
Contribution analysis aims to reduce uncertainty about a program's contribution, and its product is a contribution story: an evidence-based account of how the program contributed to the outcomes, which links in the theory of change are well supported, and which rival explanations have been examined. Mayne (2001) first presented the approach for use with performance measurement data in the Canadian Journal of Program Evaluation, and he later described it in six steps (Mayne, 2008) and reviewed its development (Mayne, 2012). The tabs apply the six steps to the first six months of the Cedar Valley program.
Set out the cause-effect issue to be addressed. The evaluation team frames the question as "Has the Connector program contributed to lower loneliness among participants in the first-wave clinics, and through which links?" It agrees with the health authority that the question concerns outcomes below the accountability ceiling, and it records what level of confidence the authority needs before the second-wave decision.
Develop the postulated theory of change and the risks to it. The team uses the theory of change in Figure 3.1 and lists the rival explanations that could produce the same pattern of results. These include regression to the mean, since only people scoring 6 or higher were referred and high scores tend to drift toward the average on remeasurement; natural recovery after events such as bereavement; seasonal change, if many people were referred in winter and followed up in spring; other seniors' programs that started in the same period; selective loss to follow-up; and connectors administering the follow-up questionnaire themselves, which may encourage favourable answers.
Gather the existing evidence on the theory of change. Program records show that 312 older adults were referred and 241 attended a first meeting (77.2 percent). Of these 241, 188 had both a baseline and a twelve-week loneliness score (78.0 percent), and their mean score fell from 7.1 to 6.3, a decline of 0.8 points on the 3 to 9 scale. The remaining 53 participants have no follow-up score. Meeting logs, connection plans and linkage records describe delivery. The wider literature on social prescribing is supportive but limited in quality: a systematic review by Bickerdike and colleagues (2017) concluded that the evidence base was weak and called for better evaluations.
Assemble and assess the contribution story and the challenges to it. The team judges that the early links are well supported by records: referrals arrived, most referred people attended a first meeting, and delivery broadly followed the protocol. The links through continued participation and meaningful contact (A3 and A4) have little direct evidence yet. The 0.8-point fall cannot by itself be credited to the program, because regression to the mean alone would be expected to produce some decline in a group selected for high scores, and the 53 participants without follow-up may differ from those with it. The story at this stage is plausible but weak.
Seek out additional evidence. The team plans data that would strengthen or weaken the story where it is weakest. It adds a six-month question on continued participation, interviews with participants who did and did not keep attending, and a comparison of loneliness change across levels of participation, since a larger fall among people who kept attending is what the theory predicts. It also proposes measuring loneliness twice before referral in a sample of patients, or using older adults in second-wave clinics who are screened but not yet served, to estimate how much decline occurs without the program. Lesson 7 develops the second option into a comparison-group design.
Revise and strengthen the contribution story. With the new evidence, the team rewrites the story, states which links are supported and which remain uncertain, and reports its level of confidence. If loneliness fell most among people who kept attending a chosen group, and fell less in the comparison group, the story becomes considerably stronger. If the fall was similar regardless of participation, the team would conclude that regression to the mean or another factor probably explains much of it.
When is a contribution claim credible?
Mayne (2008, 2012) argued that a reasonable contribution claim can be made when four conditions hold. The program is based on a reasoned theory of change, with plausible links and assumptions that are at least partly supported. The activities were implemented as set out in the theory of change. The theory of change is supported by evidence: the expected chain of results occurred and the key assumptions held. Other influencing factors have been assessed and either shown not to have made a significant contribution or their relative role has been recognized.
"In its first six months, the Connector program received 312 referrals from twelve clinics, and 241 older adults attended a first meeting. Connectors delivered meetings and plans broadly as designed. Among the 188 participants with both measurements, mean loneliness fell from 7.1 to 6.3 on a scale from 3 to 9. This decline is consistent with the program's theory of change, but it cannot yet be credited to the program. Participants were selected for high loneliness scores, so some decline would be expected without the program, and 53 participants have no follow-up score. Evidence on whether participants keep taking part after twelve weeks, and on whether their new contacts are meaningful to them, is not yet available. We therefore judge the program's contribution to lower loneliness as plausible but unconfirmed, and we have planned additional evidence on these links before the second-wave decision."
Strengths and limits of contribution analysis
Contribution analysis suits programs in settings where many factors influence outcomes and an experimental counterfactual is unavailable. It makes the reasoning behind a causal claim explicit, uses evidence the program already collects, and directs new data collection to the weakest links. It does not estimate an effect size, so it cannot say how much of the change the program caused, and its conclusions depend on the evaluator's judgement about rival explanations, which can lean toward confirming the program's theory. Mayne (2012) encouraged combining it with other methods; Befani and Mayne (2014), for example, combined it with process tracing, which applies formal tests to the evidence for each causal link. Contribution analysis is best seen as complementary to the designs in Lessons 6 to 8, and at Cedar Valley the second-wave rollout offers a comparison group that can turn the contribution story into a stronger causal claim.
Summary of Section 3
A theory of change explains how and why a program is expected to produce its long-term outcome. It is built by backward mapping, read forward as a test, and annotated with rationales and assumptions for each link, with an accountability ceiling separating outcomes the program is judged on from outcomes it contributes to. Contribution analysis uses the theory of change to build and test a contribution story when attribution is not possible.
Reflection
A public health unit runs a school-based vaping prevention program in grade 8 classrooms. It has three components: a four-session curriculum delivered by teachers, a peer-leader component in which trained grade 11 students lead small-group discussions, and a letter to parents with conversation tips. Its long-term goal is lower vaping among students when they reach grade 10. (a) Using backward mapping, identify one intermediate outcome and at least three preconditions, and name the program component that produces each early precondition. (b) Place an accountability ceiling and justify your choice. (c) State two assumptions with a rationale for each. (d) Name one rival explanation that a contribution analysis would need to examine if grade 10 vaping fell.
(a) Working backward from lower vaping in grade 10, the intermediate outcome is that students who are offered a vape in grades 9 and 10 decline it. For that to happen, students need accurate beliefs about the harms and addictiveness of nicotine, the confidence and skills to refuse an offer, and a perception that most of their peers do not vape. Further down, parents need to talk with their children about vaping. The curriculum produces accurate beliefs, the peer-leader groups produce refusal skills and corrected peer norms, and the parent letter produces parent conversations.
(b) I would place the accountability ceiling between refusing offers and grade 10 vaping prevalence. Prevalence depends on product availability, marketing, enforcement and household vaping, which the program does not control, so the program should be judged on beliefs, skills, norms and refusal.
(c) The first assumption is that teachers deliver all four sessions as designed; the rationale is that curriculum effects depend on delivery, which varies across schools. The second assumption is that grade 8 students regard grade 11 peer leaders as credible; the rationale is that norm-based programs work through perceived similarity to the messenger.
(d) A contribution analysis would need to examine a provincial change in vaping product rules during the same period, which could reduce vaping in all schools regardless of the program.
Minimum 20 characters required.
Question 1: What distinguishes backward mapping from the way logic models are usually drafted?
Question 2: In a theory of change, what is the accountability ceiling?
Question 3: In Mayne's contribution analysis, which condition is part of a credible contribution claim?
Question 4: Among the 188 Cedar Valley participants with both measurements, mean loneliness fell from 7.1 to 6.3. Why can this fall not yet be credited to the program?
Realist Program Theory and a Worked Example
Learning Objectives for this section
- Explain the realist account of causation and the question realist evaluation asks: what works, for whom, in what circumstances, and why.
- Write context-mechanism-outcome configurations that separate the resources a program offers from the reasoning they trigger.
- Describe how an initial program theory is developed from documents, interviews, literature and formal theory, and how it is refined.
- Produce a logic model and a one-page theory of change narrative with stated assumptions for a program, following the Cedar Valley worked example.
4.1 The Realist Account of Causation
The program theories in Sections 1 to 3 describe a single pathway that is expected to apply to everyone the program serves. Ray Pawson and Nick Tilley (1997) argued that this expectation is usually wrong. In Realistic Evaluation, they observed that the same program produces different results for different people in different settings, and that evaluations reporting only an average effect conceal this variation. They proposed that the useful question is what works, for whom, in what circumstances, and why. Later realist writing extends the question to ask in what respects and to what extent a program works (Pawson, 2013).
The approach rests on scientific realism. Its central claim for evaluation is that programs do not produce outcomes directly: they offer resources, such as information, support, opportunities or material help, and outcomes depend on how people respond. Pawson and Tilley called the response a mechanism and argued that mechanisms operate only in certain contexts. This generative account of causation explains an outcome by identifying the mechanism that produced it and the conditions that allowed it to operate. It differs from the successionist account behind experimental designs, in which a cause is inferred from a regular association between program and outcome under controlled conditions. Realist evaluators accept that experiments can estimate average effects, and they argue that an average effect tells decision-makers little about where and for whom to deliver a program.
Lesson 1 introduced realist evaluation as one of the approaches on the methods branch of Alkin and Christie's evaluation theory tree, and it presented two Cedar Valley configurations. This section teaches how to write such configurations and how to assemble them into a program theory that an evaluation can test.
4.2 Context-Mechanism-Outcome Configurations
The analytic unit of realist evaluation is the context-mechanism-outcome configuration, usually abbreviated in writing as CMOC. It states that in a given context, a given mechanism is triggered, which produces a given outcome. Pawson and Tilley (1997) expressed the relationship in a short formula, and Dalkin and colleagues (2015) refined it by dividing the mechanism into two parts.
Two forms of the realist formula
Pawson and Tilley (1997): Context + Mechanism = Outcome.
Dalkin et al. (2015): Mechanism (resource) + Context → Mechanism (reasoning) = Outcome.
In the second form, the resource is what the program offers, the reasoning is how participants think, feel or decide in response, and the context determines whether the resource triggers that reasoning.
Each element has a precise meaning. The context is the set of conditions that determines whether a mechanism is triggered. It includes characteristics of participants, such as their history, relationships, beliefs and circumstances, and features of the setting, such as available services, organizational culture and community norms. A demographic label such as "rural" is a context only if the evaluator can say why it matters; "rural participants who no longer drive and live beyond walking distance of any group" states the condition that makes transportation help important. The mechanism is the participant's response to what the program offers. It is usually hidden, it operates at the level of reasoning or feeling, and it must be inferred from evidence. The outcome is the result, intended or unintended, and it may be an early outcome such as attendance or a later one such as lower loneliness.
The table below sets out four configurations for Cedar Valley. The first two develop the examples given in Lesson 1. The third describes a configuration in which the program may fail or cause harm, which realist evaluation treats as equally informative. The fourth illustrates a mechanism that operates through a volunteer role and a sense of purpose.
| Context | Resource | Reasoning | Outcome |
|---|---|---|---|
| Older adults who were recently widowed and have no existing link to community groups. | The connector offers to go with the person to the first session of a group. | Anxiety about entering a room of strangers eases, and the person feels expected. | The person keeps attending and reports less loneliness at twelve weeks. |
| Rural older adults who no longer drive and live beyond walking distance of any group. | The transport fund pays for rides to and from a chosen activity. | Attendance becomes feasible and predictable, so the person commits to a regular schedule. | Regular participation is established and sustained after twelve weeks. |
| Older adults whose loneliness accompanies low mood or long-standing social anxiety. | The connector encourages the person to join a group early in the twelve weeks. | The person experiences the suggestion as pressure, and an awkward first visit confirms a fear of rejection. | The person withdraws from the program, and loneliness does not improve or worsens. |
| Recently retired adults with skills they value who feel they have lost a role. | The connector arranges a volunteer role matched to the person's skills. | The person feels needed and useful, and the role gives structure to the week. | Sustained participation, a wider network and less loneliness. |
The third configuration has direct implications for the program. If the evaluation finds evidence for it, Cedar Valley might change its protocol for participants with signs of low mood, for example by pacing the move toward group activity more slowly or by linking them with mental health support first. This kind of finding is what realist evaluation is designed to produce: specific guidance about how to adapt a program to the circumstances of different participants.
Writing configurations well
Three errors are common in configurations written by evaluators new to the approach. The first is to write a program activity as the mechanism, as in "the mechanism is accompaniment"; accompaniment is the resource, and the mechanism is the change in the participant's reasoning that accompaniment produces. The second is to write a context that does not explain anything, such as "women" or "people in the city", without stating what about that group makes the mechanism more or less likely to fire. The third is to state an outcome that cannot be observed, such as "empowerment", without saying how it would be recognized. A useful drafting device is the "if-then-because" statement: if the program offers this resource to people in this context, then this outcome will follow, because people will respond in this way. The "because" clause contains the mechanism.
4.3 Developing an Initial Program Theory
A realist evaluation begins with an initial program theory: a set of tentative configurations that state how the program is expected to work, for whom and in what circumstances. The evaluation then tests and refines this theory, and its main product is a refined program theory. Initial program theories are built from several sources. Program documents and interviews with designers and staff reveal the reasoning behind the program. Realist interviews, in which the evaluator presents candidate theories to an interviewee and asks them to confirm, refute or refine them, are a distinctive tool for this purpose; Pawson (1996) called this the teacher-learner relationship, and Manzano (2016) gave practical guidance. The research literature, including realist reviews of social prescribing (Husk et al., 2020; Tierney et al., 2020), offers configurations that other researchers have proposed and tested. Formal social science theories, which realists call middle-range theories after Robert Merton, offer general explanations of mechanisms, such as theories of how loneliness affects attention to social threat (Hawkley and Cacioppo, 2010). People with lived experience, such as the older adults on the Cedar Valley steering committee, can say which mechanisms ring true and which do not.
The evaluator collects the explanations offered in documents, interviews, literature and formal theory, without yet judging them. At Cedar Valley this would include the connectors' accounts of why some participants keep attending and others do not, and the explanations offered in published reviews of social prescribing.
Each candidate theory is written as a statement with a context, a resource, a response and an outcome. Statements that cannot be written in this form are usually too vague to test and need further discussion with the people who proposed them.
The evaluator groups related statements, removes duplicates and arranges the configurations along the program's pathway, often using the theory of change as a frame. The four Cedar Valley configurations in Section 4.2 sit at different points on the pathway in Figure 3.1.
An evaluation cannot test every configuration. The team selects those that are most important to decisions, most uncertain and most feasible to examine. At Cedar Valley, the configuration about low mood would be a priority because, if supported, it would change how connectors work with some participants.
Realist evaluations use mixed methods. Quantitative data show patterns of outcomes across contexts, such as whether loneliness fell more among bereaved participants who were accompanied. Qualitative data, especially realist interviews, explain why those patterns occur. Realists call recurring but imperfect patterns of this kind demi-regularities.
The evaluator revises each configuration in light of the evidence, by a process realists call retroduction: reasoning from the observed patterns back to the mechanisms and contexts that best explain them. The refined theory is reported with the evidence for each configuration, following the RAMESES II reporting standards for realist evaluations (Wong et al., 2016).
The configurations in an initial program theory should be developed with the people whose programs they describe. For the land-based connection pathway being co-designed with one First Nation, the theory of how land-based activities support connection belongs first to that Nation and its knowledge holders. An evaluator's role would be to help articulate and test configurations in ways the partners consider appropriate, under the governance arrangements discussed in Lesson 4.
4.4 What Works, for Whom, in What Circumstances and Why
Realist evaluation changes the practical design of an evaluation in several ways. Sampling for interviews is purposive across contexts, so that the evaluation hears from participants in the circumstances where different mechanisms are expected, such as bereaved and non-bereaved participants, urban and rural participants, and those who stopped attending as well as those who continued. Quantitative analysis is planned around the configurations, for example by comparing outcomes across the contexts they specify. Findings take the form of refined configurations, which tell decision-makers where and for whom the program works and what to adapt.
The approach is especially useful when an average result conceals opposing effects. Suppose the Cedar Valley evaluation finds little average change in loneliness. A realist analysis might show that bereaved participants who were accompanied improved substantially, while participants with low mood who were pushed toward groups early withdrew. The average would mislead the health authority, which might abandon a program that works well for a large group of participants. The realist finding instead points to a targeted change in protocol.
Realist evaluation also has limits. Context and mechanism are hard to separate in practice, and different evaluators may classify the same factor differently. The approach demands time and skill, particularly for realist interviewing and retroductive analysis, and its findings are harder to summarize for decision-makers who want a single answer. Configurations developed after the data are collected can become stories that fit the evidence too easily, a risk reduced by stating the initial program theory in advance and reporting evidence that contradicts it.
Choosing among the three forms of program theory
| Form | Question it answers | Main strength | Main limit |
|---|---|---|---|
| Logic model | What does the program do, and what does it expect to achieve? | A compact, shared picture of the program that identifies what to measure at each stage. | Shows sequence without explaining why links hold or for whom. |
| Theory of change | How and why is the program expected to produce its long-term outcome? | Explains the pathway and attaches testable assumptions to each link. | Usually describes one pathway that is assumed to apply to everyone. |
| Realist program theory | What works, for whom, in what circumstances, and why? | Explains variation in outcomes and guides adaptation and targeting. | Demands time and skill, and context and mechanism are hard to separate. |
The three forms are complementary. Most evaluation plans include a logic model and a theory of change, and many add realist configurations for the links where variation across participants is expected. The Cedar Valley worked example in Section 4.5 shows the first two.
4.5 Worked Example: The Cedar Valley Logic Model and Theory of Change Narrative
The worked example below sets out a logic model and a theory of change for the Cedar Valley Connector program. The first part is the logic model, which for Cedar Valley is Figure 2.1 in Section 2. The second part is a one-page theory of change narrative with stated assumptions, which follows. The narrative refers to the pathway in Figure 3.1, and a narrative of this kind is normally accompanied by a diagram like it.
Long-term outcome and accountability ceiling. The Cedar Valley Connector program aims to improve self-rated health among older adults referred from participating clinics and to reduce their emergency department visits and primary care visits made mainly for social reasons. These outcomes depend on many factors beyond the program, so they sit above the accountability ceiling. The program holds itself accountable for the intermediate outcome of lower loneliness and higher social participation among participants within twelve months of referral.
Pathway. Clinicians in the twelve first-wave clinics screen adults aged 65 and older with the three-item UCLA Loneliness Scale and refer those who score 6 or higher, or whom they judge to be isolated. A connector contacts each person within ten business days and meets them up to six times over twelve weeks to agree a connection plan based on what matters to them. Connectors accompany participants to a first activity where helpful and arrange transport through the $40,000 transport fund, while small grants to community partners support welcoming practices such as a buddy at a newcomer's first session. These early preconditions together lead participants to attend a first activity or role they chose. Participants who keep taking part after the twelve weeks end and who form contacts they find meaningful are expected to report lower loneliness and to take part more in community life. Lower loneliness is in turn expected to contribute to better self-rated health and less use of emergency and primary care for social reasons.
Assumptions. The pathway rests on five assumptions. (A1) Clinicians screen consistently during routine visits; the pilot tested the referral pathway in two clinics, but its consistency across twelve clinics is untested. (A2) Suitable groups and roles exist within reach of participants, including rural participants who no longer drive; supply is strongest in the city. (A3) Participation continues after the connector's support ends; accompaniment and welcoming practices are designed to make this likely, but evidence is limited. (A4) New contacts address the kind of loneliness participants feel; research suggests that opportunity for contact helps less when loneliness is maintained by negative expectations of others (Masi et al., 2011), so the program may work less well for some participants. (A5) Lower loneliness leads to better health and less use of care for social reasons; this is supported by observational associations, and causal evidence is limited.
Implications for the evaluation. Assumptions A3 and A4 are the most important and least supported, and they sit on the links just below the accountability ceiling, so the evaluation will give them the most attention. Attendance at a first chosen activity is the point through which every pathway passes, and it will serve as an early indicator of whether the theory is working. The evaluation will also examine whether the program works differently for participants with low mood, for whom early encouragement to join groups may be counterproductive.
What makes the example work
The narrative names a specific long-term outcome and places the accountability ceiling explicitly, so readers know what the program should be judged on. It describes the pathway in the order of the diagram and uses the program's real details (clinic numbers, the referral threshold, the meeting schedule and the transport fund) consistently with the program description from Lesson 1. Each assumption is attached to a link, states its evidence, and is phrased so that the evaluation can test it. The final paragraph turns the theory into priorities for the evaluation, which prepares the ground for the evaluation questions in Lesson 4. At about 480 words, it fits on one page.
Checking a logic model and theory of change
Together, a logic model and a theory of change narrative set out the causal argument an evaluation will test, so that its questions, indicators and design can be tied to specific links and assumptions. A draft can be checked against five points.
- Statements are placed in the correct components, and every outcome describes a change in people or systems.
- The pathway is plausible, has no long leaps, and is consistent with the program description and objectives.
- Assumptions are attached to specific links, are supported by stated evidence, and can be tested.
- The accountability ceiling is placed and justified.
- The program’s interest holders can read and recognize the model and narrative.
In later lessons, the links and assumptions of the Cedar Valley theory become evaluation questions (Lesson 4), the boxes of its logic model become indicators (Lesson 5), and the most important causal links shape the choice of design (Lessons 6 to 8).
Reflection
The fictional Cedar Valley Connector program refers adults aged 65 and older who screen as lonely to a community connector. The connector meets each person up to six times over twelve weeks, co-develops a connection plan, and links the person to community groups, volunteer roles, transportation help and services, sometimes going with them to a first activity. Write two context-mechanism-outcome configurations for this program: one in which the program is likely to work and one in which it is likely to fail or cause harm. Do not reuse the configurations about bereavement, rural transport, low mood or retirement given in the lesson. For each configuration, separate the resource from the reasoning, explain why the context matters, and state one piece of evidence the evaluation could collect to test it.
Configuration 1 (likely to work). Context: older adults who care for a spouse with dementia and have stopped seeing friends because they cannot leave the house for long. Resource: the connector links the caregiver to a caregiver support group that offers on-site respite for the spouse during meetings. Reasoning: the caregiver feels it is safe and acceptable to take time for themselves, and recognizes others in the group as people who understand their situation. Outcome: regular attendance and lower loneliness at twelve weeks. The context matters because the usual barrier is the caregiving role, which the respite resource directly addresses. Evidence: compare attendance and loneliness change among caregivers linked to groups with and without respite, and ask caregivers in interviews what made attending possible.
Configuration 2 (likely to fail). Context: older adults with untreated hearing loss. Resource: the connector refers them to a large, noisy social group. Reasoning: they struggle to follow conversation, feel embarrassed and exhausted, and conclude that groups are not for them. Outcome: they stop attending after one or two visits, and loneliness may increase. The context matters because hearing loss changes how a group setting is experienced. Evidence: add a hearing question at intake and compare continued attendance by hearing status, and interview participants who stopped attending.
Minimum 20 characters required.
Question 1: In Dalkin and colleagues' refinement of the realist formula, which element is the reasoning part of the mechanism in the Cedar Valley accompaniment configuration?
Question 2: Which statement is a well-formed realist context?
Question 3: What is an initial program theory in realist evaluation?
Question 4: A realist analysis finds that bereaved participants who were accompanied improved, while participants with low mood who were encouraged to join groups early withdrew. What does this finding imply?
Final Assessment
Bringing It All Together
This lesson has treated program theory as the argument a program makes about how change happens, and it has shown four ways of making that argument explicit. Chen's change and action models divide the theory into the causal process the program relies on and the arrangements needed to deliver it, and Weiss's distinction between implementation failure and theory failure shows why an evaluation that measures only final outcomes cannot explain what it finds. Logic models summarize the theory as a sequence from inputs to outcomes, and their value depends on precise definitions, especially the line between outputs and outcomes.
Theories of change explain the sequence. They are built backward from the long-term outcome, read forward as a test, and annotated with the assumptions and rationales that make each link testable, with an accountability ceiling that separates what a program is judged on from what it contributes to. Contribution analysis uses a theory of change to build a credible account of a program's contribution when attribution is not possible, and the Cedar Valley example showed how regression to the mean and loss to follow-up limit what a pre-post change can show. Realist program theory adds the question of for whom and in what circumstances a program works, expressed as context-mechanism-outcome configurations.
For the fictional Cedar Valley Connector program, these tools converge on the same priorities: first attendance at a chosen activity as an early indicator, and the assumptions about continued participation and meaningful connection as the links most in need of evidence.
Key Takeaways from this lesson
- Program theory is the explicit account of how a program is expected to produce its outcomes, and it combines a causal account with an operational account.
- Chen's change model consists of the intervention, the determinants and the goals and outcomes, while his action model describes the six elements that must be organized for delivery.
- Evaluations that measure only final outcomes cannot distinguish implementation failure from theory failure, and they cannot say which components carried an effect.
- A logic model presents inputs, activities, outputs and outcomes in three time frames, together with assumptions and external factors.
- Outputs count delivery and reach and are largely under the program's control, while outcomes describe changes in people or systems that depend on how they respond.
- Linear, nested and outcome-chain formats suit different programs and audiences, and outcome chains show the specific links an evaluation can test.
- Backward mapping builds a theory of change from the long-term outcome through necessary preconditions to interventions, and forward reading tests the result.
- Assumptions and rationales should be attached to specific links, rated by importance and strength of evidence, and tested where they are important and weakly supported.
- Contribution analysis assembles and tests a contribution story when attribution is not possible, and its claims are credible only after rival explanations have been assessed.
- Realist evaluation explains variation in outcomes through context-mechanism-outcome configurations, and it begins with an initial program theory that the evaluation refines.
Core Concepts Reviewed
Section 1: program theory, social science and evaluation theory, espoused theory and theory-in-use, Chen's change and action models, determinants, black-box evaluation, and implementation and theory failure.
Section 2: logic model components, the output-outcome distinction, if-then reading, linear, nested and outcome-chain formats, common errors, and the Cedar Valley logic model.
Section 3: theory of change, backward mapping, preconditions, the accountability ceiling, assumptions and rationales, attribution and contribution, and Mayne's six steps of contribution analysis.
Section 4: realist causation, context-mechanism-outcome configurations with resource and reasoning, initial program theory, realist interviews, refinement, and the Cedar Valley theory of change narrative.
The final reflection asks you to apply the lesson's tools to a program you have not seen before.
Reflection
A health authority plans to evaluate a community paramedicine program in which paramedics visit adults aged 65 and older at home within seven days of a hospital discharge. At each visit, the paramedic reviews medications, checks home safety and refers the person to home care where needed. The program operates from every hospital in the region, so no comparison group is available. Using concepts from all four sections of this lesson, describe (a) one change model determinant and one action model element the evaluation should measure; (b) the long-term outcome, one intermediate outcome and two preconditions of a theory of change for the program, with an accountability ceiling; (c) two assumptions on specific links; and (d) how contribution analysis would be used, including one rival explanation it must address. You may add one context-mechanism-outcome configuration if it strengthens your answer.
(a) A key determinant in the change model is the number of medication problems identified and resolved, since the program expects better medication use to prevent complications. An action model element to measure is the program implementers: whether paramedics have the training and time to complete a full medication review at every visit.
(b) The long-term outcome is fewer emergency department visits and readmissions within ninety days of discharge. An intermediate outcome is that patients take their medications as prescribed and have home care in place within two weeks. Two preconditions are that medication problems are identified and resolved with the patient's pharmacist or physician, and that home care accepts referrals promptly. I would place the accountability ceiling between the intermediate outcome and readmissions, because readmission depends on illness severity and hospital practices beyond the program's control.
(c) On the link from referral to home care in place, the program assumes that home care has capacity to accept referrals within two weeks. On the link from medication review to correct use, it assumes that patients act on the paramedic's advice.
(d) Because no comparison group exists, contribution analysis would set out this theory of change, gather evidence on each link from visit records and home care data, and assess rival explanations. One rival is a simultaneous change in hospital discharge planning, which could reduce readmissions on its own. A configuration would add that patients living alone may benefit most because no family member checks their medications.
Minimum 30 characters required.
Final Knowledge Assessment
Question 1: A program's logic model shows activities leading to outputs and outcomes. Which addition would turn it into a theory of change in the sense used in this lesson?
Question 2: Which pairing correctly matches Chen's two models with the evidence most often used to evaluate them?
Question 3: At one Cedar Valley clinic, connectors delivered the full protocol and participants joined groups, but loneliness did not fall, and interviews show the groups produced no close relationships. This pattern is best described as:
Question 4: Which of the following statements describes an outcome?
Question 5: Where should "fewer emergency department visits among participants" sit in the Cedar Valley theory of change, and why?
Question 6: During backward mapping, a team proposes "participants receive a monthly newsletter" as a precondition of "participants attend a first chosen activity". Which test from the method best checks the proposal?
Question 7: Which situation most clearly calls for contribution analysis?
Question 8: Which rival explanation is especially relevant to the Cedar Valley pre-post change in loneliness because of the program's referral rule?
Question 9: In a realist configuration about rural participants who no longer drive, the transport fund is best classified as:
Question 10: A logic model draws an arrow directly from "connector meetings" to "better self-rated health". Which error is this, and what is the correction?
Question 11: Why does a realist evaluation sample interviewees purposively across contexts?
Question 12: Which source is most distinctive to developing an initial program theory in realist evaluation?
Question 13: Which Cedar Valley assumption is both important and weakly supported, and so deserves the most evaluation attention?
Question 14: Which statement correctly contrasts the contributions of Carol Weiss and of Ray Pawson and Nick Tilley?
Question 15: A draft theory of change narrative lists its assumptions in a final paragraph without linking them to the pathway. Which change would most strengthen it?
Glossary: Key Terms, People & Frameworks
📚 Reference page, available throughout the lesson
This glossary defines the terms, frameworks and people introduced in Lesson 3; use the search box to filter entries.