HSCI 826 · Lesson 10

Economic Evaluation, Reporting and Evaluation Use

Program Planning & Evaluation

Learning objectives for this lesson:

  • Distinguish the financial cost of a program from its economic cost, and explain how the perspective of an analysis determines which costs are counted.
  • Calculate a program's cost per participant with gross costing and time-driven activity-based costing, and explain how the use of capacity changes it.
  • Distinguish cost-minimization, cost-effectiveness, cost-utility, cost-benefit and cost-consequence analysis, and calculate the QALYs gained from utility measurements.
  • Calculate an incremental cost-effectiveness ratio and net monetary benefit, place the result on the cost-effectiveness plane, and interpret scenario and probabilistic sensitivity analyses.
  • Explain the cautions that apply to social return on investment, and identify the reporting items in CHEERS 2022.
  • Construct an evaluative rubric and synthesize mixed evidence into judgements of merit and worth with a stated level of confidence.
  • Describe a program with the TIDieR checklist so that others can understand, cost and replicate it.
  • Plan reports, executive summaries, data visualizations and knowledge translation products for decision-makers and communities, and explain the conditions under which evaluations are used.
  • Describe the structure of a complete evaluation proposal, including its work plan, budget and risk register, its costing, synthesis and reporting plans, and an executive summary for decision-makers.

This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University, drawing on Rossi, P. H., Lipsey, M. W., & Henry, G. T. (2019). Evaluation: A Systematic Approach (8th ed.). SAGE; and Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and Quasi-Experimental Designs for Generalized Causal Inference. Houghton Mifflin.

Lesson 10 · HSCI 826

Economic Evaluation, Reporting and Evaluation Use

A short guided orientation before you work through the final lesson of the course.

Program Planning & Evaluation
Where we are

Completing the evaluation plan

Lessons 1 to 9

The plan describes the program, its theory, its questions, its indicators, its design and its implementation framework.

Lesson 10

The plan adds a costing, an economic evaluation, a method of synthesis and a plan for reporting and use.

Lesson map

Four sections

1. Costing a program

Perspectives, resources and costing methods.

2. Economic evaluation

QALYs, the ICER, the plane and uncertainty.

3. Synthesis and judgement

Rubrics, mixed evidence and TIDieR.

4. Reporting and use

Reports, charts, knowledge translation and use.

Running case

Cedar Valley in this lesson (fictional)

$840,000First-wave annual budget
640Referrals in year one (illustrative)
500Adults attending a first meeting (illustrative)
How to work through it

The final worked example

  • Each section builds on the results of the section before it.
  • A worked example in Section 2 reports scenario analyses and a probabilistic sensitivity analysis for the ICER.
  • Section 4 shows a complete evaluation proposal for Cedar Valley with an executive summary.
Section 1 of 5

Costing a Program

⏱ Estimated reading time: 40 minutes
Section 1 of 5

Costing a Program

What a program really costs, from whose point of view, and how to measure it.

Opportunity cost

Financial and economic cost

Financial cost

The money a program spends, as recorded in budgets and accounts.

Economic cost

The value of all resources a program uses, including donated space, clinician time and volunteer hours.

Perspective

Whose costs count

Program

The program's own budget of $840,000 a year.

Health system

The budget plus overhead, clinician time, clinic space and changes in other service use.

Societal

All of the above plus volunteer, participant and community organization costs.

Three steps

Identify, measure and value

Total cost
\[ \text{Total cost} = \sum_i \, q_i \times p_i \]

Here q is the quantity of each resource and p is its unit cost.

Time-driven activity-based costing

The cost of connector time

$50.00Capacity cost rate per hour
10.5 hConnector time per participant
$525Connector cost per participant
Capacity

Average cost falls as caseloads fill

Year one: 500 participants

The budget of $840,000 gives a gross cost of $1,680 per participant.

Full capacity: 840 participants

The same budget gives a gross cost of $1,000 per participant.

Worked costing

The first year, by perspective

Program

$840,000, or $1,680 per participant.

Health system

$972,000, or $1,944 per participant.

Societal (valued items)

$1,022,000, or $2,044 per participant.

Carry forward

From cost to value

  • The full health system cost is about 16 percent higher than the budget.
  • Time-driven costing shows the cost of direct service and makes unused capacity visible.
  • Section 2 combines the cost of $1,944 per participant with the program's effects.

Learning Objectives for this section

  • Distinguish the financial cost of a program from its economic cost, and explain why a program budget usually understates the resources a program consumes.
  • Explain how the perspective of an analysis (program, health system or societal) determines which costs are counted.
  • Identify, measure and value the resources a program consumes, including in-kind, volunteer and participant resources.
  • Compare gross costing, micro-costing and time-driven activity-based costing, and calculate a cost per participant from connector time data.
  • Distinguish fixed, variable, average and marginal costs, and explain how the use of capacity changes the cost per participant.

This lesson completes the evaluation plan that the course has built since Lesson 1. Decision-makers ask what a program costs and whether the money would do more good elsewhere. This section teaches the first question, and Section 2 uses the answer to teach economic evaluation. Costing also supports budgets for scale-up and documents implementation cost, one of the implementation outcomes of Proctor and colleagues (2011) taught in Lesson 9.

The running example is the fictional Cedar Valley Connector program, a community connector (social prescribing) program for adults aged 65 and older run by the fictional Cedar Valley Health Authority in British Columbia. Its first wave operates in 12 of the region's 24 primary care clinics on an annual budget of $840,000: seven connector positions ($595,000), a coordinator ($105,000), a half-time analyst ($55,000), a transport fund ($40,000), partner grants ($30,000), and training, travel and data systems ($15,000). For this lesson, assume that in the program's first full year clinicians made 640 referrals and 500 older adults attended a first meeting with a connector (312 referrals and 241 first meetings in the first six months, and 328 referrals and 259 first meetings in the second six months). These year-one figures, and the unit costs introduced below, are illustrative assumptions for this lesson.

1.1 Costs as Opportunity Costs

Economic evaluation rests on the idea of opportunity cost: the cost of using a resource for one purpose is the value of the best alternative use that is given up (Drummond et al., 2015). When the health authority pays a connector, the money could instead have funded home-care hours or a nurse in a long-term care home. When a clinic lends a meeting room to the program, the room could have been used for a foot-care clinic. The budget records the first of these choices, because money changes hands, and misses the second, because nothing is paid. An evaluator who wants to know the full cost of a program therefore distinguishes the financial cost, which is the money spent and recorded in accounts, from the economic cost, which is the value of all resources used, whether or not anyone paid for them.

Opportunity costClick to explore
Financial costClick to explore
Economic costClick to explore
In-kind contributionsClick to explore
Sunk costsClick to explore
Transfer paymentsClick to explore

Three related ideas follow. In-kind contributions are counted at their opportunity cost, sunk costs are left out of forward-looking decisions, and transfer payments are left out of societal costs because they move money without using resources. The distinction between financial and economic cost matters most for programs that rely on partners. Social prescribing programs link people to community groups and volunteer roles that the health system does not pay for, so a costing that reports only the budget makes them appear cheaper than they are and hides costs that community organizations may not be able to absorb as referrals grow.

1.2 The Perspective of the Analysis

The perspective of an economic analysis defines whose costs and consequences are counted. The program perspective counts costs to the organization delivering the program. The health system perspective, called the publicly funded health care payer perspective in Canadian guidance, counts all costs to the public health system, including changes in the use of other services. The societal perspective counts all costs, whoever bears them.

The guidelines for economic evaluation published by CADTH (2017), the organization now called Canada's Drug Agency, specify the publicly funded health care payer perspective for the reference case and allow a broader societal perspective as an additional analysis when important costs fall outside the health system. The Second Panel on Cost-Effectiveness in Health and Medicine recommended that analysts report two reference cases, one from the health care sector perspective and one from the societal perspective, and that they complete an impact inventory listing the consequences of an intervention inside and outside the health sector, even those that are not valued (Sanders et al., 2016). For a program such as Cedar Valley, whose resources and effects extend into the community, the impact inventory makes visible what a health-system analysis leaves out.

The cost of the first wave is its budget, $840,000 a year. This perspective answers the program manager's question about how much money the program needs, and it omits corporate overhead, clinician referral time, donated rooms and effects on other services.

The cost adds corporate overhead, clinician screening and referral time and clinic space, together with any change in emergency department and primary care use, which may offset part of the program cost. This is the reference perspective in Canadian guidance.

The cost also includes volunteer time, participants' and caregivers' time and travel, and costs that community organizations and the First Nations health centre bear beyond what the program pays them. This perspective reveals cost-shifting from the health system to the community.

Cost item (Cedar Valley first wave)ProgramHealth systemSocietal
Connector, coordinator and analyst salaries and benefitsYesYesYes
Transport fund, partner grants, training, travel and data systemsYesYesYes
Health authority corporate overheadNoYesYes
Clinician time for screening and referralNoYesYes
Clinic meeting rooms provided in kindNoYesYes
Changes in emergency department and primary care visitsNoYesYes
Volunteer time in community groupsNoNoYes
Participants' and caregivers' time and travelNoNoYes
Costs to community organizations beyond partner grantsNoNoYes

1.3 Identifying, Measuring and Valuing Resources

Drummond and colleagues (2015) describe costing as three steps. The evaluator first identifies the resources the program uses, then measures the quantity of each resource, and then values each quantity with a unit cost. The total cost is the sum, across resources, of the quantity multiplied by the unit cost.

The costing identity

Total cost = Σ (quantity of resourcei × unit cost of resourcei)

For example, 640 referrals that each take clinician time valued at $45 cost 640 × $45 = $28,800.

Identification starts from the inputs and activities in the program's logic model (Lesson 3). The evaluator walks through each activity, from screening in the clinic to the last follow-up call, and asks what staff time, space, equipment, travel and partner resources it uses. The inventory should be agreed with staff and partners, who know about resources that never appear in the accounts, such as the volunteer who drives participants to their first group meeting.

Measurement records quantities in natural units, such as staff hours, referrals, rides funded and visits to other services, from payroll records, activity logs, case management systems, administrative data (Lesson 5) and participant surveys.

Valuation attaches a unit cost to each quantity. Staff time is valued at salary plus benefits per hour of work. Purchased items are valued at their market price. Physician services are often valued with the provincial fee schedule, which in British Columbia is the Medical Services Plan payment schedule, and hospital and emergency care can be valued with national cost estimates from the Canadian Institute for Health Information. Resources without a market price need a shadow price. Volunteer time can be valued at the wage of a paid worker doing the same task (the replacement cost method) or at what the volunteer would otherwise earn, and the choice changes the result. The value of retired participants' time is contested, and many analyses report their hours without a dollar value.

Start-up and recurring costsv

Start-up costs are incurred once, such as hiring, initial training, building the resource directory and co-designing the land-based pathway. Recurring costs are incurred every year. Report them separately, because continuing a program depends on recurring costs and adopting it elsewhere depends on both.

Capital costsv

Equipment that lasts several years, such as laptops or a vehicle, is spread over its useful life as an equivalent annual cost, which accounts for depreciation and for the return the money could have earned.

Overhead and shared costsv

Corporate services such as human resources, finance and information technology are allocated to the program, often as a percentage of direct costs or by staff numbers. State the allocation method, because it can change the result.

Protocol-driven costsv

Costs that occur only because a program is being evaluated, such as research interviews or extra questionnaires, are excluded from the program cost and belong in the evaluation budget.

Price year and inflationv

Costs from different years are expressed in the prices of one price year with an index such as the consumer price index from Statistics Canada, and reports state the currency and price year.

1.4 Gross Costing, Micro-Costing and Time-Driven Activity-Based Costing

Costing methods differ in how finely they measure resource use. Gross costing, also called top-down costing, divides a total expenditure by a measure of output. For the first full year of Cedar Valley, $840,000 divided by 500 participants who attended a first meeting gives a gross cost of $1,680 per participant. Gross costing is quick and uses data that already exist, but it treats every participant as identical and cannot explain why one clinic or one participant costs more than another.

Micro-costing, or bottom-up costing, measures each resource used to deliver the program to each participant and values it separately (Frick, 2009). It is more demanding, and it is the method of choice when a program is new and no unit costs exist, when costs are expected to vary across participants or sites, or when the evaluation needs to show which components drive the total. Micro-costing data can come from time-and-motion studies, activity logs kept by staff, case management records, or structured interviews with staff about how long tasks take.

Time-driven activity-based costing is a form of micro-costing developed in management accounting by Kaplan and Anderson (2004) and applied to health care by Kaplan and Porter (2011). It needs two estimates for each type of staff: the capacity cost rate, which is the cost of supplying that staff capacity divided by the practical capacity in hours, and the time each activity in the process requires. The cost of serving a participant is the time each activity takes multiplied by the capacity cost rate, summed along a process map of the participant's path through the program.

Capacity cost rate for a Cedar Valley connector

Cost of one connector position (salary and benefits) = $595,000 ÷ 7 = $85,000 a year

Paid hours = 37.5 hours × 52 weeks = 1,950 hours a year

Practical capacity, after vacation, statutory holidays, sick leave, training and team meetings = 1,700 hours a year (assumed)

Capacity cost rate = $85,000 ÷ 1,700 hours = $50.00 per hour

The Cedar Valley analyst asked connectors to keep an activity log for four weeks and reviewed case records for a sample of closed cases. Participants attended a mean of four meetings (the first plus three follow-ups) of the six allowed. The table and process map summarize mean connector time per participant.

ActivityMean connector time (hours)Cost at $50.00 per hour
Referral triage and first telephone contact0.5$25
First meeting and co-development of the connection plan1.5$75
Follow-up meetings (mean of 3.0 meetings of 1.0 hour)3.0$150
Linkage work: calls to groups, arranging transport, accompanied first visits2.0$100
Travel to homes and community venues2.0$100
Documentation and case review1.5$75
Total per participant10.5$525
Connector time along one participant's path Triage0.5 h$25 Firstmeeting1.5 h$75 Follow-upmeetings3.0 h$150 Linkagework2.0 h$100 Travel2.0 h$100 Recordsand review1.5 h$75 Total: 10.5 hours × $50.00 per hour = $525 Teal: contact with the participant. Red: supporting time.
A process map for time-driven activity-based costing of the fictional Cedar Valley Connector program. Each activity's time is multiplied by the connector capacity cost rate of $50.00 per hour, and the products are summed along the participant's path.

Direct connector time costs $525 per participant, which is less than a third of the gross cost of $1,680. The difference has three sources. The first is the shared costs that the process map does not include: the coordinator, the analyst, training, travel and data systems. The second is the transport fund and the partner grants, which pay for participants' links to the community and lie outside the connector's own work. The third is unused capacity. Seven connectors supply 7 × 1,700 = 11,900 practical hours a year, and 500 participants at 10.5 hours each used 5,250 hours, or 44.1 percent of that capacity. A time study would show how the remaining hours were spent, for example on partner development, clinic liaison, building the resource directory and supporting referred adults who never attended. Kaplan and Anderson (2004) argued that one advantage of time-driven costing is that it makes unused capacity visible, where a gross cost per participant hides it inside a single average.

1.5 Fixed, Variable, Average and Marginal Costs

Costs behave differently as the number of participants changes. Fixed costs, such as connector and coordinator salaries, the analyst and the data system, do not change with the number of participants in the short run. Variable costs, such as transport vouchers and clinician time for referrals, rise with each participant. Step costs are fixed over a range and then jump, as when a full caseload requires an eighth connector. The average cost is the total cost divided by the number of participants, and the marginal cost is the additional cost of serving one more participant.

When a program has spare capacity, the marginal cost of an extra participant is small, because the connectors are already paid. The average cost then falls as the caseload grows, since the fixed costs are spread over more people. Suppose a connector working at full caseload can serve 120 participants a year, which uses 120 × 10.5 = 1,260 hours and leaves 440 of the 1,700 practical hours for partner development, liaison and other work that does not belong to a single participant. Seven connectors could then serve 840 participants a year, and the gross cost per participant at full capacity would be $840,000 ÷ 840 = $1,000.

Average budget cost per participant, $840,000 a year $0$1,000$2,000$3,000 300400500600700800 Participants attending a first meeting per year Capacity of seven connectors: 840 Year one: 500, $1,680 Full capacity: 840, $1,000
With a fixed budget, the average cost per participant falls as the caseload grows toward the capacity of the seven connectors. Beyond 840 participants a year, an eighth connector would be needed, which is a step cost.

Costs measured in a start-up year overstate the steady-state cost per participant, while an analysis that assumes full capacity overstates efficiency if referrals never reach that level. A careful costing reports the observed cost, states the program's capacity, and presents the cost at a plausible steady-state caseload as a scenario, which Section 2 does for Cedar Valley.

Try it: Cost per participant in year two

Suppose that in year two the first-wave budget is unchanged at $840,000 and 588 older adults attend a first meeting. Calculate the gross cost per participant and the share of the 840-participant capacity in use. Then explain what the marginal cost of a 589th participant would mostly consist of, and at what caseload the health authority would face a step cost. (The gross cost is $840,000 ÷ 588 = $1,428.57, and 588 ÷ 840 = 70 percent of capacity. The marginal cost would consist mainly of variable items such as transport support and the clinician's referral time, since the connectors are already paid. A step cost would arise once demand exceeded about 840 participants a year, when an eighth connector would be needed.)

1.6 The Full Economic Cost of the First Wave

The table below brings the elements of this section together for the first full year of the first wave. The unit costs for overhead, clinician time, rooms and volunteer time are assumptions chosen for teaching, and a real costing would document the source of each one. The health authority applies a corporate overhead rate of 10 percent, clinicians spend time valued at $45 on each screening and referral, each of the 12 clinics provides meeting space valued at $1,600 a year, and community group volunteers contribute about 2,000 hours a year to welcoming and accompanying participants, valued at a replacement wage of $25 an hour.

ResourceCalculationAnnual cost
Program budgetSeven connectors, coordinator, analyst, transport fund, partner grants, training, travel and data systems$840,000
Health authority overhead10% × $840,000$84,000
Clinician screening and referral time640 referrals × $45$28,800
Clinic meeting rooms provided in kind12 clinics × $1,600$19,200
Health system totalPer participant: $972,000 ÷ 500 = $1,944$972,000
Volunteer time in community groups2,000 hours × $25$50,000
Societal total (valued items)Per participant: $1,022,000 ÷ 500 = $2,044$1,022,000

The full health-system cost is about 16 percent higher than the budget ($972,000 ÷ $840,000 = 1.157). Some societal items were identified and not valued: participants' and caregivers' time and travel, the costs that community organizations bear beyond their partner grants, and any costs that the First Nations health centre bears in hosting a connector beyond the funded position. The evaluation plan proposes to measure participants' time and travel with a short questionnaire and to estimate partner costs in discussion with each partner, including the First Nations health centre, in a way that respects the partner's authority over its own information. Items that cannot be valued are listed in the impact inventory so that readers know what the totals omit.

Case: The budget request for the second wave

The executive asks what the second wave will add to the annual budget. This is a question of affordability, which a budget impact analysis answers for a specific budget holder over one to five years (Sullivan et al., 2014), and it is separate from the question of value that Section 2 addresses. The team assumes that the 12 new clinics need seven more connectors ($595,000), a second coordinator ($105,000), and transport, partner grant, training and data funds at first-wave levels ($85,000), while the half-time analyst is shared. The second wave would add $785,000, for an annual budget across 24 clinics of $1,625,000. The team also notes the first-year start-up costs of hiring and training, including the $64,000 implementation support package for the second-wave clinics that Lesson 9 specified, and the new referrals that community organizations in the second-wave areas would receive.

With the cost of the program established from more than one perspective, the evaluation can ask whether the outcomes justify it. Section 2 combines the health-system cost of $1,944 per participant with estimates of the program's effects to calculate an incremental cost-effectiveness ratio and to judge how confident a decision-maker can be in the result.

Reflection

A health authority runs a community paramedicine program in which paramedics visit frail older adults at home. Its annual budget is $650,000: four paramedic positions ($440,000), a supervisor ($120,000), leased vehicles ($60,000), and equipment and supplies ($30,000). Other resources are used but not paid for by the program: dispatch staff in the provincial ambulance service schedule the visits (time valued at $25,000 a year), family physicians spend time valued at $40 per patient preparing care plans, the municipality provides space in a fire hall at no charge (valued at $8,000 a year), and family caregivers spend time at home visits. In the year, 400 patients received at least one visit. Calculate the cost per patient from the program perspective and from the publicly funded health system perspective, state which further items a societal perspective would add, and explain one reason the program-perspective cost could mislead a decision to expand the program to a second region.

Model answer

From the program perspective, the cost is the budget, $650,000 ÷ 400 = $1,625 per patient. The health system perspective adds the dispatch staff time ($25,000), because the ambulance service is publicly funded, and the physicians' care-planning time (400 × $40 = $16,000). The total is $650,000 + $25,000 + $16,000 = $691,000, or $1,727.50 per patient.

A societal perspective would add the fire hall space ($8,000), because the municipality is outside the health system but the space has an opportunity cost, bringing the valued total to $699,000, or $1,747.50 per patient. It would also add family caregivers' time at visits, and patients' own time, which should at least be listed in an impact inventory even if they are not valued.

The program-perspective figure could mislead an expansion decision because it hides resources that a second region may not have. A new region might have no spare dispatch capacity and no free municipal space, so its budget would need to cover $33,000 of resources that the first region receives in kind. A strong answer might also note that the first-year cost includes start-up effects, or that cost per patient depends on whether the paramedics' caseloads are full.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: A clinic lends a meeting room to the Cedar Valley program at no charge. How should a health-system costing treat the room?

Economic evaluation values resources at their opportunity cost, so donated space is included at the value of its best alternative use. Option a describes a financial costing, which omits in-kind resources. Transfer payments (option d) move money without using up resources, whereas the room is a real resource that could have been used for another clinic service.

Question 2: Which perspective would count the time of volunteers in community groups and participants' own travel time?

The societal perspective counts all costs regardless of who bears them, including volunteers' and participants' time. The program and health system perspectives (options a, b and c) count only costs that fall on the program or on the publicly funded health system, and who recruited the volunteers does not change that.

Question 3: A Cedar Valley connector position costs $85,000 a year and provides 1,700 practical hours. A participant uses 10.5 hours of connector time. What is the connector cost per participant?

The capacity cost rate is $85,000 ÷ 1,700 = $50.00 per hour, and 10.5 × $50.00 = $525. Using paid hours (option b) understates the rate, because connectors cannot spend vacation, training and team meeting time with participants. Options c and d are gross costs per participant for the whole budget, which include much more than connector time.

Question 4: Why is the marginal cost of one more Cedar Valley participant in year one much lower than the average cost of $1,680?

In year one the seven connectors used about 44 percent of their practical capacity, so serving one more participant would add mainly variable costs such as transport support and referral time. The average cost spreads fixed costs, such as salaries, over all participants. Option b confuses the issue: the gap arises from fixed costs and spare capacity, whether or not start-up costs are present.
Section 2 of 5

Economic Evaluation

⏱ Estimated reading time: 50 minutes
Section 2 of 5

Economic Evaluation

Comparing costs and consequences, and representing uncertainty.

Types of evaluation

How consequences are measured

Cost-effectiveness

A single natural unit, such as an older adult no longer lonely.

Cost-utility

Quality-adjusted life years, comparable across health areas.

Cost-benefit

All consequences valued in money.

Cost-consequence

Several outcomes reported side by side.

Quality-adjusted life years

QALYs as the area between utility profiles

Base case
\[ \Delta E = \tfrac{1}{2} \times 1 \text{ year} \times 0.03 = 0.015 \text{ QALYs} \]

If the twelve-week difference persisted to twelve months, the gain would be 0.0265 QALYs.

The ICER

Incremental cost-effectiveness ratio

Cedar Valley base case
\[ \text{ICER} = \frac{\Delta C}{\Delta E} = \frac{\$1{,}868}{0.015} = \$124{,}533 \text{ per QALY} \]
The plane

Four quadrants and a threshold

North-east: trade-off

The program costs more and gains health, as in the Cedar Valley base case.

South-east: dominant

The program saves money and gains health.

North-west: dominated

The program costs more and achieves less.

South-west: trade-off

The program saves money and loses health.

Net monetary benefit at $100,000 per QALY is $1,500 minus $1,868, or −$368.

Uncertainty

Scenarios and probabilistic analysis

$124,533Base-case ICER per QALY
$41,617Persistent effects and full caseloads
0.282Probability cost-effective at $100,000 per QALY
Money values and reporting

Cost-benefit analysis, SROI and CHEERS 2022

Social return on investment

Ratios depend on financial proxies and deadweight, and ratios from different analyses cannot be compared.

CHEERS 2022

The 28-item reporting checklist adds items on analysis plans, distributional effects and engagement.

Carry forward

From an ICER to a judgement

  • The base-case ICER is about $125,000 per QALY, and plausible scenarios fall below $50,000.
  • An economic result is one line of evidence among several.
  • Section 3 combines economic and other evidence into a judgement.

Learning Objectives for this section

  • Distinguish partial from full economic evaluations, and describe cost-minimization, cost-effectiveness, cost-utility, cost-benefit and cost-consequence analysis.
  • Explain how quality-adjusted life years combine length and quality of life, and calculate QALYs gained as the area between utility profiles.
  • Calculate an incremental cost-effectiveness ratio and net monetary benefit, and place a result on the cost-effectiveness plane.
  • Describe how scenario analyses, probabilistic sensitivity analysis and cost-effectiveness acceptability curves represent uncertainty.
  • Explain the cautions that apply to social return on investment analyses, and identify the main reporting items in CHEERS 2022.

Section 1 established that the first wave of the fictional Cedar Valley Connector program costs the health system $1,944 per participant in its first full year. A cost on its own cannot tell a decision-maker whether the program is worth funding. Economic evaluation compares the costs and consequences of a program with those of an alternative, and it asks whether the additional health gained justifies the additional resources used.

2.1 Full and Partial Economic Evaluation

Drummond and colleagues (2015) define a full economic evaluation by two features: it examines both costs and consequences, and it compares two or more alternatives. Studies that lack one of these features are partial evaluations. A cost description of one program, an outcome description of one program, or a comparison of the costs of two programs without their outcomes are all partial evaluations, which can be useful steps toward a full one. Full economic evaluations differ in how they measure consequences.

TypeHow consequences are measuredSummary resultWhen it is used
Cost-minimization analysisConsequences shown or assumed to be equivalentDifference in costRarely, because equivalence must be demonstrated
Cost-effectiveness analysisA single natural unit, such as an older adult no longer lonelyCost per unit of effect gainedComparing programs with the same main outcome
Cost-utility analysisQuality-adjusted life years (QALYs)Cost per QALY gainedComparing programs across health areas
Cost-benefit analysisMonetary value of consequencesNet benefit or benefit-cost ratioComparing health and non-health investments
Cost-consequence analysisSeveral outcomes reported side by side without aggregationA table of costs and outcomesComplex programs with many outcomes

Cost-consequence analysis is common in public health, where programs affect several outcomes that do not combine easily into one measure. It leaves the weighing of outcomes to the reader, which is transparent but provides no single decision rule. Section 3 shows how an evaluative rubric can structure that weighing.

2.2 The Incremental Cost-Effectiveness Ratio

Economic evaluation compares a program with what would happen without it, which is usually usual care. The key quantity is the incremental cost-effectiveness ratio (ICER), the difference in cost divided by the difference in effect (Weinstein & Stason, 1977).

Incremental cost-effectiveness ratio

ICER = (C1 − C0) ÷ (E1 − E0) = ΔC ÷ ΔE

Here C1 and E1 are the mean cost and effect per person with the program, and C0 and E0 are the mean cost and effect with the comparator. The ICER is the additional cost of each additional unit of effect.

The ratio is incremental because the decision concerns the additional money and the additional effect. An average ratio, such as total cost divided by the number of people whose loneliness improved, attributes every improvement to the program, including improvements that would have happened anyway.

Both increments come from the evaluation design. For this lesson, assume that a comparison-group evaluation of the first wave (Lesson 7) produced the following illustrative results per participant over twelve months. Emergency department visits were 0.12 lower and primary care visits 0.40 lower than in the comparison group. These figures count participants’ own visits; the Lesson 7 exercise that found a rise in recorded primary care visits across all older patients of first-wave clinics measured a different quantity, which may include visits that connectors arranged or recorded. At assumed unit costs of $450 per emergency visit and $55 per primary care visit, these differences offset 0.12 × $450 + 0.40 × $55 = $54 + $22 = $76 of the program cost. The incremental cost is therefore ΔC = $1,944 − $76 = $1,868 per participant.

Worked example: A cost-effectiveness analysis in natural units

At twelve weeks, 34 percent of program participants and 22 percent of comparison participants scored below the referral threshold of 6 on the three-item UCLA Loneliness Scale (illustrative values). The incremental effect is 0.34 − 0.22 = 0.12, or 12 additional older adults no longer screening as lonely for every 100 participants. For 100 participants, the incremental cost is 100 × $1,868 = $186,800, so the ICER is $186,800 ÷ 12 = $15,567 per additional older adult no longer screening as lonely.

This result is easy to explain, but a decision-maker cannot tell from it whether $15,567 is good value, because no other program reports its results in the same unit and the measure ignores effects on health beyond loneliness.

2.3 Cost-Utility Analysis and Quality-Adjusted Life Years

Cost-utility analysis addresses the problem of incomparable units by measuring consequences in a common unit, the quality-adjusted life year (QALY). A QALY weights each period of life by a utility value that represents health-related quality of life on a scale on which 1 is full health and 0 is a state equivalent to death. States judged worse than death have negative values. One year in full health is one QALY, and one year at a utility of 0.70 is 0.70 QALYs (Weinstein et al., 2009).

Utilities are elicited from people's preferences between health states. In the standard gamble, a respondent chooses between a certain health state and a gamble between full health and death. In the time trade-off, the respondent chooses between a longer life in the health state and a shorter life in full health. Torrance and colleagues developed the time trade-off at McMaster University in the early 1970s, and Torrance (1986) reviewed the methods. Most evaluations use a generic instrument with a preference-based scoring algorithm, such as the EQ-5D-5L, which has a Canadian value set derived from the time trade-off (Xie et al., 2016), the Health Utilities Index, or the SF-6D.

QALYs gained

QALYs = Σ (utility in each period × length of the period in years)

QALYs gained by a program are the area between the utility profiles of the program and comparison groups over the time horizon. With measurements at a few time points, the area is usually computed with the trapezoid rule, assuming that utility changes in a straight line between measurements.

Suppose Cedar Valley participants complete the EQ-5D-5L at intake, twelve weeks and twelve months. In the illustrative analysis, the groups have the same mean utility at intake, the program group's mean utility is 0.03 higher at twelve weeks, and the difference has returned to zero by twelve months. The area between the profiles is a triangle with a base of one year and a height of 0.03, so the incremental QALYs are ΔE = 0.5 × 1 × 0.03 = 0.015 per participant. If the difference of 0.03 were instead sustained to twelve months, the area would be 0.5 × 0.03 × (12 ÷ 52) + 0.03 × (40 ÷ 52) = 0.0035 + 0.0231 = 0.0265 QALYs.

00.010.020.030.04 01252Weeks since intake Utility difference Base case: area = 0.015 QALYs If the effect persists: 0.0265 QALYs
Incremental QALYs as the area between utility profiles for the illustrative Cedar Valley analysis. The teal triangle is the base case, in which the difference fades by twelve months, and the shaded red area is added if the twelve-week difference persists.

The base-case ICER is ΔC ÷ ΔE = $1,868 ÷ 0.015 = $124,533 per QALY gained. The figure shows why the assumption about persistence matters: no one measured utility between twelve weeks and twelve months, so the shape of the profile in that interval is a modelling choice that Section 2.5 tests.

QALYs have known limits for programs such as Cedar Valley. Generic instruments such as the EQ-5D-5L describe mobility, self-care, usual activities, pain and anxiety or depression, and they may miss changes in social connection that participants value. Capability measures such as the ICECAP-O, developed for older people (Coast et al., 2008), assess attachment, security, role, enjoyment and control, and some evaluations of social programs report them alongside QALYs. QALYs also raise equity questions, since a QALY is valued equally whoever gains it, and the Cedar Valley evaluation reports results for subgroups so that readers can see how gains and costs are distributed.

2.4 The Cost-Effectiveness Plane and Net Monetary Benefit

The cost-effectiveness plane plots the incremental effect on the horizontal axis and the incremental cost on the vertical axis, with the comparator at the origin (Black, 1990). Its four quadrants classify results.

North-east: more costly, more effectiveClick to explore
South-east: less costly, more effectiveClick to explore
North-west: more costly, less effectiveClick to explore
South-west: less costly, less effectiveClick to explore

A line through the origin with a slope equal to the decision-maker's willingness to pay per QALY, written λ (lambda), divides the north-east quadrant. Results below the line are cost-effective at that threshold. Canada has no official threshold. The CADTH (2017) guidelines ask analysts to show results across a range of willingness-to-pay values, and Canadian studies often report results against reference values such as $50,000 and $100,000 per QALY. The figure places the illustrative Cedar Valley results on the plane.

$50,000 per QALY $100,000 per QALY Base case $124,533 per QALY Optimistic scenario $41,617 per QALY North-west: dominated North-east: trade-off South-east: dominant South-west: trade-off −0.02−0.0100.010.020.030.04 Incremental QALYs per participant −$1,000$0$1,000$2,000$3,000
The illustrative Cedar Valley base case ($1,868 for 0.015 QALYs) lies above both threshold lines. The optimistic scenario, with persistent effects and full caseloads ($1,104 for 0.0265 QALYs), lies below the $50,000 line.

Ratios are awkward to analyze, because the same ICER can arise in the north-east and south-west quadrants and because a ratio becomes unstable when ΔE is near zero. Net monetary benefit avoids these problems by converting health gains into money at the threshold value (Stinnett & Mullahy, 1998).

Net monetary benefit

NMB = λ × ΔE − ΔC

At λ = $50,000, the Cedar Valley base case gives NMB = $50,000 × 0.015 − $1,868 = $750 − $1,868 = −$1,118. At λ = $100,000, NMB = $1,500 − $1,868 = −$368. A program is cost-effective at a threshold when its NMB is positive, which is the same as an ICER below λ in the north-east quadrant.

2.5 Representing Uncertainty

Every input to the Cedar Valley analysis is uncertain. Deterministic sensitivity analysis changes one input, or a set of inputs that form a scenario, and recalculates the result. Probabilistic sensitivity analysis assigns each uncertain parameter a probability distribution, draws a value for every parameter many times, and recalculates the result for each draw, so that the spread of results reflects the joint uncertainty in all parameters (Briggs et al., 2006). The share of draws with a positive NMB at each threshold is plotted as a cost-effectiveness acceptability curve (van Hout et al., 1994). Probabilistic analysis handles uncertainty in parameter values. Uncertainty about structure, such as whether effects persist, is better shown with scenarios, because the answer changes the model itself.

Analyses with time horizons longer than one year also discount future costs and QALYs to present values, and the CADTH (2017) guidelines specify a rate of 1.5 percent a year for the reference case. The Cedar Valley analysis uses a one-year horizon, so it needs no discounting.

Worked example: The Cedar Valley ICER, scenarios and a probabilistic sensitivity analysis

Scenario analyses. The base case uses the inputs from Section 1 and this section. The QALY gain of 0.015 is the area between the utility curves, computed with the trapezoid rule from a utility difference that rises to 0.03 at twelve weeks and returns to zero at twelve months. With an incremental cost of $1,868, the ICER is $124,533 per QALY, and the net monetary benefit is −$1,118 at $50,000 per QALY and −$368 at $100,000 per QALY. Five scenarios change one assumption, or a pair of assumptions, and recalculate the result.

ScenarioIncremental costIncremental QALYsICER (dollars per QALY)
Base case$1,8680.0150$124,533
Effect persists to 52 weeks$1,8680.0265$70,388
Full caseloads (840 a year)$1,1040.0150$73,630
No health care offsets$1,9440.0150$129,600
Societal perspective$1,9680.0150$131,200
Persistence and full caseloads$1,1040.0265$41,617

The scenarios show which assumptions matter. Removing the health care offsets or adding valued volunteer time changes the ICER only modestly ($129,600 and $131,200 per QALY). Persistence of the effect ($70,388) and full caseloads ($73,630, with an incremental cost of $1,104) each bring the ICER below $100,000, and together they bring it to $41,617, below $50,000. The program's value depends mainly on two things the first year cannot show, whether benefits last and whether caseloads fill.

Probabilistic sensitivity analysis. The probabilistic analysis draws 5,000 sets of inputs: the cost per participant from a gamma distribution (mean $1,944, standard deviation $120), the differences in emergency and primary care visits from normal distributions, and the twelve-week utility difference from a normal distribution with mean 0.03 and standard deviation 0.012. Across the draws, the mean incremental cost is $1,865.60 and the mean QALY gain is 0.0151. The table gives the share of draws with a positive net monetary benefit at each willingness to pay, which is the cost-effectiveness acceptability curve.

Willingness to pay per QALYProbability cost-effective
$25,0000.000
$50,0000.000
$75,0000.062
$100,0000.282
$125,0000.511
$150,0000.665
$200,0000.830

Almost all draws (99.4 percent) lie in the north-east quadrant, and 0.6 percent lie in the north-west, where the program is dominated. The probability that the program is cost-effective is 0.000 at $50,000 per QALY, 0.282 at $100,000 and 0.511 at $125,000, close to the base-case ICER. Under base-case assumptions, a decision-maker willing to pay $100,000 per QALY would face about a 28 percent chance that the program is good value, and the scenario analyses show that this chance would rise considerably if effects persist or caseloads fill.

2.6 Cost-Benefit Analysis and Social Return on Investment

Cost-benefit analysis values all consequences in money and reports the net benefit (benefits minus costs) or the benefit-cost ratio. Health effects can be valued by willingness to pay, elicited through contingent valuation or discrete choice experiments, or with a monetary value per QALY. If the Cedar Valley health gain is valued at $100,000 per QALY, the benefits per participant are 0.015 × $100,000 + $76 of averted health care costs = $1,576. Against the gross health-system cost of $1,944, the net benefit is $1,576 − $1,944 = −$368 and the benefit-cost ratio is $1,576 ÷ $1,944 = 0.81. The net benefit equals the NMB at λ = $100,000, which shows that a cost-utility analysis with a threshold is a restricted form of cost-benefit analysis.

Social return on investment (SROI) is a form of cost-benefit analysis developed in the social enterprise sector and set out in a guide first published by the United Kingdom Cabinet Office in 2009 and revised by the SROI Network in 2012 (Nicholls et al., 2012). It maps outcomes with interest holders, assigns financial proxies to each outcome, and adjusts for deadweight (what would have happened anyway), attribution (the share due to others), displacement and drop-off. The result is a ratio of social value to investment. SROI has appeal for community programs because it involves interest holders in deciding which outcomes count. A systematic review of SROI in public health found wide variation in methods and in the quality of the proxies and adjustments used (Banke-Thomas et al., 2015).

Worked example: How assumptions drive an SROI ratio

Suppose a hypothetical SROI of Cedar Valley assigns a financial proxy of $4,000 to each participant whose loneliness score falls by at least one point, and 300 of the 500 participants meet that criterion. The claimed social value is 300 × $4,000 = $1,200,000, and the ratio to the $840,000 budget is 1.43 to 1. If deadweight is set at 50 percent, because the comparison group suggests that half of the improvement would have happened anyway, the value falls to $600,000 and the ratio to 0.71 to 1. Using the full health-system cost of $972,000 as the investment, the ratio falls further to 0.62 to 1. The proxy itself is a judgement that no data in the evaluation can confirm.

Financial proxiesv

Proxies are often drawn from value banks or from unrelated contexts, and a ratio is only as credible as its least defensible proxy. Reports should give the source of each proxy and test alternatives.

Counterfactual and deadweightv

Many SROI analyses estimate deadweight from general statistics or judgement instead of a comparison group, which can inflate the claimed value. The designs of Lessons 6 to 8 provide a stronger basis for the adjustment.

Double countingv

Counting both a reduction in loneliness and the improved wellbeing that follows from it values the same change twice. Outcome maps should separate final outcomes from intermediate ones.

Comparability of ratiosv

Because methods vary so widely, a ratio from one SROI cannot be compared with a ratio from another, and a ratio above one does not show that a program is a better use of money than alternatives.

2.7 Reporting with CHEERS 2022

The Consolidated Health Economic Evaluation Reporting Standards 2022 (CHEERS 2022) replaced the 2013 statement with a 28-item checklist for reporting any health economic evaluation (Husereau et al., 2022). Items cover the title and abstract, the setting and comparators, the perspective, time horizon and discount rate, the selection, measurement and valuation of outcomes and resources, the currency and price date, the rationale and description of any model, and the characterization of uncertainty. CHEERS 2022 added items asking authors to state whether a health economic analysis plan was prepared, to describe how distributional effects were examined, and to report how patients, the public and others affected by the study were engaged and what difference that engagement made. Distributional cost-effectiveness analysis, which examines how costs and health gains fall across social groups, is one way to address the distributional item (Cookson et al., 2017).

Try it: Check the Cedar Valley analysis against CHEERS 2022

For each of the following CHEERS topics, write one sentence stating what the illustrative Cedar Valley analysis reports or omits: perspective, comparator, time horizon, discount rate, outcome measure and its valuation, currency and price year, characterization of uncertainty, distributional effects, and engagement with older adults and First Nations partners. (The analysis reports a health-system perspective with a societal scenario, usual care in comparison clinics, a one-year horizon with no discounting, EQ-5D-5L utilities with the Canadian value set, scenario and probabilistic analyses, and subgroup results. It has not yet stated a price year, and it should describe how the steering committee's older adults and First Nations representatives shaped the outcomes and the interpretation.)

An ICER, a probability of cost-effectiveness and an SROI ratio are each one line of evidence about a program. Section 3 turns to the task of combining such evidence with evidence on reach, cultural safety and effectiveness into an overall judgement.

Reflection

A health authority is considering a home-based exercise program to prevent falls among adults aged 75 and older. Compared with usual care, the program costs $900 per participant and reduces hospital admissions for falls enough to save $300 per participant over one year. The program gains 0.008 quality-adjusted life years (QALYs) per participant over the same year. A probabilistic sensitivity analysis found that the probability the program is cost-effective is 0.35 at a willingness to pay of $50,000 per QALY and 0.70 at $100,000 per QALY. A community agency has separately published a social return on investment (SROI) analysis claiming $5 of social value for every $1 invested, without a comparison group. Calculate the incremental cost, the incremental cost-effectiveness ratio (ICER) and the net monetary benefit at $50,000 and at $100,000 per QALY, state where the result lies on the cost-effectiveness plane, interpret the probabilistic results, and explain how you would treat the SROI claim in advice to the health authority.

Model answer

The incremental cost is $900 − $300 = $600 per participant, and the ICER is $600 ÷ 0.008 = $75,000 per QALY gained. The result lies in the north-east quadrant, because the program costs more and gains health. The net monetary benefit is $50,000 × 0.008 − $600 = $400 − $600 = −$200 at $50,000 per QALY, and $800 − $600 = $200 at $100,000 per QALY. The program is therefore cost-effective at $100,000 per QALY and not at $50,000.

The probabilistic results agree with this. At $50,000 per QALY only 35 percent of simulations show a positive net benefit, while at $100,000 the figure is 70 percent, so a decision-maker willing to pay $100,000 would face about a 30 percent chance that the program is poor value. I would present these probabilities together with the scenario analyses that drive them.

I would not treat the SROI ratio as comparable evidence. Without a comparison group, its deadweight adjustment is a judgement, and the ratio depends on financial proxies whose sources should be checked. I would ask for the proxies, the deadweight and attribution assumptions and a test of alternatives, and I would note that a ratio above one does not show that the program is better value than other uses of the money.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: What distinguishes a full economic evaluation from a partial one?

Drummond and colleagues define a full economic evaluation as one that examines both costs and consequences for two or more alternatives. QALYs (option b) define cost-utility analysis, which is one type of full evaluation, and neither probabilistic modelling nor a societal perspective is required.

Question 2: Compared with usual care, a program costs $1,868 more per participant and gains 0.015 QALYs. What is the ICER, and where does the result lie on the cost-effectiveness plane?

The ICER is $1,868 ÷ 0.015 = $124,533 per QALY. The program costs more and gains health, which places it in the north-east quadrant. The south-east quadrant (option b) is for programs that save money while gaining health, and options c and d contain arithmetic errors.

Question 3: At a willingness to pay of $100,000 per QALY, a program has an incremental cost of $1,868 and an incremental effect of 0.015 QALYs. What is its net monetary benefit, and what does it imply?

Net monetary benefit is $100,000 × 0.015 − $1,868 = $1,500 − $1,868 = −$368, so the program is not cost-effective at that threshold. Option a omits the cost, option c uses a threshold of $50,000, and a negative net benefit in the north-east quadrant does not mean the program is dominated.

Question 4: Which statement about social return on investment (SROI) is most accurate?

SROI results depend on financial proxies and on adjustments for deadweight, attribution, displacement and drop-off, and reviews have found wide variation in how these are done. A ratio above one (option a) says nothing about alternatives, and SROI ratios from different analyses cannot be compared. SROI uses proxies precisely because many outcomes have no market price.
Section 3 of 5

Synthesis and Judgement

⏱ Estimated reading time: 40 minutes
Section 3 of 5

Synthesis and Judgement

Combining mixed evidence into transparent judgements of merit and worth.

Evaluative conclusions

Merit and worth

Merit

The intrinsic quality of a program, judged on criteria such as effectiveness and delivery.

Worth

The value of a program in its context, which brings in cost, need and alternatives.

Rubrics

The Cedar Valley year-one rubric

ReachEquity of reachCultural safety (hurdle)EffectivenessValue for money

Each criterion has four levels (excellent, good, adequate and poor), descriptors, evidence sources and an importance rating.

Synthesis approaches

Combining ratings across criteria

Qualitative weight and sum

Ratings are compared within importance categories, with the reasoning shown.

Numerical weight and sum

Strong ratings can offset a harmful one, and the weights imply false precision.

Hurdles

A minimum level on one criterion is required before others can raise the rating.

Profiles

Ratings are reported by criterion when audiences weigh criteria differently.

Mixed evidence

Convergence, complementarity and dissonance

Convergence

Sources point to the same conclusion.

Complementarity

Sources describe different aspects of the same change.

Dissonance

Sources disagree, and the disagreement is investigated.

Worked synthesis

Good merit, adequate worth

Merit: good

Effectiveness is good, delivery is culturally safe, and reach is good.

Worth at year-one caseloads: adequate

Equity of reach and value for money are adequate, with low confidence in value for money.

TIDieR

Describing the program as delivered

NameWhyMaterialsProceduresWho providedHowWhereWhen and how muchTailoringModificationsHow well: plannedHow well: actual
Carry forward

From judgement to use

  • A rubric agreed in advance makes the basis of a judgement transparent.
  • Dissonant evidence is investigated and explained.
  • Section 4 asks how judgements reach the people who will use them.

Learning Objectives for this section

  • Distinguish descriptive findings from evaluative conclusions about the merit and worth of a program.
  • Construct an analytic rubric with criteria, performance levels, descriptors, evidence sources and importance for a program.
  • Compare qualitative weight-and-sum, numerical weight-and-sum, hurdle and profile approaches to synthesis, and explain the risks of numerical weighting.
  • Synthesize convergent, complementary and dissonant evidence into a judgement with a stated level of confidence.
  • Describe a program with the TIDieR checklist so that others can understand, cost and replicate it.

Sections 1 and 2 produced a cost and an economic result for the fictional Cedar Valley Connector program. The evaluation has also produced evidence on reach, cultural safety and changes in loneliness. Decision-makers asked a single evaluative question, whether the program is good enough and good value enough to extend to the second wave, and answering it requires a judgement that brings these lines of evidence together. Lesson 1 introduced evaluative reasoning as a movement from criteria to standards to evidence to synthesis. This section develops the tools that make that synthesis explicit.

3.1 From Findings to Evaluative Conclusions

A descriptive finding states what happened: 78.1 percent of referred adults attended a first meeting. An evaluative conclusion states how good that is: reach was good against the standard the steering committee agreed in advance. Davidson (2005) argues that evaluation is distinguished from other applied research by its obligation to draw conclusions of this second kind, and that an evaluation which reports only descriptive findings leaves the hardest part of the work to readers who have less information than the evaluator.

Scriven's distinctions from Lesson 1 shape the conclusions an evaluation can reach. Merit is the intrinsic quality of a program, judged against criteria such as effectiveness and quality of delivery. Worth is its value in a particular context, which brings in cost, need and the alternatives available to the decision-maker. A program can have good merit and modest worth, as when it works well but costs more than an equally useful alternative. Keeping the two separate lets an evaluation report that the Cedar Valley program delivers what it promises while being candid about its value for money at current caseloads.

Synthesis happens at two levels. Within a criterion, several pieces of evidence are combined into a rating, as when a difference-in-differences estimate, participant interviews and connector records together determine the rating for effectiveness. Across criteria, the ratings are combined into an overall judgement. Each level needs a stated rule, and the rules should be agreed with the primary intended users (Lesson 4) before the evidence arrives.

3.2 Building an Evaluative Rubric

An evaluative rubric sets out the criteria in rows and describes what performance at each level looks like for each criterion. King and colleagues (2013) describe rubrics as a method for surfacing the values of interest holders and for making the basis of a judgement transparent, and Oakden (2013) gives a practical account of developing them with program partners. An analytic rubric rates each criterion separately, while a holistic rubric describes overall levels of performance in a single set of descriptors. Analytic rubrics are more common in program evaluation because they show where a program is strong and where it is weak.

CriteriaClick to explore
Performance levelsClick to explore
DescriptorsClick to explore
Evidence sourcesClick to explore
Importance and hurdlesClick to explore

Rubrics are developed in steps. The evaluator drafts criteria from the logic model and the evaluation questions, holds a workshop in which interest holders revise the criteria and draft the descriptors, specifies the evidence for each criterion, agrees the importance of each criterion and any hurdles, and tests the rubric on hypothetical results to check that the descriptors discriminate. For Cedar Valley, the steering committee, including its four older adults with lived experience and its two First Nations representatives, built on the six-month rubric of Lesson 1 to agree the year-one rubric below. The value-for-money criterion follows King (2017), who argues that economic results should enter an evaluation as evidence rated against agreed standards, alongside the other criteria.

Criterion and importanceExcellentGoodAdequatePoor
Reach: percentage of referred adults who attend a first meeting (very important)80 percent or more65 to 79 percent50 to 64 percentBelow 50 percent
Equity of reach across sex, language, rurality and Indigenous identity (very important)All groups within 5 percentage pointsAll groups within 10 pointsGaps over 10 points, with an agreed plan to close themGaps over 10 points, with no plan
Cultural safety, as judged by Indigenous participants and partners (hurdle and very important)Safe and respectful, with no unresolved concernsGenerally safe, with concerns resolved promptlyMixed reports, with a plan in placeDisrespect or harm without a response
Effectiveness: program-attributable reduction in mean loneliness at twelve weeks (extremely important)0.5 points or more, with a confidence interval excluding zero0.3 to 0.49 points, with a confidence interval excluding zero0.1 to 0.29 points, or a larger estimate with an interval including zeroBelow 0.1 points
Value for money (very important)Dominant, or an ICER below $50,000 per QALY with a probability of at least 0.7 of being cost-effective at that valueBase-case ICER below $100,000 per QALYBase-case ICER above $100,000, with plausible scenarios below $100,000Dominated, or an ICER above $100,000 in all plausible scenarios

The committee agreed two synthesis rules. Cultural safety is a hurdle: the program cannot be rated better than adequate overall if cultural safety is rated poor. The overall judgement of merit cannot be higher than the rating for effectiveness, because the program exists to reduce loneliness. The committee also agreed to report merit and worth separately, with value for money informing worth.

3.3 Approaches to Synthesis Across Criteria

Several approaches exist for combining criterion ratings into an overall judgement, and they can produce different conclusions from the same ratings. Scriven (1991) and Davidson (2005) recommend qualitative approaches for most program evaluations.

Each criterion is assigned an importance category, such as extremely important, very important or important, and each is rated on the rubric. The evaluator then compares the pattern of ratings within each importance category, giving most attention to the most important criteria, and states the overall judgement in words with the reasoning shown. The method keeps the reasoning visible and avoids arithmetic on ordinal ratings (Scriven, 1991; Davidson, 2005).

Each rating is converted to a number (for example, excellent 4, good 3, adequate 2 and poor 1), multiplied by a weight, and summed. With weights of 0.30 for effectiveness, 0.20 each for cultural safety and value for money, and 0.15 each for reach and equity, a hypothetical program rated excellent on every criterion except a poor rating for cultural safety scores 0.30 × 4 + 0.15 × 4 + 0.15 × 4 + 0.20 × 1 + 0.20 × 4 = 3.40, which reads as better than good. The method lets strong ratings compensate for a harmful one, treats ordinal ratings as if they were measurements, and gives an impression of precision that the weights do not support.

A hurdle, or bar, is a minimum level that a program must reach on a criterion before other criteria can raise the overall rating. Cedar Valley's cultural safety hurdle records the committee's view that a program which harms some participants cannot be rescued by good results for others. Hurdles are combined with one of the other approaches.

Some evaluations report a profile of ratings without an overall judgement. A profile is appropriate when audiences legitimately weigh the criteria differently, for example when a ministry, a health authority and a community partner will each make their own decision. The evaluator then explains the implications of the profile for each audience.

Two levels of synthesis Evidence Outcome estimatesInterviews andtalking circlesProgram recordsEconomic results Within criteria One rating percriterion, fromthe rubric, witha confidencelevel Hurdle check Is culturalsafety ratedabove poor?If not, overallis capped Across criteria Merit: qualityand effectivenessWorth: merit incontext and cost Rules for each step are agreed with primary intended users before the evidence arrives.
The synthesis used for the fictional Cedar Valley evaluation, in which evidence is first combined within each criterion and the ratings are then combined across criteria into separate judgements of merit and worth.

3.4 Synthesizing Mixed Evidence

Within a criterion, evidence from different sources can relate in three ways (Greene, 2007). It can converge, when sources point to the same conclusion. It can be complementary, when sources describe different aspects of the same phenomenon, as when a survey measures how much loneliness changed and interviews describe how. It can be dissonant, when sources disagree. Lesson 5 introduced joint displays, which set quantitative and qualitative results side by side for each question and make these relationships visible.

Dissonance deserves investigation. The evaluator asks whether the sources measure the same thing, cover the same people and time, and are equally credible, and then uses the program theory of Lesson 3 to ask whether a mechanism might operate in some contexts and not others. Averaging dissonant evidence into a middle rating hides the most useful information the evaluation has produced.

Every rating should carry a statement of confidence. Confidence depends on the strength of the design, the consistency of evidence across sources, the precision of estimates, how directly the evidence addresses the criterion, and the risk of bias in each source. This reasoning resembles the GRADE approach to the certainty of evidence (Guyatt et al., 2008), applied here to one program. Many evaluations rate confidence as high, moderate or low and give the main reason for each rating.

Worked example: The year-one synthesis for Cedar Valley

The following year-one results are illustrative. Reach was 500 ÷ 640 = 78.1 percent, which is good, with high confidence because it comes from complete program records. Equity of reach showed attendance of 64 percent among referred adults whose first language is neither English nor French, compared with 80 percent among English or French speakers, and 71 percent in rural areas compared with 80 percent in urban areas. The gaps exceed 10 points, and the committee has agreed a plan of interpreter-supported first meetings and a rural transport pilot, so the rating is adequate, with high confidence.

Cultural safety was rated good: talking circles and the review with First Nations partners described the program as generally safe, and a concern about meeting locations was resolved within a month. Confidence is moderate, because the talking circles reached a small number of participants.

Effectiveness drew on three sources. The difference-in-differences estimate was a 0.4-point reduction in mean loneliness at twelve weeks (95 percent confidence interval 0.1 to 0.7), which is good. Interviews converged with this result for most participants, who described new routines and group memberships, and connector records complemented it by showing which linkages were made. The evidence was dissonant for rural participants, who were satisfied with their connectors and yet showed little change in loneliness. Program records explained the dissonance: the transport fund was exhausted in the ninth month, and the groups to which rural participants were linked were often a long drive away, so the mechanism of attending community groups was blocked by context. Effectiveness is rated good with moderate confidence, because the design is non-randomized, although trends in emergency and primary care visits were parallel before launch and the sources mostly agree.

Value for money was rated adequate: the base-case ICER is $124,533 per QALY, plausible scenarios range from $41,617 to $131,200, and the probability of cost-effectiveness at $100,000 per QALY is 0.282. Confidence is low, because persistence of the effect beyond twelve weeks has not been measured.

The cultural safety hurdle is passed. The extremely important criterion, effectiveness, is good, and so the committee rated the program's merit as good. Among the very important criteria, reach and cultural safety are good and equity of reach and value for money are adequate, so the committee rated the program's worth at year-one caseloads as adequate. The recommendation to the executive was to proceed with the second wave on three conditions: fill caseloads before adding positions, replenish the transport fund and pilot rural transport, and measure loneliness and utility at twelve months to test whether effects persist.

Moving the standards after the results arrivev

Revising descriptors once results are known, so that a disappointing result earns a better rating, undermines the credibility of the whole evaluation. If a standard proves unrealistic, the change and its reason should be reported.

Letting the most precise number dominatev

An ICER or an effect estimate looks more authoritative than interview evidence, but precision is a different property from relevance. A precise estimate of a narrow outcome should not override strong evidence on a criterion it does not address.

Treating absence of evidence as poor performancev

When a criterion cannot be rated because data are missing, the rubric should record it as not rated, with the reason, instead of assigning the lowest level.

Ignoring who the evidence describesv

A good average can conceal poor results for a subgroup. Ratings for equity criteria, and subgroup results within other criteria, protect against this.

3.5 Describing the Intervention with TIDieR

A judgement applies to a program as it was actually delivered, and readers can use the judgement only if they know what that program was. Glasziou and colleagues (2008) found that descriptions of interventions in published trials and reviews were often too incomplete for the intervention to be replicated. The Template for Intervention Description and Replication (TIDieR) is a 12-item checklist for describing interventions in reports and protocols (Hoffmann et al., 2014), and TIDieR-PHP adapts it for population health and policy interventions (Campbell et al., 2018). A full description also serves the other parts of this lesson: costing depends on the dose and the staff involved, and an adopter needs both the description and the cost to judge whether the program could work in their setting.

TIDieR itemCedar Valley Connector program (fictional)
1. Brief nameCedar Valley Connector program, a community connector (social prescribing) program for older adults.
2. WhyLoneliness and social isolation harm health, and many isolated older adults face barriers to joining community activities that one-to-one support and practical help can reduce.
3. What: materialsScreening card with the three-item UCLA Loneliness Scale, connection plan template, community resource directory, transport fund and partner grants.
4. What: proceduresClinician screening and referral, triage call, first meeting with a co-developed plan, linkage to groups and services, follow-up meetings and closure.
5. Who providedSeven community connectors, one hosted by a First Nations health centre, supported by a coordinator and trained in person-centred planning and cultural safety.
6. HowFace to face, one to one, with telephone follow-up.
7. WhereTwelve first-wave primary care clinics, participants' homes, community venues and the First Nations health centre.
8. When and how muchUp to six meetings over twelve weeks; a mean of four meetings and 10.5 hours of connector time per participant in year one.
9. TailoringEach plan is co-developed with the participant, and a land-based connection pathway is being co-designed with one First Nation.
10. ModificationsChanges during delivery are documented with the FRAME approach described in Lesson 9.
11. How well: plannedFidelity is monitored with a checklist of core components and monthly case review.
12. How well: actualReported from the process evaluation: reach, meetings attended, linkages made and fidelity to core components.
Try it: Find the gaps in a program description

A report describes a program as follows: "Participants received peer support sessions from trained volunteers at community centres." Using the 12 TIDieR items, list the information a reader would need in order to cost and replicate the program. (A strong answer notes the missing rationale, the materials and procedures of a session, the training and background of the volunteers, whether sessions were individual or group and in person or remote, the number, length and frequency of sessions, tailoring, modifications, and planned and actual fidelity.)

A judgement of merit and worth, a statement of confidence and a clear description of the program are the core content of an evaluation report. Section 4 asks how that content should be written, shown and shared so that decision-makers and communities can use it.

Reflection

A peer support program for new parents was evaluated against a rubric agreed with its advisory group. Criteria, importance and levels were as follows. Effectiveness (extremely important): reduction in mean Edinburgh Postnatal Depression Scale score compared with a comparison group, rated excellent at 2.0 points or more, good at 1.0 to 1.9, adequate at 0.5 to 0.9 and poor below 0.5. Reach among low-income parents (very important): share of participants with low income, rated excellent at 60 percent or more, good at 50 to 59 percent, adequate at 35 to 49 percent and poor below 35 percent. Safety (hurdle and very important): good if all safeguarding incidents were handled promptly under the protocol, and poor if any was not; a poor rating caps the overall rating at adequate. Value for money (very important): excellent if the ICER is below $50,000 per QALY with a probability of at least 0.7 of being cost-effective at that value, good if below $100,000, adequate if above $100,000 with plausible scenarios below it, and poor otherwise. The group also agreed that overall merit cannot be rated higher than effectiveness. Evidence: a non-randomized comparison with similar baseline scores found a reduction of 1.5 points (95 percent confidence interval 0.6 to 2.4); 45 percent of participants had low income; one safeguarding incident occurred and was handled under the protocol within a week; the ICER was $38,000 per QALY with a probability of 0.75 of being cost-effective at $50,000. In interviews with 30 parents, most described feeling less alone, but parents who joined more than three months after birth described little benefit. Rate each criterion with a level of confidence, state overall judgements of merit and worth, and explain how you would handle the interview evidence from late joiners.

Model answer

Effectiveness is good (1.5 points, with an interval that excludes zero), with moderate confidence because the comparison is non-randomized although baseline scores were similar and the interviews mostly converge. Reach among low-income parents is adequate (45 percent), with high confidence because it comes from program records. Safety is good, so the hurdle is passed, with moderate confidence because rare incidents may go unreported. Value for money is excellent ($38,000 per QALY with a probability of 0.75 at $50,000), with moderate confidence because it inherits the uncertainty of the effect estimate.

Merit is good, the highest rating the effectiveness rule allows. Worth is good: the program is good value and safe, and its main weakness is reach among the parents who may need it most.

The late-joiner evidence is dissonant with the overall effect, and I would investigate it before combining it with the other evidence. I would check whether late joiners differ in baseline scores or circumstances, and whether the program theory, in which peer support reduces isolation during the early weeks of parenthood, implies that timing matters. If program records show smaller score changes among late joiners, I would report the subgroup pattern and recommend earlier referral, for example at the postnatal discharge visit, together with outreach to low-income parents.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: In Scriven's terms, what is the difference between the merit and the worth of a program?

Merit is the intrinsic quality of a program, judged on criteria such as effectiveness and quality of delivery, while worth is its value in a particular context, which brings in cost, need and alternatives. The distinction does not depend on who judges (option a) or on the type of evidence (option c).

Question 2: A rubric converts ratings to numbers and sums weighted scores across criteria. What is the main risk of this approach?

Numerical weight and sum lets strong ratings compensate for a poor one, treats ordinal ratings as measurements, and implies more precision than the weights support. Agreeing criteria in advance (option c) is good practice for any rubric, and the method can be applied to ratings based on qualitative evidence.

Question 3: Rural Cedar Valley participants were satisfied with their connectors, yet their loneliness scores changed little. What is the best response in the synthesis?

Dissonant evidence is a finding to investigate. In the Cedar Valley example, records showed that the transport fund ran out and rural groups were far away, which explained why satisfaction did not translate into lower loneliness. Averaging (option a) hides that explanation, and precision (option b) is a different property from relevance.

Question 4: What does the TIDieR item "when and how much" ask for in a description of the Cedar Valley program?

Item 8 of TIDieR asks for the number of sessions, their schedule, duration and intensity, which for Cedar Valley is up to six meetings over twelve weeks. Option b belongs to the item on who provided the intervention, option c to modifications, and option d to planned fidelity.
Section 4 of 5

Reporting and Use

⏱ Estimated reading time: 45 minutes
Section 4 of 5

Reporting and Use

Reports, messages, charts and the conditions under which evaluations are used.

The report

Findings, conclusions and recommendations

Finding

Attendance was 64 percent among adults whose first language is neither English nor French.

Conclusion

Equity of reach was adequate against the agreed standard.

Recommendation

Offer interpreter-supported first meetings before the second wave.

Executive summaries

The 1:3:25 format

1Page of main messages
3Pages of executive summary
25Pages of report, at most
Data visualization

Charts for decision-makers

  • Each chart carries one message, and its title states that message.
  • Position on a common scale supports the most accurate comparison.
  • Categories are sorted, points are labelled directly, and the standard is shown.
  • Economic results for executives are given as scenarios and plain probabilities.
Knowledge translation

Different products for different audiences

Executive: briefing before the decisionSteering committee: interpretation workshopFirst Nations partners: joint review before releaseOlder adults: plain-language summaryClinicians: one-page summaryResearchers: report and article
Evaluation use

Types of use and the conditions for it

Instrumental

Findings directly inform a decision.

Conceptual

Findings change how people understand a program.

Symbolic

An evaluation legitimizes a decision already made.

Process

People learn by taking part in the evaluation.

Worked example

The Cedar Valley evaluation proposal

6,000Words in the proposal
11Parts, drawn from Lessons 1 to 10
2Pages of executive summary, at most
Carry forward

Toward the final assessment

  • Findings, conclusions and recommendations are kept distinct.
  • Products, messengers and timing are planned for each audience.
  • Use is most likely when an evaluation is relevant, credible, timely and well communicated.

Learning Objectives for this section

  • Plan an evaluation report that separates findings, conclusions and recommendations and answers the evaluation questions in order.
  • Write main messages and an executive summary for decision-makers, using the 1:3:25 format as a guide.
  • Apply principles of data visualization to present evaluation findings to decision-makers.
  • Plan knowledge translation for different audiences, including communities that hold rights over their data.
  • Describe types of evaluation use and the conditions under which evaluations are used or misused.
  • Describe an evaluation work plan, budget and risk register, and explain how an evaluation proposal assembles the parts of an evaluation plan behind an executive summary for decision-makers.

An evaluation that reaches a sound judgement has done most of its work, and it can still fail if the judgement never reaches the people who need it, arrives after the decision, or is written in a form they cannot use. This section covers the report, the executive summary, the presentation of data, knowledge translation and the conditions under which evaluations are used. It ends with the work plan, budget and risk register of an evaluation and a worked example of a complete evaluation proposal for the fictional Cedar Valley Connector program.

4.1 The Evaluation Report

The full report is the reference document of an evaluation. It records what was evaluated, the questions asked, the methods used, the evidence and its limits, and the judgements reached, in enough detail that a reader can check the reasoning. The Program Evaluation Standards introduced in Lesson 1 apply directly: the accuracy standards call for conclusions that are justified by the evidence and reasoning, the utility standards call for timely and appropriate communication, and the propriety standards call for full and fair disclosure of findings, including unwelcome ones (Yarbrough et al., 2011).

Front matterv

The title page, acknowledgements (including the contributions of participants, partners and the steering committee), main messages and the executive summary. Main messages and the executive summary are often the only parts decision-makers read, so they are written last and with the most care.

Program and contextv

A description of the program as delivered, following TIDieR (Section 3), with its logic model and theory of change (Lesson 3) and the context in which it operates.

Purpose, questions and approachv

The purpose of the evaluation, its primary intended users, the evaluation questions in priority order (Lesson 4), and the evaluative criteria and rubric agreed with interest holders.

Methodsv

The design and its justification (Lessons 6 to 8), data sources and indicators (Lesson 5), implementation measures (Lesson 9), the costing and economic methods (Sections 1 and 2), analysis and synthesis methods, and the ethics review and data governance arrangements.

Findings by evaluation questionv

Findings organized by question, in the same order as the questions, so that a reader looking for the answer to one question finds all the relevant evidence in one place.

Conclusions, recommendations and limitationsv

Evaluative conclusions for each question with a level of confidence, recommendations that follow from the conclusions, and the limitations that affect how far the conclusions can be trusted.

Appendicesv

The evaluation matrix, instruments, detailed tables, the technical economic appendix with a completed CHEERS 2022 checklist, and the rubric with the evidence used for each rating.

Three kinds of statement should be kept distinct throughout. A finding reports evidence: attendance was 64 percent among adults whose first language is neither English nor French. A conclusion makes a judgement against a standard: equity of reach was adequate. A recommendation proposes an action that follows from one or more conclusions: offer interpreter-supported first meetings in all clinics before the second wave. Readers can then trace each recommendation back to its evidence. Good recommendations are specific, feasible, and linked to the conclusions that justify them. They name who would act, and they are developed with the primary intended users so that they fit the resources and authority those users have. Patton (2008) advises that recommendations focus on actions within the control of intended users and state the costs, benefits and challenges of carrying them out.

4.2 Main Messages and Executive Summaries

Background: the 1:3:25 format and Lavis's five questions

Two planning tools from the knowledge translation literature recur in this section and the next. The Canadian Health Services Research Foundation, whose work continues in Healthcare Excellence Canada, promoted a reader-friendly format for reports to decision-makers known as 1:3:25: one page of main messages, a three-page executive summary, and a report of no more than 25 pages, with technical detail in appendices. The proportions matter less than the principle that each layer must stand on its own for a reader who goes no further. Lavis and colleagues (2003) proposed five questions for planning the transfer of research knowledge to decision-makers: what should be transferred, to whom, by whom, how, and with what effect. Applied to an evaluation, the answers produce a plan with different products and messengers for different audiences, as the table in Section 4.4 shows.

HSCI 241 Lesson 12, Sections 3.1 and 3.3 (Reporting and Translating Review Findings), applies both tools to the findings of a systematic review and is optional fuller reading.

Main messages are the conclusions and their implications, written as statements a decision-maker could act on. They are distinct from a summary of the contents of the report. The executive summary then gives the context, the questions, the approach in a few sentences, the answer to each question with its confidence, and the recommendations. Writing for decision-makers involves leading with the answer, using plain language and consistent terms, giving numbers with their context (for example, 12 more older adults out of every 100 no longer screening as lonely), and stating uncertainty in words that a non-specialist can interpret.

Evaluation adds two requirements to this general advice. First, the main messages of an evaluation report are its evaluative conclusions from Section 3, each with its level of confidence, so the rubric and the synthesis rules agreed with the steering committee determine what the messages can claim. A message about worth also names the standard behind it, such as the willingness-to-pay value or the rubric's descriptor for good reach, so that the reader can see the basis of the judgement. Second, the executive summary answers the evaluation questions in the order the primary intended users agreed them, so that each user can find the answer to the question they asked. The executive summary of an evaluation proposal, as in Section 4.7, has a different job: it states what the evaluation will deliver, how and when, so that a funder can decide whether to support it.

Weaker main messageStronger main message
This report presents the results of a mixed-methods evaluation of the Connector program's first year.The Connector program reduced loneliness modestly among the older adults it served in its first year, and the evaluation team has moderate confidence in this result.
Cost-effectiveness results were sensitive to assumptions.At current caseloads the program costs about $125,000 for each year of full health gained, and the cost would fall below $50,000 if caseloads fill and benefits last.
Some groups had lower attendance.Older adults whose first language is neither English nor French were less likely to attend a first meeting, and interpreter-supported first meetings could close this gap before the second wave.

4.3 Data Visualization for Decision-Makers

Charts in evaluation reports communicate a message to a reader who has little time. Cleveland and McGill (1984) ranked visual encodings by how accurately people judge the quantities they show, and their experiments confirmed the upper part of the ranking: positions along a common scale are judged most accurately, lengths and angles less accurately, and, in their ranking, areas and colour shading least accurately. Dot plots and bar charts therefore support more accurate comparison than pie charts or bubble charts. Tufte (2001) urged designers to maximize the share of ink that shows data and to remove decoration that carries no information, and Evergreen (2017) adapted these ideas for evaluators with practical guidance on choosing and formatting charts for reports.

One message per chartClick to explore
A title that states the messageClick to explore
Use position on a common scaleClick to explore
Sort and label directlyClick to explore
Show uncertainty and standardsClick to explore
Design for every readerClick to explore
Attendance was lowest among adults whose first language is neither English nor French Good: 65% Overall: 78% WomenEnglish or French speakersUrban clinicsMenRural clinicsOther first language 81%80%80%72%71%64% 50%60%70%80%90% Referred adults who attended a first meeting, year one
Illustrative year-one attendance for the fictional Cedar Valley Connector program. The title states the finding, categories are sorted, values are labelled directly, and the overall rate and the rubric threshold give the reader a standard for comparison.

Economic results need particular care. The cost-effectiveness plane and the acceptability curve are useful for technical readers and are hard for most decision-makers to read. For an executive audience, a short scenario table and a plain statement of probability work better, for example: "If the health authority is willing to pay $100,000 for each year of full health gained, there is about a 28 percent chance that the program is good value at current caseloads."

4.4 Knowledge Translation

The Canadian Institutes of Health Research (CIHR) define knowledge translation as a dynamic and iterative process of synthesis, dissemination, exchange and ethically sound application of knowledge to improve health and strengthen the health system. CIHR distinguishes integrated knowledge translation, in which knowledge users take part in the work from the start, from end-of-grant knowledge translation, in which findings are shared once the work is complete. A utilization-focused evaluation with an engaged steering committee is integrated knowledge translation by design. Lesson 9's Knowledge-to-Action framework (Graham et al., 2006) describes what happens after knowledge reaches users.

In an evaluation, the steering committee is the usual vehicle for integrated knowledge translation. When knowledge users sit on the committee that chooses the questions, agrees the rubric and interprets draft findings, they shape the products and their timing and carry the findings back to their own organizations, which answers the Lavis question of who should transfer the knowledge. A finding is more likely to be acted on when the messenger is someone the audience trusts, which for clinicians may be a respected peer and for a community may be one of its own members, so committee members often present findings to their own constituencies. The table applies these ideas to Cedar Valley, using the five questions introduced in Section 4.2.

AudienceMain interestProduct and messengerTiming
Health authority executiveWhether to fund the second wave, and on what conditionsTwo-page briefing and a presentation by the evaluation lead and the program directorBefore the second-wave budget decision
Steering committeeInterpretation, recommendations and program improvementData interpretation workshop on draft findingsBefore the report is finalized
First Nations partnersFindings about Indigenous participants and the land-based pathwayJoint review and co-interpretation, and products the partners choose, under agreed data governanceBefore any release, as agreed
Older adults and familiesWhat the program offers and what it achievedPlain-language summary in large print and at community meetings, presented with the committee's older adult membersAfter the executive briefing
Referring cliniciansWhether referral helps their patientsOne-page summary and a short talk at clinic meetings by a clinician championBefore second-wave clinics begin referring
Other health authorities and researchersTransferability and methodsFull report and a journal article reported with CHEERS 2022 and TIDieRAfter the report is released

The First Nations row reflects commitments made in Lessons 4 and 5. Under the First Nations principles of OCAP® (ownership, control, access and possession), the partners decide how information about their members is interpreted and shared, and findings are reviewed with them before release. Reporting back to communities and participants is an obligation in its own right, separate from its value as a means of influencing decisions.

4.5 Evaluation Use

Research on evaluation use began when evaluators noticed that many evaluations had no visible effect on decisions. Weiss (1979) showed that research is used in many ways besides the direct application of findings, and reviews of evaluation use (Leviton & Hughes, 1981; Cousins & Leithwood, 1986) organized these into types that remain in use today.

Instrumental useClick to explore
Conceptual useClick to explore
Symbolic useClick to explore
Process useClick to explore
Evaluation influenceClick to explore

Instrumental use is the most visible type, but conceptual use and process use often matter as much over time, because they change how managers and staff think about a program. Cousins and Leithwood (1986) found that use depended on characteristics of the evaluation, such as its relevance, credibility, quality of communication and timeliness, and on characteristics of the decision setting, such as information needs, the political climate, competing information and the receptiveness of users. A later review of the empirical literature by Johnson and colleagues (2009) emphasized the engagement of interest holders throughout the evaluation. Patton (2008) called the presence of an identifiable person or group who cares about the findings the personal factor, and his utilization-focused approach (Lesson 1) builds the evaluation around such primary intended users.

Timelinessv

An evaluation that reports after the decision has little chance of instrumental use. The Cedar Valley plan schedules an interim briefing before the second-wave budget decision, with the full report to follow.

Relevance and engagementv

Questions chosen with primary intended users (Lesson 4) produce answers they want. Engagement throughout also builds the trust that makes unwelcome findings easier to accept.

Credibilityv

Users act on findings they believe. Credibility comes from a defensible design, transparent synthesis with a rubric agreed in advance, independence where it matters, and candour about limitations.

Communicationv

Products matched to each audience, clear main messages and good charts make findings usable by people who will never read the full report.

Decision contextv

Budgets, political commitments and competing priorities limit what an evaluation can change. Evaluators who understand the decision context can frame recommendations that are feasible within it.

Use also has an ethical side. Misuse occurs when findings are suppressed, distorted or selectively reported, or when an evaluation is commissioned only to justify a decision already made. Evaluators can reduce the risk by agreeing in advance, in the evaluation contract or terms of reference, how findings will be released, who owns the report, and how disagreements about interpretation will be handled, and by keeping Indigenous partners' data governance rights separate from any power to suppress unfavourable findings about the program.

4.6 Work Plan, Evaluation Budget and Risk Register

An evaluation proposal closes with the practical commitments that let a funder judge whether the evaluation can be delivered on time and within its means: a work plan, a budget and a risk register. They apply to the evaluation itself the program work plan and budget outline described in Lesson 2, Section 4.6 (Needs Assessment and Planning Models).

Background: timelines, roles and budgets

A Gantt chart lays the tasks of a project along a timeline, with a bar for the duration of each task, markers for key dates and arrows for dependencies between tasks. A roles table names one accountable person for each task, and a budget lists each cost line with the calculation behind it (quantity, unit cost and total). HSCI 207 Lesson 6, Sections 4.1 to 4.3 (Choosing an Approach and Managing a Research Project), introduces these tools for a research project and is optional reading.

The evaluation work plan

The work plan lists each evaluation task with the person or role responsible, its start and end dates and the output that marks its completion, and a Gantt chart shows the same tasks on a timeline. Its milestones, the dated checkpoints on which later work depends, include ethics approval, the data access agreement, the end of baseline data collection, the interim briefing and the final report. For Cedar Valley, the plan works backward from the second-wave budget decision: the interim briefing must reach the executive before that decision, so baseline data collection and the first analysis are scheduled to finish in time for it. Tasks that depend on outside bodies, such as research ethics review, review under the First Nations data governance agreement and requests for linked administrative data, often take longest and are started first.

The evaluation budget

The evaluation budget is separate from the program budget. Its largest line is usually staff time for the evaluation lead, analysts, research coordinators and interviewers, each costed as hours or full-time equivalents multiplied by a rate that includes benefits. Other common lines are incentives or honoraria for participants and for community members who advise the evaluation, transcription of interviews and focus groups, data access and linkage fees charged by data stewards, translation and printing of instruments, travel to clinics and community meetings, and the production of knowledge translation products. Each line shows its calculation, so that a reviewer can check it and the team can revise it when a quantity changes.

The risk register

A risk register lists what could stop the evaluation from answering its questions on time. For each risk it records the likelihood, the impact, the mitigation and the owner, the person responsible for watching the risk and acting on it. Likelihood and impact are often rated low, medium or high, so that the team attends first to risks that are both likely and serious. The table shows three illustrative entries for Cedar Valley.

RiskLikelihoodImpactMitigationOwner
Approval of the request for linked administrative data arrives after the interim briefing is dueMediumHighSend the request in the first month, and prepare the briefing from survey data if the linked data are lateEvaluation lead
Follow-up survey response is low among older adults who did not attend a first meetingHighMediumTelephone follow-up by trained interviewers, a modest honorarium, and an analysis of attritionResearch coordinator
Second-wave clinics launch earlier than planned, which shortens the comparison periodLowHighAgree launch dates with the steering committee in advance and record any changeProgram director, with the evaluation lead

4.7 Worked Example: The Cedar Valley Evaluation Proposal

An evaluation proposal brings the parts of an evaluation plan together in one document that a funder can review. The worked example is the 6,000-word proposal that the Cedar Valley evaluation team wrote at the start of the first wave, before any of the illustrative results in Sections 2 and 3 existed. A proposal describes a plan, so its executive summary states what the evaluation will do, why, and how its findings will be used. The table shows how the team allocated the 6,000 words and which lesson covers each part.

Proposal sectionSource in the courseWords
1. Program description, need and objectivesLessons 1 and 2700
2. Program theory: logic model and theory of changeLesson 3600
3. Interest holders, engagement and evaluation questionsLesson 4600
4. Evaluation matrix, indicators and data systemsLesson 5700
5. Design and its justificationLessons 6 to 81,000
6. Implementation evaluationLesson 9500
7. Costing and economic evaluationLesson 10, Sections 1 and 2600
8. Synthesis: criteria, rubric and confidenceLesson 10, Section 3400
9. Ethics, Indigenous data governance and equityLessons 1, 4 and 5300
10. Reporting, knowledge translation and useLesson 10, Section 4300
11. Timeline, evaluation budget and risksLesson 10, Section 4.6, and Lesson 2, Section 4.6300
Total6,000
Worked example: Executive summary of the Cedar Valley evaluation proposal (excerpt)

Main messages. The Cedar Valley Health Authority will decide in one year whether to extend the Connector program to its remaining 12 primary care clinics. This evaluation will tell the executive, before that decision, whether the program reaches the older adults who need it, whether it reduces loneliness, what it costs, and whether it is good value compared with usual care. It will also tell program staff how to improve delivery for groups who are not being reached.

Questions. The steering committee agreed five questions: how equitably the program reaches referred older adults; whether it is delivered as intended and in a culturally safe way; what effect it has on loneliness, social participation, self-rated health and use of emergency and primary care at twelve weeks and twelve months; what it costs and whether it is good value; and which contexts and mechanisms explain differences in results.

Approach. Because the first-wave clinics were chosen for readiness, the evaluation will compare them with the 12 second-wave clinics before those clinics launch, using a difference-in-differences design with checks of trends before launch. An implementation evaluation will track reach, fidelity and adaptations. A costing will combine program accounts with a time study of connector work, and a cost-utility analysis from the health-system perspective, with a societal scenario, will use the EQ-5D-5L at intake, twelve weeks and twelve months. Results will be judged against a rubric agreed in advance with the steering committee, with cultural safety as a hurdle.

Partnership and use. Older adults with lived experience and First Nations representatives on the steering committee will shape the instruments, the interpretation and the reports. Information about First Nations participants will be governed by an agreement with the partners consistent with OCAP®. The executive will receive an interim briefing before the budget decision, and older adults, clinicians and partners will each receive products designed for them.

What an evaluation proposal contains

An evaluation proposal of this kind, about 6,000 words excluding references and appendices, is organized as in the worked example. Beyond the parts developed in earlier lessons, it includes a costing and economic evaluation plan that states the perspective, the resources to be measured and the outcome measure, a draft rubric with synthesis rules, a reporting and use plan for at least three audiences, and the work plan, budget and risk register of Section 4.6. An executive summary of no more than two pages, written for decision-makers, precedes it.

A sound proposal connects the program's need, theory, questions, design and methods in a consistent line of reasoning. It justifies the design and states its assumptions and threats to validity with the checks that will address them, names a perspective for the costing and identifies resources (including in-kind and partner resources), and makes the rubric and synthesis rules explicit, developed or planned with primary intended users. Its plans for ethics, equity, Indigenous data governance, reporting and use are specific to the program's interest holders, and its executive summary can be read on its own and tells a decision-maker what the evaluation will deliver and when.

Reflection

A provincial ministry commissioned an evaluation of a youth mental health drop-in program. The evaluator worked independently and met the ministry once at the start. The 140-page report arrived three months after the ministry had set the next year's budget. It contained 60 tables, every chart was titled "Results", and it had no executive summary. It was sent only to the ministry, so the program's youth advisory council and its staff never saw it. A year later, staff said nothing in the program had changed. Using what is known about the conditions under which evaluations are used, identify three reasons this evaluation had little use. Then redesign the reporting and knowledge translation plan for three audiences (the ministry, program staff, and young people), naming for each a product, a messenger and a timing.

Model answer

Three conditions for use were missing. Timeliness: the report arrived after the budget decision, so instrumental use was impossible. Engagement and relevance: the evaluator met the ministry once and did not involve staff or young people, so no primary intended users were invested in the findings, which Patton calls the absence of the personal factor. Communication: a 140-page report without an executive summary or message titles is unusable for busy readers, and sending it only to the ministry excluded the people able to change practice.

For the ministry, I would provide a two-page briefing with main messages and recommendations, presented by the evaluator with the program director, at least one month before the budget decision, followed by a report of about 25 pages with technical appendices. For program staff, I would hold a data interpretation workshop on the draft findings, led by the evaluator and a respected program manager, before the report is finalized, so that staff shape the recommendations they will carry out. For young people, I would co-produce a short visual summary and a social media version with the youth advisory council, presented by council members at drop-in sites after the ministry briefing. Throughout, the evaluator would meet these users regularly so that the questions and products reflect their decisions.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: Which of the following statements is a conclusion, as distinct from a finding or a recommendation?

A conclusion makes a judgement against a standard, as option a does. Option b reports evidence and is a finding, option c proposes an action and is a recommendation, and option d describes a method.

Question 2: According to Cleveland and McGill (1984), which visual encoding supports the most accurate comparison of quantities?

Cleveland and McGill found that people judge quantities most accurately from positions along a common scale, and less accurately from angles, areas and shading. This is why dot plots and bar charts are preferred to pie and bubble charts for comparisons.

Question 3: The executive funds the second wave of Cedar Valley on the conditions the evaluation recommended. Which type of use is this?

Instrumental use occurs when findings directly inform a decision or action. Conceptual use changes understanding without an immediate decision, symbolic use legitimizes a decision already taken, and process use arises from taking part in the evaluation.

Question 4: Which plan follows the 1:3:25 format for a report to decision-makers?

The format promoted by the Canadian Health Services Research Foundation consists of one page of main messages, a three-page executive summary and a report of no more than 25 pages, with technical detail in appendices. Each layer must stand on its own for a reader who goes no further.
Section 5 of 5

Final Assessment

⏱ Estimated time: 30 minutes

Bringing It All Together

This lesson completed the evaluation plan by asking what a program costs, whether it is good value, how evidence is combined into a judgement, and how that judgement reaches the people who will use it. Section 1 showed that the budget of the fictional Cedar Valley Connector program ($840,000) understates its economic cost, which rises to $972,000 from the health-system perspective and to $1,022,000 when volunteer time is valued, and that time-driven activity-based costing puts the connector time for each participant at $525. It also showed how spare capacity in a start-up year inflates the average cost per participant. Section 2 used these costs in an illustrative cost-utility analysis, which gave an incremental cost of $1,868, a gain of 0.015 QALYs and an ICER of $124,533 per QALY, with scenarios ranging from $41,617 to $131,200 and a probability of 0.282 of being cost-effective at $100,000 per QALY.

Section 3 placed the economic result beside evidence on reach, equity, cultural safety and effectiveness in a rubric agreed in advance, and reached separate judgements of merit (good) and worth at year-one caseloads (adequate), each with a stated level of confidence. Section 4 turned to reports, executive summaries, data visualization, knowledge translation and the conditions for use, and ended with the work plan, budget and risk register of an evaluation and a worked example of a complete evaluation proposal. The lesson's main argument is that economic and other evidence serve decisions best when they are combined openly, with the standards agreed before the results arrive and the findings delivered in time and in forms that each audience can use.

Key Takeaways from this lesson

  • Economic evaluation values resources at their opportunity cost, so donated space, clinician time and volunteer time count even when no money changes hands.
  • The perspective of an analysis determines whose costs count, and the CADTH reference case uses the publicly funded health care payer perspective, with a societal analysis when important costs fall outside the health system.
  • Costing proceeds by identifying, measuring and valuing resources, and time-driven activity-based costing multiplies the time each activity takes by a capacity cost rate.
  • Average cost per participant falls as caseloads approach capacity, so start-up costs overstate steady-state costs and full-capacity assumptions overstate efficiency when demand is uncertain.
  • A full economic evaluation compares the costs and consequences of two or more alternatives, and cost-utility analysis expresses health gains in QALYs so that programs in different areas can be compared.
  • The ICER divides incremental cost by incremental effect, the cost-effectiveness plane classifies results by quadrant, and net monetary benefit converts health gains into money at a stated willingness to pay.
  • Scenario analyses show the effect of structural assumptions such as the persistence of benefits, while probabilistic sensitivity analysis and acceptability curves show the effect of uncertainty in parameter values.
  • Social return on investment ratios depend heavily on financial proxies and deadweight adjustments, and ratios from different analyses cannot be compared.
  • Evaluative rubrics make the reasoning behind a judgement explicit, and synthesis rules, hurdles and confidence statements should be agreed with primary intended users before the evidence arrives.
  • Evaluations are used when they answer questions their users care about, arrive before decisions are made, are credible, and are communicated in products designed for each audience, including communities that hold rights over their data.

Core Concepts Reviewed

Section 1: opportunity cost, financial and economic cost, program, health system and societal perspectives, identification, measurement and valuation, gross costing, micro-costing, time-driven activity-based costing, fixed, variable, average and marginal costs, and budget impact.

Section 2: full and partial economic evaluation, cost-effectiveness, cost-utility, cost-benefit and cost-consequence analysis, QALYs and utilities, the ICER, the cost-effectiveness plane, net monetary benefit, scenario and probabilistic sensitivity analysis, acceptability curves, social return on investment and CHEERS 2022.

Section 3: merit and worth, evaluative rubrics, qualitative and numerical weight and sum, hurdles and profiles, convergent, complementary and dissonant evidence, confidence in judgements, and TIDieR.

Section 4: findings, conclusions and recommendations, main messages and the 1:3:25 format, principles of data visualization, integrated knowledge translation, instrumental, conceptual, symbolic and process use, conditions for use, evaluation work plans, budgets and risk registers, and the structure of an evaluation proposal.

The final reflection asks you to bring the lesson together by turning the evidence from an evaluation into main messages, a recommendation and an explanation of an economic result for a decision-maker.

Reflection

A regional health authority must decide whether to renew a community diabetes self-management program for three years. The program ran for two years in 10 communities with an annual budget of $1.2 million and served 1,000 adults a year. The evaluation found the following. The health-system cost, including overhead and physician time, was $1.38 million a year. Compared with similar communities without the program, average HbA1c (a measure of blood glucose control) fell by an additional 0.3 percentage points at 12 months (95 percent confidence interval 0.1 to 0.5). The incremental cost-effectiveness ratio (ICER) was $45,000 per quality-adjusted life year (QALY), with a probability of 0.62 of being cost-effective at $50,000 per QALY. Participation was 40 percent of eligible adults in remote communities and 65 percent in towns. In interviews, participants valued group sessions led by peers, and remote participants said virtual sessions were hard to join because of poor internet service. Write three main messages for the executive summary and one recommendation, and explain in two or three sentences how you would present the economic result to an executive who is not an economist.

Model answer

Main messages. First, the program improved blood glucose control modestly among the adults it served, with an additional fall in HbA1c of 0.3 percentage points compared with similar communities, and we have moderate confidence that the program caused this improvement. Second, the program costs about $1,380 per participant a year to the health system and is likely to be good value at commonly used reference values, although there remains a meaningful chance that it is not. Third, adults in remote communities took part much less often than adults in towns (40 percent compared with 65 percent), mainly because virtual sessions depend on internet service that many remote households lack.

Recommendation. Renew the program for three years, and use part of the renewal to offer in-person, peer-led sessions in remote communities, with participation monitored against a target agreed with those communities.

Presenting the economic result. I would say that each year of full health gained costs about $45,000, which is below the $50,000 reference value often used in Canada, and that when the analysis accounts for uncertainty there is roughly a six-in-ten chance that the program is good value at that level. I would show a short table of the scenarios that most change the result, instead of a cost-effectiveness plane, and say which assumption the decision depends on most.

Minimum 30 characters required.

✓ Reflection saved

Final Knowledge Assessment

Final Assessment, this lesson: Economic Evaluation, Reporting and Evaluation Use (15 Questions)

Question 1: The fictional Cedar Valley program's health-system cost ($972,000) is higher than its budget ($840,000). Which items account for the difference?

The health-system cost adds overhead ($84,000), clinician screening and referral time ($28,800) and clinic rooms ($19,200) to the budget, for a total of $972,000. Volunteer and participant time (option a) belong to the societal perspective, and sunk costs and transfer payments are excluded.

Question 2: Why should an economic evaluation of a social prescribing program report a societal perspective alongside the health-system perspective?

Social prescribing programs mobilize community resources outside the health system, so a health-system analysis alone would hide costs shifted to community organizations and volunteers. The CADTH reference case is the publicly funded health care payer perspective (option b), and adding societal costs usually raises the ICER, as it did for Cedar Valley.

Question 3: Compared with usual care, the utility difference rises from 0 at intake to 0.04 at twelve weeks and returns to 0 at fifty-two weeks. What are the QALYs gained over the year?

The area between the profiles is a triangle with a base of one year and a height of 0.04, so the QALYs gained are 0.5 × 1 × 0.04 = 0.02. Option b counts only the first twelve weeks, and option a treats a utility difference as if it lasted the full year.

Question 4: A program costs $600 more per participant than usual care and gains 0.010 QALYs. At a willingness to pay of $50,000 per QALY, which statement is correct?

The ICER is $600 ÷ 0.010 = $60,000 per QALY, and the net monetary benefit is $50,000 × 0.010 − $600 = −$100, so the program is not cost-effective at $50,000. A program is dominant (option d) only when it also saves money.

Question 5: In the Cedar Valley probabilistic sensitivity analysis, the probability of cost-effectiveness at $100,000 per QALY was 0.282. What does this mean?

The probability on a cost-effectiveness acceptability curve is the share of simulated parameter sets for which net monetary benefit is positive at that willingness to pay. It reflects parameter uncertainty about the average result, not variation across clinics (option b).

Question 6: Which assumption in the Cedar Valley scenario analysis brought the ICER below $50,000 per QALY?

Only the combined scenario of persistent effects and full caseloads gave an ICER below $50,000 ($41,617). Full caseloads alone (option d) gave $73,630, while removing offsets and adding volunteer time raised the ICER to $129,600 and $131,200.

Question 7: Why does the Cedar Valley analysis test the persistence of effects with a scenario instead of only through the probabilistic analysis?

Whether the effect persists between twelve weeks and twelve months is an assumption about the shape of the model, since no one measured utility in that period. Scenarios make the consequences of such structural choices visible. Persistence changes QALYs within one year (option a), from 0.015 to 0.0265, and probabilistic analysis can include effect parameters.

Question 8: A cost-utility analysis that uses a threshold of $100,000 per QALY is a restricted form of which type of analysis?

Net monetary benefit converts QALYs into money at the threshold value, so a cost-utility analysis with a threshold is a cost-benefit analysis in which only health gains are valued. In the Cedar Valley example, the net benefit at $100,000 per QALY equalled the net monetary benefit of −$368.

Question 9: The Cedar Valley committee agreed that the program cannot be rated better than adequate overall if cultural safety is rated poor. What is this rule called, and why is it used?

A hurdle, or bar, is a minimum requirement that a program must meet before other criteria can raise its overall rating. It records the committee's view that harm to some participants cannot be rescued by good results for others. Weights (option b) allow compensation, which is what a hurdle prevents.

Question 10: At year-one caseloads, Cedar Valley was rated good on effectiveness and adequate on value for money. Which conclusion follows Scriven's distinction between merit and worth?

Merit reflects intrinsic quality, such as effectiveness and delivery, which were good. Worth reflects value in context, including cost, which was adequate at year-one caseloads. Cost bears on worth, so option b reverses the distinction, and option a applies a cost criterion to merit.

Question 11: Why does a complete TIDieR description matter for the economic evaluation of Cedar Valley?

A cost estimate applies to a program as delivered, so readers need the dose, providers and materials that TIDieR records (for Cedar Valley, a mean of four meetings and 10.5 hours of connector time) to interpret or transfer the cost. TIDieR complements CHEERS 2022 and does not replace it (option a).

Question 12: An evaluation team learns that the second-wave budget decision will be made in month ten, and its full report is due in month fourteen. What would most improve the chance of instrumental use?

Timeliness is one of the strongest conditions for instrumental use, and a briefing before the decision is the only option that puts findings in front of decision-makers when they decide. Asking to delay a budget decision (option c) is rarely feasible, and an acceptability curve is hard for non-specialists to read.

Question 13: Which practice best respects the rights of First Nations partners under OCAP® when the Cedar Valley evaluation reports its findings?

OCAP® affirms First Nations ownership, control, access and possession of information about their members, so findings about them are reviewed and interpreted with the partners before release, under the agreed data governance. Removing Indigenous participants (option c) would erase their experience from the evaluation.

Question 14: Which reporting guideline applies to the Cedar Valley cost-utility analysis, and what did its 2022 version add?

CHEERS 2022 is the 28-item reporting standard for health economic evaluations, and it added items on health economic analysis plans, distributional effects, and engagement with patients and others affected by the study. TIDieR describes interventions, CONSORT reports trials and StaRI reports implementation studies.

Question 15: A consultant reports that Cedar Valley returns $1.43 in social value for every dollar invested. Which question should an evaluator ask first?

SROI ratios depend most on the financial proxies and on the deadweight adjustment. In the Cedar Valley example, a deadweight of 50 percent reduced the ratio from 1.43 to 0.71. Ratios from different SROI analyses are not comparable (option a), and counting volunteer time as a benefit could double count value already reflected in the outcomes.
✦ Complete the final reflection above before submitting

Congratulations!

You have successfully completed this lesson: Economic Evaluation, Reporting and Evaluation Use.

You can now cost a program from more than one perspective, calculate and interpret an incremental cost-effectiveness ratio and net monetary benefit, represent the uncertainty in an economic result, judge the claims of a social return on investment analysis, synthesize mixed evidence with a rubric into judgements of merit and worth, describe a program with TIDieR, and plan reports and knowledge translation that give an evaluation a good chance of being used.

This lesson closes HSCI 826. Across ten lessons the course has described the Cedar Valley Connector program and its need, built its logic model and theory of change, engaged interest holders and framed evaluation questions, specified indicators and data systems, chosen among randomized and quasi-experimental designs, planned an implementation evaluation, and now costed the program and planned how its evaluation will reach a judgement and be used. The evaluation proposal in Section 4.7 draws these parts together into a single document with an executive summary for decision-makers, the kind of document a British Columbia health authority could fund and act on.

Back to Course Home →
Reference

Glossary: Key Terms, People & Frameworks

📚 Reference page, available throughout the lesson

This glossary defines the terms, tools and people introduced in Lesson 10, grouped by the lesson's main themes.

Core Concepts
Opportunity cost The value of the best alternative use of a resource. Economic evaluation values resources at their opportunity cost, so the economic cost of a program includes donated and absorbed resources as well as its financial cost.
Perspective The point of view that defines whose costs and consequences an economic analysis counts, such as the program, the publicly funded health system or society as a whole.
Gross costing A top-down method that divides a total expenditure by a measure of output, such as a budget divided by the number of participants. It is quick but cannot show why costs vary.
Micro-costing A bottom-up method that measures and values each resource used to deliver a program to each participant. It is used when a program is new, when costs vary, or when the drivers of cost must be shown.
Time-driven activity-based costing A micro-costing method that multiplies the time each activity takes by a capacity cost rate, the cost of supplying staff capacity divided by practical capacity in hours (Kaplan & Anderson, 2004).
Marginal cost The additional cost of serving one more participant. When a program has spare capacity, its marginal cost is usually much lower than its average cost.
Cost-effectiveness analysis A full economic evaluation that measures consequences in a single natural unit, such as an additional older adult no longer screening as lonely, and reports the cost per unit gained.
Cost-utility analysis A full economic evaluation that measures consequences in quality-adjusted life years, allowing comparison of programs across health areas.
Cost-benefit analysis A full economic evaluation that values all consequences in money and reports net benefit or a benefit-cost ratio.
Cost-consequence analysis An analysis that reports costs alongside several outcomes without combining them into one measure, leaving the weighing of outcomes to the reader.
Quality-adjusted life year (QALY) A measure of health that weights each period of life by a utility value between 0 (equivalent to death) and 1 (full health). One year in full health equals one QALY.
Incremental cost-effectiveness ratio (ICER) The difference in cost between two alternatives divided by the difference in their effects, giving the additional cost of each additional unit of effect.
Cost-effectiveness plane A graph with incremental effect on the horizontal axis and incremental cost on the vertical axis, whose quadrants classify a program as dominant, dominated or involving a trade-off (Black, 1990).
Net monetary benefit The health gain valued at a willingness to pay per unit, minus the incremental cost. A positive value means the program is cost-effective at that willingness to pay (Stinnett & Mullahy, 1998).
Probabilistic sensitivity analysis A method that assigns probability distributions to uncertain parameters, draws values many times, and recalculates the result for each draw to show the joint effect of parameter uncertainty.
Cost-effectiveness acceptability curve A plot of the probability that a program is cost-effective across a range of willingness-to-pay values, derived from a probabilistic sensitivity analysis (van Hout et al., 1994).
Social return on investment (SROI) A form of cost-benefit analysis that assigns financial proxies to outcomes identified with interest holders and adjusts for deadweight, attribution, displacement and drop-off to give a ratio of social value to investment.
Deadweight In SROI, the share of an outcome that would have occurred without the program. Its estimate has a large effect on the resulting ratio.
Merit and worth Scriven's distinction between the intrinsic quality of a program (merit) and its value in a particular context, including cost and need (worth).
Evaluative rubric A table of criteria and performance levels with descriptors that state what performance at each level looks like, used to make evaluative reasoning explicit.
Hurdle A minimum level a program must reach on a criterion before other criteria can raise its overall rating. Also called a bar.
Instrumental use The direct use of evaluation findings to inform a decision or action.
Conceptual use A change in how people understand a problem or program as a result of an evaluation, without an immediate decision.
Symbolic use The use of an evaluation to legitimize a decision already made or to signal accountability, with little attention to its findings.
Process use Changes in thinking and behaviour that result from taking part in an evaluation, such as staff learning to use data (Patton, 2008).
Knowledge translation In the definition of the Canadian Institutes of Health Research, an iterative process of synthesis, dissemination, exchange and ethically sound application of knowledge to improve health and the health system.
Frameworks & Tools
CHEERS 2022 The Consolidated Health Economic Evaluation Reporting Standards, a 28-item checklist for reporting health economic evaluations, which added items on analysis plans, distributional effects and engagement (Husereau et al., 2022).
TIDieR The Template for Intervention Description and Replication, a 12-item checklist for describing interventions (Hoffmann et al., 2014). TIDieR-PHP adapts it for population health and policy interventions.
EQ-5D-5L A generic preference-based measure of health with five dimensions and five levels, which has a Canadian value set derived from the time trade-off (Xie et al., 2016).
ICECAP-O A capability measure for older people covering attachment, security, role, enjoyment and control (Coast et al., 2008), sometimes reported alongside QALYs for social programs.
1:3:25 format A reader-friendly reporting format promoted by the Canadian Health Services Research Foundation: one page of main messages, a three-page executive summary and a report of no more than 25 pages.
Qualitative weight and sum A synthesis method in which criteria are given importance categories and the pattern of ratings within each category is compared, without arithmetic on ordinal ratings (Scriven, 1991).
Key People
Michael F. Drummond Health economist at the University of York and lead author of Methods for the Economic Evaluation of Health Care Programmes, whose definitions of full and partial economic evaluation this lesson follows.
George W. Torrance Health economist at McMaster University who, with colleagues, developed the time trade-off method for measuring health state utilities and contributed to the Health Utilities Index.
Milton C. Weinstein Harvard health decision scientist who, with William Stason, set out the foundations of cost-effectiveness analysis for health and medical practices in 1977.
Robert S. Kaplan Harvard Business School accounting scholar who developed activity-based costing and, with Steven Anderson, time-driven activity-based costing.
E. Jane Davidson Evaluator and author of Evaluation Methodology Basics (2005), which sets out practical methods for evaluative rubrics and for synthesizing evidence into judgements.
Tammy C. Hoffmann Researcher in evidence-based practice who led the development of the TIDieR checklist for describing interventions (Hoffmann et al., 2014).
Carol H. Weiss Evaluation scholar at Harvard who showed that research and evaluation are used in many ways besides the direct application of findings (Weiss, 1979).
J. Bradley Cousins Evaluation scholar at the University of Ottawa whose review with Kenneth Leithwood (1986) identified the characteristics of evaluations and decision settings that influence use.
No matching entries. Try a different search term.