Implementation Science
Program Planning & Evaluation
Learning objectives for this lesson:
- Distinguish efficacy, effectiveness and implementation research, and explain how the research-to-practice gap and the voltage drop arise.
- Define the eight implementation outcomes of Proctor and colleagues and distinguish them from service and client outcomes.
- Use Nilsen's taxonomy to classify implementation theories, models and frameworks by the purpose they serve.
- Apply CFIR 2.0 to assess the determinants of a program's implementation, using the second-wave rollout of the Cedar Valley Connector program.
- Apply RE-AIM, interpreted with PRISM, to calculate and interpret indicators for the evaluation of a program rollout.
- Select implementation strategies from the ERIC compilation, specify them with the seven dimensions of Proctor, Powell and McMillen, and document adaptations with FRAME.
- Explain the three types of hybrid effectiveness-implementation studies, including the 2022 update, and choose a type that fits the state of evidence for a program.
- Identify the reporting requirements of StaRI for an implementation study.
- Select an implementation framework for a program and specify the implementation outcomes to be measured.
This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University, drawing on Rossi, P. H., Lipsey, M. W., & Henry, G. T. (2019). Evaluation: A Systematic Approach (8th ed.). SAGE; and Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and Quasi-Experimental Designs for Generalized Causal Inference. Houghton Mifflin.
The Research-to-Practice Gap
Learning Objectives for this section
- Describe the research-to-practice gap and explain why programs that work in trials often lose benefit when they are delivered at scale.
- Distinguish efficacy, effectiveness and implementation research by the question each asks, the conditions it studies and the outcomes it measures.
- Explain the difference between intervention failure and implementation failure, and why an evaluation needs evidence to tell them apart.
- Define the eight implementation outcomes of Proctor and colleagues and distinguish them from service and client outcomes.
- Write an indicator for each implementation outcome for the Cedar Valley Connector program and calculate penetration from referral data.
1.1 Why Evidence Does Not Reach Practice on Its Own
The first eight lessons of this course asked whether a program works and how an evaluator can tell. This lesson asks what determines whether a program with evidence behind it is taken up, delivered well and kept going in the settings that are supposed to use it. The question matters because the movement of evidence into routine practice is slow and uneven. A frequently quoted estimate holds that it takes about 17 years for research findings to reach routine clinical practice (Balas & Boren, 2000). Morris, Wooding and Grant (2011) reviewed the studies behind that figure and found that estimates of the time lag vary widely and depend on where the start and end points of the lag are placed. The figure is best treated as an indication that translation is slow, with considerable uncertainty about how slow.
A second problem concerns what happens to programs once they are delivered. Interventions tested under controlled conditions often produce smaller benefits when they are delivered by ordinary staff, to more varied populations, with fewer resources and less supervision. Chambers, Glasgow and Stange (2013) discussed this loss of benefit, often called a voltage drop, together with the related idea of program drift, the gradual departure of everyday practice from a tested protocol. They questioned the assumption that a program can be optimized once, at the end of a trial, and then delivered unchanged in settings that themselves keep changing. Section 3 returns to their argument when it turns to adaptation and sustainability.
Implementation science is the field that studies these problems. Eccles and Mittman (2006), in the opening editorial of the journal Implementation Science, defined it as the scientific study of methods that promote the systematic uptake of research findings and other evidence-based practices into routine practice, with the aim of improving the quality and effectiveness of health services. In Canada the broader enterprise is usually called knowledge translation. The Canadian Institutes of Health Research (CIHR) describe knowledge translation as a dynamic and iterative process that includes the synthesis, dissemination, exchange and ethically sound application of knowledge to improve health, provide more effective services and strengthen the health system (Straus, Tetroe & Graham, 2009). CIHR's planning guidance also distinguishes integrated knowledge translation, in which knowledge users take part throughout a study, from end-of-grant knowledge translation, which follows it. Implementation science is the part of this enterprise that tests and explains how evidence-based practices become part of routine care.
Dissemination and implementation
The literature separates two activities that are often run together. Dissemination is the targeted distribution of information and intervention materials to a specific public health or clinical audience. Implementation is the use of strategies to adopt and integrate evidence-based health interventions and to change practice patterns within specific settings. A guideline can be mailed to every family physician in a health authority while practice stays the same. These definitions follow the funding announcements of the United States National Institutes of Health.
Social prescribing programs illustrate both sides of the gap. Community connector programs have spread across the United Kingdom, Canada and other countries faster than evidence about their effectiveness has accumulated, and early systematic reviews judged that evidence to be limited (Bickerdike et al., 2017). Evaluators of these programs therefore often face the reverse of the usual problem: a program that is already being implemented widely while questions about its effects remain open. Where a connector program does help people, its effect in a population depends on whether clinicians refer, whether referred people engage, whether connectors deliver the core functions of the role, and whether community resources exist to connect people to. Each of those conditions is a question for implementation science.
The Cedar Valley Connector program is a fictional community connector (social prescribing) program run by the fictional Cedar Valley Health Authority in British Columbia. Primary care clinicians refer adults aged 65 and older who score 6 or higher on the three-item UCLA Loneliness Scale (scored 3 to 9), or whom they judge to be isolated, to a community connector. The connector meets each person up to six times over twelve weeks, co-develops a plan with them, and links them to community groups, volunteer roles, transportation help and services. The program launched in 12 of the region's 24 primary care clinics. As Lesson 7 established, those first-wave clinics were chosen for their readiness. In this lesson the 12 second-wave clinics, which start a year later, are on average less ready: they include four rural clinics, two clinics that depend heavily on locum physicians, and several clinics with no spare room for a connector.
In its first six months the program received 312 referrals, held first meetings with 241 older adults, and recorded both a baseline and a follow-up loneliness score for 188 of them, whose mean score fell from 7.1 to 6.3. Lesson 7 showed why that pre-post change cannot by itself be read as the program's effect. This lesson asks why some first-wave clinics referred many patients and others very few, what the second-wave clinics will need, which strategies the health authority should use to support them, and how an evaluation could study the program's effects and its implementation at the same time.
1.2 Efficacy, Effectiveness and Implementation
Flay (1986) distinguished efficacy trials, which test whether an intervention does more good than harm when it is delivered under optimum conditions, from effectiveness trials, which test whether it does more good than harm when it is delivered under real-world conditions. Lesson 6 developed the same distinction in the vocabulary of explanatory and pragmatic trials and showed how the PRECIS-2 wheel places a trial between the two poles. Glasgow, Lichtenstein and Marcus (2003) argued that the conventional sequence, in which efficacy trials come first and effectiveness studies follow, slows translation. Efficacy trials select motivated participants, expert staff and well-resourced settings, so their results say little about the conditions in which programs will actually be used.
Implementation research asks a third kind of question. It treats an intervention with evidence of effectiveness as given and studies the methods used to get it adopted, delivered and sustained. Its independent variable is usually an implementation strategy, such as training, practice facilitation, audit and feedback, or a change to the electronic medical record. Its dependent variables are implementation outcomes, such as the proportion of clinics that adopt the program or the proportion of eligible patients who are referred. Health outcomes remain relevant, and Section 4 shows how a single study can measure both, but the primary test of an implementation strategy is its effect on implementation outcomes. Table 1.1 sets the three kinds of research side by side.
| Feature | Efficacy research | Effectiveness research | Implementation research |
|---|---|---|---|
| Question | Can the intervention work under ideal conditions? | Does the intervention work when delivered in usual practice? | Which methods get the intervention adopted, delivered well and sustained? |
| What is compared | The intervention and a control condition | The intervention and usual care | Two or more implementation strategies, or a strategy and usual rollout |
| Typical unit | Individual participants | Individuals or clusters | Clinicians, clinics, organizations or regions |
| Primary outcomes | Health outcomes and mechanisms | Health outcomes in routine conditions | Implementation outcomes such as adoption, fidelity and penetration |
| Cedar Valley example | A trial in which research staff deliver the connector model to screened volunteers | A comparison of loneliness in clinics with and without the program | A comparison of two ways of supporting second-wave clinics to begin referring |
Curran (2020) proposed a plain-language way of keeping these questions apart, which many implementation scientists now use in teaching. The tabs below apply it to the Cedar Valley program.
Curran calls the intervention, practice or innovation "the thing". For Cedar Valley, the thing is the connector program: screening in primary care, referral, and up to six meetings with a connector who co-develops a plan and links the person to community resources, as the Lesson 3 logic model describes.
Effectiveness research asks whether the thing works. For Cedar Valley, this is the question of Lessons 6 to 8: whether referral to a connector reduces loneliness and changes health service use compared with what would have happened without it.
Implementation strategies are the activities used to help people and places do the thing. For Cedar Valley, they include training clinicians to ask about loneliness, adding a referral template to the electronic medical record, appointing a clinic champion, and sending each clinic a regular report on its referrals.
Implementation outcomes describe how much and how well people and places do the thing. For Cedar Valley, they include the number of clinics that start referring, the share of eligible older adults who are referred, whether connectors deliver the core functions of the role, and whether referrals continue once the launch support ends.
1.3 Intervention Failure and Implementation Failure
When an outcome evaluation finds no effect, there are two broad explanations. The program may not work, because its theory of change is wrong or its activities are too weak to produce the intended change. This is intervention failure, sometimes called theory failure. Alternatively, the program may work when it is delivered but may not have been delivered, or delivered to too few people, or delivered without its essential components. This is implementation failure. Basch and colleagues (1985) called the evaluation of a program that was never adequately implemented a Type III error: the evaluator draws a conclusion about a program that, in practice, did not take place.
Proctor and colleagues (2011) argued that the distinction between the two failures is one of the main reasons to measure implementation outcomes explicitly. Without them, a null result cannot be interpreted. The process evaluation methods of Lesson 5, including measures of reach, dose delivered, dose received and fidelity, supply much of the evidence needed. Table 1.2 shows the four combinations an evaluator can encounter.
| Program and delivery | Delivered well | Delivered poorly |
|---|---|---|
| Program works when delivered | Benefits are observed, and the program is a candidate for spread. | Benefits are lost through implementation failure, so the remedy lies in implementation strategies. |
| Program does not work | The null result reflects intervention failure, so the program theory needs revision. | The result cannot be interpreted, because the evaluation cannot say whether the program would have worked. |
In the Cedar Valley first wave, referral rates differed several-fold between clinics, which invites a comparison of outcomes between well and poorly implemented clinics. Such an analysis is observational even within a randomized study, because clinics that implement well may differ in staffing or community resources that also affect loneliness, so the confounding logic of Lesson 7 applies.
1.4 Implementation Outcomes
Proctor and colleagues (2009) proposed a conceptual model that separates three kinds of outcomes. Implementation outcomes are the effects of deliberate actions to implement a program. Service outcomes describe the quality of the care or services delivered, using the six quality domains of the United States Institute of Medicine: efficiency, safety, effectiveness, equity, patient-centredness and timeliness. Client outcomes describe changes in the people served, such as satisfaction, function and symptoms. In the model, implementation strategies act on implementation outcomes, and implementation outcomes are preconditions for service and client outcomes, as Figure 1.2 shows.
Proctor and colleagues (2011) then defined eight implementation outcomes, reviewed how each had been measured, and proposed that each is most salient at a particular stage of implementation. Acceptability, adoption, appropriateness and feasibility matter most early, when settings decide whether to take a program up. Fidelity and implementation cost matter most during delivery. Penetration and sustainability matter most later, when the question is whether the program has become part of routine operations. The cards below define each outcome and give an indicator for the Cedar Valley second wave.
Several of these outcomes can be measured with brief instruments. Weiner and colleagues (2017) developed and tested the Acceptability of Intervention Measure, the Intervention Appropriateness Measure and the Feasibility of Intervention Measure, each with four items scored on a five-point scale, and found evidence for their reliability and for the distinctness of the three constructs. Other outcomes are measured from administrative records (adoption and penetration), program logs and observation (fidelity), financial and time records (implementation cost), or repeated follow-up of settings (sustainability). Ten years after the original taxonomy, Proctor and colleagues (2023) reviewed the studies that had used it and reported that some outcomes, notably acceptability, fidelity and feasibility, had been studied far more often than others, and that the links between implementation outcomes and client outcomes had rarely been tested.
Implementation outcomes and process evaluation
Several implementation outcomes overlap with the process evaluation components of Lesson 5: fidelity appears in both, and penetration is close to reach. In a process evaluation these measures help explain an outcome evaluation's findings, while in implementation research they are the dependent variables against which a strategy is judged.
1.5 Worked Example: First-Wave Penetration and Its Consequences
The Cedar Valley referral rule is a UCLA Loneliness Scale score of 6 or higher, or clinician judgement. To calculate penetration, the evaluator needs a denominator: the number of eligible older adults attached to the first-wave clinics. Lesson 2 estimated that the 12 first-wave clinics have about 23,000 attached patients aged 65 and older. The regional survey described in Lessons 1 and 2 found that 24.5 percent of respondents aged 65 and older scored 6 or higher, which gives an estimated 5,635 eligible older adults.
Penetration of the referral step, first wave, first six months
Penetration = referrals ÷ estimated eligible older adults = 312 ÷ 5,635 = 0.055, or 5.5 percent.
Share of the eligible population who attended a first meeting = 241 ÷ 5,635 = 4.3 percent.
Share of referred older adults who attended a first meeting = 241 ÷ 312 = 77.2 percent.
Share of the eligible population with both loneliness measurements = 188 ÷ 5,635 = 3.3 percent.
Across the 12 first-wave clinics, penetration ranged from about 1 percent to 10 percent, and 52 of the 88 clinicians in those clinics (59.1 percent) made at least one referral in the six months. These figures change how the pre-post loneliness result should be read. Even if the program reduces loneliness among participants, the reduction applies to fewer than one in twenty of the eligible older adults in those clinics, and the 188 people with two measurements are about one in thirty. A program's population impact is the product of how many eligible people it reaches and how much it helps them, which is the reasoning behind the RE-AIM framework in Section 2.
The denominator deserves scrutiny. It applies a survey prevalence to clinic panels, which assumes that the panels resemble the survey population. Rural clinics, clinics with older panels, or clinics serving communities with more people living alone may have higher prevalence, so a single regional figure could understate need in some clinics and overstate it in others. A better denominator would count patients who were actually screened and scored 6 or higher, which requires a structured screening field in the electronic medical record. The second-wave plan in Section 4 adds that field.
Classify each of the following Cedar Valley measures as an implementation outcome, a service outcome or a client outcome, and name the specific outcome where you can. (1) The proportion of clinicians in a clinic who made at least one referral. (2) The mean change in UCLA Loneliness Scale scores among participants. (3) The median number of working days between referral and a connector's first contact. (4) Clinicians' scores on the Acceptability of Intervention Measure. (5) Participants' self-rated health at twelve weeks. (6) The proportion of participants whose plan was co-developed by the third meeting. (7) Whether referral rates in older adults who speak languages other than English match those in English speakers. Then open the answers below.
Item 1 is an implementation outcome (adoption at the clinician level). Item 2 is a client outcome (symptoms, in Proctor's terms). Item 3 is a service outcome (timeliness). Item 4 is an implementation outcome (acceptability). Item 5 is a client outcome (function or health status). Item 6 is an implementation outcome (fidelity). Item 7 is best treated as a service outcome (equity), although it can also be read as the representativeness of penetration across groups, which shows that some measures sit on the boundary between categories. What matters for an evaluation plan is that each measure is placed deliberately and that its role in the analysis is stated.
Section 2 turns to the frameworks used to plan, explain and evaluate this kind of work, each applied to the Cedar Valley second wave.
Reflection
A regional health authority introduced a community paramedicine program in which paramedics visit frail older adults at home every two weeks for six months to review medications, check blood pressure and arrange services. The program ran in 10 communities. After one year, an outcome evaluation found no difference in emergency department visits between participants and a matched comparison group. The process data show the following. All 10 communities started the program. In 4 of the 10 communities, staffing shortages meant that paramedics completed fewer than half of the planned visits. Of the older adults whom family physicians identified as eligible, 35 percent were enrolled. Paramedics rated the program as a good fit for the needs of their communities, but they disliked the program's documentation software and found it unpleasant to use. In the 6 communities with full staffing, audits found that the visit protocol was delivered as written in about 90 percent of visits.
Proctor and colleagues define eight implementation outcomes: acceptability (the program is agreeable or satisfactory to those involved), adoption (the decision or action to use it), appropriateness (its perceived fit for the setting, provider or problem), feasibility (the extent to which it can be carried out in the setting), fidelity (delivery as intended), implementation cost, penetration (the share of eligible people who receive it) and sustainability (its maintenance in routine operations).
(a) Identify which implementation outcomes the process data describe, and characterize each as favourable or unfavourable. (b) Explain whether the null result is better described as intervention failure, as implementation failure, or as not yet interpretable, and state what additional analysis or data would help you decide.
(a) Adoption was favourable, because all 10 communities started the program. Appropriateness was favourable, since paramedics saw the program as a good fit for their communities. Acceptability was mixed and unfavourable for one component, because paramedics disliked the documentation software. Feasibility was unfavourable in 4 communities, where staffing made the planned schedule impossible. Fidelity was low in those 4 communities, where fewer than half of the visits took place, and high in the other 6, where about 90 percent of visits followed the protocol. Penetration was modest, at 35 percent of eligible older adults.
(b) The null result is not yet interpretable. It pools 4 communities with clear implementation failure and 6 with adequate delivery, and it applies to about a third of the eligible population. I would first estimate the effect separately in the 6 fully staffed communities, recognizing that this comparison is observational because those communities may differ in other ways. I would examine whether participants who received more visits had fewer emergency visits, again with attention to confounding, and check whether the study had enough power to detect a plausible effect. I would also compare enrolled and non-enrolled eligible adults. If the program shows no effect even where it was delivered as intended, intervention failure becomes the more likely explanation; if an effect appears there, the priority is staffing and implementation support.
Minimum 20 characters required.
Question 1: Which of the following is an implementation research question about the Cedar Valley Connector program?
Question 2: A clinic has an estimated 180 eligible older adults and made 27 referrals to the connector program in six months. What is the penetration of the referral step?
Question 3: An evaluation finds no effect of a connector program on loneliness, but process data show that only a third of referred older adults ever met a connector. What does this most likely illustrate?
Question 4: Clinicians at a rural clinic say the connector program fits their patients' needs well and that a referral takes under a minute, but they dislike the program's scripted screening questions and find them disagreeable to use. Using Proctor's definitions, how should these findings be described?
Implementation Frameworks: Process, Determinant and Evaluation
Learning Objectives for this section
- Use Nilsen's taxonomy to classify implementation theories, models and frameworks by the purpose they serve.
- Describe the Knowledge-to-Action framework and the EPIS framework as process frameworks and apply them to a program rollout.
- Explain the five domains of CFIR 2.0 and apply them in a pre-implementation assessment of the Cedar Valley second wave.
- Explain when the Theoretical Domains Framework is a better fit than CFIR 2.0 for diagnosing an implementation problem.
- Calculate RE-AIM indicators for the Cedar Valley second wave and explain how PRISM adds context to their interpretation.
2.1 Theories, Models and Frameworks
Implementation science has produced a large number of theories, models and frameworks, and new users often find it hard to tell which ones answer which questions. Nilsen (2015) proposed a taxonomy that sorts them by purpose. He distinguished a theory, a set of analytical principles or statements that structure observation, understanding and explanation, from a model, which is a deliberate simplification of a phenomenon, and from a framework, which is a structure of descriptive categories and the relations between them that are presumed to account for a phenomenon. Frameworks name and organize the factors that matter without necessarily explaining how those factors produce their effects.
Nilsen grouped the approaches into five categories that serve three aims. Process models describe or guide the process of moving research into practice, usually as a sequence of steps or phases. Determinant frameworks, classic theories and implementation theories share the aim of understanding and explaining what influences implementation outcomes. Determinant frameworks list barriers and enablers across several levels; classic theories, such as diffusion of innovations, come from psychology, sociology and organizational science; and implementation theories, such as Normalization Process Theory, were developed within the field. Evaluation frameworks specify what to measure to judge whether implementation succeeded. Figure 2.1 arranges the five categories under their aims.
The taxonomy has a practical use. An evaluator who needs to plan a rollout, diagnose why it is struggling, and judge whether it worked needs one approach for each purpose, and confusion arises when a framework built for one purpose is stretched to serve another. Some frameworks sit across categories; EPIS, for example, describes phases and also names determinants within them.
2.2 Process Frameworks: Knowledge-to-Action and EPIS
Process frameworks lay out the steps through which evidence moves into use. Two are widely used in health services and fit the Cedar Valley program well.
The Knowledge-to-Action framework was developed by Graham and colleagues (2006) at the University of Ottawa and is used by CIHR to describe knowledge translation. It has two parts. The knowledge creation funnel refines knowledge from primary studies (knowledge inquiry) through syntheses (knowledge synthesis) to guidelines, decision aids and other tools (knowledge products). The action cycle describes the steps of applying that knowledge: identify a problem and select the relevant knowledge, adapt the knowledge to the local context, assess barriers to knowledge use, select, tailor and implement interventions, monitor knowledge use, evaluate outcomes, and sustain knowledge use. The phases can overlap and recur, and the funnel can feed the cycle at any point.
For Cedar Valley, the problem was identified through the regional survey in which about one in four older adults scored 6 or higher on the UCLA scale. Knowledge was adapted to the local context through the 18-month pilot in two clinics. The second wave now sits at the step of assessing barriers to knowledge use, which is where the CFIR assessment in Section 2.3 belongs, followed by the selection and tailoring of interventions, which in this framework means implementation strategies (Section 3).
The EPIS framework of Aarons, Hurlburt and Horwitz (2011) names four phases: Exploration, in which a system or organization considers its needs and identifies a program; Preparation, in which it plans, identifies barriers and facilitators and prepares for adaptation; Implementation, in which delivery begins and is monitored; and Sustainment, in which the program continues with whatever adaptation is needed. Within each phase, EPIS identifies factors in the outer context (policy, funding, inter-organizational networks) and the inner context (organizational characteristics, leadership, staff). A systematic review of its use by Moullin and colleagues (2019) gave greater emphasis to bridging factors, the relationships and arrangements that link outer and inner contexts, and to innovation factors, the characteristics of the program itself.
EPIS is useful for Cedar Valley because the program's clinics are in different phases at the same time. The first-wave clinics are moving from Implementation toward Sustainment, while the second-wave clinics are in Preparation. Bridging factors include the steering committee with older adults and First Nations representatives, the agreement under which a First Nations health centre hosts a connector, and data-sharing agreements between the clinics and the health authority.
2.3 Determinant Frameworks: CFIR 2.0 and the Theoretical Domains Framework
The Consolidated Framework for Implementation Research
The Consolidated Framework for Implementation Research (CFIR) was published by Damschroder and colleagues in 2009 as a consolidation of constructs drawn from earlier implementation theories. It became one of the most widely used determinant frameworks in the field. Damschroder, Reardon, Widerquist and Lowery (2022b) published an updated version, CFIR 2.0, based on a review of published applications and a survey of authors who had used it. The update has five domains containing 48 constructs and 19 subconstructs. It uses the general term "innovation" for the program or practice being implemented, gives more attention to the people who receive an innovation, and adds constructs related to equity. A companion paper, the CFIR Outcomes Addendum (Damschroder et al., 2022a), recommends that users define the implementation outcomes they are trying to explain before they assess determinants, so that the assessment explains something specific.
Using CFIR 2.0 starts with three definitions: the innovation (what exactly is being implemented), the inner setting (where it is implemented) and the outer setting (the setting in which the inner setting exists). The accordion below summarizes the five domains with examples of their constructs.
The characteristics of the innovation itself. Constructs include innovation source, evidence-base, relative advantage, adaptability, trialability, complexity, design and cost. For Cedar Valley, the innovation is defined as screening for loneliness in primary care, referral, and connector support, so the complexity of the referral process belongs here.
The setting in which the inner setting exists, such as the health system, the community and the policy environment. Constructs include critical incidents, local attitudes, local conditions, partnerships and connections, policies and laws, financing and external pressure. For Cedar Valley, the availability of community groups and transportation in each clinic's area are outer setting conditions.
The setting in which the innovation is implemented. Constructs include structural characteristics, relational connections, communications, culture, tension for change, compatibility, relative priority, incentive systems, mission alignment, available resources, and access to knowledge and information. Several have subconstructs, such as space and funding under available resources, and recipient-centredness under culture. For Cedar Valley, each primary care clinic is an inner setting.
The roles and characteristics of the people involved. The roles include high-level leaders, mid-level leaders, opinion leaders, implementation facilitators, implementation leads, implementation team members, other implementation support, innovation deliverers and innovation recipients. Their characteristics are described with four constructs drawn from the COM-B model introduced in Lesson 2: need, capability, opportunity and motivation. For Cedar Valley, clinicians and connectors are innovation deliverers, clinic managers are mid-level leaders, and older adults are innovation recipients.
The activities and strategies used to implement the innovation. Constructs include teaming, assessing needs, assessing context, planning, tailoring strategies, engaging, doing, reflecting and evaluating, and adapting. For Cedar Valley, whether clinics receive feedback on their own referrals falls under reflecting and evaluating.
CFIR data are usually collected through semi-structured interviews built from the framework's construct definitions, sometimes supplemented by surveys and documents. Transcripts are coded deductively to constructs, and each construct is then rated for each site. Damschroder and Lowery (2013), in an evaluation of a weight management program across Veterans Affairs medical centres, rated each construct for its valence (positive or negative influence on implementation) and strength, on a scale from −2 to +2, with 0 for a neutral influence and X for mixed evidence. Comparing ratings between sites with high and low implementation identified the constructs that distinguished them. The same approach can be used prospectively, before launch, to anticipate barriers.
Three months before the second wave, the evaluation team interviewed the manager and one clinician in each of the 12 second-wave clinics (24 interviews) and the seven first-wave connectors. The interview guide drew questions from CFIR 2.0 constructs chosen with the steering committee. The team defined the innovation as screening, referral and connector support; the inner setting as each clinic; and the outer setting as the clinic's community and the health authority. The outcome the assessment was meant to explain was penetration of the referral step, which had ranged from about 1 to 10 percent across first-wave clinics. Two analysts coded the transcripts deductively, rated constructs, and resolved disagreements by discussion. Table 2.1 shows a selection of the results.
| Domain and construct | Evidence from interviews | Anticipated rating | Implication for the second wave |
|---|---|---|---|
| Innovation: complexity | First-wave clinicians described a referral that required a separate form outside the medical record, and three second-wave clinics use a different record system. | −1 | Build the referral into each record system before launch. |
| Innovation: relative advantage | Clinicians described having had nowhere to send lonely patients before the program. | +2 | Use clinicians' accounts of this advantage in engagement. |
| Outer Setting: local conditions | The four rural clinics serve communities without public transit and at long distances from group activities. | −2 (rural clinics) | Plan transport support and telephone or video meetings. |
| Outer Setting: partnerships and connections | Urban clinics have links with seniors' centres, while rural clinics have few organized groups to connect people to. | X (mixed) | Map and build community partnerships before launch. |
| Inner Setting: available resources (space) | Seven of the 12 clinics have no room for a connector to meet patients. | −1 | Arrange meeting space in community sites. |
| Inner Setting: relative priority | Managers of the two locum-dependent clinics said access to physicians outweighs every other priority. | −2 (two clinics) | Reduce clinician effort and appoint a non-physician champion. |
| Individuals: innovation deliverers (capability) | Clinicians were unsure how to ask about loneliness without causing offence. | −1 | Provide brief, practice-based skills training. |
| Individuals: mid-level leaders | Ten of 12 managers supported the program and offered a staff member as champion. | +1 | Formalize champion roles. |
| Implementation Process: reflecting and evaluating | First-wave clinics never received data on their own referral numbers. | −1 | Provide regular audit and feedback reports. |
The ratings in Table 2.1 are illustrative and summarize the evidence across clinics, noting where a rating applies only to some. A real assessment would report ratings by clinic so that strategies can be tailored to each. The ratings are useful mainly because of the discipline they impose: every barrier must be tied to evidence, placed in a domain, and linked to a decision. Section 3 takes these barriers forward and matches them to implementation strategies.
The Theoretical Domains Framework
CFIR 2.0 spans many levels, from policy to individual clinicians. When the implementation problem is a specific behaviour of health professionals, a framework focused on behaviour can give a sharper diagnosis. The Theoretical Domains Framework (TDF) was developed by Michie and colleagues (2005) through a consensus process that synthesized constructs from behaviour change theories, and Cane, O'Connor and Michie (2012) validated a revised version with 14 domains and 84 constructs. The domains are knowledge; skills; social or professional role and identity; beliefs about capabilities; optimism; beliefs about consequences; reinforcement; intentions; goals; memory, attention and decision processes; environmental context and resources; social influences; emotion; and behavioural regulation. The domains map onto the COM-B model from Lesson 2, which allows a TDF diagnosis to lead into the Behaviour Change Wheel's intervention functions. Atkins and colleagues (2017) published a guide to using the TDF in interviews and analysis.
For Cedar Valley, the behaviour of interest is a clinician asking an older patient about loneliness and making a referral. A TDF analysis might find gaps in knowledge (clinicians unsure of the referral rule), skills (how to ask), beliefs about consequences (fear that the question will offend), memory, attention and decision processes (forgetting during short visits), and social or professional role (uncertainty whether loneliness is a medical concern). CFIR 2.0 would capture some of these under the capability and motivation of innovation deliverers, but with less detail. Birken and colleagues (2017) reviewed studies that combined the two frameworks and described how they complement each other, with CFIR covering the organizational context and the TDF the behaviour of individuals.
2.4 Evaluation Frameworks: RE-AIM and PRISM
RE-AIM
RE-AIM was proposed by Glasgow, Vogt and Boles (1999) to evaluate the public health impact of health promotion programs. Its central argument is that the impact of a program depends on more than its effect on participants: a program that helps people a great deal but reaches few of them, is taken up in few settings, is delivered inconsistently, or disappears when funding ends has limited impact. The framework has five dimensions, assessed at the individual level, the setting level, or both, as Table 2.2 shows.
| Dimension | Level | Definition |
|---|---|---|
| Reach | Individual | The absolute number, proportion and representativeness of eligible individuals who participate. |
| Effectiveness | Individual | The impact of the program on important outcomes, including negative effects, quality of life, and variation across subgroups. |
| Adoption | Setting and staff | The absolute number, proportion and representativeness of settings and staff who are willing to initiate the program. |
| Implementation | Setting and staff | Fidelity to the program's key functions and components, consistency of delivery, adaptations made, and the time and cost of delivery. |
| Maintenance | Individual and setting | At the individual level, the long-term effects six months or more after the last program contact; at the setting level, the extent to which the program becomes part of routine practice. |
The emphasis on representativeness is what distinguishes RE-AIM from a simple count. Reach and adoption ask who takes part and which settings deliver, and whether they differ from those who do not, which makes the framework a natural tool for examining equity. In a review of the framework's first 20 years, Glasgow and colleagues (2019) encouraged pragmatic use, in which evaluators select the dimensions that matter for a particular decision without an obligation to measure all five, and gave greater weight to qualitative methods, adaptations, costs and equity.
PRISM
RE-AIM says what to measure but says little about why the results take the values they do. PRISM, developed by Feldstein and Glasgow (2008), adds contextual domains that are expected to influence RE-AIM outcomes: the intervention as seen from the organizational and the patient perspective; the characteristics of recipients, both organizations and patients; the implementation and sustainability infrastructure; and the external environment. Figure 2.2 shows the structure. In practice, PRISM plays a role similar to a determinant framework within an evaluation, which is why some teams use it in place of CFIR when they want a single integrated approach.
Worked example: RE-AIM for the Cedar Valley second wave at six months
The 12 second-wave clinics have about 21,000 attached patients aged 65 and older. Applying the regional estimate of 24.5 percent gives about 5,145 eligible older adults, with the same caution about the denominator that Section 1.5 raised. Eleven of the 12 clinics signed a participation agreement and made at least one referral within three months; the twelfth, a two-physician clinic that lost a physician shortly before launch, deferred. In the first six months the eleven clinics referred 263 older adults, and 198 attended a first meeting. Table 2.3 sets out the indicators.
| Dimension | Indicator | Calculation | Value |
|---|---|---|---|
| Reach | Eligible older adults who attended a first meeting | 198 ÷ 5,145 | 3.8% |
| Reach | Referred older adults who attended a first meeting | 198 ÷ 263 | 75.3% |
| Reach (representativeness) | Men among participants, compared with men among attached patients aged 65 and older | 31% compared with 45% | Men under-represented |
| Reach (representativeness) | Participants from the four rural clinics, which hold 30% of the eligible population | 36 ÷ 198 | 18.2% |
| Effectiveness | Mean UCLA score among the 151 participants with baseline and twelve-week scores | 7.2 at baseline, 6.5 at follow-up | Pre-post change of −0.7 |
| Adoption (setting) | Clinics that made at least one referral within three months | 11 ÷ 12 | 91.7% |
| Adoption (staff) | Clinicians in adopting clinics who referred at least once | 41 ÷ 74 | 55.4% |
| Implementation | Participants with a co-developed written plan by the third meeting | 167 ÷ 198 | 84.3% |
| Implementation | Participants with at least one active linkage to a community resource | 137 ÷ 198 | 69.2% |
| Maintenance | Clinics still referring six months after launch support ends; loneliness six months after the last meeting | Planned for month 12 and later | Not yet measured |
Read against the first wave, these results show that the less ready second-wave clinics came close to the first wave on reach (3.8 percent of the eligible population compared with 4.3 percent) and on clinician adoption (55.4 percent compared with 59.1 percent). The representativeness indicators show where the shortfall lies. Men and older adults served by rural clinics were under-represented among participants, and the rural clinics contributed 18.2 percent of participants while holding 30 percent of the estimated eligible population. The effectiveness indicator is a pre-post change and carries the limitations Lesson 7 described; RE-AIM asks the evaluator to report effectiveness from the best available design, which for Cedar Valley is the first-wave comparison, and to check whether effects differ across subgroups such as rural participants. PRISM directs attention to contextual explanations for the rural gap, in the external environment (distance and transport) and in recipient characteristics.
2.5 Choosing and Combining Frameworks
No single framework serves every purpose, and evaluations commonly combine one framework from each of Nilsen's aims. Birken and colleagues (2018) developed the Implementation Theory Comparison and Selection Tool (T-CaST), which helps teams compare candidate frameworks against criteria such as usability, testability and applicability to their setting. Two cautions apply. A framework should shape data collection from the start; when it is applied only to label findings at the end, it adds vocabulary without adding explanation.
The Proctor outcomes and RE-AIM overlap in ways that are worth stating explicitly. Table 2.4 maps them onto each other.
| Proctor implementation outcome | Closest RE-AIM dimension | Comment |
|---|---|---|
| Penetration | Reach | Both divide participants by the eligible population; RE-AIM adds representativeness. |
| Adoption | Adoption | Both apply to settings and staff. |
| Fidelity and implementation cost | Implementation | RE-AIM also includes adaptations and consistency of delivery. |
| Sustainability | Maintenance (setting level) | RE-AIM adds maintenance of individual outcomes. |
| Acceptability, appropriateness and feasibility | No direct equivalent | These perceptions can be measured within implementation or treated as PRISM context. |
For the Cedar Valley second wave, a coherent combination would use EPIS to describe the phases and the bridging arrangements, CFIR 2.0 to assess determinants before launch and again at six months, and RE-AIM, with PRISM to interpret context, to evaluate the rollout, with the Proctor perception measures (acceptability, appropriateness and feasibility) added to the implementation dimension. Section 4 assembles this combination into a worked example of the implementation component of the Cedar Valley evaluation plan.
For each question, name the category in Nilsen's taxonomy and a framework that fits: (1) the phases the second-wave clinics will pass through; (2) why some clinicians in a well-resourced clinic rarely ask about loneliness; (3) the share of eligible older adults who took part, and whether they were representative; (4) the features of rural communities that will hold back referrals. Suggested answers: (1) a process model, EPIS; (2) a determinant framework focused on behaviour, the TDF; (3) an evaluation framework, RE-AIM; (4) a determinant framework, CFIR 2.0, or PRISM's contextual domains.
Reflection
A health authority launched a falls-prevention referral program in community pharmacies. Pharmacists screen adults aged 65 and older for falls risk and refer those at risk to a free community exercise class. Twenty pharmacies were invited; 15 signed participation agreements, and 12 made at least one referral within three months. In those 12 pharmacies, 38 of 60 pharmacists referred at least once. The pharmacies serve an estimated 3,000 eligible older adults. In six months, 420 were referred and 260 attended at least one class. Interviews produced five findings: (i) pharmacists said the screening tool has too many steps to complete during busy dispensing; (ii) two rural pharmacies are 40 kilometres from the nearest class and there is no bus service; (iii) pharmacy managers said dispensing volume always comes first and the program ranks low among their priorities; (iv) pharmacists who attended the training felt confident raising falls with customers; (v) no pharmacy had received any information about its own referral numbers.
CFIR 2.0 has five domains: Innovation (characteristics of the program itself, including its complexity), Outer Setting (the community and system around the organization, including local conditions), Inner Setting (the organization where the program is delivered, including relative priority), Individuals (the roles and characteristics of people involved, including the capability of innovation deliverers), and Implementation Process (the activities used to implement, including reflecting and evaluating). RE-AIM defines reach as the number, proportion and representativeness of eligible individuals who participate, and adoption as the number, proportion and representativeness of settings and staff who initiate the program.
(a) Assign each of the five findings to a CFIR 2.0 domain and construct. (b) Calculate reach, the share of referred older adults who attended, setting adoption and staff adoption. (c) Name one further piece of information you would need to judge representativeness, and explain why.
(a) Finding (i) belongs to the Innovation domain under complexity. Finding (ii) belongs to the Outer Setting under local conditions, since distance and the absence of transit lie outside the pharmacy. Finding (iii) belongs to the Inner Setting under relative priority. Finding (iv) belongs to the Individuals domain under the capability of innovation deliverers, and it is a facilitator. Finding (v) belongs to the Implementation Process under reflecting and evaluating.
(b) Reach is 260 divided by 3,000, or 8.7 percent of eligible older adults. Of those referred, 260 of 420, or 61.9 percent, attended a class, and referral penetration was 420 of 3,000, or 14.0 percent. Setting adoption, defined by making a referral, is 12 of 20 invited pharmacies, or 60 percent; signing an agreement alone would give 15 of 20, or 75 percent, so the definition should be stated. Staff adoption is 38 of 60 pharmacists in adopting pharmacies, or 63.3 percent.
(c) I would need the characteristics of participants compared with all eligible older adults, such as age, sex, language and rural residence. Reach of 8.7 percent could hide large gaps, for example if almost no participants came from the rural pharmacies. I would also compare adopting and non-adopting pharmacies, because if rural or independent pharmacies adopt less, the program's reach will be unequal before any older adult is screened.
Minimum 20 characters required.
Question 1: In Nilsen's taxonomy, CFIR 2.0 is best classified as which kind of approach?
Question 2: Which of the following is a RE-AIM adoption indicator for the Cedar Valley second wave?
Question 3: In a CFIR 2.0 assessment, where does the finding that the four rural clinics serve communities without public transit belong?
Question 4: What does PRISM add to RE-AIM?
Implementation Strategies and Adaptation
Learning Objectives for this section
- Define an implementation strategy and distinguish discrete, multicomponent and blended strategies from the program they support.
- Describe the ERIC compilation and its nine clusters, and select strategies that match barriers identified with CFIR 2.0.
- Specify an implementation strategy using the seven dimensions recommended by Proctor, Powell and McMillen.
- Distinguish a program's core functions from its forms and document adaptations with FRAME, including whether each is fidelity-consistent.
- Distinguish horizontal from vertical scale-up and describe how sustainability is defined, assessed and planned.
3.1 What an Implementation Strategy Is
Proctor, Powell and McMillen (2013) defined implementation strategies as the methods or techniques used to enhance the adoption, implementation and sustainability of a clinical program or practice. Powell and colleagues (2012) distinguished three kinds. A discrete strategy is a single action or process, such as a reminder in the electronic medical record. A multicomponent strategy combines two or more discrete strategies, such as training plus audit and feedback. A blended strategy interweaves several discrete strategies into a protocolized package that is often given a name and tested as a unit. Most implementation efforts in health services use multicomponent or blended strategies, because the barriers they face sit at several levels at once.
A practical difficulty is deciding where the program ends and the strategy begins. For Cedar Valley, the connector meetings, the co-developed plan and the linkage to community resources are clearly part of the program. Training clinicians to ask about loneliness, adding a referral template to the medical record, and sending clinics reports on their referral rates are clearly implementation strategies. Screening is harder to classify: it is part of the program's referral pathway, yet a prompt that reminds clinicians to screen is a strategy. The boundary matters because an evaluation needs to know which components it is testing and which it is holding constant. The logic model from Lesson 3 should state the boundary explicitly, and the CFIR 2.0 definition of the innovation in Section 2 should match it.
3.2 The ERIC Compilation
To give the field a common vocabulary, Powell and colleagues (2015) ran the Expert Recommendations for Implementing Change (ERIC) project, a modified Delphi process with implementation scientists and clinical experts that produced a compilation of 73 discrete implementation strategies, each with a name and a definition. Waltz and colleagues (2015) then asked experts to sort the strategies by similarity and to rate their importance and feasibility, and used concept mapping to group them into nine clusters. The cards below describe each cluster with examples of its strategies and a possible use in the Cedar Valley second wave. Two cluster labels are paraphrased here because the originals use an older term for interest holders.
The compilation is a vocabulary, and it carries no evidence ratings. Evidence about the effectiveness of particular strategies comes from other sources, such as systematic reviews of audit and feedback or of practice facilitation. Baskerville, Liddy and Hogg (2012), for example, reviewed trials of practice facilitation in primary care and found that facilitation improved the adoption of evidence-based guidelines, which supports its use where clinic workflows are the main barrier.
3.3 Selecting and Tailoring Strategies
Strategies should be chosen to address the determinants that stand in the way of implementation, which is why a determinant assessment usually comes first. Waltz and colleagues (2019) asked implementation experts to name the ERIC strategies they would use to address each CFIR barrier and published the results as the CFIR-ERIC matching tool. Experts' recommendations varied considerably, so the tool's output is best used as a starting list for a team's own judgement. Powell and colleagues (2017) reviewed structured methods for selecting and tailoring strategies, including concept mapping, group model building, conjoint analysis and intervention mapping, the last of which Lesson 2 introduced as a planning framework. All of these methods share three elements: they identify barriers, they link each barrier to a strategy through an explicit rationale, and they involve the people who will deliver the program.
The evaluation team took the barriers in Table 2.1 to a workshop with the steering committee, two first-wave champions, three second-wave managers and the connectors. For each barrier, the group reviewed the strategies suggested by the CFIR-ERIC matching tool, discussed their feasibility in Cedar Valley, and agreed on a set. Table 3.1 shows the result.
| Barrier (CFIR 2.0 construct) | ERIC strategy | Form in Cedar Valley |
|---|---|---|
| Complexity of referral | Change record systems; remind clinicians | A one-step referral template and a screening prompt in each record system |
| Capability of clinicians to ask about loneliness | Conduct educational meetings; make training dynamic | A one-hour session with role-plays, delivered in each clinic |
| Relative priority in locum-dependent clinics | Identify and prepare champions; revise professional roles | A nurse or office assistant champion, with screening moved to office staff |
| No feedback on referrals | Audit and provide feedback | A quarterly clinic report of referrals per estimated eligible patient |
| No space for connectors | Change service sites | Meetings in libraries, community centres and participants' homes |
| Rural transport and few groups | Access new funding; build a coalition; promote adaptability | A larger rural transport fund, a rural partners' network, and telephone or video meetings |
| Several barriers in each clinic | Facilitation | A practice facilitator who supports each clinic through launch |
3.4 Specifying Strategies
Reports of implementation studies have often described strategies too vaguely to be replicated or compared, with labels such as "training" or "support" that could describe almost anything. Proctor, Powell and McMillen (2013) recommended that authors name each strategy, define it, and specify it on seven dimensions: the actor who delivers it; the action taken; the action target, meaning the people, unit or determinant it is meant to change; the temporality, meaning when it is used; the dose, meaning its frequency and intensity; the implementation outcome it is expected to affect; and the justification for its use, drawn from theory, evidence or practical reasoning. Lewis and colleagues (2018) added that the justification should name the mechanism through which the strategy is expected to work, so that a study can test whether the strategy succeeded or failed for the hypothesized reason.
| Dimension | Specification of practice facilitation for the Cedar Valley second wave |
|---|---|
| Name and definition | Practice facilitation: a trained facilitator helps each clinic team adapt its workflow to screen and refer, solve problems, and use referral data to improve. |
| Actor | A half-time practice facilitator employed by the program, with primary care experience and training in quality improvement, supervised by the program coordinator. |
| Action | Meets the clinic champion and manager to map the referral workflow, agree changes, review the quarterly audit and feedback report, and set next steps. |
| Action target | The clinic team (clinicians, office staff and manager), and the determinants of complexity, relative priority, deliverer capability, and reflecting and evaluating. |
| Temporality | Begins two months before a clinic launches and continues for six months after launch. |
| Dose | One 60-minute meeting before launch, six fortnightly visits of 30 minutes in the first three months, and three monthly visits of 30 minutes in the next three months: 5.5 hours of contact per clinic, or 60.5 hours across eleven clinics. |
| Implementation outcomes affected | Clinician-level adoption and penetration of the referral step (primary); completeness of referral information (secondary). |
| Justification and mechanism | The CFIR assessment found barriers in workflow, priority, capability and feedback; facilitation has improved guideline adoption in primary care (Baskerville et al., 2012); the hypothesized mechanism is that facilitation increases the team's capability and embeds referral in routine workflow. |
Specification also makes the strategy's cost visible. The second-wave implementation support package consists of the half-time facilitator ($52,000 for the year), clinician training sessions ($8,000) and the production of audit and feedback reports ($4,000), for a total of $64,000. Across the eleven adopting clinics, this is about $5,818 per clinic, which is the implementation cost indicator proposed in Section 1. The cost is separate from the cost of delivering the program, the connectors' salaries and the transport fund, and Lesson 10 shows how both enter a program costing and a budget impact analysis. When a strategy changes during a study, for example when visits are moved online, the change should itself be documented; Miller and colleagues (2021) extended the FRAME approach described next to modifications of implementation strategies, as FRAME-IS.
3.5 Adaptation: Core Functions, Forms and FRAME
Programs are almost always changed when they move into new settings. Some changes improve fit and preserve the program's effects; others remove the components that make it work. The older view treated any departure from the protocol as a threat to fidelity. A more useful view, set out by Hawe, Shiell and Riley (2004) for complex interventions, holds that what should be standardized is the function of each component, the purpose it serves in the program's theory of change, while its form, the specific activity that serves the purpose, can vary with context. Jolles, Lengnick-Hall and Mittman (2019) developed this distinction of core functions and forms for health services interventions. An adaptation that changes a form while preserving the function is fidelity-consistent; one that removes or undermines a function is fidelity-inconsistent.
| Core function | First-wave form | Permitted second-wave forms |
|---|---|---|
| A person-centred conversation that identifies what matters to the older adult and what stands in the way of connection | A one-to-one meeting, usually in a clinic room | A one-to-one meeting at home, in a community site, or by telephone or video |
| A co-developed, written connection plan | A paper plan in English | A plan in the participant's preferred language, on paper or by email |
| Active linkage to at least one community resource | An introduction or accompaniment to a group | An introduction, accompaniment, volunteer driver, or land-based activity within the co-designed pathway |
| Follow-up to address barriers and revise the plan | Up to six meetings over twelve weeks | Up to six contacts over twelve weeks, in any of the permitted formats |
Wiltsey Stirman, Baumann and Miller (2019) published FRAME, an expanded framework for reporting adaptations and modifications to evidence-based interventions, which built on an earlier coding system by Stirman and colleagues (2013). FRAME records eight aspects of each modification: when and how in the implementation process it was made; whether it was planned and proactive or unplanned and reactive; who decided to make it; what was modified (content, context, training or evaluation); at what level of delivery it was made, such as an individual participant, a clinic or the whole program; the type of change, such as adding, removing, shortening, substituting or reordering elements, or changing the format, setting or personnel; whether it was fidelity-consistent; and the reasons for it, including its goal and the contextual factors behind it. The accordion below logs three second-wave adaptations with FRAME.
When and how: during implementation, in month two, after several rural participants missed meetings. Planned or unplanned: unplanned and reactive, then formalized as a permitted form. Who decided: the rural connectors with the program coordinator. What was modified: context, specifically the format of delivery. Level: participants in the four rural clinics. Type: a change of format from mainly in-person to mainly remote meetings, with the first meeting kept in person where possible. Fidelity: consistent, because all four core functions are preserved. Reasons: to increase reach and retention; the contextual factors were distance, the lack of public transit and winter road conditions.
When and how: during implementation, in month four, at one large urban clinic. Planned or unplanned: unplanned and reactive. Who decided: one connector, without consultation. What was modified: content and context. Level: new referrals from that clinic. Type: substitution of a group session for the first one-to-one meeting. Fidelity: probably inconsistent, because a group session is unlikely to deliver a person-centred conversation about what matters to each person. Reasons: to manage a waiting list; the contextual factor was a caseload above the connector's capacity. Response: the coordinator kept the group session as an optional addition, restored the one-to-one first meeting, and redistributed referrals across connectors, an example of how FRAME records can prompt a timely correction.
When and how: co-designed before the second wave and introduced during it. Planned or unplanned: planned and proactive. Who decided: the First Nation's health centre and community members, with the connector hosted by the health centre and the program team in a supporting role. What was modified: content and context. Level: participants from that community who choose the pathway. Type: addition of land-based and cultural activities as linkage options, delivered with Elders and community members. Fidelity: consistent, because the pathway delivers the core functions through culturally grounded forms. Reasons: cultural safety, appropriateness and self-determination. The documentation of this adaptation, and any data about it, follow the agreements made with the Nation, consistent with the OCAP® principles and relational accountability discussed in Lesson 4.
A FRAME log is cheap to keep if it is built into routine work, for example as a short form that connectors complete when they change how they deliver the program, reviewed monthly by the coordinator. The log supports two analyses. It shows the fidelity-inconsistent changes that need correction, and it allows the evaluation to ask whether adaptations were associated with reach or outcomes, keeping in mind that such associations are observational.
3.6 Scale-Up and Sustainability
Scale-up
Scale-up refers to deliberate efforts to increase the impact of health innovations that have been tested successfully, so that they benefit more people and become part of policy and programs on a lasting basis. The World Health Organization and ExpandNet (2010) distinguished two directions. Horizontal scale-up, also called expansion or replication, extends a program to more sites or populations. Vertical scale-up, also called institutionalization, embeds the program in policy, budgets, regulations and organizational structures. The Cedar Valley second wave is horizontal scale-up, from 12 clinics to 24; moving the connector positions from time-limited funding into the health authority's base budget would be vertical scale-up. Figure 3.1 shows the two directions.
Before scaling up, a program's scalability can be assessed. Milat and colleagues (2013) described scalability as the ability of an intervention shown to be efficacious on a small scale or under controlled conditions to be expanded under real-world conditions to reach a greater proportion of the eligible population while retaining its effectiveness, and their group later developed the Intervention Scalability Assessment Tool (Milat et al., 2020) to structure the judgement around the problem, the intervention, the strategic and political context, the evidence of effectiveness, the costs, and the infrastructure needed. For Cedar Valley, the main scalability questions are whether the rural communities have enough groups and services to connect people to, and whether connector caseloads remain manageable as referrals grow.
Sustainability
Moore, Mascarenhas, Bain and Straus (2017) reviewed definitions of sustainability and proposed one with five elements: after a defined period of time, the program, clinical intervention or implementation strategies continue to be delivered, or individual behaviour change is maintained; and the program and behaviour change may evolve or adapt while continuing to produce benefits for individuals or systems. The definition accepts change as part of sustainability, which is consistent with the Dynamic Sustainability Framework of Chambers and colleagues (2013). That framework argues that programs should be expected to change over time, through continuous measurement and refinement of their fit with changing settings, and that sustainment should be planned from the start. Shelton, Cooper and Wiltsey Stirman (2018) reviewed the determinants of sustainability in public health and health care and identified factors at the levels of the program, the organization, the outer context and the implementation process, many of which overlap with CFIR 2.0 constructs.
Sustainability can be assessed with the Program Sustainability Assessment Tool (PSAT) of Luke and colleagues (2014), which rates a program's capacity for sustainability across eight domains: environmental support, funding stability, partnerships, organizational capacity, program evaluation, program adaptation, communications and strategic planning. For Cedar Valley, the steering committee completed the PSAT at the end of the first year. Partnerships and program evaluation scored well, while funding stability scored poorly, because the connector positions are funded from a time-limited allocation. The sustainability plan therefore centres on vertical scale-up: presenting the first-wave evaluation and the second-wave RE-AIM results to the health authority's executive with a proposal to move the positions into the base budget, and building connector referral into regional primary care policy.
The program coordinator proposes three changes for the third year. (1) Replace the written connection plan with a verbal agreement, because writing plans takes time. (2) Allow trained volunteers to make the follow-up calls in the last six weeks, with connectors reviewing each call. (3) Offer the first meeting in a group format to anyone who prefers it, while keeping a one-to-one option. For each change, state which core function it touches, whether it is likely to be fidelity-consistent, and what you would record under each FRAME element. A strong answer will treat change 1 as likely fidelity-inconsistent, change 2 as a change of personnel that may be consistent if follow-up still revises the plan, and change 3 as consistent only if every participant still has a one-to-one conversation.
Section 4 turns to study design and shows how a single study can test a program's effectiveness and its implementation at the same time.
Reflection
A health authority is implementing a tobacco cessation program in 16 hospital units: nurses ask every admitted patient about tobacco use and offer smokers a referral to a provincial quitline. Interviews identified three barriers: (1) nurses forget to ask during busy admissions; (2) unit managers rank the program low among their priorities; (3) nurses are unsure how to raise smoking without sounding judgemental. The program's core function is that every admitted patient is asked about tobacco use and every smoker is offered a referral to cessation support.
Available ERIC strategies: remind clinicians (prompts at the point of care); audit and provide feedback (performance data returned to staff); identify and prepare champions (people who promote the program within a unit); conduct educational meetings (training sessions); change record systems (new fields or templates in the record); facilitation (a trained facilitator helps units solve problems); alter incentive structures; and mandate change.
Proctor, Powell and McMillen's seven dimensions are actor (who delivers the strategy), action (what is done), action target (who or what it aims to change), temporality (when), dose (how often and how intensely), implementation outcome affected, and justification. FRAME records when the modification was made, whether it was planned or unplanned, who decided, what was modified, the level of delivery, the type of modification, whether it is fidelity-consistent, and the reasons for it.
(a) Choose one strategy for each barrier and justify it. (b) Specify one of your strategies on all seven dimensions. (c) Two months after launch, one unit replaced nurse questioning with a tablet questionnaire that patients complete at admission, with nurses offering referrals to smokers it identifies. Record this modification using the FRAME elements and judge whether it is fidelity-consistent.
(a) For forgetting, I would change record systems to add a mandatory tobacco question to the admission template, which acts as a reminder at the point of care. For low priority, I would identify and prepare a champion on each unit and pair this with audit and feedback, so that managers see their unit's performance. For uncertainty about raising smoking, I would conduct brief educational meetings with practice in non-judgemental wording.
(b) Audit and feedback: the actor is the health authority's quality analyst; the action is a monthly one-page report showing each unit's share of admissions asked and smokers referred, compared with other units; the action targets are unit managers and nurses, and the determinants are relative priority and reflecting on performance; temporality runs from the first month after launch for twelve months; the dose is one report and a ten-minute discussion at a monthly unit meeting; the outcomes affected are fidelity (share asked) and penetration (share of smokers referred); the justification is that managers ranked the program low and had no data on it.
(c) When: during implementation, at two months. Unplanned and reactive. Decided by the unit manager. What: context, the format and personnel of the question. Level: one unit. Type: substitution of a patient questionnaire for a nurse's question. Reasons: workload. It is likely fidelity-consistent, because the core function, asking everyone and offering referral to every smoker, is preserved, provided patients who cannot use a tablet are still asked by a nurse.
Minimum 20 characters required.
Question 1: Which of the following is an implementation strategy for the Cedar Valley program?
Question 2: A specification states that practice facilitation begins two months before a clinic launches and continues for six months after launch. Which of Proctor, Powell and McMillen's dimensions does this describe?
Question 3: A connector replaces each new participant's first one-to-one meeting with a group information session because her caseload is full. Which judgement about this modification is most defensible under FRAME?
Question 4: Moving the Cedar Valley connector positions from time-limited funding into the health authority's base budget is an example of what?
Hybrid Effectiveness-Implementation Designs
Learning Objectives for this section
- Explain why hybrid effectiveness-implementation studies were proposed and define hybrid types 1, 2 and 3.
- Describe the changes recommended in the 2022 update by Curran and colleagues, including how a study's type is determined and labelled.
- Design a version of the Cedar Valley evaluation as each hybrid type, stating its aims, outcomes, design and unit of analysis.
- Calculate the approximate number of clinics needed for a hybrid type 3 study powered on a clinic-level implementation outcome.
- Identify the reporting requirements of StaRI and select an implementation framework and implementation outcomes for a program.
4.1 Why Combine Effectiveness and Implementation Questions
Section 1 described the conventional sequence of research, in which efficacy trials are followed by effectiveness trials and then, often years later, by implementation studies. Curran, Bauer, Mittman, Pyne and Stetler (2012) argued that this sequence is slow, and that it produces effectiveness evidence under conditions that may not resemble those of later implementation. They proposed effectiveness-implementation hybrid designs, studies that take a dual focus on effectiveness and implementation, as a way to speed translation, to learn about implementation while effectiveness is still being established, and to generate evidence that is more useful to the people who decide whether to adopt a program. Effectiveness-implementation hybrid designs are unrelated to the hybrid observational designs of epidemiology, such as the case-cohort and nested case-control designs, which combine features of cohort and case-control sampling.
The idea is especially relevant to programs like Cedar Valley. Health authorities rarely wait for a full sequence of trials before rolling out a social prescribing program, and the evaluator is usually asked to learn as much as possible from a rollout already under way. A hybrid study makes that learning systematic, by stating in advance which questions about effectiveness and which questions about implementation the study will answer, and with what priority.
4.2 The Three Hybrid Types
Curran and colleagues (2012) defined three types along a continuum, shown in Figure 4.1.
A hybrid type 1 study tests the effects of an intervention on relevant outcomes as its primary aim, while observing and gathering information on implementation as a secondary aim. The secondary aim typically examines barriers and facilitators, the acceptability and feasibility of the intervention in its setting, and the potential for wider implementation. Type 1 suits interventions whose effectiveness is not yet established in the target setting, where the intervention carries little risk and there is reason to prepare for implementation if it proves effective.
A hybrid type 2 study has co-primary aims: it tests the intervention's effects and also tests or studies an implementation strategy. Type 2 suits interventions with some evidence of effectiveness, perhaps in related populations or settings, where there is a reasonable expectation that a particular implementation strategy is feasible and supportable in the setting.
A hybrid type 3 study tests an implementation strategy as its primary aim, while observing and gathering information on the intervention's clinical or health outcomes as a secondary aim. Type 3 suits interventions with strong effectiveness evidence, where the main uncertainty concerns how best to implement them and whether their effects hold up when they are delivered in new settings, the voltage drop of Section 1.
| Type | Primary aim | Secondary aim | Suited to |
|---|---|---|---|
| Hybrid type 1 | Test the intervention's effects on health outcomes | Gather information on implementation, such as barriers, facilitators and acceptability | Limited effectiveness evidence in the target setting, low risk |
| Hybrid type 2 | Test the intervention's effects and test or study an implementation strategy (co-primary) | Further implementation and outcome questions | Some effectiveness evidence and a strategy ready to be tested |
| Hybrid type 3 | Test an implementation strategy on implementation outcomes | Gather information on the intervention's health outcomes | Strong effectiveness evidence, with uncertainty about implementation |
4.3 The 2022 Update
A decade after the original paper, Curran, Landes, McBain, Pyne, Smith, Fernandez, Chambers and Mittman (2022) reviewed how the types had been used and recommended several changes. They proposed describing hybrid studies in place of hybrid designs, because the term "design" had caused confusion and a hybrid can use many research designs, including non-randomized ones. A study's type is set primarily by its aims and research questions, and secondarily by its outcomes, and the research design is a separate choice. They therefore recommended that authors name both, as in "a parallel cluster randomized hybrid type 3 effectiveness-implementation study".
The update clarified each type. In type 1 studies, power is usually calculated for participant-level health outcomes, implementation outcomes are often not measured formally, and the implementation aim typically explores the intervention's potential for implementation. Type 2 studies should genuinely test an implementation strategy that is deployed in real-world conditions; dual randomization of both the intervention and the strategy is ideal but uncommon, and many type 2 studies randomize individuals to the intervention while applying a strategy to all sites, or study the strategy in a single-arm pre-post design. Type 3 studies should be powered on implementation outcomes, which are usually measured at the level of the site, the implementer or the system, and their strategies should be delivered by or built into the local system wherever possible.
The authors also offered four considerations for choosing a type: the strength of the existing evidence for the intervention's effectiveness; how much adaptation the intervention is expected to need in the new setting; how much is known about the determinants of its implementation; and whether the team is ready to test implementation strategies. Strong evidence and little expected adaptation point toward type 3; weak evidence and little knowledge of determinants point toward type 1. They noted that a hybrid study is not indicated when an intervention has not yet been shown to be safe, when no question about its effectiveness remains for the setting, or when a study examines only the determinants of implementation. Because the types lie on a continuum, authors should justify their choice and explain why the neighbouring types were not selected.
4.4 The Cedar Valley Program as Each Hybrid Type
The same program can be studied as any of the three types at different points in its history, depending on what is known and what decision is being made. The tabs below describe a version of the Cedar Valley evaluation for each type.
Situation. At the first wave, the only local evidence came from the 18-month pilot in two clinics, and the international evidence for social prescribing was limited. The program carries little risk, and the health authority had already committed to a second wave.
Aims. The primary aim is to estimate the program's effect on loneliness at twelve weeks. As Lesson 7 described, the first-wave clinics were chosen for readiness, so the comparison is between older adults in first-wave clinics and similar older adults in second-wave clinics during the first year, analyzed with the comparison-group and difference-in-differences methods of that lesson. The secondary aim is to gather implementation information for the second wave: CFIR 2.0 interviews with clinicians and connectors, the Acceptability, Appropriateness and Feasibility measures, and referral data by clinic.
Label. A non-randomized comparison-group hybrid type 1 effectiveness-implementation study. The CFIR findings in Table 2.1 are the product of its secondary aim.
Situation. By the second wave, the first-wave evaluation has produced moderate, non-randomized evidence of benefit. The second-wave clinics are less ready, and a larger share of their eligible older adults are served by rural clinics, so substantial adaptation is expected, and the CFIR assessment has identified determinants and a tailored strategy bundle (Table 3.1) that the health authority wants to test.
Aims. The first co-primary aim is to test effectiveness in the new setting. Because connector capacity in the second wave is limited and a waiting list is likely in any case, referred older adults could be randomized to start connector support immediately or after twelve weeks, the waitlist design discussed in Lesson 6, with loneliness at twelve weeks compared between groups. The second co-primary aim is to test the strategy bundle, which every clinic receives, against a pre-specified benchmark: clinic-level penetration of the referral step of at least 5.5 percent at six months, the first-wave level. Adoption, fidelity and implementation cost are further implementation outcomes, reported with RE-AIM as in Table 2.3.
Label. An individually randomized hybrid type 2 effectiveness-implementation study with a single-arm implementation aim. Its main weakness is that the implementation aim has no comparison group, so a rise or fall in penetration cannot be attributed to the bundle with confidence. A dual-randomized version would also randomize clinics to the bundle or a standard launch, but with 12 clinics such a comparison would be underpowered, as the calculation below shows.
Situation. Suppose that, after both waves, the program's effectiveness is well established and the model is offered to 40 primary care clinics in other regions of the province. The decision-makers now need to know which implementation strategy to fund.
Aims. The primary aim is to compare two strategies: a standard launch (training, a referral template and a toolkit) and an enhanced launch (the standard launch plus practice facilitation and audit and feedback), with clinic-level penetration of the referral step at six months as the primary outcome. Clinics are randomized, 20 to each strategy. Secondary aims observe loneliness change among participants, to check that effects hold in the new regions, along with fidelity and cost.
Label. A parallel cluster randomized hybrid type 3 effectiveness-implementation study. Because penetration is a clinic-level outcome, the number of clinics determines the study's power.
How many clinics does a type 3 study need?
Suppose first-wave clinic-level penetration at six months had a standard deviation of about 2 percentage points across clinics, and the enhanced strategy is expected to raise penetration by 2.5 percentage points (for example, from 5.5 to 8.0 percent). Treating each clinic's penetration as one observation, the approximate number of clinics per arm for a two-sided test at the 5 percent level with 80 percent power is:
n per arm = 2 × (z0.975 + z0.80)2 × SD2 ÷ difference2 = 2 × (1.96 + 0.84)2 × 22 ÷ 2.52 = 2 × 7.84 × 4 ÷ 6.25 = 10.04, which rounds up to 11 clinics per arm, or 22 in total.
The calculation uses the normal approximation, which understates the requirement slightly when the number of clusters is small, and it ignores variation in clinic size. With 40 clinics, the type 3 study has room for these corrections and for some clinics to withdraw. A comparison within the 12 second-wave clinics alone, at 6 per arm, would fall well short, which is why the type 2 tab treats the strategy bundle as a single-arm aim. Lesson 6 explained the related design effect for individual-level outcomes in cluster trials.
Other designs can test implementation strategies. Brown and colleagues (2017) reviewed the options, which include parallel and stepped-wedge cluster randomized trials, interrupted time series, and adaptive designs. Sequential multiple assignment randomized trials, introduced in Lesson 6, can test adaptive implementation strategies: Kilbourne and colleagues (2014) described a trial in which sites that did not respond to a standard implementation strategy were randomized to receive additional facilitation of different intensities. In a Cedar Valley version, clinics with penetration below 3 percent at three months could be randomized to add practice facilitation or facilitation plus a funded champion, so that the more intensive and costly support goes only to the clinics that need it. Where randomization is not possible, the interrupted time series methods of Lesson 8, applied to monthly referral counts, can test whether a strategy changed the level or trend of referrals.
4.5 Reporting Implementation Studies with StaRI
Implementation studies report two things at once, the implementation strategy and the intervention it supports, and readers need to know about both. The Standards for Reporting Implementation Studies (StaRI) statement of Pinnock and colleagues (2017) was developed through a systematic review and an international e-Delphi process. It contains 27 items, and for many of them it asks authors to report separately on the implementation strategy and on the intervention, which the checklist sets out as two parallel columns. The items cover the title and abstract, the rationale and context, the specification of the strategy and the intervention, the outcomes for each, their fidelity and adaptation, resource use, and the interpretation and generalizability of findings. StaRI applies to implementation studies with a wide range of designs and is used alongside design-specific guidelines: the CONSORT statement and its cluster extension for randomized studies (Lesson 6), SQUIRE 2.0 for quality improvement studies (Lesson 8), and TIDieR for describing interventions (Lesson 10). Proctor's specification dimensions and FRAME logs supply much of what StaRI asks for.
Two columns in practice
For the Cedar Valley type 2 study, the methods section of a StaRI-compliant report would describe the strategy bundle (facilitation, audit and feedback, champions, training, record changes, community sites and transport support) with its specification, and separately describe the connector program with its core functions and permitted forms. The results would report adoption, penetration and the cost of the bundle in one strand, and loneliness outcomes and fidelity to the core functions in the other, with the FRAME log summarizing adaptations to each.
4.6 Worked Example: An Implementation Framework and Outcomes for Cedar Valley
The implementation component of an evaluation plan selects an implementation framework for the program and specifies the implementation outcomes to be measured. The Cedar Valley version follows.
Framework choice and justification. The evaluation of the second wave will use CFIR 2.0 as its determinant framework and RE-AIM, interpreted with PRISM, as its evaluation framework. CFIR 2.0 was chosen because the first-wave experience showed that barriers sit at several levels at once (the referral process, clinic priorities and space, clinicians' skills, and rural community resources), and because its Outcomes Addendum ties the determinant assessment to the outcomes it is meant to explain. RE-AIM was chosen because the health authority's question concerns population impact across 24 clinics, which depends on reach and adoption as much as on effectiveness, and because its representativeness indicators allow the evaluation to track rural older adults and men, the groups the first wave reached least. EPIS will be used descriptively to report the phase of each clinic. A process model alone was not chosen because the steps of the rollout are already set; the open questions concern determinants and impact.
Study type. The second-wave evaluation is a hybrid type 2 study, with co-primary aims for loneliness and for clinic-level penetration, as described in Section 4.4.
| Implementation outcome | Indicator and data source | Timing and level |
|---|---|---|
| Adoption | Clinics making at least one referral within three months, divided by clinics invited; clinicians referring at least once, divided by clinicians in adopting clinics (referral records) | Months 3 and 6; clinic and clinician |
| Penetration (co-primary) | Referrals divided by patients aged 65 and older who screened 6 or higher, using the new screening field in the medical record | Monthly, with the primary test at month 6; clinic |
| Fidelity | Participants with a co-developed written plan by the third meeting and at least one active linkage, divided by participants (connector logs, with a 10 percent audit of records) | Quarterly; participant and connector |
| Acceptability, appropriateness and feasibility | Mean scores on the four-item AIM, IAM and FIM among clinicians, office staff and connectors | Month 3; individual deliverers |
| Implementation cost | Cost of the strategy bundle per adopting clinic (finance records and facilitator time logs) | Month 12; program |
| Sustainability | Clinics referring at or above their month-6 rate twelve months after facilitation ends; PSAT ratings by the steering committee | Month 18 and later; clinic and program |
The plan keeps the implementation outcomes separate from the effectiveness outcomes (loneliness, social participation, self-rated health and service use), which remain in the evaluation matrix from Lesson 5. It also states the data each indicator needs before launch, which is why the screening field in the medical record and the connector log appear in the plan. The CFIR 2.0 assessment will be repeated at six months, and its ratings compared between clinics above and below the penetration benchmark, to explain the implementation results that RE-AIM reports.
What an implementation component contains
An implementation component, often about two pages of an evaluation plan, names one or two implementation frameworks, classifies each with Nilsen's taxonomy, and justifies the choice in a paragraph that refers to the program's setting and evaluation questions. It then specifies four to six implementation outcomes in a table that gives each outcome's indicator (with numerator and denominator), data source, timing and level of measurement, and it states whether the evaluation is, or could become, a hybrid study, and if so which type and why.
A sound component chooses a framework that fits the purpose it is asked to serve and justifies it with reference to the program's context. It defines each implementation outcome in the terms of Proctor and colleagues and keeps it distinct from the effectiveness outcomes, states the numerator and denominator of each indicator, and relies on data sources that exist or can be created before launch. At least one outcome addresses representativeness or equity across the groups the program serves, and any hybrid type is justified with reference to the strength of the effectiveness evidence and the expected adaptation.
Reflection
Suppose that one randomized trial in another province found that a peer-led diabetes self-management program improved blood glucose control among South Asian adults. A health authority in British Columbia wants to deliver the program in 16 community centres, with peer leaders recruited locally and with substantial changes to language, food examples and session timing. Little is known about the barriers to delivering it in British Columbia community centres. The health authority has designed a facilitation strategy for the centres and wants to know whether it works.
Hybrid effectiveness-implementation studies come in three types. In type 1, the primary aim tests the intervention's effects on health outcomes and a secondary aim gathers information on implementation. In type 2, co-primary aims test the intervention's effects and test an implementation strategy. In type 3, the primary aim tests an implementation strategy on implementation outcomes and a secondary aim gathers information on health outcomes. Curran and colleagues (2022) suggest four considerations when choosing a type: the strength of existing effectiveness evidence, how much adaptation is expected, how much is known about implementation determinants, and whether the team is ready to test an implementation strategy.
(a) Recommend a hybrid type and justify it using all four considerations. (b) Name the primary outcome or outcomes and the unit of analysis for each. (c) Give one argument that a different type could be defended.
(a) I recommend a hybrid type 2 study. The effectiveness evidence is moderate: one trial in another province supports the program, but it has not been tested in this population and setting. Substantial adaptation is expected, which makes it unsafe to assume that the original effect will carry over, so effectiveness needs to remain a primary aim. Little is known about implementation determinants in British Columbia community centres, which argues for studying implementation closely, and the health authority has a specified facilitation strategy that it is ready to test. Together these point to co-primary aims.
(b) The effectiveness outcome would be change in glycated hemoglobin at six months, measured on individual participants and analyzed with a model that accounts for clustering by centre. The implementation outcomes would be reach (the share of eligible adults in each centre's area who attend at least one session) and fidelity to the program's core sessions, measured at the centre level. With only 16 centres, randomizing centres to facilitation or no facilitation would have little power, so I would test the strategy against a pre-specified benchmark and describe its delivery in detail, acknowledging this limitation.
(c) A type 1 study could be defended if the adaptations are so extensive that the program is effectively new, because then the first question is whether the adapted program works at all, and implementation questions would be exploratory.
Minimum 20 characters required.
Question 1: A study randomizes 40 clinics to two implementation strategies, is powered on clinic-level referral penetration, and records participants' loneliness scores as a secondary outcome. How should it be classified?
Question 2: According to the 2022 update by Curran and colleagues, what primarily determines a study's hybrid type?
Question 3: When is a hybrid type 1 study most appropriate?
Question 4: What is distinctive about the StaRI reporting standard?
Final Assessment
Bringing It All Together
This lesson moved the evaluation from the question of whether a program works to the question of how it comes to work in practice. Section 1 distinguished efficacy, effectiveness and implementation research, explained why benefits shrink as programs move into routine settings, and defined the eight implementation outcomes of Proctor and colleagues. The Cedar Valley first wave showed that a program with an encouraging pre-post result had reached fewer than one in twenty eligible older adults in its clinics, and that penetration varied several-fold between clinics.
Sections 2 and 3 provided the tools for acting on that finding. Nilsen's taxonomy sorted frameworks by purpose. CFIR 2.0 identified the barriers facing the less ready second-wave clinics, and RE-AIM, interpreted with PRISM, showed that the second wave nearly matched the first on reach and adoption while under-serving men and rural older adults. The ERIC compilation, the specification dimensions of Proctor, Powell and McMillen, and FRAME turned those findings into specified strategies and documented adaptations, with core functions held constant and forms allowed to vary.
Section 4 brought effectiveness and implementation questions into a single study. The Cedar Valley first wave fits a hybrid type 1 study, the second wave a hybrid type 2 study, and a provincial spread a hybrid type 3 study powered on clinic-level penetration, reported with StaRI.
Key Takeaways from this lesson
- Efficacy research asks whether a program can work under ideal conditions, effectiveness research asks whether it works in usual practice, and implementation research asks which methods get it adopted, delivered well and sustained.
- Benefits often shrink as programs move into routine settings, and evaluations need implementation data to tell intervention failure from implementation failure.
- Proctor and colleagues defined eight implementation outcomes (acceptability, adoption, appropriateness, feasibility, fidelity, implementation cost, penetration and sustainability), which are distinct from service and client outcomes.
- Nilsen's taxonomy groups implementation approaches into process models, determinant frameworks, classic theories, implementation theories and evaluation frameworks, according to the purpose each serves.
- CFIR 2.0 organizes determinants into innovation, outer setting, inner setting, individuals and implementation process domains, and its constructs can be rated for each site to guide strategy selection.
- RE-AIM judges a program's public health impact through reach, effectiveness, adoption, implementation and maintenance, with representativeness at its centre, and PRISM adds the context that explains the results.
- Implementation strategies should be selected to address identified barriers, named with the ERIC compilation, and specified by actor, action, target, timing, dose, outcome and justification.
- Adaptations are expected; FRAME documents them, and a program's core functions should be preserved while their forms vary with context.
- Hybrid type 1, 2 and 3 studies place effectiveness and implementation aims in different priority, and the 2022 update ties the type to a study's aims and asks authors to name both the design and the type.
- StaRI asks implementation studies to report the implementation strategy and the intervention as two parallel strands.
Core Concepts Reviewed
Section 1: the research-to-practice gap, the voltage drop, efficacy, effectiveness and implementation research, intervention and implementation failure, the Type III error, and the eight implementation outcomes of Proctor and colleagues.
Section 2: Nilsen's taxonomy, the Knowledge-to-Action and EPIS process frameworks, CFIR 2.0 and the Theoretical Domains Framework as determinant frameworks, and RE-AIM and PRISM as evaluation frameworks.
Section 3: implementation strategies and the ERIC compilation, the CFIR-ERIC matching approach, the specification of strategies, core functions and forms, FRAME, horizontal and vertical scale-up, and sustainability.
Section 4: hybrid effectiveness-implementation types 1, 2 and 3, the 2022 update, sample size for clinic-level implementation outcomes, designs for testing strategies, and reporting with StaRI.
The final reflection asks you to plan the implementation component of an evaluation for a hospital-to-home transition program, drawing on all four sections.
Reflection
A health authority will roll out a hospital-to-home transition program for older adults discharged from medical units at 6 hospitals over 18 months. A transition nurse visits each patient before discharge, reconciles medications, and telephones the patient twice in the following two weeks. One randomized trial in a single large hospital found fewer readmissions within 30 days. The smaller hospitals in the rollout have fewer staff and expect to change some parts of the program, and the health authority wants to test a facilitation strategy.
Frameworks available: CFIR 2.0, a determinant framework with five domains (innovation, outer setting, inner setting, individuals and implementation process); the Theoretical Domains Framework, a determinant framework of 14 domains focused on the behaviour of individual health professionals; EPIS, a process framework of exploration, preparation, implementation and sustainment phases; RE-AIM, an evaluation framework of reach, effectiveness, adoption, implementation and maintenance; and PRISM, which adds contextual domains to RE-AIM. Proctor and colleagues' implementation outcomes are acceptability, adoption, appropriateness, feasibility, fidelity, implementation cost, penetration and sustainability. Hybrid type 1 has effectiveness as its primary aim; type 2 has co-primary effectiveness and implementation aims; type 3 has an implementation strategy as its primary aim. StaRI is a 27-item reporting standard for implementation studies.
(a) Choose one determinant framework and one evaluation framework and justify each. (b) Specify four implementation outcomes, each with an indicator that states its numerator and denominator, a data source and a time of measurement. (c) Choose a hybrid type and justify it. (d) Name the reporting standard you would use and one thing it requires.
(a) I would use CFIR 2.0 as the determinant framework, because the likely barriers sit at several levels, including hospital staffing, discharge workflows and community services, and the smaller hospitals differ from the trial site in ways the inner setting domain captures. I would use RE-AIM as the evaluation framework, because the health authority needs to know how many eligible patients and which hospitals and nurses are reached, and whether the program persists; PRISM would help interpret differences between hospitals.
(b) Adoption: units with at least one transition visit in the first two months, divided by participating medical units, from program logs at month 2. Penetration: eligible discharged patients who received a pre-discharge visit, divided by all eligible discharges, from discharge records, monthly. Fidelity: patients who received the visit, medication reconciliation and both calls, divided by patients enrolled, from nurse documentation with a record audit, quarterly. Implementation cost: the cost of facilitation per hospital, from finance records and facilitator time logs at month 18.
(c) A hybrid type 2 study fits. One single-site trial gives moderate evidence, adaptation is expected in smaller hospitals, and a facilitation strategy is ready to test, so effectiveness (30-day readmissions) and implementation (penetration) should be co-primary aims. With 6 hospitals, the strategy would be tested against a benchmark or with a stepped rollout.
(d) I would report with StaRI, which requires parallel description of the implementation strategy and the transition program, including fidelity and adaptations to each.
Minimum 30 characters required.
Final Knowledge Assessment
Question 1: Flay distinguished efficacy trials from effectiveness trials primarily by which feature?
Question 2: Second-wave Cedar Valley clinics produce a smaller reduction in loneliness than the two-clinic pilot did. Which term describes this pattern?
Question 3: Which RE-AIM dimension corresponds most closely to Proctor's implementation outcome of penetration?
Question 4: A health authority invites 20 clinics. Sixteen agree to take part, and 14 make at least one referral within three months. In those 14 clinics, 60 of 96 clinicians refer at least once. Which statement reports RE-AIM adoption correctly, using a referral as the criterion?
Question 5: In Nilsen's taxonomy, which purpose does EPIS mainly serve?
Question 6: Clinicians in a well-resourced clinic with supportive managers still rarely ask older patients about loneliness. Which framework is best suited to diagnosing why?
Question 7: In the CFIR rating approach of Damschroder and Lowery, what does a rating of −2 for a construct indicate?
Question 8: First-wave Cedar Valley clinics never received data on their own referral numbers. Which ERIC strategy most directly addresses this barrier?
Question 9: Why do Proctor, Powell and McMillen recommend specifying strategies by actor, action, action target, temporality, dose, implementation outcome and justification?
Question 10: Which statement reflects the distinction between core functions and forms?
Question 11: Which element belongs to the definition of sustainability proposed by Moore and colleagues (2017)?
Question 12: A second-wave evaluation has co-primary aims: estimating the program's effect on loneliness and testing whether a strategy bundle achieves a pre-specified referral penetration. Which hybrid type is it?
Question 13: Clinic-level penetration has a standard deviation of 4 percentage points, and two strategies are expected to differ by 5 points. Using n per arm = 2 × (1.96 + 0.84)2 × SD2 ÷ difference2, about how many clinics are needed per arm?
Question 14: Which description follows the labelling recommended by Curran and colleagues in 2022?
Question 15: An evaluator wants to identify barriers before launch and to report the program's public health impact across individuals and settings afterward. Which pairing fits these two purposes?
Glossary: Key Terms, People & Frameworks
📚 Reference page, available throughout the lesson
Definitions of the implementation science terms, frameworks and people introduced in this lesson.