# Lesson 8: Quasi-Experimental Designs II: Time Series, Discontinuities and Natural Experiments

*Companion-podcast transcript, Sarah and Kiffer*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This week we finish the block on evaluation designs that began in Lesson six. After the randomized designs of Lesson six and the comparison-group designs of Lesson seven, this lesson adds interrupted time series, regression discontinuity, natural experiments and instrumental variables, and it closes with quality improvement designs.

**Sarah:** That is a lot of designs for one week.

**Kiffer:** It is, but they share a theme. Each one uses some feature of how a program was delivered to build a comparison. A launch date gives you a before and an after. An eligibility rule gives you people just above and just below a cut-off. A policy decided by someone else can split a population into exposed and unexposed groups.

**Sarah:** And we are still working with the Cedar Valley Connector program.

**Kiffer:** We are. For anyone new, it is a fictional community connector program run by the fictional Cedar Valley Health Authority in British Columbia. Primary care clinicians refer adults aged sixty-five and older who screen as lonely or isolated to a community connector, who meets them up to six times over twelve weeks and links them to groups, volunteering and transportation help. Twelve clinics started in the first wave and twelve more a year later. All the Cedar Valley numbers in this episode are illustrative, and the data in the worked examples are simulated.

**Sarah:** Let's start with Section one, the interrupted time series. What is it, in plain terms?

**Kiffer:** You measure an outcome at regular intervals, say every month, for a good while before a program starts and for a while after. The months before the launch tell you the level of the outcome, its trend and how much it bounces around. Then you ask whether the series changed at the launch by more than that earlier pattern would predict.

**Sarah:** So the earlier trend is the counterfactual.

**Kiffer:** Exactly. If emergency department visits among older adults were already creeping upward, the counterfactual for the program period is that the creep continued. A simple before and after comparison would miss that. In the Cedar Valley worked example, the naive comparison showed only a tiny drop in visits, because it ignored the fact that visits had been rising for four years.

**Sarah:** What can go wrong?

**Kiffer:** The big one is history. Anything else that happens at the same time as the launch, a new provincial policy, a bad respiratory virus season, a hospital staffing change, can produce a change in the series. Instrumentation is the second, meaning a change in how the outcome is recorded, like a new emergency department information system. And if the population itself changes, for example many older people moving into new housing in the region, a rate can change without anyone's risk changing.

**Sarah:** And the analysis is segmented regression.

**Kiffer:** Yes. You fit one line before the launch and another after it, in a single model. There is a term for time, a term called post that switches from zero to one at the launch, and a term called time after that counts the months since the launch. The coefficient on post is the level change, the immediate step at the launch month compared with where the old trend would have put you. The coefficient on time after is the slope change.

**Sarah:** And the effect of the program at some later month?

**Kiffer:** It is the level change plus the slope change multiplied by the number of months since the launch. So if you want the effect two years in, you add both pieces, and because you are adding two estimates, the confidence interval needs their covariance as well as their variances. The worked example shows how to do that.

**Sarah:** The reading makes a point about deciding the shape of the effect in advance. Why does that matter?

**Kiffer:** Because if you look at the data first and then pick the shape that fits best, you can fit random wiggles and call them an effect. Lopez Bernal and colleagues call the expected shape the impact model, and they recommend specifying it beforehand from the program theory. For Cedar Valley, enrolment builds gradually, a few hundred older adults in the first months and more as the second wave of clinics joins, so the theory predicts a slope change. A sudden region-wide step at launch would actually make me suspicious.

**Sarah:** Suspicious of what?

**Kiffer:** Of some other event. A program that has reached a few hundred people cannot plausibly move a regional rate overnight.

**Sarah:** Let's talk about seasonality.

**Kiffer:** Emergency visits by older adults rise every winter. If a program launches just before the winter peak, or just after it, the seasonal swing can look like an effect. So the model has to include seasonality. You can add an indicator for each calendar month, which costs eleven parameters for monthly data, or you can add Fourier terms, which are pairs of sine and cosine waves with a twelve-month period. One pair describes a smooth annual cycle with just two parameters.

**Sarah:** And autocorrelation is the other problem.

**Kiffer:** Right. A busy month tends to be followed by another busy month. The residuals from the model are correlated over time. The coefficients are still fine if the model is right, but the ordinary standard errors are too small, so confidence intervals are too narrow and p-values too small. You find it by plotting the residuals, by looking at the autocorrelation function, and with tests like Durbin-Watson.

**Sarah:** Remind me how to read the Durbin-Watson statistic.

**Kiffer:** It sits near two when there is no first-order autocorrelation and falls toward zero as positive autocorrelation grows. In the Cedar Valley worked example it was about one point two six, which is clear positive autocorrelation. The lag-one correlation of the residuals was about zero point three seven.

**Sarah:** And the fix?

**Kiffer:** There are a few. Newey-West standard errors keep the same coefficients and replace the standard errors with ones that allow for autocorrelation up to a chosen lag. Generalized least squares with autoregressive errors, such as the Prais-Winsten estimator, models the correlation directly. In the worked example, Newey-West standard errors raised the standard error of the slope change enough that its p-value went from about zero point zero one five to just over zero point zero five.

**Sarah:** So the slope change became borderline.

**Kiffer:** It did, on its own. The level change stayed clearly different from zero, and the effect at month seventy-two, which combines the two, had a confidence interval well away from zero.

**Sarah:** How long does a series need to be?

**Kiffer:** There is no single rule. Cochrane's Effective Practice and Organisation of Care group requires at least three points before and three after for inclusion in its reviews, which is only a floor. For monthly data, twelve before and twelve after is often cited as a working minimum, partly because you need a full year on each side to separate seasonality from the effect. Power depends on the variability of the series, the autocorrelation, the size and shape of the effect, and where the launch falls. Zhang and colleagues showed how to estimate power by simulation.

**Sarah:** You also talk about dilution.

**Kiffer:** This one matters a lot for programs like Cedar Valley. If you measure emergency visits for every older adult in the region, and the program reaches only a few percent of them, then even a real effect among participants barely moves the regional rate. In the worked example I did a plausibility check. Suppose about one in twenty older residents had taken part by month seventy-two, which is more than the program's attendance figures suggest. The controlled estimate would then require each participant to avoid about a third of a visit per year, which is roughly two thirds of the emergency visits an average older resident makes.

**Sarah:** That seems implausible.

**Kiffer:** It does. The simulated effect was set large so the methods are easy to see. In a real evaluation, that arithmetic would tell me the regional rate is an insensitive outcome, and I would add a series restricted to referred older adults using linked records.

**Sarah:** Now the controlled interrupted time series.

**Kiffer:** A single series cannot rule out history, so we add a comparison series that should feel the same outside events but did not get the program. In the worked example it is a fictional neighbouring region called Juniper Ridge. The key assumption is that, without the program, the difference between the two series would have kept following its pre-launch path.

**Sarah:** And what happened in Juniper Ridge?

**Kiffer:** Its rate dropped at the launch month too, by about zero point seven seven visits per thousand, with no change in slope. Juniper Ridge has no connector program, so that drop points to something shared, which in the simulation stands in for a province-wide change. When I modelled the difference between the two regions, the controlled effect at month seventy-two was about one point three eight visits per thousand, compared with about two point zero three from Cedar Valley alone.

**Sarah:** So about a third of the apparent effect was shared.

**Kiffer:** Yes. Most of the step at launch was shared with the comparison region, and the remaining evidence for an effect rests on Cedar Valley's slope bending downward afterward.

**Sarah:** Is that the same as difference-in-differences from Lesson seven?

**Kiffer:** It is a close relative. You can think of it as difference-in-differences with many periods and an explicit trend on each side of the launch.

**Sarah:** Let's move to Section two, regression discontinuity.

**Kiffer:** Many programs decide eligibility with a rule on a score. Adults become eligible at a certain age, patients get a treatment when a lab value crosses a cut-off, and in Cedar Valley, clinicians refer older adults who score six or higher on the three-item University of California, Los Angeles Loneliness Scale. The score is called the running variable, and the cut-off is the threshold. People just below and just above the threshold are similar in everything except eligibility, so a jump in the outcome right at the threshold tells you the effect for people near it.

**Sarah:** Where did the design come from?

**Kiffer:** Thistlethwaite and Campbell introduced it in nineteen sixty to study merit awards for students whose test scores crossed a cut-off. A nice Canadian example is Ontario's school-based human papillomavirus vaccine program, which began in grade eight in two thousand seven, so eligibility depended on birth date. Smith and colleagues compared girls on either side of the cut-off and reported no evidence that vaccination increased clinical indicators of sexual behaviour.

**Sarah:** And the drinking age example.

**Kiffer:** The minimum legal drinking age is eighteen in Alberta, Manitoba and Quebec and nineteen elsewhere. Callaghan and colleagues compared hospital admissions just below and just above the legal age and found increases in alcohol-related admissions above it, especially among young men.

**Sarah:** What is the assumption underneath all this?

**Kiffer:** Continuity. Without the program, the average outcome would change smoothly as the score crosses the threshold. Lee showed that if people cannot precisely control their own score, assignment near the threshold behaves almost like a lottery.

**Sarah:** The reading distinguishes sharp and fuzzy designs.

**Kiffer:** In a sharp design the rule is followed exactly, so the probability of getting the program jumps from zero to one. In a fuzzy design it jumps by less than one. Cedar Valley is fuzzy on purpose, because clinicians can refer someone who scores five if they judge the person isolated, and some people who score six or more are never referred.

**Sarah:** How do you get an effect from a fuzzy design?

**Kiffer:** You measure two jumps. The jump in the outcome at the threshold is called the reduced form. The jump in the probability of getting the program is called the first stage. You divide the first by the second. That ratio is the Wald estimator. In the simulated Cedar Valley data, the probability of referral jumped by about zero point four one at a score of six, and the six-month loneliness score dropped by about zero point two. The ratio is about minus zero point four nine points.

**Sarah:** Whose effect is that?

**Kiffer:** The effect for compliers at the threshold, meaning the older adults who were referred because their score reached six and would not have been referred at five. And that is exactly the group that matters for the steering committee's question about lowering the threshold.

**Sarah:** The worked example had two other comparisons that looked very different.

**Kiffer:** It did, and they are worth remembering. Referred older adults had six-month scores about one point two higher than people who were not referred, because referral selects lonelier people. And among referred older adults, scores fell by about zero point seven four from screening to six months, which mixes any program effect with regression to the mean. The discontinuity design avoids both problems, at the price of describing only the margin of eligibility.

**Sarah:** How wide a window do you use around the threshold?

**Kiffer:** That is the bandwidth question. The usual estimator is local linear regression, a straight line on each side within the window. A narrow window compares very similar people but uses fewer of them, so it is noisy. A wide window is more precise but leans harder on the straight-line approximation. Imbens and Kalyanaraman developed a bandwidth that balances the two, and Calonico, Cattaneo and Titiunik showed how to correct the confidence intervals. Their methods are available in standard statistical software. Gelman and Imbens warned against fitting high-order polynomials to the whole range.

**Sarah:** And the Cedar Valley score has only seven possible values.

**Kiffer:** Which is awkward. There is nothing between five and six, so the line from below has to be extrapolated a full point, and the window cannot shrink below one point. I report results in more than one window and I am cautious about precise confidence intervals. In the worked example, a narrower window gave about minus zero point five four, similar to the main estimate.

**Sarah:** What is the biggest threat to the design?

**Kiffer:** Manipulation. If a clinician who wants to secure a referral records a borderline five as a six, the people just above the threshold are no longer comparable to those just below. That leaves a fingerprint: too many people just above the cut-off. McCrary's density test, and the newer rddensity method, look for that. You also check that characteristics fixed before screening, like age, do not jump, and that no jumps appear at placebo thresholds where no rule applies.

**Sarah:** And in the simulated data?

**Kiffer:** The distribution of scores declined smoothly with no pile-up at six, and age showed no clear jump.

**Sarah:** On to Section three. What makes an experiment natural?

**Kiffer:** The Medical Research Council guidance, written by Craig and colleagues in twenty twelve, describes natural experiments as events or interventions that researchers do not control but that divide a population into exposed and unexposed groups. The word natural describes how exposure was assigned.

**Sarah:** So what makes one good?

**Kiffer:** An assignment process unrelated to the outcome, or as Dunning puts it, as if random. The classic case is John Snow in London in the eighteen fifties. Two water companies supplied houses on the same streets, one drawing cleaner water from upstream, and households generally had not chosen their supplier. That made the comparison of cholera deaths unusually convincing.

**Sarah:** What else does the guidance say?

**Kiffer:** That natural experimental studies are most informative when the intervention is expected to have an effect large enough to detect, when good data cover large populations, and when the assignment process is understood well enough to support a credible comparison. It also encourages testing assumptions with more than one method.

**Sarah:** The reading has several Canadian examples.

**Kiffer:** Provinces make policy at different times, so Canada is full of natural experiments. Saskatchewan introduced public hospital insurance and then medical care insurance before the other provinces, and Hanratty used the staggered introduction of national health insurance to study infant health. Baker, Gruber and Milligan compared Quebec with other provinces after Quebec introduced low-fee child care in nineteen ninety-seven. Stockwell and colleagues studied minimum alcohol prices in British Columbia and Saskatchewan. And Forget linked health records decades after the Mincome guaranteed income experiment in Dauphin, Manitoba, and reported fewer hospitalizations during the experiment.

**Sarah:** Let's get to instrumental variables. Give me the intuition first.

**Kiffer:** Suppose attendance at the connector program is voluntary, and the people who attend are more motivated or healthier. Comparing attenders with non-attenders is confounded. An instrument is something that nudges some people into the program for reasons unrelated to their outcomes. If you look only at the variation in attendance caused by that nudge, you get around the confounding.

**Sarah:** And the conditions?

**Kiffer:** There are four. Relevance means the instrument really shifts participation. Independence means it shares no causes with the outcome. The exclusion restriction means it affects the outcome only through the program. And monotonicity means nobody is pushed away from the program by it. Only relevance can be checked directly in the data.

**Sarah:** The reading uses a capacity example from Cedar Valley.

**Kiffer:** In the second year, after the second wave of clinics joined, connectors' caseloads were sometimes full. If a referral came in when no slot was open, the person went on a waiting list, and many never started. Whether a slot was open depended on when earlier participants finished their twelve weeks, which the evaluator argues has nothing to do with the person being referred.

**Sarah:** And the numbers?

**Kiffer:** These are illustrative. Eighty percent of people referred when a slot was open started within four weeks, compared with thirty-five percent when caseloads were full. Six-month loneliness scores averaged five point six and five point nine. So the Wald estimate is minus zero point three divided by zero point four five, about minus two thirds of a point.

**Sarah:** For whom?

**Kiffer:** For compliers, the people who started only because a slot happened to be open. Imbens and Angrist called this the local average treatment effect. People who would have started anyway, or who would never start, are unaffected by the instrument, so they contribute nothing to the estimate.

**Sarah:** I want to push on the independence assumption. Couldn't caseloads be fuller at certain times of year?

**Kiffer:** That is exactly the right worry. If caseloads fill in winter, when loneliness and illness are higher, the instrument is tangled up with season. Or if some clinics serve more isolated neighbourhoods and are always full, the instrument is tangled up with place. So you compare referrals within the same clinic and season, and you check whether baseline scores and age are balanced between open-slot and full-caseload referrals.

**Sarah:** And the exclusion restriction?

**Kiffer:** If waitlisted people got a welcome call and a list of community groups, or clinicians sent them somewhere else, the waiting list could affect loneliness directly. Then the instrument has a second path to the outcome.

**Sarah:** What about weak instruments?

**Kiffer:** If the instrument barely moves participation, you are dividing by a small number. The estimate becomes very imprecise, it drifts toward the confounded estimate, and any small violation of the exclusion restriction gets magnified. Staiger and Stock suggested treating a first-stage F statistic below about ten as a warning. Hernán and Robins also warned that the key assumptions cannot be verified.

**Sarah:** How does this connect to the regression discontinuity?

**Kiffer:** A fuzzy regression discontinuity is an instrumental variable analysis in which being above the threshold is the instrument. Same Wald estimator, same complier interpretation. But the two Cedar Valley estimates answer different questions. The threshold estimate is about the effect of referral at the margin of eligibility. The capacity estimate is about the effect of starting promptly among people already referred.

**Sarah:** And how do you judge a natural experiment overall?

**Kiffer:** Start with who decided who was exposed, and why. Look for a comparison that shares the same outside events. Run falsification tests, like a negative control outcome the program should not change. For Cedar Valley, emergency visits for fractures could play that role. And compare designs with different weaknesses, which Lawlor and colleagues call triangulation. If a controlled time series, a comparison of clinic waves and a threshold design all point the same way, a single bias is less likely to explain all three.

**Sarah:** Section four is quality improvement. Why is that in an evaluation course?

**Kiffer:** Because programs use it to improve their own delivery, and evaluators need to know what those data can and cannot show. Quality improvement has been defined as systematic, data-guided work to bring about immediate improvement in care in a particular setting. Its purpose is local and practical. Evaluation judges merit and worth, and research aims for knowledge that generalizes.

**Sarah:** And the Model for Improvement?

**Kiffer:** It comes from Langley and colleagues and asks three questions. What are we trying to accomplish? How will we know that a change is an improvement? What change can we make that will result in improvement? The changes are then tested in Plan-Do-Study-Act cycles, which grew out of Walter Shewhart's work and were developed by W. Edwards Deming.

**Sarah:** Walk me through a cycle in Cedar Valley.

**Kiffer:** The coordinator notices that many referred older adults wait more than two weeks for a first meeting, and some lose interest. The aim is to raise the share with a first meeting within fourteen days from about fifty-four percent to seventy percent within six months. In the first cycle, one connector at one clinic phones every new referral within two business days for two weeks, predicting that eight of ten will be reached. Six of nine are reached, and several people say they would prefer a text first. So the second cycle adds a text message. Later cycles spread the change to all the wave-one clinics.

**Sarah:** That prediction step seems small.

**Kiffer:** It is what turns a test into learning. Taylor and colleagues reviewed published applications and found that few showed the key features: linked cycles, explicit predictions, small-scale testing and data over time. Reed and Card argued that the method looks simple but needs real skill and support.

**Sarah:** Then run charts.

**Kiffer:** A run chart plots the measure over time with the median as the centre line. Perla, Provost and Murray describe four rules. A shift is six or more consecutive points on one side of the median. A trend is five or more points all rising or all falling. Too few or too many runs suggest a non-random pattern. And an astronomical point is a value that is obviously different from the rest.

**Sarah:** What did the Cedar Valley chart show?

**Kiffer:** The baseline median was fifty-four percent, and all ten weeks after the change lay above it. That meets the shift rule, so the team can say the process changed. What it cannot say, from the chart alone, is that the phone and text change caused it. A new coordinator or a change in how meetings were recorded could produce the same shift. That is the history threat again, in a new setting.

**Sarah:** And statistical process control charts?

**Kiffer:** Shewhart distinguished common-cause variation, the ordinary noise of a stable process, from special-cause variation, which comes from identifiable sources. A control chart adds limits, usually three standard deviations from the centre line. You pick the chart to match the data: a p-chart for proportions, a u-chart for rates, a c-chart for counts and an individuals chart for single values.

**Sarah:** Give me the p-chart example.

**Kiffer:** With a centre line of zero point five four and fifty-two referrals in a month, the limits run from about zero point three three to about zero point seven five. A month at seventy percent sits inside the limits, so on its own it is consistent with ordinary variation. A month at seventy-five percent would sit just above the upper limit and signal a special cause.

**Sarah:** Do the time series problems from Section one apply here?

**Kiffer:** They do. Standard control limits assume independent observations, so autocorrelation creates false signals. And if you put Cedar Valley's monthly emergency visit rate on a u-chart, every winter peak would look like a special cause unless you modelled the season first.

**Sarah:** And how do teams report this work?

**Kiffer:** With the Standards for Quality Improvement Reporting Excellence, SQUIRE two point oh, published by Ogrinc and colleagues. It has eighteen items. The ones I would point evaluators to are the rationale, which asks for the theory behind the change, the context, and the study of the intervention, which is the design used to judge whether the change caused the results.

**Sarah:** The section ends by comparing all the designs. What is the big lesson there?

**Kiffer:** That every design trades one assumption for another. A randomized trial relies on randomization being done and kept. Difference-in-differences relies on parallel trends. A time series relies on the old trend continuing and on a comparison series that shares outside events. A discontinuity relies on continuity at the threshold. An instrument relies on independence and exclusion. And quality improvement designs put speed and local learning ahead of causal certainty.

**Sarah:** So no design wins outright.

**Kiffer:** None does. The choice follows from the evaluation question, from how the program was assigned, from the data and from what is ethical and feasible. Strong plans often combine designs whose weaknesses differ.

**Sarah:** Which brings us to writing a design justification.

**Kiffer:** Yes. A justification sets out the chosen design for the people who will use the evaluation. State the evaluation questions and the causal contrast each one needs. Describe how the program was or will be assigned. Name the design you chose and at least one alternative you rejected, with reasons. Then list each key assumption in plain language, pair it with a check planned in advance, and say what you will conclude if the check fails. Include a short table of assumptions and checks.

**Sarah:** And the Cedar Valley version?

**Kiffer:** It uses difference-in-differences between the two waves of clinics for loneliness, a controlled interrupted time series with Juniper Ridge for emergency visits, with a plausibility check for dilution, and a fuzzy regression discontinuity for the threshold question. Each assumption is paired with a check, like an event-study plot, a negative control outcome or the distribution of scores at five and six.

**Sarah:** Any last advice?

**Kiffer:** Write the justification for a steering committee. If the people who fund and run the program can see what you are assuming and how you will know if you are wrong, they can weigh the result sensibly.

**Sarah:** Thanks, Kiffer. That's it for this week's Office Hours.

**Kiffer:** Thanks, Sarah, and thanks, everyone, for listening.
