# Lesson 8: Mediation, Moderation and Path Analysis

*Companion-podcast transcript • Sarah & Kiffer*  
*Office Hours episode to listen to after working through the lesson*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This episode goes with Lesson eight, Mediation, Moderation and Path Analysis.

**Sarah:** It's meant for after you've worked through the lesson. We'll bring in some perspective and some critique, work through the questions that students tend to find thorny with this material, and add a few worked examples of our own, including one that's harder than anything on the lesson page.

**Kiffer:** At a few points we'll ask you to work something out before we give the answer. When that happens, you'll hear a few seconds of quiet. Pause the audio if you'd like more time.

**Sarah:** Here are the four questions. Why does a mediation analysis still need a no-confounding assumption when the exposure is randomized? How do you read the coefficients of an interaction model with a continuous moderator? Can an interaction be present and absent in the same data? And can an exposure be linked to an outcome through a mediator when its total association with that outcome is close to zero? We'll finish with a question Kiffer and I see differently, which is whether mediation analysis belongs in cross-sectional data at all.

**Kiffer:** Let's take the first one.

**Sarah:** Question one. Randomizing the exposure is supposed to take care of confounding. So why does the list of conditions for a causal indirect effect still include no unmeasured confounding of the mediator and the outcome, even in a trial?

**Kiffer:** Randomization protects only the arrows that start at the exposure. It makes the exposure independent of everything that came before it, so the total effect and path a are protected. The mediator is another matter. People end up with high or low values of the mediator for many reasons besides the exposure, and some of those reasons can also affect the outcome.

**Sarah:** So those shared causes confound path b, and the indirect effect along with it. And there's a second problem when we hold the mediator constant to get the direct effect.

**Kiffer:** Right. Holding the mediator constant means comparing people who have the same value of it. Within that group, the exposure and the other causes of the mediator become linked, because if the exposure didn't produce someone's value of the mediator, something else probably did. In DAG terms, the mediator is a collider between the exposure and its other causes, and conditioning on it opens a path from the exposure to the outcome through those other causes.

**Sarah:** So the collider idea turns up inside a mediation analysis.

**Kiffer:** It does, and there's a well-known real example. In 2006, Hernández-Díaz, Schisterman and Hernán looked at births in the United States in 1991. Babies whose mothers smoked were more likely to be born with low birth weight, and more likely to die in infancy. Among the low-birth-weight babies, though, infant mortality was lower for the babies of smokers, with a relative rate of 0.79.

**Sarah:** That's about twenty-one percent lower. So here's a question for everyone. Smoking doesn't protect babies. Using the idea we just described, what could produce a lower death rate among the low-birth-weight babies of smokers? Take a few seconds.

*(Pause)*

**Kiffer:** Low birth weight has several causes. Smoking is one of them. Others, such as serious birth defects, are far more dangerous to the baby. Among low-birth-weight babies whose mothers didn't smoke, the low weight is more likely to come from one of those other causes. So within that group, having a mother who smoked goes with the less dangerous reasons for being small.

**Sarah:** So birth weight lies on the pathway from smoking to infant death, and it shares causes with infant death. Restricting the analysis to small babies is the same as holding the mediator constant.

**Kiffer:** Exactly. The authors used causal diagrams to show that this structure can produce the inverse association even when smoking has no protective effect at all. They also pointed out that an estimate of the direct effect of smoking, adjusted for birth weight, can be biased when low birth weight and mortality share an unmeasured cause. That's the fourth condition from the lesson, in a real data set.

**Sarah:** And in a study of social support and depressive symptoms, the hidden common cause might be something like a history of depression, which can raise both loneliness and later symptoms.

**Kiffer:** Right. So a trial that randomizes the exposure still leaves the mediator part of the analysis observational. A careful mediation analysis in a trial measures the likely common causes of the mediator and the outcome at baseline, adjusts for them, and runs a sensitivity analysis for the ones it couldn't measure.

**Sarah:** Question two is about interaction models, and I'd like to try this one myself.

**Kiffer:** Here's a made-up example. A survey of twelve hundred adults regresses a loneliness score on weekly hours of in-person contact with friends and family, on age, and on the product of the two. Both predictors are centred, contact at its mean of eight hours a week and age at its mean of forty-five years. The centred contact coefficient is minus 0.15, and the interaction coefficient is minus 0.004 per year of age. What's the simple slope of contact at age seventy-five? Take a few seconds.

*(Pause)*

**Sarah:** Okay. The simple slope is the contact coefficient plus the interaction coefficient times age. So I take minus 0.004 times seventy-five, which is minus 0.30, and add it to minus 0.15. That gives a slope of minus 0.45 loneliness points per weekly hour of contact.

**Kiffer:** Here's a check worth doing. What does your method give at age forty-five?

**Sarah:** Minus 0.004 times forty-five is minus 0.18, and adding that to minus 0.15 gives minus 0.33. That can't be right. The centred coefficient is the slope at the mean age, so at forty-five the answer has to be minus 0.15.

**Kiffer:** So where did it go wrong?

**Sarah:** I multiplied by the age itself. The model was fitted with centred age, so I need the distance from the mean. Seventy-five minus forty-five is thirty, and minus 0.004 times thirty is minus 0.12. Adding that to minus 0.15 gives minus 0.27.

**Kiffer:** That's right. At seventy-five, each extra weekly hour of contact goes with a loneliness score about 0.27 points lower. What about age twenty-five?

**Sarah:** Twenty-five minus forty-five is minus twenty, and minus 0.004 times minus twenty is plus 0.08. So the slope is minus 0.15 plus 0.08, which is minus 0.07. In this made-up example, contact is barely associated with loneliness among young adults and more strongly associated at older ages.

**Kiffer:** The check you used works for any centred model. Put in the centring value, and you should get the main effect back. It's also worth remembering that a single survey compares people of different ages. These slopes describe differences between younger and older people at one point in time. Seeing how the association changes as one person ages would take repeated measurements of the same people.

**Sarah:** And the age coefficient in that model would be the slope of age for someone with eight hours of contact a week.

**Kiffer:** Right. Every main effect in an interaction model belongs to a particular value of the other variable. Centring lets you choose a value that actually occurs in the data, and it leaves the interaction coefficient and its test unchanged.

**Sarah:** Question three. Can an interaction be present and absent in the same data?

**Kiffer:** It can, and a classic real example comes from occupational health. Asbestos is also part of Canada's history, since it was mined in Quebec for more than a century, and a federal ban on asbestos and products containing it took effect at the end of 2018. The study I have in mind was published in 1979 by Hammond, Selikoff and Seidman, and it compared lung cancer deaths among asbestos insulation workers with those in a much larger group of workers who hadn't been exposed to asbestos.

**Sarah:** And they looked at smoking as well.

**Kiffer:** They did. Compared with non-smokers who hadn't been exposed to asbestos, the lung cancer death rate was about five times as high in non-smokers who had been exposed, about ten times as high in smokers who hadn't, and about fifty times as high in smokers who had.

**Sarah:** Let me work through that. I'll call the rate in unexposed non-smokers one unit. Without asbestos, smoking takes the rate from one to ten, a ratio of ten. With asbestos, smoking takes it from five to fifty, which is also a ratio of ten. So on the ratio scale, smoking does the same thing in both groups.

**Kiffer:** And on the difference scale?

**Sarah:** Without asbestos, smoking adds nine units. With asbestos, it adds forty-five units. That's an excess five times as large among the asbestos workers.

**Kiffer:** So there's almost no interaction on the multiplicative scale and a large one on the additive scale. Here's another way to see it. If the excess rates of the two exposures simply added up, smokers exposed to asbestos would be at one unit, plus nine for smoking, plus four for asbestos, which makes fourteen. The observed figure was about fifty.

**Sarah:** So if I fitted a Poisson or logistic regression with a smoking-by-asbestos product term to data like these, its coefficient would be close to zero.

**Kiffer:** And if you read that as no interaction, you'd miss a finding that matters a great deal for public health. The excess death rate that goes with smoking was about five times larger among the workers exposed to asbestos, and on the additive scale, which counts cases, that's where smoking prevention would matter most. I'd show the rates in all four groups, give the interaction on both scales, and lead with the additive scale when the question is where an intervention would prevent the most cases.

**Sarah:** Question four. Can an exposure be linked to an outcome through a mediator when its total association with that outcome is close to zero?

**Kiffer:** It can, when two pathways point in opposite directions. This is the harder example, and it's made up. A study of four hundred family caregivers compares those who joined a weekly peer support group with those who didn't. The mediator is a perceived support score, and the outcome is a distress score. Joining goes with a support score 0.80 points higher. Among caregivers with the same group status, each point of support goes with 0.50 points less distress. And among caregivers with the same support score, those who joined report 0.30 points more distress.

**Sarah:** Let me make sure I have the paths. Path a is 0.80, path b is minus 0.50, and the direct path, c prime, is plus 0.30.

**Kiffer:** That's it. So here's the question for everyone. Using the decomposition from the lesson, what's the total association between joining the group and distress? Take a few seconds.

*(Pause)*

**Sarah:** The indirect effect is a times b, which is 0.80 times minus 0.50, or minus 0.40. The total is the direct effect plus the indirect effect, and adding 0.30 to minus 0.40 gives minus 0.10.

**Kiffer:** So the total association is small. Now suppose the confidence interval for that total includes zero. What would the Baron and Kenny steps conclude?

**Sarah:** The first step requires a significant total effect, so they'd stop there and conclude there's nothing to mediate, even though the indirect path is four times the size of the total.

**Kiffer:** And suppose the study were large enough for the total to be significant. What would the last step, the shrinking of the exposure coefficient, show?

**Sarah:** The exposure coefficient is minus 0.10 in the total model and plus 0.30 once support is added. It changes sign and becomes three times as large in size, so the shrinking step fails as well.

**Kiffer:** And what proportion mediated would the software print?

**Sarah:** Minus 0.40 divided by minus 0.10, which is four. That's four hundred percent. A share has to lie between zero and one, so the number has no sensible reading here.

**Kiffer:** This is inconsistent mediation. The two parts have opposite signs and largely cancel. A sensible report gives the indirect and direct associations separately, each with a bootstrap interval, and leaves the proportion mediated out.

**Sarah:** What about the plus 0.30? It's tempting to say that the group itself makes caregivers more distressed.

**Kiffer:** The direct path holds everything the model doesn't name. One explanation is the time and effort the meetings take. Another, in a study where caregivers chose whether to join, is confounding, since the caregivers under the most strain may be the ones who look for a group. Those explanations call for different responses from the program, and this decomposition can't choose between them.

**Sarah:** So a total close to zero is a poor reason to stop looking for pathways.

**Kiffer:** That's my view. When there's a reason to expect opposing pathways, I'd test the indirect effect directly, whatever the total shows, and I'd say in advance which pathways I expected.

**Sarah:** That brings us to the question Kiffer and I see differently. Should anyone run a mediation analysis on a single cross-sectional survey?

**Kiffer:** Before we argue, here's one more question for everyone, because it bears on the answer. Take a three-variable model with no direct path, in which support predicts loneliness and loneliness predicts depressive symptoms. Now reverse both arrows, so that depressive symptoms predict loneliness and loneliness predicts support. Which version fits the lesson's survey data better? Think about what each model says about support and depressive symptoms once loneliness is held constant. Take a few seconds.

*(Pause)*

**Sarah:** They fit equally well. Each model has one degree of freedom, and each says the same thing about the data, which is that support and depressive symptoms are unrelated once loneliness is held constant. They allow exactly the same patterns of covariances, so they get the same chi-square.

**Kiffer:** A third model fits identically too, one in which loneliness affects both support and depressive symptoms. In DAG terms, the chain, the reversed chain and the fork imply the same conditional independence. In the survey data all three fit badly by exactly the same amount, because the link between support and depressive symptoms that each one leaves out is clearly different from zero. The fit says nothing about which way the arrows point. These are called equivalent models. In 1993, MacCallum and colleagues reviewed fifty-three published applications and found that equivalent models existed routinely, often in large numbers.

**Sarah:** That's my point. The direction of the arrows has to come from the design or from knowledge outside the data, and for these three variables a single survey leaves it open. In 2011, Maxwell, Cole and Mitchell showed that a cross-sectional analysis can suggest a substantial indirect effect even when the true longitudinal indirect effect is zero. And for most readers, the word mediation means a causal pathway, whatever the limitations paragraph says. I'd keep mediation analysis for designs with temporal order, and describe cross-sectional results as patterns of partial association.

**Kiffer:** I agree that the risk is real, and that the word carries a lot of weight in a title. Where I differ is on giving up the decomposition. Many public health questions start with survey data, because a cohort with well-timed waves is expensive. A decomposition shows whether the pattern of associations is consistent with a proposed pathway, and how large the indirect association is in those data. That helps when deciding whether a longitudinal study is worth running.

**Sarah:** But the same work showed that very different longitudinal processes can produce almost the same cross-sectional correlations. So the decomposition could point that decision in the wrong direction as well.

**Kiffer:** That's fair, and it's why I wouldn't let one cross-sectional estimate settle anything. I'd treat it as one input, with the sensitivity analysis and the equivalent models reported right beside it.

**Sarah:** I'm still more skeptical than Kiffer, but we agree on a practical rule. If you decompose an association from a single survey, say in the first sentence of the results that it's a statistical decomposition. Name the reversed-arrow explanation and the most likely confounder of the mediator and the outcome, report the sensitivity analysis, and describe the longitudinal design that would test the pathway.

**Kiffer:** Let's pull it together with three things to take away.

**Sarah:** First, randomizing the exposure protects the total effect and path a. The mediator still varies for other reasons, which can confound path b, and holding it constant can open a path through those other causes, as the birth weight example shows.

**Kiffer:** Second, every main effect in an interaction model belongs to a particular value of the other variable, and every interaction belongs to a particular scale. Check a simple slope by putting in the centring value, and say whether an interaction is on the additive or the multiplicative scale.

**Sarah:** And third, report the direct and indirect associations separately, with their intervals. A total close to zero can hide two opposing pathways, and the proportion mediated works as a share only when both parts have the same sign and the total is well away from zero.

**Kiffer:** If you'd like more practice, rework the caregiver support group example with the direct path changed to minus 0.30, and see what happens to the total and the proportion mediated. Then try it with a direct path of zero.

**Sarah:** Lesson eight is the last lesson, so this is the last Office Hours episode for this course. Thanks for working through all eight lessons with us.

**Kiffer:** Take care, everyone.

**Sarah:** Good luck with your next analysis.
