# Lesson 3: Causal Webs and Directed Acyclic Graphs

*Companion-podcast transcript, Sarah and Kiffer*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This week we are on Lesson three of Health Sciences two-oh-seven, which is about causal webs and directed acyclic graphs.

**Sarah:** Those phrases may make some students nervous. Before the definitions, what problem is this lesson solving?

**Kiffer:** In Lesson two, the Cedar Valley team wrote a structured research question. A question names an exposure and an outcome, but it does not tell you what else your study needs to measure. Every health outcome has many causes, tangled together, and if you measure only the exposure and the outcome, you cannot tell whether an association between them is a real effect or the work of something else.

**Sarah:** And presumably the opposite mistake is to measure everything.

**Kiffer:** Yes. A survey that asks about everything becomes long and expensive, fewer people finish it, and it collects personal information that nobody needed. So researchers decide what to measure before they collect any data, and they make that decision by drawing what they believe about causes.

**Sarah:** Let's anchor this in the running case. Remind listeners what the Cedar Valley study is.

**Kiffer:** The Cedar Valley Social Connection Study is fictional. It is a study planned by a university research team, led by Doctor Maya Hart, working with a fictional regional health authority in British Columbia. The team wants to understand loneliness among adults aged sixty-five and older. Their question, written in Lesson two in the population, exposure, comparison and outcome format, asks whether loneliness is associated with the number of emergency department visits in the twelve months after a regional survey.

**Sarah:** And loneliness is measured how?

**Kiffer:** With the three-item loneliness scale developed at the University of California, Los Angeles. Each item is scored from one to three, so totals run from three to nine, and the team counts a score of six or higher as lonely. The survey received one thousand six hundred completed responses, and three hundred and ninety-two people, or about one in four, scored as lonely.

**Sarah:** Section one is about the causal web. What exactly is a causal web?

**Kiffer:** It is a broad diagram of the factors that evidence and experience suggest are connected to the outcome, and to one another. The idea goes back to a nineteen sixty textbook by Brian MacMahon, Thomas Pugh and Johannes Ipsen, who described the causes of disease as a web of causation, at a time when a single-cause model did not fit heart disease or cancer very well.

**Sarah:** Why does the image of a web help?

**Kiffer:** It captures four things at once. Most outcomes have several causes acting together. Each cause has causes of its own, so living alone may follow the death of a spouse. Causes sit at different levels, from the individual to relationships to whole communities. And factors can affect each other in loops. Loneliness can lead to depression, and depression can lead people to withdraw from others.

**Sarah:** The lesson also describes a critique of the web idea. Who made it?

**Kiffer:** Nancy Krieger, a social epidemiologist, in a nineteen ninety-four paper whose title asks whether anyone has seen the spider. If disease arises from a web, something has to spin it. She argued that the webs drawn in epidemiology put individual behaviours and biology in the foreground and treated social, economic and political conditions as background, or left them out altogether.

**Sarah:** Is that critique still relevant to someone drawing a first web today?

**Kiffer:** It is, and in a practical way. If your web contains only individual risk factors, your study will measure only individual risk factors. So when students draw their webs, I ask them to look deliberately at income, housing, transport and access to services. Those outer levels are the easiest to forget.

**Sarah:** Let's talk about how you actually build a web. Where do students start?

**Kiffer:** With the question. Write the exposure and the outcome at the top of a page. Then build an evidence table. For each source, record what kind of evidence it is and what it suggests about causes. Lesson two showed how to find three to five key papers, so those papers are the raw material.

**Sarah:** What did the Cedar Valley team's evidence table contain?

**Kiffer:** It had a mix. A large meta-analysis by Julianne Holt-Lunstad and colleagues found that loneliness and social isolation were associated with higher mortality. A review linked loneliness to heart disease and stroke. Work by John Cacioppo and colleagues found that loneliness predicted later depressive symptoms, another study found that it predicted reduced physical activity, and a study of older adults in the United States found it was associated with more physician visits.

**Sarah:** Those are all published studies. Was anything else in the table?

**Kiffer:** Two rows came from people. The team's advisory group of six older adults said that bereavement often leads to living alone, that losing a driver's licence or living far from town cuts people off, and that hearing loss makes conversation tiring. Staff at partner clinics said that patients with several chronic conditions use the emergency department more, especially where after-hours primary care is scarce.

**Sarah:** Once the table is built, what is the next step?

**Kiffer:** Rewrite each finding as a link statement. A link statement has the form A affects B, with the source in brackets. So instead of writing that loneliness and depression are related, you write that loneliness leads to later depressive symptoms, and you cite Cacioppo and colleagues. The link statement forces you to commit to a direction.

**Sarah:** What if the evidence only shows an association?

**Kiffer:** Then the direction is an assumption you are bringing to the evidence, and you should say so. Most of the findings in a first evidence table are associations. Writing the direction down makes the assumption visible, so that someone else can question it.

**Sarah:** And then the drawing itself?

**Kiffer:** First, sort the factors into rough levels and check whether any level looks thin. The Cedar Valley team found its first list said little about transport and income, so it went back to the advisory group's notes. Then draw the web, with the exposure near the centre and the outcome on the right, and show the draft to people who know the problem.

**Sarah:** Describe the Cedar Valley web for someone who can't see it.

**Kiffer:** On the left are causes of loneliness: bereavement, living alone, hearing loss, limited transport and rural residence. Low income and chronic conditions also point into loneliness. Loneliness sits in the centre, and emergency department visits sit on the right. There is a direct arrow from loneliness to visits, and two pathways run through depression and through physical inactivity. Chronic conditions also point straight to visits. The web contains two loops as well. Loneliness and depression are drawn affecting each other, and there is a longer loop in which chronic conditions raise loneliness, loneliness reduces physical activity, and inactivity worsens chronic conditions. Over years, these factors really do feed each other.

**Sarah:** So why not stop there? The web seems to describe the situation well.

**Kiffer:** It describes it well, but it does not tell the team what to measure. Its boxes can be broad, its arrows can run both ways, and it does not say what happened first. That is where the directed acyclic graph comes in. I will call it a DAG from here on.

**Sarah:** Spell out what the three words mean.

**Kiffer:** Directed means every arrow points one way, from cause to effect. Acyclic means there are no loops, so if you start anywhere and follow the arrows you can never come back to where you started. And graph is the mathematical word for a set of points joined by lines. In a DAG, every box is called a node, and each node is a single variable that could be measured at a particular time.

**Sarah:** That brings us to section two. Let's go through the parts of a DAG.

**Kiffer:** An arrow from A straight to B represents a direct effect, meaning one that does not pass through any other variable in the diagram. Two more words are useful. A node's ancestors are all the nodes from which you can reach it by following arrows forward. Its descendants are all the nodes you can reach from it. The program we use, DAGitty, puts those words in its legend.

**Sarah:** The lesson makes a point about arrows that surprised me. It says an arrow is a modest claim.

**Kiffer:** That is right. An arrow from A to B says only that A might directly affect B. It does not say how large the effect is, or whether it is harmful or protective. Adding an arrow is the cautious choice when the evidence is unclear, because it allows for an effect without assuming one.

**Sarah:** And a missing arrow is the strong claim.

**Kiffer:** Yes. If there is no arrow from A to B, the diagram asserts that A has no direct effect on B at all. Any influence has to travel through other variables in the diagram. The arrows students leave out are often their most important assumptions.

**Sarah:** Is there a Cedar Valley example of that?

**Kiffer:** There is a good one. The team's first small DAG had four nodes: age, hearing loss, loneliness and emergency department visits. Hearing loss pointed into loneliness, and there was no arrow from hearing loss to visits. That missing arrow claimed that hearing loss affects visits only by making people lonelier.

**Sarah:** And someone challenged it.

**Kiffer:** A member of the advisory group did. She pointed out that poor hearing can make it harder to hear traffic, alarms or a clinician's instructions, and that several people she knew had been injured in falls they connected with their hearing. That is a route to the emergency department that has nothing to do with loneliness. So the team added the arrow. As we will see in section three, that one arrow changed the role hearing loss plays in the study.

**Sarah:** The lesson gives four conventions for drawing a first DAG. Can you take us through them?

**Kiffer:** The first is that one node is one variable at one time. Social factors fails that test. Living alone at the time of the survey passes it. The second is that arrows point forward in time, so you arrange the diagram from left to right, with earlier things on the left. The third is that loops are unrolled into time.

**Sarah:** Explain that third one, because the web had loops and the DAG cannot.

**Kiffer:** You split the factor into time-stamped nodes. Depression before the survey can affect loneliness at the survey. Loneliness at the survey can affect depression during the following year. And depression before the survey can affect depression afterwards, because symptoms often persist. The loop becomes a sequence, and every arrow points forward.

**Sarah:** And the fourth convention?

**Kiffer:** If a variable causes two or more variables already in the diagram, it belongs in the diagram even if you cannot measure it. Frailty is the Cedar Valley example. A frail older adult may find it hard to leave the house, which raises loneliness, and may be more likely to be injured, which raises emergency visits. The survey does not measure frailty, so the team keeps it in the diagram and marks it as unobserved.

**Sarah:** Now the practical part. How does a student actually draw this in DAGitty?

**Kiffer:** DAGitty is free software developed by Johannes Textor and colleagues. You go to dagitty dot net and choose the link that says Launch DAGitty online in your browser. There is nothing to install and no account to create. From the Model menu you choose New model, and the program asks for the names of the exposure and the outcome. It draws them with an arrow between them, and you keep that arrow, because it is the effect your question asks about.

**Sarah:** And then you add the other variables.

**Kiffer:** You double-click an empty spot on the canvas, or move the pointer to an empty spot and press the n key. A box appears, you type the name, and you press Enter. To add an arrow, you double-click the cause, which becomes highlighted, and then double-click the effect. To remove an arrow, you repeat the same two double-clicks.

**Sarah:** What comes after the arrows are in place?

**Kiffer:** Drag the nodes so time runs from left to right. Then set each variable's status. You click a node and use the checkboxes in the panel labelled Variable, or you hover over it and press e for exposure, o for outcome or u for unobserved. Then you look at the box labelled Model code. DAGitty writes your whole diagram there as text, with each arrow on its own line.

**Sarah:** And that text is how you save your work.

**Kiffer:** Yes. Copy everything in the Model code box into a plain text file with a version number, something like Cedar Valley DAG version one. To reopen the diagram, you paste the text back in and click Update DAG. You can also export a picture from the Model menu, as a PNG file for a document or as a PDF.

**Sarah:** DAGitty also colours things and suggests what to adjust for. What should students make of that?

**Kiffer:** For now, they should leave it alone. Those features apply formal rules for reading a DAG, which Health Sciences two-thirty introduces in its Lesson seven and three forty-one states in full in its own Lesson seven. In this course DAGitty is a drawing and record-keeping tool.

**Sarah:** The section ends with something called a justification table.

**Kiffer:** It is a simple table with three columns: the arrow or missing arrow, a one-line reason, and a source. The source can be a study, the knowledge of interest holders, or the order of events. For example, there is no arrow from emergency visits to loneliness, because the visits are counted after the survey. Tennant and colleagues recommended that researchers report their full diagrams and explain how they were built, and a justification table is a simple way of doing that in a small study.

**Sarah:** Let's move to section three, which I think is the heart of the lesson. Variable roles.

**Kiffer:** The first point is that roles are defined by the question. A variable has no role on its own. The exposure is the variable whose effect you are asking about, and the outcome is where you measure that effect. In Cedar Valley, loneliness is the exposure. In a different study asking whether a connector program reduces loneliness, loneliness would be the outcome. Three more roles describe the other variables. A confounder is a common cause of the exposure and the outcome. A mediator lies on a pathway from the exposure to the outcome. A collider is a common effect of two variables in the diagram.

**Sarah:** Start with the confounder. What is the example you like to use?

**Kiffer:** Carrying matches and lung cancer. People who carry matches get lung cancer more often than people who do not. Matches do not cause cancer. Smoking causes people to carry matches, and smoking causes lung cancer. So smoking is a confounder. A study that ignored smoking would conclude that matches are dangerous.

**Sarah:** And in Cedar Valley?

**Kiffer:** Chronic conditions diagnosed before the survey. Conditions such as heart failure or arthritis can keep people at home, which raises loneliness, and they produce acute episodes that bring people to the emergency department. If the team compared lonely and non-lonely respondents without accounting for chronic conditions, part of the difference in visits would really be a difference in illness.

**Sarah:** You said earlier that hearing loss changed roles. Explain that now.

**Kiffer:** In the first draft, hearing loss pointed only into loneliness, so it was a cause of the exposure and nothing more. It was not a confounder. Once the team added the arrow from hearing loss to emergency visits, hearing loss caused both the exposure and the outcome, and it became a confounder. Its role changed because one arrow changed.

**Sarah:** Now the mediator.

**Kiffer:** A mediator carries part of the effect. Picture a school breakfast program. Breakfast may improve children's attention in the morning, and better attention may raise test scores. Attention is a mediator of the program's effect on scores. In Cedar Valley, depressive symptoms that develop after the survey are a mediator. Loneliness can lead to depression, and depression can lead to emergency visits.

**Sarah:** So what goes wrong if someone treats a mediator as a confounder?

**Kiffer:** They remove part of the very effect they set out to estimate. If loneliness raises visits partly by causing depression, that route is part of the answer to the team's question, and removing it leaves a narrower answer. Studying how much of an effect travels through a mediator is called mediation analysis, which is taught in Health Sciences four-ten.

**Sarah:** The lesson says students confuse mediators and confounders more than any other pair. Is there evidence for that?

**Kiffer:** There is. In a recent survey of students in Health Sciences three forty-one in this series, about one in four, twenty-three percent, chose a variable on the causal pathway as the definition of a confounder. That is the definition of a mediator. These are upper-year students, so the confusion clearly persists.

**Sarah:** Why is it so persistent?

**Kiffer:** Because both are third variables associated with the exposure and with the outcome. In a dataset they can look exactly alike. The difference lies in the direction of the arrow between the third variable and the exposure, and that comes from knowledge about timing and mechanism. The data alone cannot tell you.

**Sarah:** So how does a student tell them apart?

**Kiffer:** I give three questions. The first is about direction. Does the arrow point from the third variable into the exposure, or from the exposure into the third variable? Into the exposure fits a confounder. Out of the exposure fits a mediator. The second is about timing. Did the variable take its value before the exposure or after it? A confounder comes first. A mediator comes after the exposure and before the outcome. The third question is about change. If you could change the exposure, would this variable change as a result? If it would, the variable is downstream of the exposure and cannot be a confounder.

**Sarah:** Can you run depression through those three questions?

**Kiffer:** Depression before the survey points into loneliness, it was present before the survey, and changing loneliness could not change it, so it is a confounder. Depression after the survey is pointed to by loneliness, develops afterwards, and would change if loneliness changed, so it is a mediator. The condition is the same in both cases, and the only thing that separates the two roles is time.

**Sarah:** Which means the team has to know when depression was measured.

**Kiffer:** It does. A single measure of depression with no date cannot be given a role.

**Sarah:** Let's turn to colliders, which I find the hardest of the three.

**Kiffer:** Most people do. A collider is a variable caused by two other variables in the diagram, so two arrowheads meet, or collide, at it. Colliders cause trouble only when a study restricts its sample to one level of the collider, or adjusts for it.

**Sarah:** The lesson uses a worked example. Walk us through it.

**Kiffer:** Imagine one thousand older adults, of whom two hundred have poor balance and two hundred have poor vision, and the two problems are unrelated. Twenty percent of people with poor balance have poor vision, and twenty percent of people with good balance have poor vision. A falls-prevention program enrols anyone with poor balance or poor vision. Enrolment is caused by both, so it is a collider. The program ends up with three hundred and sixty people. Among those with good balance, every single one has poor vision, because poor vision was their only reason to enrol. A researcher who studied only the program would wrongly conclude that good balance goes with poor vision. Selecting on the collider created the association.

**Sarah:** Is there a collider in Cedar Valley?

**Kiffer:** Referral to the health authority's community connector program. Clinicians refer older adults who seem lonely, and emergency department staff refer patients who keep coming back. So loneliness and visits both cause referral. If the team analyzed only referred adults, or adjusted for referral, it could produce a misleading association, just as the falls program did.

**Sarah:** So what does the final Cedar Valley DAG look like?

**Kiffer:** It has twelve nodes. There are eight confounders: age, income, living alone, rural residence, hearing loss, chronic conditions, earlier depression and frailty, which is unmeasured. There is one mediator, later depression, and one collider, connector referral, alongside the exposure and the outcome.

**Sarah:** That brings us to section four, turning the DAG into something the team can use to collect data.

**Kiffer:** The bridge is a variable list. It is a table with one row for every node in the DAG. Each row records the role, the time at which the variable takes its value, the source of the data, and the planned use.

**Sarah:** Why bother with a separate table? The DAG already shows the roles.

**Kiffer:** The DAG does not say where the data will come from or when, and the variable list does three jobs. It turns the diagram into a practical plan for the survey and the data request. It makes unmeasured nodes visible early, so they can be discussed as limitations. And it justifies every piece of personal information the team wants to collect.

**Sarah:** That last one sounds like it connects to ethics.

**Kiffer:** It does, and Lesson five develops it. Research ethics boards expect researchers to explain why they need each piece of personal information. The team collects income, for example, because income is a common cause of loneliness and emergency visits, and it can point to the arrows that say so.

**Sarah:** Walk me through how the list is built.

**Kiffer:** There are five steps. List every node, using the exact names from the model code, including unobserved ones like frailty. Record each role. Record the timing, which in Cedar Valley means before the survey, at the survey, or during the twelve months of follow-up. Identify a source. And state the planned use in plain words.

**Sarah:** What are the sources for Cedar Valley?

**Kiffer:** Loneliness, age, income, living alone, hearing loss and rural residence come from the survey. Chronic conditions and both depression variables come from physician billing and hospital records linked through Population Data British Columbia for respondents who consent, and emergency visits come from linked emergency department records. Frailty has no source, and connector referral is deliberately not collected.

**Sarah:** And the planned uses?

**Kiffer:** For confounders, account for as a confounder. For the mediator, do not treat as a confounder, and keep for a later mediation analysis. For the collider, do not restrict the sample to referred adults and do not adjust for referral. For frailty, record as a limitation. Health Sciences three forty-one teaches the formal rules for deciding adjustments, so at this stage the plan is provisional.

**Sarah:** Then comes the data dictionary.

**Kiffer:** A data dictionary describes every variable in the dataset so that anyone using the data knows what each column contains. Each entry has a short name, a readable label, a definition, the allowed values and codes, the type of variable, the source and timing, and any notes.

**Sarah:** The lesson distinguishes two kinds of definition.

**Kiffer:** A conceptual definition says what the idea means in words. Daniel Perlman and Letitia Anne Peplau defined loneliness as the unpleasant experience that occurs when a person's social relationships fall short in quantity or quality. An operational definition says exactly how the idea is measured in your study, which for Cedar Valley is a score of six or higher on the three-item scale.

**Sarah:** The last step is a gap check. What is that?

**Kiffer:** It compares the DAG with the data the team can obtain, and the Cedar Valley check found four things. Frailty has no source, so it goes down as a limitation, and the team asked the advisory group whether one question on difficulty leaving the home could give a partial measure. Second, later depression and emergency visits are measured over the same twelve months, so the data cannot always show which came first. The team set a rule that any later mediation analysis would use depression in the first six months and visits in the following six. Third, several confounders are measured at the same moment as loneliness, so the DAG assumes they were in place beforehand. And fourth, physical inactivity was dropped as a mediator because the survey measured it at the same time as loneliness.

**Sarah:** Are there variables that appear in the dataset but not in the DAG?

**Kiffer:** Yes. Gender and the language spoken at home describe who took part and help readers judge to whom the findings apply. They go in the data dictionary with the role descriptive, and Lesson eleven shows how they appear in a Table one.

**Sarah:** Let's finish by putting it together. What does the whole chain look like for Cedar Valley?

**Kiffer:** It starts with the PECO question from Lesson two. The team built an evidence table from its key papers and from what its advisory group and clinic staff told it, and wrote each finding as a link statement. It drew a causal web that includes structural causes, then a first DAG in DAGitty, saved the model code as a versioned text file, and exported a picture. It wrote a justification table for the arrows and for the important missing arrows, labelled every node with its role, and turned the labelled diagram into a variable list, a data dictionary and a gap check.

**Sarah:** What changes when the question is qualitative, in the SPIDER format?

**Kiffer:** The causal web becomes a map of the factors that shape the experience being studied. The DAG is then drawn for one related quantitative question, which is exactly what the Cedar Valley team did alongside its own qualitative question. And the variable list still matters, because it grows into the data dictionary that guides data collection.

**Sarah:** If students remember one thing from this lesson, what should it be?

**Kiffer:** When a third variable puzzles you, ask about direction, timing and change. Those three questions resolve most of the confusion between mediators and confounders, and they will serve students well in Health Sciences three forty-one and four-ten.

**Sarah:** Thanks, Kiffer. Next week is Lesson four, on interest holder mapping and engagement.

**Kiffer:** Thanks, Sarah. See you all next week.
