HSCI 207 · Lesson 3

Causal Webs and Directed Acyclic Graphs

Research Methods in Health Sciences

Learning objectives for this lesson:

  • Explain what a causal web is, where the idea came from, and why researchers draw one before deciding what data to collect.
  • Turn a literature summary and the knowledge of interest holders into link statements and a causal web that includes community and structural causes.
  • Describe how a directed acyclic graph (DAG) differs from a causal web, and explain what an arrow and a missing arrow each claim.
  • Draw, save and export a DAG in DAGitty, redrawing any feedback loop as a sequence of time-stamped nodes.
  • Identify the exposure, outcome, confounders, mediators and colliders in a DAG by working through examples.
  • Distinguish a mediator from a confounder by the direction of arrows, the order of events and what a change in the exposure could change.
  • Convert a DAG into a variable list and a first data dictionary, and record unmeasured variables and timing problems in a gap check.

This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University. It is the applied research methods course of the Public Health Assessment and Analysis series.

Lesson 3 · HSCI 207

Causal Webs and Directed Acyclic Graphs

This short walkthrough orients you before you work through the lesson at your own pace.

Research Methods in Health Sciences
Why this lesson

Deciding what to measure

Too little

A study that measures only the exposure and outcome cannot separate a real effect from the influence of other causes.

Too much

A study that measures everything is long, costly and intrusive for the people who take part.

Drawing the causes first turns a research question into a measurement plan.

The running case

The fictional Cedar Valley Social Connection Study

1,600adults aged 65 and older completed the regional survey.
392respondents (24.5 percent) scored 6 or higher on the UCLA scale.
12 monthsof emergency department visits form the outcome.
The plan

Four sections

1. Causal webs

Evidence becomes link statements and a web.

2. Drawing a DAG

Nodes, arrows and missing arrows are drawn in DAGitty.

3. Variable roles

Each variable is named by its role.

4. Variable list

The DAG becomes a data collection plan.

By the end of the lesson

What you will be able to build

An evidence tableA causal webA first DAG in DAGittyA justification tableA starter variable list

Each section closes with a knowledge check and a reflection.

Section 1 of 5

From a Literature Summary to a Causal Web

⏱ Estimated reading time: 40 minutes
Section 1 of 5

From a Literature Summary to a Causal Web

This section covers the web of causation, link statements and the step from a web to a DAG.

The idea

The web of causation

Many causes

Most outcomes arise from several causes acting together.

Causes have causes

Living alone may follow bereavement, for example.

Several levels

Causes belong to individuals, relationships and communities.

Loops

Two factors can affect each other over time.

The idea comes from MacMahon, Pugh and Ipsen (1960).

A critique

Has anyone seen the spider?

Epidemiology and the web of causation: has anyone seen the spider?Title of Krieger, 1994

The concern

Webs tended to omit the social and economic conditions behind individual risk.

The practical step

Check that income, housing, transport and services appear in your web.

The process

Five steps from evidence to web

1. Start from the question2. Build an evidence table3. Write link statements4. Group factors by level5. Draw and review with others

A link statement names a direction and a source, such as "loneliness leads to later depressive symptoms (Cacioppo et al., 2006)".

The running case

The Cedar Valley causal web

Causes of loneliness

Bereavement, living alone, hearing loss, limited transport, rural residence, low income and chronic conditions are on the left.

Pathways

Loneliness may raise ED visits directly and through depression and physical inactivity.

Loops

Loneliness and depression affect each other, and inactivity worsens chronic conditions.

Comparison

Causal web and DAG

Causal web

Boxes may be broad ideas, arrows may run both ways, and loops are allowed.

Directed acyclic graph

Each node is one measurable variable, each arrow runs one way, and loops are not allowed.

In a DAG, a missing arrow is a deliberate claim of no direct effect.

Carry forward

From web to graph

  • A causal web gathers possible causes from evidence and experience.
  • Each arrow should trace back to a link statement with a source.
  • Section 2 turns the web into a DAG drawn in DAGitty.

Learning Objectives for this section

  • Explain what a causal web is and why researchers draw one before deciding what data to collect.
  • Describe the origins of the web of causation and Nancy Krieger's critique of it.
  • Turn a short literature summary into a list of cause-and-effect links, each tied to its source.
  • Draw and read a causal web for a research question, and describe how a causal web differs from a directed acyclic graph.

1.1 Why Researchers Draw Causes Before Collecting Data

Lesson 2 turned a broad topic into a structured research question. A question names an exposure and an outcome, but it does not tell you what else to measure. Every outcome in health has many causes, and those causes are tangled together. A study that measures only the exposure and the outcome cannot tell whether an association between them reflects a real effect or the influence of something else, and a study that measures everything becomes long, expensive and intrusive. Researchers therefore need a way to decide, before any data are collected, which variables belong in the study and why.

This lesson teaches two drawings that support that decision. The first is a causal web, a broad diagram of the causes that the literature and experience suggest are connected to the outcome. The second is a directed acyclic graph (DAG), a more disciplined diagram that states exactly which variable is assumed to affect which. In this lesson the word cause has a plain meaning: a variable is a cause of an outcome if changing it would change how likely, or how large, the outcome is. Smoking is a cause of lung cancer in this sense, because a person who smokes more has a higher chance of developing it than the same person would have had after smoking less.

Case: the Cedar Valley Social Connection Study (fictional)

Throughout this course we follow the Cedar Valley Social Connection Study, a fictional mixed-methods study planned by a research team at a British Columbia university with the fictional Cedar Valley Health Authority. The region serves about 210,000 residents, about 46,000 of whom are aged 65 and older. The team is led by Dr. Maya Hart and includes a graduate research assistant, a community research associate and an advisory group of six older adults.

In Lesson 2 the team wrote a PECO question: among adults aged 65 and older in Cedar Valley (population), is loneliness, measured as a score of 6 or higher on the three-item UCLA Loneliness Scale (exposure), compared with a score of 3 to 5 (comparison), associated with the number of emergency department visits in the 12 months after the survey (outcome)? Before building its survey, the team must decide what else to measure. This lesson follows the team as it makes that decision.

1.2 The Web of Causation

The idea that disease has a single cause served public health well in the era of infectious disease, when identifying the organism behind cholera or tuberculosis pointed directly to prevention. It served less well for heart disease, cancer and other chronic conditions, which arise from many factors acting together over long periods. In their 1960 textbook Epidemiologic Methods, Brian MacMahon, Thomas Pugh and Johannes Ipsen described the causes of disease as a web of causation, in which many factors are linked to the outcome and to one another (MacMahon et al., 1960). Each cause in a web has causes of its own, and an intervention can act at any strand that is open to change.

Many causesClick to explore
Causes have causesClick to explore
Causes act at several levelsClick to explore
Webs can contain loopsClick to explore

The web of causation has been criticized as well as used. In a widely cited paper, Nancy Krieger (1994) asked whether anyone had seen the spider: if disease arises from a web, what force spins it? She argued that webs drawn in epidemiology tended to place individual behaviours and biological factors in the foreground, while the social, economic and political conditions that shape those factors were left out or treated as background. Krieger went on to develop ecosocial theory. For a student drawing a first causal web, her critique has a practical consequence. A web that contains only individual risk factors will lead to a study that measures only individual risk factors, so the outer levels of the web, such as income, housing, transport and access to services, deserve deliberate attention.

HSCI 341 Lesson 1 Section 3 returns to the web of causation to separate direct from indirect causes and to link the web with the component-cause model.

1.3 From a Literature Summary to a List of Links

A causal web is built from evidence and from the experience of people who know the problem. Lesson 2 showed how to find three to five key papers on a topic. Those papers, together with the knowledge of people affected by the problem, supply the raw material. The process has five steps, and each step produces a record that the team keeps with its study files.

Step 1: Start from the question

Write the exposure and the outcome at the top of a page. Everything that follows is chosen because it may be connected to one or both of them. For Cedar Valley, the exposure is loneliness and the outcome is the number of emergency department (ED) visits in the 12 months after the survey.

Step 2: Build an evidence table

For each source, record what kind of evidence it is and what it suggests about causes. Write each finding in your own words beside its citation. The table below shows part of the Cedar Valley team's evidence table. The published studies are real and their general findings are summarized briefly; the last two rows record knowledge from the team's advisory group and partner clinics.

SourceType of evidenceWhat it suggestsLinks for the web
Holt-Lunstad et al. (2015)Meta-analysis of prospective studiesLoneliness and social isolation were associated with higher mortality.Loneliness affects physical health.
Valtorta et al. (2016)Systematic review and meta-analysisLoneliness and social isolation were associated with a higher risk of coronary heart disease and stroke.Loneliness leads to chronic conditions.
Cacioppo et al. (2006)Cross-sectional and longitudinal analyses of middle-aged and older adultsLoneliness was associated with depressive symptoms and predicted later depressive symptoms.Loneliness leads to depression.
Hawkley et al. (2009)Cross-sectional and longitudinal analyses of middle-aged and older adultsLoneliness predicted reduced physical activity.Loneliness leads to physical inactivity.
Gerst-Emerson & Jayawardhana (2015)Longitudinal survey of older adults in the United StatesLoneliness was associated with more physician visits.Loneliness affects health service use.
Cedar Valley advisory groupLived experience of six older adultsBereavement often leads to living alone; losing a driver's licence or living far from town reduces contact; hearing loss makes conversation tiring.Bereavement leads to living alone; rural residence and limited transport lead to loneliness; hearing loss leads to loneliness.
Partner clinic staffPractice knowledgePatients with several chronic conditions use the ED more, especially where after-hours primary care is scarce.Chronic conditions lead to ED visits.

Step 3: Write each link as a short statement

Rewrite the last column as a list of link statements of the form "A affects B", one link per line, with the source in brackets. A link statement makes the direction explicit. "Loneliness and depression are related" is too vague to draw, whereas "loneliness leads to later depressive symptoms (Cacioppo et al., 2006)" names a direction and a source. Where sources disagree about direction, write both statements and note the disagreement, because Section 2 shows how to handle a two-way relationship.

Step 4: Group the factors by level

Sort the factors into rough levels: individual health (chronic conditions, hearing loss, depression), relationships (living alone, bereavement), and community or structural conditions (income, rural residence, transport). Then ask, with Krieger's critique in mind, whether any level is thin. The Cedar Valley team found its first list thin on transport and income and returned to the advisory group's notes.

Step 5: Draw the web and review it with others

Place the exposure near the centre and the outcome to its right, add each factor, and draw an arrow for each link statement. Then show the draft to people who know the problem. Lesson 4 develops interest holder engagement in detail; at this stage, a conversation with a few people who live with the problem or work with it often reveals missing factors. A more formal and reproducible version of this process, called evidence synthesis for constructing directed acyclic graphs (ESC-DAGs), has been published by Ferguson and colleagues (2020). This lesson uses a simplified form suited to a first project, and systematic searching itself is taught in HSCI 241.

Try it: extract the links

Read the two sentences below, which could come from the results section of a study. Write each causal claim as a link statement of the form "A affects B", and name the level each factor belongs to.

"Older adults who had stopped driving reported fewer weekly social contacts and higher loneliness scores. Those living in households with lower income were more likely to have stopped driving."

A good answer contains three statements: lower income affects whether a person stops driving; stopping driving affects the number of weekly social contacts; and fewer contacts affect loneliness. The factors run from the structural level (income) through an individual circumstance (driving) to relationships (contacts) and the individual experience of loneliness. Because the sentences report associations, the direction of each arrow is an assumption the researcher brings to the evidence.

1.4 The Cedar Valley Causal Web

The figure below shows the team's first causal web. Read it from left to right. The factors on the left (bereavement, living alone, hearing loss, limited transport and rural residence) are drawn as causes of loneliness. Loneliness sits in the centre, and the outcome, ED visits, sits on the right. Chronic conditions and low income appear as factors linked to loneliness, and two pathways from loneliness toward ED visits pass through depression and physical inactivity.

Bereavement Living alone Hearing loss Limited transport Rural residence Low income Loneliness Chronic conditions Depression Physical inactivity ED visits
The Cedar Valley team's first causal web. Teal arrows show assumed causal links. The two red arrows close feedback loops: loneliness and depression affect each other, and physical inactivity worsens chronic conditions, which in turn raise loneliness. ED stands for emergency department.

Three features of the web deserve attention. First, every arrow traces back to a row of the evidence table, so a reader can ask where each link came from. Second, the web contains two loops, shown in red. Loneliness and depression are drawn as affecting each other, and a second loop runs from chronic conditions to loneliness to physical inactivity and back to chronic conditions. Third, the web does not say whether chronic conditions come before loneliness or after it, or which factors the team can measure. A causal web is a thinking tool for gathering possibilities, and it is deliberately broad.

1.5 From a Causal Web to a DAG

A causal web is good at showing how much is going on. It is less good at telling the team exactly what to measure, because its boxes can be broad, its arrows can run in both directions, and it does not distinguish what happened first. A directed acyclic graph imposes three disciplines. Each box, called a node, is a single variable that could in principle be measured at a particular time. Each arrow points one way only, from cause to effect, which is what "directed" means. And the diagram contains no loops, which is what "acyclic" means: if you start at any node and follow the arrows, you can never return to where you started. "Graph" is the mathematical name for a set of nodes joined by lines.

DAGs grew out of path diagrams, which the geneticist Sewall Wright developed in the 1920s, and were given a formal footing by the computer scientist Judea Pearl in the 1990s (Pearl, 1995). Sander Greenland, Judea Pearl and James Robins introduced them to epidemiologists in an influential paper (Greenland et al., 1999), and they are now widely used in health research to plan which variables to measure and how to analyze them (Hernán & Robins, 2020; Tennant et al., 2021). The table below summarizes how the two diagrams differ.

FeatureCausal webDirected acyclic graph
PurposeIt gathers the possible causes suggested by evidence and experience.It states the team's assumptions precisely enough to plan data collection and analysis.
BoxesA box may be a broad idea, such as social factors.A node is one variable that could be measured at a stated time.
ArrowsArrows may run both ways between two factors.Each arrow runs one way, from cause to effect.
LoopsLoops are allowed.Loops are not allowed; a two-way relationship is drawn as a sequence over time.
Missing arrowsA missing arrow usually means nobody drew it.A missing arrow is a deliberate claim that one variable has no direct effect on another.

The web and the DAG are two stages of the same work. The web keeps the team honest about the breadth of causes, including the structural ones that Krieger warned are easily dropped. The DAG then forces the team to decide, for each pair of variables, whether one affects the other and in what order. Section 2 shows how to make those decisions and draw the result in DAGitty.

Problem 1: Too many boxesv

A web with thirty or more factors is hard to read and harder to turn into a DAG. Keep the factors that the evidence links to the exposure, the outcome or both, and record the others in a separate list in case they become relevant later.

Problem 2: Arrows without sourcesv

An arrow that cannot be traced to a study, a theory or the knowledge of people affected by the problem is a guess. Label guesses as such so that reviewers can question them.

Problem 3: Vague boxesv

Boxes such as "social factors" or "lifestyle" cannot be measured and hide several distinct causes. Split them into specific factors, such as living alone, frequency of contact with family, and physical activity.

Problem 4: Only individual factorsv

A web made entirely of individual behaviours and diagnoses repeats the pattern Krieger (1994) criticized. Check that community and structural conditions, such as income, housing, transport and access to services, have been considered.

Problem 5: Arrows that point backward in timev

An arrow from a later event to an earlier one cannot be causal. An ED visit in the year after the survey cannot cause loneliness measured at the survey, even though a past ED visit might. Section 2 shows how to label variables by time to avoid this error.

Where this goes next

Section 2 turns the causal web into a first DAG. The evidence table and link statements built with the five steps above are the working materials for that DAG, and each arrow in it should trace back to one of them.

Reflection

A student team asks whether food insecurity is associated with poorly controlled type 2 diabetes among adults in a Canadian city. Its evidence table contains four findings: (1) adults with lower incomes were more likely to be food insecure; (2) food-insecure adults more often skipped meals and bought cheaper foods high in refined carbohydrates; (3) diets high in refined carbohydrates were associated with higher blood glucose; and (4) clinic staff reported that some patients facing food insecurity ration their diabetes medication to afford food.

A link statement has the form "A affects B" and names its source. Nancy Krieger (1994) argued that causal webs in epidemiology tend to leave out the social, economic and political conditions that shape individual risk factors. (a) Write the link statements contained in the four findings. (b) Sort the factors into individual, relationship and community or structural levels. (c) Using Krieger's critique, name one structural factor the team should consider adding to its web and state the link it would add.

Model answer

(a) The findings contain five link statements: lower income affects food insecurity (finding 1); food insecurity affects meal skipping and the purchase of foods high in refined carbohydrates (finding 2); a diet high in refined carbohydrates affects blood glucose control (finding 3); food insecurity affects rationing of diabetes medication (finding 4, clinic staff); and medication rationing affects blood glucose control (finding 4, by clinical reasoning). The findings report associations, so each direction is an assumption the team should state.

(b) Diet, meal skipping, medication rationing and blood glucose control are individual-level factors. Food insecurity describes a household's circumstances and sits between the individual and relationship levels. Income is a structural factor.

(c) The web says nothing about the cost of medication or about drug coverage, both of which are shaped by public policy and bear directly on rationing. I would add the link statement "lack of drug coverage affects medication rationing" and look for Canadian studies of cost-related non-adherence to support it. Housing costs would be another strong addition, since rent paid from a fixed income reduces the money left for food.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: Which authors described the causes of disease as a "web of causation" in a 1960 textbook?

MacMahon, Pugh and Ipsen introduced the web of causation in Epidemiologic Methods (1960). Krieger criticized the idea in 1994, and Greenland, Pearl and Robins introduced DAGs to epidemiology in 1999.

Question 2: In asking whether anyone had seen the spider, what did Nancy Krieger (1994) argue about the web of causation?

Krieger argued that webs placed individual behaviours and biology in the foreground and left out the social, economic and political conditions that spin the web. She did not call for single-cause models.

Question 3: Which of the following is a link statement that can be drawn directly as an arrow in a causal web?

A link statement names a direction (A affects B) and a source. "Closely related" gives no direction, and "social factors" and "a vulnerable group" are too vague to draw as a single arrow.

Question 4: Which feature is allowed in a causal web but not in a directed acyclic graph?

A DAG is acyclic, so it cannot contain a loop; a two-way relationship must be redrawn as time-stamped nodes. Cited arrows, structural factors and large diagrams are all allowed in a DAG.
Section 2 of 5

Drawing a DAG in DAGitty: Nodes, Arrows and the Absence of Arrows

⏱ Estimated reading time: 40 minutes
Section 2 of 5

Drawing a DAG in DAGitty

This section covers nodes, arrows, the absence of arrows and the DAGitty workflow.

Vocabulary

The parts of a DAG

Node

A variable that could be measured at a stated time.

Arrow

It points from an assumed cause to its effect.

Direct effect

It does not pass through another node.

Ancestors and descendants

They are the nodes upstream and downstream of a node.

Claims

Arrows and missing arrows

An arrow

A may directly affect B, by any amount and in either direction.

A missing arrow

A has no direct effect on B, which is a strong assumption to defend.

The advisory group's comment led the team to add an arrow from hearing loss to ED visits.

Conventions

Four drawing conventions

One variable, one time

Each node can be measured at a stated time.

Forward in time

Time runs from left to right across the diagram.

Unroll loops

A two-way relationship becomes time-stamped nodes.

Include shared causes

Unmeasured shared causes such as frailty stay in, marked as unobserved.

Convention 3

Unrolling a loop into time

Loneliness Depression Earlier depression Loneliness Later depression Causal webDAG, ordered in time
DAGitty, steps 1 to 5

Building the diagram

  • Open dagitty.net and choose "Launch DAGitty online in your browser".
  • Choose Model, then New model, and enter the exposure and the outcome.
  • Double-click empty canvas, or press n, to add a variable.
  • Double-click the cause and then the effect to add an arrow.
  • Repeat the double-clicks to remove an arrow; press d over a node to delete it.
DAGitty, steps 6 to 10

Arranging, saving and exporting

Arrange and label

Drag nodes into time order and set exposure, outcome and unobserved status.

Save and export

Save the Model code as a versioned text file and export a PNG or PDF.

Age -> Loneliness is how one arrow appears in the model code.

Carry forward

From drawing to roles

  • A missing arrow is a claim that needs a justification.
  • Loops are unrolled into time-stamped nodes.
  • Section 3 names the role of every node in the DAG.

Learning Objectives for this section

  • Define node, arrow, ancestor and descendant, and read a simple DAG aloud.
  • Explain what an arrow claims, what a missing arrow claims, and what a DAG does not show.
  • Apply four drawing conventions, including redrawing a feedback loop as a sequence over time.
  • Draw, save and export a DAG in DAGitty, either by pointing and clicking or by typing its model code.

2.1 The Parts of a DAG

A directed acyclic graph has two kinds of parts. A node is a variable, drawn as a box or a name. An arrow (also called an edge) joins two nodes and points from the variable assumed to be a cause to the variable assumed to be its effect. When an arrow runs directly from A to B, A is said to have a direct effect on B, meaning an effect that does not pass through any other variable shown in the diagram. Two more words help when reading larger diagrams. A node's ancestors are all the nodes from which you can reach it by following arrows forward, and its descendants are all the nodes you can reach from it by following arrows forward. DAGitty uses these words in its legend, so they are worth learning now.

The figure below shows a first, small DAG drawn by the Cedar Valley team. It contains four nodes: age, hearing loss, loneliness and ED visits. Read aloud, it says that age affects hearing loss, loneliness and ED visits directly; that hearing loss affects loneliness; and that loneliness affects ED visits. Age is therefore an ancestor of every other node, and ED visits is a descendant of every other node.

Age Hearing loss Loneliness ED visits Arrow: an assumed direct effect No arrow: no direct effect assumed
A first, four-node DAG for the Cedar Valley study. Loneliness (the exposure) is filled in teal and ED visits (the outcome) in dark grey. The dashed line is drawn here only to point out a missing arrow; in a real DAG nothing is drawn between hearing loss and ED visits.

2.2 What an Arrow Claims, and What a Missing Arrow Claims

An arrow from A to B is a modest claim. It says that A might have a direct effect on B, and that the team is not prepared to rule this out. It does not say how large the effect is, whether it raises or lowers B, or whether it is the same for everyone. Adding an arrow is the cautious choice when the evidence is unclear, because an arrow allows for an effect without assuming one.

A missing arrow is a strong claim. If there is no arrow from A to B, the diagram asserts that A has no direct effect on B at all, so that any influence of A on B must travel through other variables shown in the diagram. In the four-node DAG above, the missing arrow between hearing loss and ED visits asserts that hearing loss affects ED visits only by making people lonelier. When Dr. Hart showed this draft to the advisory group, one member pointed out that hearing loss can make it harder to hear traffic, alarms or a clinician's instructions, and that several people she knew had been injured in falls they connected with poor hearing. That is a route from hearing loss to ED visits that does not pass through loneliness. The team added an arrow from hearing loss to ED visits. Section 3 shows that this single decision changed the role hearing loss plays in the study.

What an arrow saysClick to explore
What a missing arrow saysClick to explore
What a DAG does not showClick to explore
Why time mattersClick to explore

2.3 Four Conventions for Drawing a First DAG

Experienced researchers follow a few conventions that keep a DAG honest and readable. They are drawing habits, and they do not require the formal rules for reading a DAG. HSCI 230 Lesson 7 Section 2 introduces those rules through the three basic structures and collider bias, and HSCI 341 Lesson 7 Section 3 states them in full as the backdoor criterion and a procedure for choosing an adjustment set.

Convention 1: One node is one variable at one time

A node should be something that could be measured, at least in principle, at a stated time. "Social factors" fails this test. "Living alone at the time of the survey" passes it. When a factor could be measured at two different times and the timing matters, it becomes two nodes.

Convention 2: Arrows point forward in time

Arrange the nodes so that time runs from left to right. Factors fixed early in life or measured before the survey go on the left, the exposure in the middle, and the outcome and anything measured during follow-up on the right. If an arrow points from right to left, check whether it describes something that happened earlier.

Convention 3: Unroll loops into time

The causal web in Section 1 showed loneliness and depression affecting each other. A DAG cannot contain that loop. The solution is to split the factor into time-stamped nodes: depression before the survey can affect loneliness at the survey, and loneliness at the survey can affect depression during the following year. The figure below shows the change.

Causal web: a loop DAG: the loop unrolled in time Loneliness Depression Earlier depression Loneliness Later depression Before survey Survey Follow-up
A two-way relationship in a causal web (left) becomes a sequence of time-stamped nodes in a DAG (right). The arrow from earlier to later depression reflects the fact that depressive symptoms often persist.

Convention 4: Include shared causes, even unmeasured ones

If a variable causes two or more of the variables already in the diagram, it belongs in the diagram, even if the team cannot measure it. Frailty, for example, could make it hard for an older adult to leave the house (raising loneliness) and could also raise the risk of injury (raising ED visits). The Cedar Valley survey does not measure frailty. Leaving it out of the DAG would hide a problem the team needs to discuss. DAGitty lets you mark such a node as unobserved. A variable that affects only one node, and is of no interest in itself, can usually be left out to keep the diagram readable.

2.4 Drawing the DAG in DAGitty, Step by Step

About DAGitty

DAGitty is free software for drawing and analyzing causal diagrams, developed by Johannes Textor and colleagues (Textor et al., 2016). It runs in a web browser at dagitty.net with no installation or account, and the same tool is available as a package for the R programming language. The steps below use the browser version and the labels it shows on screen.

The steps build a six-node version of the Cedar Valley DAG with the nodes Loneliness, ED_visits, Age, Chronic_conditions, Living_alone and Later_depression. Use the _ character in place of spaces in node names. DAGitty can handle names with spaces by placing them in double quotation marks in its code, but names without spaces are easier to type and can later serve as variable names in a dataset (Section 4).

  1. Open the program. Go to dagitty.net and select the link labelled "Launch DAGitty online in your browser". The screen shows a menu bar along the top, a drawing area (the canvas) in the middle, and panels of options and information at the sides.
  2. Start a new model. From the Model menu, choose New model. DAGitty asks for the name of the exposure and then the name of the outcome. Type Loneliness and then ED_visits. DAGitty draws the two nodes with an arrow from the exposure to the outcome. Keep this arrow, because it represents the effect your question asks about.
  3. Add the other variables. Double-click an empty spot on the canvas, or move the mouse pointer to an empty spot and press the n key. Type the variable's name in the box that appears and press Enter or click OK. Add Age, Chronic_conditions, Living_alone and Later_depression in this way.
  4. Add arrows. Double-click the node that is the cause; it becomes highlighted. Then double-click the node that is the effect. An arrow appears between them. As an alternative, move the pointer over the cause and press c, then move it over the effect and press c again. Add each arrow from your link list in turn.
  5. Correct mistakes. To remove an arrow, repeat the same two double-clicks that created it. To delete a variable, move the pointer over it and press d or the Delete key; its arrows disappear with it. To rename a variable, move the pointer over it and press r.
  6. Arrange the layout. Drag the nodes with the mouse so that time runs from left to right and arrows cross as little as possible. Put the exposure toward the left of centre and the outcome toward the right.
  7. Set each variable's status. Click a node and use the checkboxes in the panel labelled Variable, or move the pointer over the node and press e for exposure, o for outcome or u for unobserved. The exposure and outcome are already set from step 2. If you add Frailty, mark it as unobserved.
  8. Read the model code. Find the text box labelled Model code. DAGitty writes your diagram there as text, with each arrow on a line such as Age -> Loneliness. Compare the list with your link statements to check that nothing is missing or reversed.
  9. Save your work. Copy all the text in the Model code box and paste it into a plain text file in the study folder, with a name and version number such as CedarValley_DAG_v1.txt. To reopen the diagram later, paste the text back into the Model code box and click Update DAG.
  10. Export a figure. The Model menu offers export as a PDF, SVG, PNG or JPEG file. A PNG file is the simplest to insert into a document, and a PDF or SVG file stays sharp when enlarged.

As you work, DAGitty colours the nodes and arrows and fills a panel labelled Causal effect identification with suggestions. These features apply the formal rules for reading a DAG, which HSCI 230 Lesson 7 Section 2 introduces and HSCI 341 Lesson 7 Section 3 states in full as the backdoor criterion. For this course, treat DAGitty as a drawing and record-keeping tool and leave those panels for later.

Typing the model code directly

Some people find it faster to type the diagram than to click it, and typing also works well with a keyboard or a screen reader. The model code for the six-node DAG is shown below. Each line inside the braces either declares a variable's status in square brackets or states one arrow. Paste the code into the Model code box and click Update DAG, and DAGitty draws the diagram. The nodes may land in an untidy arrangement, and dragging them into place adds position information to the code automatically.

dag {
Loneliness [exposure]
ED_visits [outcome]
Age -> Loneliness
Age -> ED_visits
Age -> Chronic_conditions
Age -> Living_alone
Chronic_conditions -> Loneliness
Chronic_conditions -> ED_visits
Living_alone -> Loneliness
Living_alone -> ED_visits
Loneliness -> Later_depression
Later_depression -> ED_visits
Loneliness -> ED_visits
}
Try it: draw the six-node DAG

Open DAGitty and build the six-node Cedar Valley DAG by pointing and clicking, following steps 1 to 7. Then compare the text in your Model code box with the code above. Your diagram should contain six nodes and eleven arrows. If your count differs, look for an arrow drawn in the wrong direction or a node name spelled two different ways (for example, ED_visits and ED_Visits, which DAGitty treats as two separate variables).

When the counts match, save the model code as a text file and export the diagram as a PNG file. You will extend this DAG in Section 3.

A two-headed arrow appearedv

If you add an arrow in the opposite direction to one that already exists, DAGitty replaces the pair with a single two-headed arrow, which has a special meaning in causal diagrams. Remove it by repeating the double-clicks, and then decide which direction you intended. If the two variables really do affect each other, unroll the loop into time as Convention 3 describes.

Double-clicking is awkwardv

Use the keyboard alternatives: n for a new variable, c over the cause and then over the effect for an arrow, d to delete, r to rename, and e, o or u to set status. Each shortcut acts on the node under the mouse pointer. You can also type the model code directly.

The diagram looks like a tanglev

Drag the nodes into columns by time, with earlier variables on the left. If arrows still cross, try moving the shared causes above the exposure and the outcome. A readable DAG is one in which a reader can follow each arrow without tracing it with a finger.

I lost my diagramv

Save the model code to a file at the end of every work session, with a new version number whenever you make a substantial change. A text file you control is the most reliable record, and it also documents how the diagram changed over time, which is useful when the team explains its decisions in a protocol or report.

2.5 Documenting the Arrows

A DAG is easier to review when each arrow, and each important missing arrow, comes with a one-line justification. The justification can cite a study, a theory, the knowledge of interest holders, or a logical point such as the order of events. The Cedar Valley team kept a justification table beside its DAG. Part of it is shown below.

Arrow or missing arrowJustificationSource
Loneliness to Later_depressionLoneliness predicted later depressive symptoms in a study of middle-aged and older adults.Cacioppo et al. (2006)
Chronic_conditions to ED_visitsPeople with several chronic conditions have more acute episodes that bring them to the ED.Partner clinic staff
Age to Living_aloneWidowhood becomes more common with age, and many widowed adults live alone.Advisory group; logical order of events
Hearing_loss to ED_visits (added after review)Hearing loss may contribute to falls and to misunderstood instructions, routes that do not pass through loneliness.Advisory group
No arrow from ED_visits to LonelinessED visits are counted in the 12 months after the survey, so they cannot affect loneliness measured at the survey.Order of events

The last row illustrates a missing arrow that is easy to defend, because the timing makes it impossible. Most missing arrows are harder to defend, and a justification table makes the weaker ones visible. Tennant and colleagues (2021), reviewing published health studies that used DAGs, recommended that researchers report their full diagrams and explain how they were built. A justification table is a simple way to meet that recommendation in a small study. Section 3 now uses the finished diagram to name the role each variable plays.

Reflection

A researcher studies whether poor sleep (exposure) is associated with falls (outcome) over one year among adults aged 70 and older. She believes that pain disturbs sleep, that poor sleep makes pain worse the next day, and that sedating sleep medication, taken by some people who sleep poorly, causes dizziness that raises the risk of falls. She also believes that age affects sleep, pain and falls.

In a directed acyclic graph (DAG), every arrow points one way, from cause to effect, and the diagram may contain no loops. A missing arrow claims that one variable has no direct effect on another. (a) Explain why pain and sleep cannot be drawn as affecting each other in a DAG, and describe how to redraw that relationship with time-stamped nodes. (b) Describe the full DAG in words, listing each arrow. (c) Choose one missing arrow in your DAG, state what it claims, and say whether you could defend it.

Model answer

(a) Pain affecting sleep and sleep affecting pain form a loop, and a DAG cannot contain a loop because following the arrows would return you to your starting point. I would split pain into two nodes: pain at baseline, which affects sleep, and pain during follow-up, which poor sleep can worsen.

(b) The DAG contains these arrows: Age to Baseline_pain, Age to Poor_sleep, Age to Falls; Baseline_pain to Poor_sleep; Baseline_pain to Later_pain (pain tends to persist); Poor_sleep to Later_pain; Poor_sleep to Sleep_medication; Sleep_medication to Falls; Poor_sleep to Falls. I would add Later_pain to Falls, since pain can limit movement and balance.

(c) My diagram has no arrow from Baseline_pain to Sleep_medication. That missing arrow claims that pain affects the use of sleep medication only by disturbing sleep. I do not think I could defend it, because people in pain may be prescribed medications that are themselves sedating, or may take sleep medication to cope with night-time pain. I would add the arrow, and I would record the change and its reason in my justification table so that a reviewer could see why.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: In a DAG, what does a missing arrow from A to B claim?

A missing arrow is a strong claim that A has no direct effect on B. A and B may still be associated, for example through a shared cause, so the claim concerns direct effects and not associations.

Question 2: The Cedar Valley causal web shows loneliness and depression affecting each other. How should this be drawn in a DAG?

A DAG unrolls a loop into time. Earlier depression can affect loneliness at the survey, and loneliness can affect later depression. Deleting depression would hide a confounder and a mediator.

Question 3: In DAGitty, how do you add an arrow from Age to Loneliness?

Double-click the cause first (it becomes highlighted) and then the effect. Reversing the order draws the arrow the wrong way.

Question 4: Which of the following can be read from a DAG?

A DAG records which variables are assumed to have direct effects on which others. It does not show the size or direction (harmful or protective) of effects, or whether they differ between groups.
Section 3 of 5

Variable Roles: Exposure, Outcome, Confounder, Mediator and Collider

⏱ Estimated reading time: 40 minutes
Section 3 of 5

Variable Roles

This section covers the exposure, outcome, confounder, mediator and collider.

Three patterns

Roles are defined by the question

ConfounderMediatorCollider Confounder Exposure Outcome Mediator Exposure Outcome Exposure Outcome Collider
Confounder

A common cause of both

Teaching example

Smoking causes match carrying and lung cancer, so it confounds their association.

Cedar Valley

Chronic conditions raise both loneliness and ED visits.

When the team added an arrow from hearing loss to ED visits, hearing loss also became a confounder.

Mediator

A step on the pathway

Teaching example

Breakfast improves attention, and attention raises test scores.

Cedar Valley

Loneliness leads to later depression, which leads to ED visits.

Mediation analysis is taught in HSCI 410 Lesson 8.

Telling them apart

Three questions for any third variable

Direction

Does the arrow point into the exposure or out of it?

Timing

Did the variable take its value before or after the exposure?

Change

Would it change if the exposure changed?

About 1 in 4 HSCI 341 students defined a confounder as a variable on the causal pathway.

Collider

A common effect of both

20%of each balance group in the community has poor vision.
360people enrol because they have poor balance or poor vision.
100%of enrolees with good balance have poor vision.
Carry forward

The labelled Cedar Valley DAG

Eight confounders

Age, income, living alone, rural residence, hearing loss, chronic conditions, earlier depression and frailty.

One mediator

Later depression carries part of the effect.

One collider

Connector referral is caused by both loneliness and ED visits.

Learning Objectives for this section

  • Identify the exposure and the outcome in a research question and in its DAG.
  • Define a confounder, give an example, and explain why ignoring a confounder can distort an association.
  • Define a mediator and distinguish it from a confounder using the direction of arrows, the order of events and what the exposure could change.
  • Define a collider and explain, using a worked example, why restricting a study to one level of a collider can create a misleading association.
  • Label every variable in a DAG with its role.

3.1 Roles Are Defined by the Question

Once a DAG is drawn, each variable can be given a role. The roles are defined relative to the research question, so a variable has no role on its own. The exposure is the variable whose effect the question asks about. It may be a risk factor (loneliness), a treatment (a medication), a program (a community connector service) or a condition of life (living in a rural area). The outcome is the variable on which that effect is measured. In the Cedar Valley PECO question, loneliness is the exposure and the number of ED visits in the following 12 months is the outcome. The same variable can play different roles in different studies. Loneliness is the exposure in this study, and it would be the outcome in a study asking whether a connector program reduces loneliness.

Three further roles describe how other variables relate to the exposure and the outcome. A confounder is a common cause of both. A mediator lies on a causal pathway from the exposure to the outcome. A collider is a common effect of two variables in the diagram. The figure below shows the three patterns in their simplest form.

ConfounderMediatorCollider Confounder Exposure Outcome Mediator Exposure Outcome Exposure Outcome Collider A common cause of bothA step on the pathwayA common effect of both
The three basic patterns. What distinguishes them is the direction of the arrows between the third variable and the exposure and outcome.

3.2 Confounders

A confounder is a common cause of the exposure and the outcome, or a proxy for such a cause, that is not on the causal pathway between them. The word comes from the Latin for "to mix together", and the name describes the problem: a confounder mixes its own effect into the association between the exposure and the outcome. A classic teaching example concerns carrying matches and lung cancer. People who carry matches develop lung cancer more often than people who do not. Matches do not cause cancer. Smoking causes people to carry matches, and smoking causes lung cancer, so smoking is a confounder of the association between carrying matches and lung cancer. A study that compared match carriers with non-carriers and ignored smoking would wrongly conclude that matches are harmful.

In Cedar Valley, chronic conditions diagnosed before the survey are a confounder. Conditions such as heart failure, arthritis or chronic lung disease can limit how often a person leaves home and sees others, which raises loneliness. The same conditions produce acute episodes that bring people to the ED. If the team compared ED visits between lonely and non-lonely respondents without accounting for chronic conditions, some of the difference would reflect the greater burden of illness among lonely respondents. The team's DAG contains several other confounders: age, income, living alone, rural residence, earlier depression, frailty and hearing loss.

Hearing loss deserves a second look. In the first draft in Section 2, hearing loss had an arrow into loneliness and no arrow into ED visits, so it was a cause of the exposure only and was not a confounder. After the advisory group's comment, the team added an arrow from hearing loss to ED visits. Hearing loss then became a common cause of loneliness and ED visits, and therefore a confounder. Its role changed because an arrow changed, which is why the justification table in Section 2 matters.

A confounder comes firstClick to explore
Unmeasured confoundersClick to explore
What researchers do with confoundersClick to explore
A cause of the outcome onlyClick to explore

3.3 Mediators

A mediator is a variable that lies on a causal pathway from the exposure to the outcome: the exposure affects the mediator, and the mediator affects the outcome. A mediator carries part of the exposure's effect. A simple example outside the case concerns a school breakfast program. Providing breakfast may improve children's attention in morning classes, and better attention may raise test scores. Attention is a mediator of the effect of the breakfast program on test scores.

In Cedar Valley, depressive symptoms that develop after the survey are a mediator. Loneliness predicts later depressive symptoms (Cacioppo et al., 2006), and depression can lead to ED visits through self-harm, neglect of chronic conditions or physical symptoms that prompt a visit. If loneliness raises ED visits partly because it leads to depression, then depression is one of the routes by which the effect travels. That route is part of the effect the PECO question asks about.

This has a practical consequence. If the team treated later depression as though it were a confounder and removed its influence in the analysis, it would remove part of the very effect it set out to estimate. The analysis would then answer a narrower question about the effect of loneliness through routes other than depression. Studying how much of an effect travels through a mediator is called mediation analysis, which HSCI 410 Lesson 8 teaches. In HSCI 207 the task is simpler: recognize a mediator in the DAG, record it in the variable list, and avoid treating it as a confounder.

3.4 Telling a Mediator from a Confounder

Students confuse these two roles more than any others. In a recent survey of HSCI 341 students in this series, about one in four (23 percent) chose "a variable on the causal pathway" as the definition of a confounder, which is the definition of a mediator. The confusion is understandable. A confounder and a mediator are both "third variables" that are associated with the exposure and with the outcome, and in a dataset they can look alike. The difference lies in the direction of the arrow between the third variable and the exposure, and that direction comes from knowledge about timing and mechanism, which the data alone cannot supply. Three questions settle most cases.

Three questions for any third variable

Question 1, direction. Does the arrow between the third variable and the exposure point into the exposure or out of it? An arrow from the third variable into the exposure fits a confounder. An arrow from the exposure into the third variable fits a mediator.

Question 2, timing. Did the third variable take its value before the exposure or after it? A variable fixed before the exposure can be a confounder. A variable that takes its value after the exposure and before the outcome can be a mediator.

Question 3, change. If you could change the exposure, would the third variable change as a result? If yes, the variable is downstream of the exposure and cannot be a confounder. If no, it may be a confounder, provided it also affects the outcome.

Depression in Cedar Valley shows why the three questions matter. Depression can affect loneliness, and loneliness can affect depression. As Section 2 showed, a DAG handles this by splitting depression into two time-stamped nodes. The figure below places them on a time line.

Earlier depressionConfounder LonelinessExposure Later depressionMediator ED visitsOutcome Before the surveySurveyTwelve months of follow-up
The same condition plays two roles. Depression before the survey is a confounder because it affects both loneliness and later ED visits. Depression that develops after the survey is a mediator because loneliness can cause it and it can lead to ED visits.

Apply the three questions to each node. Earlier depression points into loneliness (question 1), was present before the survey (question 2), and could not be changed by changing loneliness at the survey (question 3), so it is a confounder. Later depression is pointed to by loneliness, develops after the survey, and would change if loneliness changed, so it is a mediator. The practical lesson for the team is that it must measure depression at two times, or at least know when each measurement was taken, because a single measure of "depression" with no date cannot be assigned a role.

FeatureConfounderMediator
Arrow patternThe confounder points into the exposure and into the outcome.The exposure points into the mediator, and the mediator points into the outcome.
TimingIt takes its value before the exposure.It takes its value after the exposure and before the outcome.
Plain descriptionIt is a common cause that creates an association between exposure and outcome.It is a route by which the exposure's effect reaches the outcome.
Cedar Valley exampleChronic conditions diagnosed before the survey; earlier depression.Depressive symptoms that develop after the survey.
If it is ignoredThe estimated effect of the exposure mixes in the confounder's effect.Nothing is lost for the total effect, although the route remains unexplained.
If it is treated as a confounderThis is the appropriate treatment, taught in HSCI 341.Part of the effect the question asks about is removed.
Where it is developedHSCI 230 Lesson 11 and HSCI 341HSCI 410 Lesson 8

3.5 Colliders

A collider is a variable that is caused by two other variables in the diagram, so that two arrowheads meet, or collide, at it. Colliders are less intuitive than confounders and mediators, because the problem they cause appears only when a study restricts its sample to one level of the collider or adjusts for it in the analysis. A worked example with simple numbers shows what can happen.

Worked example: a falls-prevention program

Imagine a community of 1,000 older adults. Two hundred (20 percent) have poor balance and 200 (20 percent) have poor vision, and in this community the two problems are unrelated: 20 percent of people with poor balance have poor vision, and 20 percent of people with good balance have poor vision. A falls-prevention program enrols everyone who has poor balance or poor vision, because either one raises the risk of falling. Enrolment is a collider, since both balance and vision cause it.

GroupPoor visionGood visionPercentage with poor vision
Whole community, poor balance (200)4016020 percent
Whole community, good balance (800)16064020 percent
Program enrolees, poor balance (200)4016020 percent
Program enrolees, good balance (160)1600100 percent

In the whole community, balance and vision are unrelated. Among the 360 people in the program, however, everyone with good balance has poor vision, because poor vision is the only reason they were enrolled. A researcher who studied only program enrolees would find a strong association between good balance and poor vision that does not exist in the community. Selecting on the collider created it.

In Cedar Valley, the team identified referral to the health authority's community connector program as a collider. Clinicians refer older adults who appear lonely, and ED staff refer patients who return to the ED frequently, so loneliness and ED visits both cause referral. If the team analyzed only respondents who had been referred, or adjusted for referral in its analysis, it could produce a misleading association between loneliness and ED visits, much as the falls program produced a misleading association between balance and vision. The variable list in Section 4 therefore records referral as a variable that the team will not restrict on or adjust for. HSCI 230 Lesson 7 and Lesson 8 discuss the biases that colliders produce, and HSCI 341 Lesson 7 Section 3 states the formal rules that explain them.

3.6 Labelling the Full Cedar Valley DAG

The team's final DAG for this stage contains twelve nodes. The figure below groups the eight confounders into a single box for legibility. In DAGitty each confounder is a separate node with its own arrows into loneliness and ED visits, as the model code below the figure shows.

Confounders: common causes of loneliness and ED visits Age, income, living alone, rural residence, hearing loss, chronic conditions, earlier depression, frailty (unmeasured) LonelinessExposure Later depressionMediator ED visitsOutcome Connector referralCollider
The Cedar Valley DAG with each variable labelled by role. The eight confounders are drawn as one box for legibility, and arrows among the confounders (for example, from age to living alone) are listed in the model code.
dag {
Loneliness [exposure]
ED_visits [outcome]
Frailty [latent]
Age -> Loneliness
Age -> ED_visits
Age -> Chronic_conditions
Age -> Hearing_loss
Age -> Living_alone
Age -> Frailty
Income -> Loneliness
Income -> ED_visits
Income -> Chronic_conditions
Living_alone -> Loneliness
Living_alone -> ED_visits
Rural_residence -> Loneliness
Rural_residence -> ED_visits
Hearing_loss -> Loneliness
Hearing_loss -> ED_visits
Chronic_conditions -> Loneliness
Chronic_conditions -> ED_visits
Chronic_conditions -> Frailty
Frailty -> Loneliness
Frailty -> ED_visits
Earlier_depression -> Loneliness
Earlier_depression -> ED_visits
Earlier_depression -> Later_depression
Loneliness -> Later_depression
Later_depression -> ED_visits
Loneliness -> ED_visits
Loneliness -> Connector_referral
ED_visits -> Connector_referral
}

In DAGitty's model code, the word latent in square brackets marks a variable as unobserved, which is what pressing u over a node does. Pasting this code into DAGitty reproduces the full diagram, and you can drag the nodes into the arrangement shown in the figure.

Try it: sort the variables

A research team asks whether attending a weekly walking group for older adults (exposure) reduces loneliness six months later (outcome). Assign a role to each of the following variables: (a) mobility before joining, which affects whether people can join the group and affects loneliness; (b) the number of new friendships made during the six months; (c) being chosen for a newspaper story, for which a reporter picked regular attenders who said they felt less lonely; (d) age; (e) depressive symptoms that develop during the six months as a result of attending or not attending. Write down your answers before opening the explanations below.

(a) Mobility before joiningv

Mobility before joining is a confounder. It was fixed before the exposure, it affects whether a person attends the group, and it affects loneliness. Attending the group cannot change mobility that was measured before joining.

(b) New friendships during the six monthsv

New friendships are a mediator. Attending the group creates opportunities to make friends, and new friendships reduce loneliness. The friendships take their value after the exposure and before the outcome.

(c) Being chosen for a newspaper storyv

Being chosen for the story is a collider. Both attendance and lower loneliness caused the reporter to choose a person. A study that interviewed only the people in the story would see a distorted picture of the link between attendance and loneliness.

(d) Agev

Age is most likely a confounder. Older age may reduce the chance of joining a walking group and may be related to loneliness, and attending the group cannot change a person's age.

(e) Depressive symptoms during the six monthsv

These symptoms are a mediator if they result from attending or not attending and in turn affect loneliness. Depressive symptoms measured before joining would be a confounder, which is the same timing distinction the Cedar Valley team faced.

With every node labelled, the DAG now tells the team which variables matter and why. Section 4 turns that information into the variable list and data dictionary that guide data collection.

Reflection

A team asks whether regular volunteering (exposure) reduces the number of physician visits for anxiety over two years (outcome) among adults aged 65 and older. It considers three variables: (1) self-rated health, measured before people began volunteering, which affects whether people are able to volunteer and affects anxiety; (2) the size of each person's social network, measured one year after they began volunteering; and (3) receiving a recognition certificate, which the volunteer organization awards both to people who volunteer often and to people whose physicians nominate them for improved well-being.

Definitions: a confounder is a common cause of the exposure and the outcome; a mediator lies on a causal pathway from the exposure to the outcome; a collider is caused by two variables in the diagram. Assign a role to each variable and justify each choice using the direction of the arrows and the order of events. Then explain what would go wrong if the team treated variable 2 as a confounder.

Model answer

Self-rated health is a confounder. It was measured before volunteering began, it affects who is able to volunteer, and it affects anxiety, so the arrows run from health into both the exposure and the outcome. Volunteering cannot change health that was measured before it started.

Social network size is a mediator. Volunteering brings people into contact with others, so the arrow runs from the exposure into network size, and a larger network may reduce anxiety. It is measured one year after volunteering began, which places it between the exposure and the outcome. If volunteering were changed, network size would be expected to change too.

The certificate is a collider. Two arrows point into it: one from frequent volunteering and one from improved well-being, which is closely related to the outcome. Studying only certificate recipients could create a misleading association.

If the team treated network size as a confounder and removed its influence in the analysis, it would remove part of the very effect it set out to estimate. The analysis would then describe the effect of volunteering through routes other than social contact, which may be most of what volunteering does. The team could study the network route later as a mediation question.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: Which definition describes a confounder?

A confounder is a common cause of the exposure and the outcome. Option a defines a mediator, which about one in four HSCI 341 students chose in a recent survey, and option b defines a collider.

Question 2: In the Cedar Valley DAG, depressive symptoms that develop during the 12 months after the survey are best described as:

Later depression takes its value after the exposure, can be changed by loneliness, and can lead to ED visits, so it lies on the pathway. Being associated with both variables does not make it a confounder.

Question 3: In the falls-prevention example, why did balance and vision appear strongly related among program enrolees but not in the whole community?

Enrolment is a collider caused by poor balance and by poor vision. In the community, 20 percent of both balance groups had poor vision; among enrolees, everyone with good balance had poor vision because that was their only reason for enrolling.

Question 4: In the first Cedar Valley draft, hearing loss pointed only to loneliness. After the team added an arrow from hearing loss to ED visits, hearing loss became:

With arrows into both loneliness and ED visits, hearing loss is a common cause of the exposure and the outcome. Arrows leaving a node do not make it a collider, which requires arrows pointing into it.
Section 4 of 5

From the DAG to a Variable List and Data Dictionary

⏱ Estimated reading time: 35 minutes
Section 4 of 5

From the DAG to a Variable List and Data Dictionary

This section covers the variable list, the data dictionary and the gap check.

Purpose

Why the DAG needs a variable list

A plan

Each node gets a source and a time of measurement.

Limitations

Unmeasured nodes are visible before data collection.

Justification

Each item of personal information has a stated reason.

Five steps

Building the variable list

1. List every node2. Record the role3. Record the timing4. Identify a source5. State the planned use

Frailty appears in the list with the source "none available" and the plan "record as a limitation".

Data dictionary

Definitions, values and names

Conceptual definition

Loneliness is an unpleasant experience of relationships that fall short in quantity or quality.

Operational definition

A score of 6 to 9 on the three-item UCLA Loneliness Scale counts as lonely.

Names such as ucla_total, lonely and ed_visits_12m follow these conventions.

Gap check

Comparing the DAG with the data

Unmeasured confounder

Frailty is recorded as a limitation.

Mediator timing

Depression and ED visits share one follow-up year.

Same-time measures

Some confounders are measured with loneliness.

Dropped mediator

Physical inactivity left the DAG, and the decision is recorded.

Putting it together

From a question to a data collection plan

  • An evidence table and link statements.
  • A causal web that includes structural causes.
  • A first DAG in DAGitty, saved and exported.
  • Justifications for arrows and important missing arrows.
  • A role for each node, and a variable list.
Next steps

Reflection and final assessment

Section reflection

It asks you to write a variable list and a data dictionary entry for a new study.

Final assessment

It contains an integrative reflection and 15 questions on the whole lesson.

Lesson 4 turns to interest holder mapping and engagement.

Learning Objectives for this section

  • Turn every node of a DAG into a row of a variable list that records its role, timing, data source and planned use.
  • Write data dictionary entries that give each variable a name, a label, a definition, its values and its source.
  • Check a DAG against the data that can actually be collected and record unmeasured variables and timing problems.
  • Trace the full chain from a research question to a data collection plan, using the Cedar Valley study.

4.1 Why the DAG Needs a Variable List

A DAG states what the team believes about causes. It does not yet say what the team will collect, from where, or when. The bridge between the diagram and data collection is a variable list, a table with one row for every node in the DAG. Each row records the variable's role, the time at which it takes its value, the source that will supply it and what the team plans to do with it. A variable list has three uses. It turns the DAG into a practical plan for the survey, the data request and the chart review. It shows which nodes cannot be measured, so that the team can discuss them as limitations before the study begins. And it justifies every piece of information the team proposes to collect from participants.

The third use matters for ethics as well as for efficiency. Lesson 5 shows that research ethics boards expect researchers to explain why they need each piece of personal information they collect. A variable list linked to a DAG answers that question variable by variable: the team collects income because income is a common cause of loneliness and ED visits, and it can point to the arrows that say so. Variables with no node in the DAG and no descriptive purpose are candidates for removal.

4.2 Building the Variable List

The variable list is built in five steps, working from the labelled DAG in Section 3.

  1. List every node. Copy each node name from the DAGitty model code, including unobserved nodes such as Frailty. Using the exact node names keeps the list and the diagram linked.
  2. Record the role. Write exposure, outcome, confounder, mediator or collider for each node, as in Section 3. A node that has none of these roles, such as a cause of the outcome only, is recorded as "other".
  3. Record the timing. State when the variable takes its value relative to the exposure: before the survey, at the survey or during follow-up. Timing is what separated earlier from later depression in Section 3.
  4. Identify a source. Name where the data will come from: the survey, linked administrative records, a chart review, or nowhere. Lesson 8 describes surveys, administrative data and chart reviews, and Lesson 9 describes data linkage.
  5. State the planned use. Write what the team intends to do with the variable, using plain phrases such as "account for as a confounder", "keep for a later mediation analysis", "do not restrict or adjust" or "record as a limitation". HSCI 230 Lesson 7 Section 2 introduces the structures behind the formal rules that decide which variables an analysis should adjust for, and HSCI 341 Lesson 7 Section 3 states the backdoor criterion and the procedure for choosing an adjustment set. At this stage the plan is provisional and follows from the roles.

The figure below sorts the twelve Cedar Valley nodes by when they take their values and shows the source of each. The colour of each dot shows the variable's role.

Before the surveytwo-year look-back At the surveybaseline After the surveytwelve months of follow-up Chronic conditionsLinked health records Earlier depressionLinked billing records FrailtyNot measured LonelinessSurvey, UCLA scale AgeSurvey IncomeSurvey Living aloneSurvey Hearing lossSurvey Rural residenceSurvey postal code Later depressionLinked billing records ED visitsLinked ED records Connector referralNot collected
The Cedar Valley variables arranged by when they take their values. Red-outlined dots mark confounders, the solid teal dot the exposure, the dark dot the outcome, the light teal dot the mediator and the dashed dot the collider. Linked records are obtained through Population Data BC for respondents who consent.

The completed variable list is shown below. Notice that every node appears, including the two that the team will not collect.

NodeRoleTimingSourcePlanned use
LonelinessExposureAt the surveySurvey: three-item UCLA Loneliness ScaleMain exposure; compare scores of 6 to 9 with scores of 3 to 5.
ED_visitsOutcome12 months after the surveyLinked ED records through Population Data BCMain outcome; count visits per person.
AgeConfounderAt the surveySurveyAccount for as a confounder.
IncomeConfounderAt the surveySurveyAccount for as a confounder.
Living_aloneConfounderAt the surveySurveyAccount for as a confounder.
Rural_residenceConfounderAt the surveySurvey (postal code)Account for as a confounder.
Hearing_lossConfounderAt the surveySurvey (self-reported difficulty hearing)Account for as a confounder.
Chronic_conditionsConfounderTwo years before the surveyLinked physician billing and hospital discharge recordsAccount for as a confounder.
Earlier_depressionConfounderTwo years before the surveyLinked physician billing recordsAccount for as a confounder.
FrailtyConfounder (unobserved)Before the surveyNone availableRecord as a limitation.
Later_depressionMediator12 months after the surveyLinked physician billing recordsDo not treat as a confounder; keep for a later mediation analysis.
Connector_referralCollider12 months after the surveyNot collectedDo not restrict the sample to referred adults or adjust for referral.

4.3 From the Variable List to a Data Dictionary

A data dictionary is a document that describes every variable in a dataset so that anyone using the data knows exactly what each column contains. Where the variable list is organized around the DAG, the data dictionary is organized around the dataset the team will build. It gives each variable a short name, a readable label, a definition, its possible values and codes, its type, its source and any notes a future analyst would need. Lesson 11 shows how a data dictionary becomes a codebook used for data entry and cleaning. Here the aim is a first draft written before data collection begins.

Conceptual and operational definitions

Each entry needs two definitions. A conceptual definition says what the idea means in words. Perlman and Peplau (1981) defined loneliness as the unpleasant experience that occurs when a person's network of social relationships is deficient in quantity or quality, and that is the concept the Cedar Valley team intends to study. An operational definition says exactly how the idea will be measured in this study. For Cedar Valley, loneliness is operationalized as the total score on the three-item UCLA Loneliness Scale (Hughes et al., 2004), and "lonely" means a score of 6 or higher. Lesson 7 develops conceptual and operational definitions, levels of measurement and the choice of validated instruments in detail.

Naming variables

Variable names will later be typed into statistical software, so they follow simple conventions. Use short names in lower case, start each name with a letter, join words with the _ character, and avoid spaces and special characters. Make names informative enough to be read without the dictionary, such as ed_visits_12m for ED visits in the 12 months after the survey. Keep a clear link between DAG node names and dataset names; the Cedar Valley team simply converted its node names to lower case and added detail where a node needed more than one dataset variable.

Copy these columns into a spreadsheet and add one row per dataset variable.

ColumnWhat to write
Variable nameA short name in lower case with _ between words, such as lives_alone.
LabelA readable description of up to about ten words.
DAG node and roleThe node name and its role from the variable list.
DefinitionThe operational definition, with the conceptual definition or its citation.
Values and codesEvery allowed value, the meaning of each code, and the range for numbers.
TypeCategory or number, to be refined with the levels of measurement in Lesson 7.
Source and timingWhere the data come from and the period they describe.
NotesDerivation rules, missing values, known problems.
ColumnEntry
Variable nameucla_total and lonely
LabelLoneliness score (three-item UCLA scale); lonely (score of 6 or higher)
DAG node and roleLoneliness; exposure
DefinitionThe sum of three items asking how often the respondent lacks companionship, feels left out and feels isolated from others (Hughes et al., 2004). Conceptual definition after Perlman and Peplau (1981).
Values and codesEach item: 1 = hardly ever, 2 = some of the time, 3 = often. ucla_total: 3 to 9. lonely: 1 = score of 6 to 9; 0 = score of 3 to 5.
Typeucla_total: number (whole numbers); lonely: category (two groups)
Source and timingRegional survey, at the survey date
Noteslonely is derived from ucla_total. If any item is unanswered, both variables are left blank. In the completed survey, 392 of 1,600 respondents (24.5 percent) had lonely = 1.
ColumnEntry
Variable nameed_visits_12m
LabelNumber of ED visits in the 12 months after the survey
DAG node and roleED_visits; outcome
DefinitionThe count of separate visits to any emergency department in the region during the 365 days after the respondent's survey date.
Values and codesWhole numbers from 0 upward.
TypeNumber (a count)
Source and timingLinked administrative ED records obtained through Population Data BC; 12 months after the survey
NotesAvailable only for respondents who consented to linkage. Visits to emergency departments outside the region may be missed, which the team records as a limitation.
ColumnEntry
Variable namedep_before and dep_after
LabelDepression diagnosed before the survey; depression diagnosed after the survey
DAG node and roleEarlier_depression, confounder; Later_depression, mediator
DefinitionAt least one physician visit with a depression diagnosis in the stated period: the two years before the survey date for dep_before and the 12 months after it for dep_after.
Values and codes1 = yes; 0 = no
TypeCategory (two groups)
Source and timingLinked physician billing records through Population Data BC
NotesThe two variables measure the same condition at different times and play different roles. dep_after covers the same 12 months as the outcome, so a depression diagnosis may follow an ED visit; the team records this timing problem in the gap check.

4.4 Checking the DAG Against the Data

The last step before data collection is a gap check, in which the team compares its DAG with the data it can actually obtain. The Cedar Valley gap check produced four findings, which show the kinds of problems such a check reveals.

Case: the Cedar Valley gap check

An unmeasured confounder. Frailty has no source. The team recorded it as a limitation and noted that chronic conditions and age capture part of what frailty means, so the gap is partial. It also asked the advisory group whether a short question on difficulty leaving the home could be added to the survey, which would give a partial measure.

A timing problem with the mediator. Later depression and ED visits are both measured over the same 12 months, so the data cannot always show which came first. The team decided that any later mediation analysis would use depression diagnosed in the first six months and ED visits in the following six months, and it added this rule to the data dictionary notes.

Confounders measured at the same time as the exposure. Age, income, living alone, hearing loss and rural residence come from the survey, at the same moment as loneliness. The DAG assumes that each was in place before the loneliness a respondent reports. For age this is certain. For living alone and hearing loss it is likely but not certain, so the team considered asking how long respondents had lived alone.

A dropped mediator. The causal web in Section 1 included physical inactivity as a route from loneliness to ED visits. The survey measures physical activity at the same time as loneliness, so it could not show that inactivity followed loneliness. The team left physical activity out of the DAG for this study and recorded the omission, so that a reader can see it was a decision.

The gap check also works in the other direction. Some variables the team wants to collect have no node in the DAG, such as gender and the language spoken at home. These variables are still useful, because they describe who took part and help readers judge to whom the findings apply. Lesson 11 shows how such descriptive variables appear in a Table 1. They are recorded in the data dictionary with the role "descriptive", which keeps the reason for collecting them visible.

Mistake 1: Collecting variables with no reasonv

A survey that asks about everything is longer than it needs to be, which lowers response rates and raises privacy risks. Every variable should either have a node in the DAG or a stated descriptive purpose.

Mistake 2: Losing track of timingv

A variable recorded without a date or period cannot be given a role when it could act at more than one time, as depression showed. Record the period each variable describes, especially for data drawn from administrative records.

Mistake 3: Leaving unmeasured nodes off the listv

Unmeasured confounders such as frailty belong in the variable list with the source "none available". Leaving them off hides a limitation that reviewers will want to see addressed.

Mistake 4: Inconsistent namesv

If the DAG says Living_alone, the variable list says "lives by self" and the dataset says alone2, mistakes become likely. Choose one naming scheme and use it in all three places.

Mistake 5: Treating the DAG as settled factv

A DAG records the team's current assumptions. As new evidence or interest holder input arrives, arrows may be added or removed, and roles may change, as hearing loss showed. Keep version numbers on the DAG, the variable list and the data dictionary, and update them together.

4.5 Putting It Together

The work of this lesson produces a set of linked documents: an evidence table, a causal web, a DAG with a justification table, a variable list and the start of a data dictionary. Together they explain what the team believes about causes, what it will measure and why. The Cedar Valley team will use these documents again when it maps interest holders in Lesson 4, drafts its consent form in Lesson 5 and builds its survey in Lesson 8.

The same chain applies to a qualitative question, with one adjustment. For a question in SPIDER form, the causal web serves as a map of the factors that shape the experience under study, and the DAG is drawn for a related quantitative question, as the Cedar Valley team did with its PECO question.

Reflection

A team asks whether moving into social housing with a shared common room (exposure) is associated with lower loneliness one year after moving in (outcome), among adults aged 65 and older. Loneliness is measured with the three-item UCLA Loneliness Scale, scored 3 to 9, in a resident survey. The team's DAG contains four other variables: age (a confounder, available from the tenancy application); mobility limitation (a confounder, not available in any data source); the number of weekly social activities attended during the first six months (a mediator, from a resident survey at six months); and being chosen as a tenant representative (a collider, because tenants and staff more often choose people who use the common room and who are less lonely).

A variable list records each node's role, timing, source and planned use. A data dictionary entry records the variable name, label, definition, values and codes, type, source and timing, and notes. (a) Write a variable list row for each of the four variables. (b) Write a full data dictionary entry for the outcome. (c) Name one gap between the DAG and the available data and say how you would record it.

Model answer

(a) Age: confounder; at move-in; tenancy application; account for as a confounder. Mobility limitation: confounder (unobserved); before move-in; no source; record as a limitation. Weekly social activities: mediator; first six months; six-month resident survey; do not treat as a confounder, and keep for a later mediation analysis. Tenant representative: collider; during the first year; housing office records; do not restrict the sample to representatives or adjust for this variable.

(b) Variable name: ucla3_12m. Label: loneliness score one year after moving in. Definition: the sum of three items on how often the resident lacks companionship, feels left out and feels isolated from others, each scored 1 (hardly ever) to 3 (often). Values and codes: whole numbers from 3 to 9; left blank if any item is unanswered. Type: number. Source and timing: resident survey, 12 months after the move-in date. Notes: higher scores mean greater loneliness; a derived variable could mark scores of 6 or higher.

(c) Mobility limitation has no source. I would list it in the variable list with the source "none available" and discuss it as a limitation, and I would ask whether one question on difficulty walking could be added to the move-in survey to give a partial measure.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: What is the main purpose of a variable list built from a DAG?

The variable list turns each node into a practical plan before data collection. It includes nodes that cannot be measured, so that they can be recorded as limitations.

Question 2: Which of the following is an operational definition of loneliness for the Cedar Valley study?

An operational definition states exactly how the concept is measured in the study. Option a is a conceptual definition, after Perlman and Peplau (1981).

Question 3: Which variable name follows the naming conventions described in this lesson?

Good names are short, in lower case, start with a letter and join words with the _ character. Option b starts with a number, and options a and c contain spaces, capitals or special characters.

Question 4: Frailty is a confounder in the Cedar Valley DAG, but no data source measures it. What should the team do?

An unmeasured confounder stays in the DAG, marked as unobserved, and appears in the variable list with the source "none available". Deleting or renaming it would hide a limitation.
Section 5 of 5

Final Assessment

⏱ Estimated time: 25 minutes

Bringing It All Together

This lesson followed the fictional Cedar Valley Social Connection Study from a broad literature summary to a plan for data collection. The team began with an evidence table and link statements, drew a causal web that included community and structural causes as well as individual ones, and then disciplined that web into a directed acyclic graph in which every node is a measurable variable at a stated time and every arrow points forward in time.

With the DAG drawn, the team named the role of each variable relative to its PECO question. Age, income, living alone, rural residence, hearing loss, chronic conditions, earlier depression and frailty are confounders; depression that develops after the survey is a mediator; and referral to the connector program is a collider. The contrast between earlier and later depression showed that the same condition can play different roles depending on when it is measured, which is the key to telling a mediator from a confounder.

Finally, the team turned its DAG into a variable list and a first data dictionary, and a gap check exposed an unmeasured confounder and two timing problems before any data were collected. The reflection and the 15-item assessment below ask you to apply these steps to new questions.

Key Takeaways from this lesson

  • A causal web gathers the possible causes of an outcome from evidence and experience, and Krieger's critique is a reminder to include community and structural causes.
  • Each arrow in a causal web should trace back to a link statement of the form "A affects B" with a named source.
  • A directed acyclic graph uses one node for each measurable variable at a stated time, one-way arrows from cause to effect, and no loops.
  • An arrow claims only that a direct effect may exist, while a missing arrow makes the stronger claim that no direct effect exists.
  • A feedback loop, such as loneliness and depression affecting each other, is drawn in a DAG as a sequence of time-stamped nodes.
  • DAGitty draws a DAG by pointing and clicking or from typed model code, and saving that code to a versioned text file keeps a reliable record.
  • A confounder is a common cause of the exposure and the outcome, a mediator lies on a pathway between them, and a collider is caused by two variables in the diagram.
  • The direction of the arrow between a third variable and the exposure, the order of events, and whether the exposure could change the variable together distinguish a mediator from a confounder.
  • Restricting a study to one level of a collider, or adjusting for it, can create an association that does not exist in the population.
  • A variable list and data dictionary built from the DAG justify every variable collected and reveal unmeasured variables and timing problems before data collection begins.

Core Concepts Reviewed

Section 1: the web of causation, Krieger's critique and ecosocial theory, evidence tables, link statements, levels of causes, and the differences between a causal web and a DAG.

Section 2: nodes, arrows, direct effects, ancestors and descendants, the claims made by arrows and missing arrows, four drawing conventions, and drawing, saving and exporting a DAG in DAGitty.

Section 3: exposure, outcome, confounder, mediator and collider, the three questions that separate a mediator from a confounder, and a worked example of selection on a collider.

Section 4: the variable list, conceptual and operational definitions, data dictionary entries and naming conventions, and the gap check between a DAG and the available data.

The final reflection asks you to build and label a DAG for a research question and to explain the decisions it implies for data collection.

Reflection

Use this question, or another research question that interests you: among university students in British Columbia, is working more than 20 hours a week in paid employment (exposure) associated with lower final grades in the same term (outcome)?

A confounder is a common cause of the exposure and the outcome; a mediator lies on a causal pathway from the exposure to the outcome; a collider is caused by two variables in the diagram. Describe in words a DAG with at least six nodes, listing its arrows and labelling each node's role. Identify one variable that could be mistaken for a confounder when it is really a mediator, and explain how you told them apart. Name one variable you would measure, one you would record as a limitation, and one you would not restrict or adjust for, giving a reason for each.

Model answer

Using the paid-work question, my DAG has seven nodes. Family income and year of study are confounders: lower family income makes long work hours more likely and limits money for tutoring or reduced course loads, and year of study shapes both job opportunities and course difficulty. Prior grade point average is also a confounder, because students with weaker records may choose more work and may earn lower grades. Hours of sleep during the term and hours of study during the term are mediators, because long work hours reduce the time available for both, and both affect grades. Withdrawal from a course is a collider if students withdraw both because work leaves no time and because their grades are falling.

Hours of study could be mistaken for a confounder because it is associated with work and with grades. I applied three tests: the arrow runs from work into study time, study time is measured during the term after work hours are set, and a change in work hours would change study time. It is therefore a mediator.

I would measure prior grade point average from transcripts with consent. I would record study habits before the term as a limitation if no source exists. I would not restrict the analysis to students who completed every course, because completion depends on both work and grades.

Minimum 30 characters required.

✓ Reflection saved

Final Knowledge Assessment

Final Assessment, this lesson: Causal Webs and Directed Acyclic Graphs (15 Questions)

Question 1: A student writes "social factors affect health" as a box in a first DAG. What is the main problem?

Each DAG node should be a single variable that could be measured at a stated time. The student should split the box into specific factors such as income or living alone.

Question 2: A study asks whether a home-visiting program for new parents (exposure) reduces infant ED visits (outcome). Parents' knowledge of infant care, measured three months after the program begins, is most likely:

Knowledge is measured after the program begins and can be changed by it, and it can affect ED visits, so it lies on the pathway. Association with both variables is true of mediators and confounders alike and does not settle the role.

Question 3: Which statement about DAGitty is accurate?

DAGitty is free, runs in a browser without an account, and writes each diagram as model code. Pasting saved code into the Model code box and clicking Update DAG redraws the diagram.

Question 4: Why does the Cedar Valley team record depression at two different times?

Depression before the survey can affect loneliness and ED visits, while depression after the survey can result from loneliness and lead to ED visits. A single undated measure could not be given a role.

Question 5: A first causal web for a study of asthma emergency visits among children contains only inhaler use, pet ownership and parental smoking. Applying Krieger's critique, which addition would most improve it?

Krieger argued that webs tend to omit social and economic conditions. Housing quality and income shape exposure to mould, pests and smoke, and they sit at the structural level the web currently lacks.

Question 6: Loneliness is measured at the survey, and the outcome is ED visits in the 12 months after it. A team adds ED visits in the year before the survey as a node. Which arrow is impossible?

Arrows must point forward in time. Visits that happen after the survey cannot affect loneliness measured at the survey, although past visits could.

Question 7: When DAGitty starts a new model, it draws an arrow from the exposure to the outcome. What does that arrow represent?

The arrow allows for the effect the study is designed to estimate. Keeping it makes no claim that the effect exists or is large.

Question 8: Why does the Cedar Valley variable list say not to restrict the sample to adults referred to the connector program?

Clinicians refer lonely patients and ED staff refer frequent visitors, so both arrows point into referral. Selecting on a collider can create a misleading association, as the falls-prevention example showed.

Question 9: Which question best separates a confounder from a mediator?

A variable that would change when the exposure changes is downstream of it and cannot be a confounder. Association with both the exposure and the outcome is true of confounders and mediators alike.

Question 10: How does a variable list support an application to a research ethics board?

Research ethics boards expect researchers to justify the personal information they collect. Linking each variable to a node and role provides that justification.

Question 11: A DAG contains an arrow from income to chronic conditions. Which statement does the arrow make?

An arrow is a modest claim that a direct effect may exist. It says nothing about the size or direction of the effect, and option d describes what a missing arrow plus a pathway would claim.

Question 12: Later depression and ED visits are both measured over the same 12 months. What problem does the Cedar Valley gap check identify?

A mediator must come before the outcome it affects. The team therefore planned to use depression in the first six months and ED visits in the following six months for any later mediation analysis.

Question 13: Which statement best describes the relationship between a causal web and a DAG in this lesson?

The web keeps the full range of causes in view, including structural ones, and the DAG commits to a direction and timing for each relationship so that data collection can be planned.

Question 14: In a study of whether a medication (exposure) reduces stroke (outcome), blood pressure measured after the medication starts is a mediator. What happens if the analysis treats it as a confounder?

Lowering blood pressure is one route by which the medication prevents stroke. Treating blood pressure as a confounder removes that route and leaves only the effect through other routes.

Question 15: Which item belongs in a data dictionary entry but is not shown in the DAG?

The DAG shows arrows, roles and whether a node is observed. The data dictionary adds practical details such as names, values, codes, types and sources.
✦ Complete the final reflection above before submitting

Congratulations!

You have successfully completed this lesson: Causal Webs and Directed Acyclic Graphs.

You can now turn a literature summary into a causal web, draw and document a DAG in DAGitty, name the role of each variable in it, tell a mediator from a confounder, and turn the diagram into a variable list and data dictionary.

Lesson 4 turns to the people who have an interest in a study. It shows how to map interest holders with a power-interest grid and influence maps, choose levels of engagement, and write an engagement plan. The Cedar Valley DAG will be useful there, because the people who live with a problem are often the ones who notice a missing arrow.

Continue to Lesson 4 →
Reference

Glossary: Key Terms, People & Frameworks

📚 Reference page, available throughout the lesson

This glossary defines the terms, tools and people introduced in this lesson on causal webs and directed acyclic graphs.

Core Concepts
Causal web A broad diagram of the factors that evidence and experience suggest are causally connected to an outcome and to one another. It may contain loops and broad factors, and it is used to gather possibilities before a DAG is drawn.
Directed acyclic graph (DAG) A diagram of assumed causal relationships in which each node is a variable, each arrow points one way from cause to effect, and no sequence of arrows returns to its starting point.
Node A variable in a DAG, drawn as a box or a name. Each node should be a single variable that could be measured, at least in principle, at a stated time.
Arrow (edge) A line in a DAG pointing from a variable assumed to be a cause to a variable assumed to be its effect. It allows for a direct effect of any size.
Direct effect An effect of one variable on another that does not pass through any other variable shown in the diagram.
Missing arrow The absence of an arrow between two nodes, which claims that neither has a direct effect on the other. It is a stronger assumption than drawing an arrow.
Ancestor and descendant A node's ancestors are all the nodes from which it can be reached by following arrows forward; its descendants are all the nodes that can be reached from it.
Feedback loop A relationship in which two or more factors affect one another in a cycle, such as loneliness and depression. A causal web may show one, and a DAG redraws it as time-stamped nodes.
Exposure The variable whose effect a research question asks about, such as loneliness in the Cedar Valley PECO question.
Outcome The variable on which the effect of the exposure is measured, such as the number of ED visits in the 12 months after the Cedar Valley survey.
Confounder A common cause of the exposure and the outcome (or a proxy for one) that is not on the causal pathway between them, and so mixes its own effect into their association. Chronic conditions diagnosed before the survey are a Cedar Valley example.
Mediator A variable on a causal pathway from the exposure to the outcome, which carries part of the exposure's effect. Depression developing after the Cedar Valley survey is an example.
Collider A variable caused by two other variables in a diagram, so that two arrowheads meet at it. Restricting a study to one level of a collider can create a misleading association.
Unobserved (latent) variable A variable included in a DAG that the study cannot measure, such as frailty in the Cedar Valley study. DAGitty marks it as latent.
Temporality The requirement that a cause comes before its effect. Bradford Hill listed it among his considerations, and every arrow in a DAG points forward in time.
Link statement A short sentence of the form "A affects B" with its source, used to turn an evidence table into the arrows of a causal web.
Evidence table A table recording, for each source, the type of evidence, what it suggests about causes, and the links it supports.
Justification table A table giving a one-line reason and a source for each arrow, and for each important missing arrow, in a DAG.
Variable list A table with one row per DAG node, recording its role, timing, data source and planned use.
Data dictionary A document describing every variable in a dataset, including its name, label, definition, values and codes, type, source and notes.
Conceptual definition A statement in words of what an idea means, such as Perlman and Peplau's definition of loneliness.
Operational definition A statement of exactly how an idea is measured in a particular study, such as a score of 6 or higher on the three-item UCLA Loneliness Scale.
Gap check A comparison of a DAG with the data a study can obtain, which records unmeasured variables, timing problems and variables collected for description only.
Frameworks & Tools
DAGitty Free software for drawing and analyzing causal diagrams, used in a web browser at dagitty.net or as an R package (Textor et al., 2016).
Model code DAGitty's text version of a diagram, in which each line declares a variable's status or states an arrow. It can be saved and pasted back to redraw the diagram.
Web of causation The idea, introduced by MacMahon, Pugh and Ipsen (1960), that disease arises from many interconnected causes, each with causes of its own.
Ecosocial theory Nancy Krieger's theory of disease distribution, which asks how people embody the social, economic and ecological conditions in which they live.
ESC-DAGs Evidence synthesis for constructing directed acyclic graphs, a structured method for turning published studies into a DAG (Ferguson et al., 2020).
Three-item UCLA Loneliness Scale A short loneliness measure asking how often a person lacks companionship, feels left out and feels isolated from others, scored 3 to 9 (Hughes et al., 2004).
Key People
Brian MacMahon Epidemiologist at the Harvard School of Public Health who, with Thomas Pugh and Johannes Ipsen, described the web of causation in the 1960 textbook Epidemiologic Methods.
Nancy Krieger Social epidemiologist at Harvard who asked in 1994 whether anyone had seen the spider of the web of causation and who developed ecosocial theory.
Sewall Wright American geneticist who developed path analysis and path diagrams in the 1920s, an early form of the causal diagrams used today.
Judea Pearl Computer scientist at the University of California, Los Angeles, whose work in the 1990s gave causal diagrams a formal mathematical basis.
Sander Greenland Epidemiologist and statistician who, with Judea Pearl and James Robins, introduced causal diagrams to epidemiologists in 1999.
James M. Robins Epidemiologist and biostatistician at Harvard whose work on causal inference shaped the use of DAGs in health research.
Miguel A. Hernán Epidemiologist at Harvard and co-author, with James Robins, of the textbook Causal Inference: What If (2020), which uses DAGs throughout.
Johannes Textor Computer scientist who leads the development of DAGitty, the free tool used in this lesson to draw DAGs.
Austin Bradford Hill British epidemiologist and statistician whose 1965 considerations for judging causation include temporality, the requirement that a cause precedes its effect.
No matching entries. Try a different search term.