# Lesson 3: Linear and Logistic Regression

*Companion-podcast transcript • Sarah & Kiffer*  
*Office Hours episode to listen to after working through the lesson*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This episode goes with Lesson three, Linear and Logistic Regression.

**Sarah:** It's meant for after you've worked through the lesson. We'll add some perspective on where these models came from and how they're used, and some critique of the ways they get misread. We'll also work through the questions students tend to find thorny with this material and do a few extra worked examples, including some that are harder than the ones on the lesson page.

**Kiffer:** There are a few places where we ask you to work something out before we give the answer. When we do, you'll hear a few seconds of quiet. Pause the audio if you'd like more time.

**Sarah:** Here are the four questions. If smokers have 5.8 times the odds of the outcome in the lesson's logistic model, are they 5.8 times as likely to have it? Why did the odds ratio for smoking go up when age was added, even though smokers and non-smokers had almost the same average age? How do you read the coefficients once a model has an age-by-smoking product term? And can a model rank people well and still get their risk wrong? We'll finish with the decision to turn blood pressure into a yes or no, because Kiffer and I see that one differently.

**Kiffer:** Let's start with the odds ratio.

**Sarah:** Question one. The lesson's yes-or-no outcome is a systolic reading of a hundred and twenty or more, which I'll call a high reading. Smokers have 5.8 times the odds of a high reading compared with non-smokers of the same age, gender, BMI and depression score. I often hear that turned into the claim that smokers are almost six times as likely to have one. Is that right?

**Kiffer:** It isn't, and the lesson's own predictions show why. For a fifty-year-old man with a BMI of twenty-one and a depression score of twenty-one, the predicted probability is 9.7 percent if he doesn't smoke and 38.4 percent if he does. The ratio of those probabilities is about 3.9. The odds ratio is larger because odds pull apart faster than probabilities once an outcome is common.

**Sarah:** Has that mix-up ever made the news?

**Kiffer:** A well-known case is from 1999. A study in the New England Journal of Medicine had physicians watch videos of actors describing chest pain, and it reported an odds ratio of about 0.6 for referring Black patients for cardiac catheterization compared with white patients. Much of the coverage said Black patients were forty percent less likely to be referred. The referral rates were about eighty-five percent and ninety-one percent, so Black patients were referred about seven percent less often, and most of that gap was among Black women. Schwartz, Woloshin and Welch pointed this out in the same journal later that year.

**Sarah:** Here's one I want to try myself, with made-up numbers. Suppose a model gives a non-smoker a predicted probability of 0.2, and the adjusted odds ratio for smoking is three. What's the predicted probability for an otherwise identical smoker? Take a few seconds.

*(Pause)*

**Sarah:** Okay. My first instinct is to multiply. Three times 0.2 is 0.6, so sixty percent.

**Kiffer:** Before you settle on that, use the same rule for someone whose predicted probability is 0.4.

**Sarah:** Three times 0.4 is 1.2. That's a probability above one, which is impossible. So the odds ratio has to be applied to the odds, and I need to convert first.

**Kiffer:** Go ahead.

**Sarah:** The non-smoker's odds are 0.2 divided by 0.8, which is 0.25. Three times that is 0.75. To get back to a probability, I divide the odds by one plus the odds, so 0.75 divided by 1.75, which is about 0.43.

**Kiffer:** So about forty-three percent, and the risk ratio is about 2.1, well below the odds ratio of three. The person who started at 0.4 ends up at about 0.67.

**Sarah:** And if the outcome were rare?

**Kiffer:** Then the two nearly agree. Start at one percent. The odds are just over 0.01, three times that is just over 0.03, and that converts back to about 2.9 percent. When an outcome is rare, the odds and the probability are almost the same number, so the odds ratio and the risk ratio are close. The lesson puts the line at about ten percent.

**Sarah:** So why do epidemiologists report odds ratios at all, if they're this easy to misread?

**Kiffer:** The reasons are partly historical and partly practical. David Cox set out logistic regression for binary outcomes in 1958, and in 1967 researchers on the Framingham Heart Study used a multiple logistic function to estimate the risk of coronary heart disease from several risk factors at once. The log odds can take any value, so the model never predicts an impossible probability. And in 1979, Prentice and Pyke showed that a logistic model fitted to case-control data gives valid odds ratios, which helped make it a standard tool. I'd report the odds ratio and then give predicted probabilities for a few typical people.

**Sarah:** Question two. In the lesson, the crude odds ratio for smoking was about four. Adding age alone raised it to about 5.3, yet smokers and non-smokers had almost the same average age, 45.5 and 45.1 years. If age wasn't confounding the comparison, why did the odds ratio move?

**Kiffer:** Let's build a made-up example where we know there's no confounding. Picture a population split into two halves of equal size, a lower-risk half and a higher-risk half. You can think of them as younger and older people. In each half, exactly half the people are exposed, so the exposure is spread evenly and the halves can't confound anything. In the lower-risk half, ten percent of unexposed people have the outcome, and fifty percent of exposed people do. In the higher-risk half, it's fifty percent and ninety percent.

**Sarah:** Let me get the odds ratio within each half. In the lower-risk half, the unexposed odds are 0.1 divided by 0.9, which is one ninth, and the exposed odds are 0.5 divided by 0.5, which is one. So the odds ratio is nine. In the higher-risk half, the unexposed odds are one and the exposed odds are 0.9 divided by 0.1, which is nine. The odds ratio is nine again.

**Kiffer:** So the odds ratio is nine in both halves. Now ignore the halves and pool everyone. What's the crude odds ratio for the whole population? Take a few seconds.

*(Pause)*

**Kiffer:** Among all the unexposed people, the share with the outcome is the average of ten and fifty percent, which is thirty percent. Among all the exposed, it's the average of fifty and ninety percent, which is seventy percent. The odds are 0.3 divided by 0.7, about 0.43, and 0.7 divided by 0.3, about 2.33. Dividing the second by the first gives about 5.4.

**Sarah:** So the odds ratio is nine in each half and 5.4 overall, and nothing was confounded. That seems strange to me. I'd expect the overall number to land somewhere between the numbers for the two halves.

**Kiffer:** For a risk difference or a risk ratio, it does. The risk difference is forty percentage points in each half and forty overall. The risk ratio is five in the lower-risk half, 1.8 in the higher-risk half, and about 2.3 overall, which sits between them. The odds ratio behaves differently, and epidemiologists call this non-collapsibility. When a strong predictor of the outcome is added to a logistic model, the odds ratio for the exposure tends to move further from one, even when that predictor is unrelated to the exposure.

**Sarah:** And in a linear model?

**Kiffer:** A difference in means behaves like the risk difference. If you add a strong predictor that's spread evenly across smokers and non-smokers, the smoking coefficient stays about where it was, and its confidence interval gets narrower because the residuals shrink.

**Sarah:** That matters for the change-in-estimate check. In the lesson's data, adding age changes the odds ratio by about thirty percent, so age would pass a ten percent rule easily, even though it does almost no confounding there.

**Kiffer:** Right. With odds ratios, part of any change comes from non-collapsibility, so the size of the change can make confounding look larger than it is, as it does here, or smaller when confounding pulls the other way. It also means two studies that adjusted for different variables can report different odds ratios with no confounding in either one. When you compare odds ratios across studies, it's worth checking what each one adjusted for, and I'd give more weight to predicted probabilities and risk differences when they're available.

**Sarah:** Question three is about interactions. The lesson tested an age-by-smoking product term, found no evidence for it, and kept the simpler model. Suppose the product term had been needed. How would you read the coefficients?

**Kiffer:** This is where a lot of misreading happens. Here's a hypothetical model for systolic blood pressure with age, smoking and their product. The smoking coefficient is minus four millimetres of mercury, and the product term is 0.2 millimetres per year of age. Someone reading quickly sees minus four and concludes that smokers have lower blood pressure.

**Sarah:** With a product term in the model, though, the smoking coefficient is the difference between smokers and non-smokers when the other variable equals zero. So minus four is the difference at age zero.

**Kiffer:** Exactly, and nobody in the data is zero years old. So here's the question. Using these made-up coefficients, what's the predicted difference between smokers and non-smokers at age fifty? Take a few seconds.

*(Pause)*

**Sarah:** The difference at any age is minus four plus 0.2 times the age. At fifty, 0.2 times fifty is ten, and minus four plus ten is six. So smokers are about six millimetres higher at fifty. At thirty, it's minus four plus six, which is two, and at twenty it's zero.

**Kiffer:** So the same model says smoking makes no difference at twenty and a six millimetre difference at fifty. The minus four is an extrapolation to age zero. One simple step is to centre age before fitting the model, for example by subtracting forty-five from everyone's age. The fit stays exactly the same, and so does the product term. The smoking coefficient becomes the difference at age forty-five, which is minus four plus nine, or five.

**Sarah:** And the age coefficient in that model?

**Kiffer:** It's the slope for age among non-smokers, the reference group. For smokers, the slope is that coefficient plus 0.2. I'd centre any continuous variable that goes into a product term, and I'd report the effect of the exposure at two or three ages that people in the data actually have.

**Sarah:** This sounds like part of a bigger rule. What a coefficient means depends on what else is in the model.

**Kiffer:** It is, and it applies to every row of a regression table. Take the lesson's model built from the causal diagram, with smoking, age, gender and depression. The smoking row estimates the total effect of smoking, if the diagram is right. The depression row is a different matter. The diagram has depression affecting smoking, so smoking lies on a path from depression to blood pressure, and adjusting for it removes part of the effect of depression. Nobody chose the other variables to deal with the confounders of depression, either. In 2013, Westreich and Greenland named the habit of reading every row as an effect the Table 2 fallacy, after the table where adjusted results usually appear.

**Sarah:** Question four. The lesson's logistic model had an AUC of 0.86 and a Hosmer-Lemeshow p-value of 0.59. Those check two different things. Can a model do well on one and badly on the other?

**Kiffer:** Yes, and here's a quick way to see it. Suppose you took the lesson's model and divided every predicted probability by two. What happens to the AUC? Take a few seconds.

*(Pause)*

**Sarah:** Nothing happens to it. The AUC asks whether a person with the outcome gets a higher predicted probability than a person without it. Halving everyone keeps the same order, so every pair is ranked the same way and the AUC stays at 0.86. The calibration would be badly off, though. A group in which thirty percent have the outcome would be told fifteen percent.

**Kiffer:** Right. Discrimination is about order, and calibration is about the size of the numbers. This happens with real models. In 2001, D'Agostino and colleagues tested the Framingham coronary heart disease equations in other cohorts. For Japanese American men and Hispanic men, the equations ranked people nearly as well as models built within each cohort, but they systematically overestimated the five-year risk. They overestimated it for Native American women, too. After recalibration for each group's risk factor levels and underlying rates of disease, the equations worked well in all three groups.

**Sarah:** Does that matter in Canada?

**Kiffer:** It does, because scores like this are used in practice here. The Canadian Cardiovascular Society's 2021 lipid guidelines recommend a cardiovascular risk assessment every five years for adults aged forty to seventy-five, using the Framingham Risk Score or the Cardiovascular Life Expectancy Model. A score that's poorly calibrated for a particular community will tell people their risk is higher or lower than it is, even if it sorts them in the right order.

**Sarah:** And a model can lose its calibration when it moves to new data from the same population, too.

**Kiffer:** Yes, and overfitting is a common cause. An overfitted model has learned some of the noise in its own sample, so in new data its predictions are too extreme. People it calls high risk turn out lower, and people it calls low risk turn out higher. That's one reason for the guide of about ten events per coefficient.

**Sarah:** Let me count one, with made-up numbers. A study has six hundred people, fifty-four of whom have the outcome. The predictors are age, sex, smoking, BMI, region with five categories, and education with four. That's six variables, but region needs four indicator coefficients and education needs three. So it's four coefficients for the first four variables, plus four, plus three, which is eleven. Fifty-four divided by eleven is about 4.9 events per coefficient, about half the guide.

**Kiffer:** That's right. The usual slip is to count six variables and stop there. The guide itself is rough, and a 2016 paper by van Smeden and colleagues concluded that the evidence behind it is weak. Even so, a count near five is a good reason to shorten the list of predictors or to collect more data.

**Sarah:** Okay. Now the part where we disagree.

**Kiffer:** It's about turning blood pressure into a yes or no.

**Sarah:** The lesson recodes systolic blood pressure at a hundred and twenty, so that there's a binary outcome for logistic regression. I'm uneasy about that. A reading of a hundred and twenty-one and a reading of a hundred and forty-five become the same thing, while a hundred and nineteen and a hundred and twenty-one become different things. Altman and Royston, writing in the BMJ in 2006, pointed out that splitting a variable at its median costs as much statistical power as throwing away about a third of the data.

**Kiffer:** I agree that the continuous measure carries more information. My view is that a yes-or-no outcome is often the right choice when the question is about a decision. A clinician decides whether to treat, and a health region wants to know what share of adults are above the treatment threshold. The mean can shift by two millimetres while the share above a threshold changes a lot or hardly at all, depending on what happens in the upper tail.

**Sarah:** But the thresholds themselves move. In 2025, Hypertension Canada lowered the reading at which hypertension is diagnosed in primary care to a hundred and thirty over eighty. Since 2015, the line for automated office readings had been a hundred and thirty-five over eighty-five. Their own explanation says cardiovascular risk starts to rise above about a hundred and fifteen over seventy-five, and that definitions of hypertension are pragmatic decisions. A person with the same blood pressure can change category overnight.

**Kiffer:** That's a fair point, and I read it the other way. The guideline changed because of newer evidence on treatment, and the threshold marks where clinical decisions are now made. An analysis meant to inform those decisions should use the threshold the decisions use. The cut at a hundred and twenty in the lesson was a teaching choice, and it had a practical reason. The survey's hypertension variable had only twenty-three cases, too few for a model with several predictors.

**Sarah:** Then my bigger worry is a cut-point chosen after looking at the data. Altman and Royston also warned against searching for the so-called optimal cut-point, the one with the smallest p-value, because it overstates the difference between the groups and invites false positives.

**Kiffer:** We agree on that. A cut-point found by searching the data is another form of overfitting. Where I land is that I'd keep a yes-or-no outcome when a decision turns on a threshold, and I'd analyse the continuous measure as well whenever the question is about blood pressure itself.

**Sarah:** I'd use the continuous measure more often than Kiffer would, but we agree on the practical rule. Choose any cut-point before seeing the data, take it from a guideline when one exists, say which one you used, and report the continuous analysis alongside it.

**Kiffer:** Let's pull it together with three things to take away.

**Sarah:** First, an odds ratio multiplies odds. To get a probability, convert to odds, multiply, and convert back, and when the outcome is common, report predicted probabilities alongside the odds ratio.

**Kiffer:** Second, what a coefficient means depends on the model around it. An odds ratio can move when a balanced predictor is added, a main effect next to a product term is the effect where the other variable equals zero, and only the exposure's row was built to answer the research question.

**Sarah:** And third, check discrimination and calibration separately. A model can put people in the right order and still give them the wrong numbers, especially when it moves to a new population or has too few events per coefficient.

**Kiffer:** If you'd like more practice, try the two-halves problem from this episode with your own numbers. Work out the odds ratio in each half and overall, then do the same for the risk difference and the risk ratio, and see which ones land between the two halves.

**Sarah:** Next time, it's Lesson four, Generalized Linear Models, where the same ideas carry over to counts, ordered categories and outcomes with several unordered categories.

**Kiffer:** Take care, everyone.

**Sarah:** See you in Lesson four.
