# Lesson 6: Choosing an Approach and Managing a Research Project

*Companion-podcast transcript, Sarah and Kiffer*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This week we are working on Lesson six, which covers choosing a study design and managing a research project.

**Sarah:** Where does this lesson sit in the course? So far the course has covered writing a question, drawing a causal diagram, mapping interest holders and research ethics.

**Kiffer:** Those five lessons gave students the parts of a study. This lesson turns the parts into something a team can actually carry out. The first half is about design: which kind of study answers which kind of question, and how to combine numbers and experiences in one study. The second half is about the practical work of running a project, which means the protocol, the timeline, the team, the budget, the files and the data.

**Sarah:** That second half sounds less like research methods and more like office administration.

**Kiffer:** I understand why it sounds that way, but in my experience it is where student projects and funded studies most often run into trouble. A good design does not help much if the ethics review takes three months longer than the timeline allowed, or if nobody can tell which version of the survey was actually used.

**Sarah:** That is fair. Let's start with design. Section one is called From Question to Design. What is the main idea?

**Kiffer:** The main idea is that the question chooses the design. A study design is the overall plan for how a study gathers and compares information. It says who will be studied, what will be measured, when each measurement happens, and whether the researchers change anything about people's circumstances. Designs that share the same basic logic form what we call a design family.

**Sarah:** And the questions themselves come from earlier lessons.

**Kiffer:** They do, and the structure of the question already says a lot. In Lesson two we wrote questions using population, exposure, comparison and outcome, or population, intervention, comparison and outcome. If there is a comparison, groups will be compared. If there is an intervention, someone will deliver it. A question written with the SPIDER framework, which stands for sample, phenomenon of interest, design, evaluation and research type, points toward qualitative methods.

**Sarah:** What goes wrong when people pick a design the other way around?

**Kiffer:** Beginning researchers often start with a method they know, such as a survey, and then look for a question it can answer. The findings then tend not to match what the team or its partners needed to know. The other common error is to treat designs as a ranking, with the randomized controlled trial at the top.

**Sarah:** Isn't the trial at the top, though? That is what students hear in almost every health course.

**Kiffer:** For questions about the effect of an intervention that can ethically be assigned, a well-run trial gives the strongest evidence. But many important questions are about something else. If you want to know how common loneliness is among older adults, or how they experience moving to a new town, a trial simply cannot answer that. So the most suitable design is the one that fits the question and that the team can carry out well.

**Sarah:** The lesson gives students a decision map. Walk me through it.

**Kiffer:** It is a sequence of five questions, adapted from a classification that David Grimes and Kenneth Schulz published in two thousand and two. Step one asks whether existing studies already answer the question. If they do, the right project may be a systematic or scoping review rather than new data collection.

**Sarah:** What is step two?

**Kiffer:** Step two asks whether the question is about numbers, about meanings and experiences, or about both. How many, how much, how strongly: those are quantitative. How people experience something or how a process unfolds: that is qualitative. If it is both, you are in mixed methods, which is section two.

**Sarah:** Then comes step three.

**Kiffer:** Step three asks whether the exposure is assigned on purpose, by the investigator or by another agent such as a government, and whether the assignment is random. Random assignment, such as a computer-generated allocation sequence, gives a randomized controlled trial. Deliberate assignment that is not random gives a quasi-experimental study, for example a program offered at clinics the team selects, or a policy introduced in some places and not others. If nobody assigns the exposure and the team only records exposures people already have, the study is observational.

**Sarah:** And then the observational studies split further.

**Kiffer:** They do. Step four asks whether the study compares groups to estimate an association. If there is no comparison, it is descriptive, like a prevalence survey or a case series. If there is a comparison, it is analytic, and step five asks about the unit and the timing.

**Sarah:** This is where the familiar names come in.

**Kiffer:** That's right. If the units are groups, such as health regions, the study is ecological. If the units are individuals, timing separates the designs. A cross-sectional study measures exposure and outcome at the same time. A cohort study starts with people who differ in exposure and follows them forward. A case-control study starts with people who have the outcome and people who do not, and looks back at their earlier exposures.

**Sarah:** Students will have heard those names before. Does this lesson teach each of them properly?

**Kiffer:** No, and that is deliberate. This course maps the families. The epidemiology course, Health Sciences two thirty, teaches the designs in depth in Lessons three to six. Here the goal is to know which family a question belongs to and where to go to learn more.

**Sarah:** What about predictive questions? Lesson one had four question types.

**Kiffer:** Predictive questions use the same map with a different purpose. A prediction study estimates who is at risk without asking whether each predictor causes the outcome. They usually use cohort data, because the predictors have to be measured before the outcome happens.

**Sarah:** Once the map gives you a family, are you done?

**Kiffer:** The map is only the first part. It tells you which designs could answer the question. Constraints then tell you which ones you can actually carry out. Ethics comes first. Nobody can be randomized to loneliness or poverty, so questions about the effects of those exposures use observational designs.

**Sarah:** What are the other constraints?

**Kiffer:** The frequency of the outcome and the exposure matters. If the outcome is rare, a cohort would have to follow an enormous number of people to see enough cases, so a case-control study, which starts from the cases that already exist, is far more efficient. If the exposure is rare, the logic flips, and a cohort chosen because of its exposure is more efficient. Then there is time and money. A prospective cohort can take years. And existing data can change everything, because linked administrative records can let a team assemble a cohort from time that has already passed.

**Sarah:** Let's bring in the running case. Remind listeners what it is.

**Kiffer:** The Cedar Valley Social Connection Study is a fictional mixed-methods study of loneliness and social isolation among adults aged sixty-five and older. Dr. Maya Hart's team works with the fictional Cedar Valley Health Authority, which serves about two hundred and ten thousand people, about forty-six thousand of them aged sixty-five and older.

**Sarah:** How does the map sort its questions?

**Kiffer:** The question of how common loneliness is gets a descriptive answer from the regional survey. Of sixteen hundred respondents, three hundred and ninety-two, or twenty-four and a half percent, scored six or higher on the three-item loneliness scale. The question of whether loneliness is associated with emergency department visits is analytic. And the question about how older adults living alone experience social connection after a move to a smaller town is qualitative.

**Sarah:** The lesson calls the emergency department question a cohort design. I would have guessed cross-sectional, since there is only one survey.

**Kiffer:** That is the most interesting point in the table. The team links each survey response to records of emergency department visits that happen after the survey date. Loneliness is measured first and visits are counted afterward, so the timing makes it a cohort. If the team had compared loneliness with visits in the year before the survey, the same data would only support a cross-sectional comparison, and nobody could say whether loneliness came before the visits.

**Sarah:** So to sum up section one: the question chooses the design, the five steps of the map lead to a family, constraints narrow the options, and the Cedar Valley questions land in three different families.

**Kiffer:** That last point is the bridge to section two. A study whose questions fall in different families has to decide how its strands fit together.

**Sarah:** Section two is about mixed-methods designs. What makes a study mixed methods, as opposed to a study that just has a survey and some interviews?

**Kiffer:** The defining feature is integration. John Creswell and Vicki Plano Clark describe a mixed-methods study as one that collects and analyzes both kinds of data rigorously and integrates them. Each part that uses one kind of data is called a strand. If a team runs a survey and some interviews and reports them in separate chapters, never relating one to the other, it has two kinds of data, and it has not yet done a mixed-methods analysis.

**Sarah:** Why would someone want to combine them in the first place?

**Kiffer:** Jennifer Greene and her colleagues looked at published mixed-methods evaluations in nineteen eighty-nine and identified five purposes. Triangulation looks for agreement between methods. Complementarity uses one method to elaborate the other. Development uses one method to inform the other, for example to choose a sample. Initiation looks for contradictions that raise new questions. And expansion uses different methods for different parts of a study to widen its scope.

**Sarah:** The lesson also introduces a notation with capital letters and arrows. It looked a bit like algebra.

**Kiffer:** It is simpler than it looks. Janice Morse proposed it in nineteen ninety-one. A strand written in capitals has priority, and a strand in lower case plays a supporting role. A plus sign means the strands run at the same time, and an arrow means one follows the other. So a main survey followed by a handful of supporting interviews is quan in capitals, an arrow, and qual in lower case.

**Sarah:** Then there are the three core designs.

**Kiffer:** There are three. The convergent design collects both kinds of data in the same phase, analyzes each separately, and then merges the results to compare them. It is the fastest of the three. The explanatory sequential design starts with quantitative data and then collects qualitative data to help explain what the numbers showed. The exploratory sequential design starts with qualitative data and uses what it finds to build something, like survey items or an intervention, for a later quantitative strand.

**Sarah:** When would you choose the exploratory one?

**Kiffer:** When there is no good instrument for what you want to measure, or when an intervention has to be adapted to a community. You listen first, then build, then test. The catch is time, and if you build a new questionnaire, its measurement properties have to be tested before you can trust it.

**Sarah:** And the explanatory sequential design has its own catch, I assume.

**Kiffer:** It does. Because the interviews depend on survey results, the final interview guide does not exist yet when the ethics application is written. Teams usually submit a draft guide and then send the final version to the research ethics board as an amendment. The survey also has to ask permission to contact people again, or the team has nobody to invite.

**Sarah:** The lesson talks about four ways strands can be integrated.

**Kiffer:** Those come from Michael Fetters and his colleagues. Connecting links the strands through sampling, so survey answers decide who gets interviewed. Building uses the results of one strand to write the questions of the other. Merging brings two analyzed sets of results together, and embedding links the strands at several points, often inside a larger design such as a trial.

**Sarah:** And the joint display is where merged results end up.

**Kiffer:** That is the usual tool. A joint display is a table that places the quantitative and qualitative results side by side, topic by topic, along with the conclusion drawn from reading them together. That conclusion is called a meta-inference.

**Sarah:** Give me a Cedar Valley example.

**Kiffer:** In the illustrative display, thirty-four point nine percent of respondents who live alone scored six or higher on the loneliness scale, compared with eighteen point one percent of those who live with others. In the interviews, people living alone described evenings, weekends and meals as the hardest times. The strands agree, which is called confirmation, and the interviews add something practical about when support would matter most.

**Sarah:** What happens when they do not agree?

**Kiffer:** That is called discordance, and it is a finding to report. In the Cedar Valley display, more than half of the lonely respondents said they would join a weekly social program. Yet several interview participants said they would avoid anything described as a program for lonely people, because the label felt embarrassing. Those two results pull in different directions, and together they tell the team that stated interest may overestimate attendance and that how a program is described matters.

**Sarah:** Which design did the Cedar Valley team choose?

**Kiffer:** It chose an explanatory sequential design, with both strands in capitals because the qualitative question carries as much weight as the quantitative one. The survey and the linked records come first. Then come twenty-four interviews with older adults living alone who moved to a smaller community in the past five years, and four focus groups.

**Sarah:** Why not convergent? It would have been faster.

**Kiffer:** The team considered it. It rejected the convergent design because it wanted to use survey answers to choose interviewees, and to write interview prompts that probe the patterns the survey found. It rejected the exploratory design because the loneliness scale it uses is already validated, so there was no instrument to build.

**Sarah:** And that choice ripples outward.

**Kiffer:** It does. The survey needs two extra items, one asking permission to recontact and one asking how long the person has lived at their current address. The ethics application has to describe both phases and promise an amendment for the final guide. And the interviews cannot start until the survey closes, which matters a great deal for the timeline in section four.

**Sarah:** So section two in brief: integration is the defining feature, the notation captures timing and priority, there are three core designs, and four ways of joining the strands, with the joint display as the main tool for merged results. Section three is about the research protocol. Students might confuse it with a grant proposal.

**Kiffer:** They share a lot of content, and they do different jobs. A proposal argues that a study deserves money, so it stresses the importance of the problem and the strength of the team. A protocol is the operating plan. It tells the people who carry out, approve and oversee the study exactly what will happen: who is eligible, how they are recruited, what is measured and when, how data are stored and how they will be analyzed.

**Sarah:** Who actually reads a protocol?

**Kiffer:** More people read it than students expect. The research team uses it as a manual. The research ethics board reviews it. Funders and partners check what is being asked of them. Data stewards, the organizations that hold records, read it to decide whether the requested data are needed and well protected. The advisory group checks that it reflects what was agreed. And later on, journal reviewers compare the published report with the protocol.

**Sarah:** The lesson calls the protocol a source document. What does that mean?

**Kiffer:** It means the other documents a project produces all draw on it. The grant proposal, the ethics application, the data access request, the data management plan, any registration record, and finally the methods section of the report all rest on the same decisions. If you write those decisions once in the protocol and copy from there, everything stays consistent. If each document is written from memory, small differences creep in, and someone will eventually ask you to explain them.

**Sarah:** What sections does a protocol have?

**Kiffer:** Templates vary, and the one your funder or ethics board gives you takes precedence. For randomized trials there is a checklist called the SPIRIT statement, most recently updated in twenty twenty-five. For most other studies, the lesson describes twelve sections: administrative information, background and rationale, questions and hypotheses, design and conceptual framework, setting and participants, data collection and measures, the analysis plan, ethics and engagement, data management, timeline, dissemination, and references with appendices.

**Sarah:** Is there one section you would single out?

**Kiffer:** The analysis plan. It states how the data will be analyzed before anyone has seen them. That protects a study from the temptation to try analyses until one gives a striking result and then report only that one.

**Sarah:** The worked example is a one-page summary. Why bother with a summary if you have the full document?

**Kiffer:** Partners, new staff and committees need something short. But the more useful reason is that writing it early is a test. If there is a row you cannot fill in, you have found a decision the team has not made yet.

**Sarah:** Protocols change, though. How do you keep track?

**Kiffer:** The team uses version numbers. Working drafts get decimal numbers, like zero point one and zero point two. The first version submitted to the ethics board becomes version one point zero. Later drafts become one point one, one point two, and the next submitted version becomes two point zero. Every page carries the version number and date, and the protocol keeps an amendment log.

**Sarah:** And an amendment is any change after approval?

**Kiffer:** It is any change to research that the ethics board has already approved. Under the Tri-Council Policy Statement, the team submits the change and waits for approval before making it, unless the change is needed to remove an immediate risk to participants. Adding a survey question, changing recruitment materials or adding a site all count.

**Sarah:** What does the Cedar Valley log look like?

**Kiffer:** Version one point zero went in with the ethics application in month three and was approved in month five. Version two point zero added the final interview guide. Version three point zero added telephone interviews as an option, because the advisory group pointed out that some rural participants do not have reliable internet.

**Sarah:** The section ends with registration and publication. Why would anyone publish a protocol before the results exist?

**Kiffer:** Because it creates a public, dated record of what the study planned to do. Clinical trials are expected to be registered, and many journals will only consider a trial that was registered before its first participant enrolled. For other designs, preregistration is optional, but many teams do it. The reason is a documented problem. An-Wen Chan and colleagues compared trial protocols with the published reports and found that outcomes were often reported selectively, with the ones showing clear differences more likely to appear.

**Sarah:** Let's move to section four, managing the project. Where does a timeline begin?

**Kiffer:** It begins with a work breakdown structure, which is a list of everything the project must produce, broken into tasks small enough that someone can estimate how long each will take. Then you estimate durations, and you add time to your estimates.

**Sarah:** Why add time? Shouldn't people just estimate accurately?

**Kiffer:** They try, but research by Roger Buehler and his colleagues showed that people routinely underestimate how long their own tasks will take, even when similar tasks have run late before. It is called the planning fallacy, and the practical response is to build in buffer time.

**Sarah:** Then come dependencies.

**Kiffer:** A dependency says one task cannot start until another finishes. The survey pilot cannot begin until the ethics board approves the study. Once the tasks, durations and dependencies are on a calendar, you have a Gantt chart, named after Henry Gantt, who developed these bar charts in the nineteen tens.

**Sarah:** What is the critical path?

**Kiffer:** The critical path is the longest chain of dependent tasks from start to finish. Its length sets the shortest time in which the project can be done. Any delay on that path delays the whole project. Tasks off the path have slack, meaning they can slip a bit without moving the end date.

**Sarah:** What does the Cedar Valley chart show?

**Kiffer:** It runs twenty-four months. The critical path goes from the protocol through ethics review, survey build and fieldwork, the interviews and focus groups, transcription and coding, integration, and the final report. Because interviewees come from survey respondents, both strands sit on that path. The data access request and the linkage run alongside, with some slack.

**Sarah:** Some of those steps are not really in the team's control.

**Kiffer:** Three of them depend on others: the ethics review, the data access request and the agreements with the six partner clinics. The team cannot speed them up, so it starts them as early as possible, adds buffer time after them, and plans useful work for the waiting period. In Cedar Valley, the team builds the survey during the ethics review, so the pilot can start the moment approval arrives.

**Sarah:** Next is roles. The lesson uses a RACI matrix.

**Kiffer:** The letters stand for responsible, accountable, consulted and informed. The responsible person does the work. The accountable person answers for the task and makes the final decision, and there is only ever one. Consulted people give input before decisions, and informed people are kept up to date. It sounds bureaucratic, but it prevents the situation where two people each assume the other is handling something.

**Sarah:** And authorship goes in the same document?

**Kiffer:** It goes in a short team charter that holds the matrix. The lesson uses the criteria of the International Committee of Medical Journal Editors: a substantial contribution to the design or to the data and their interpretation, drafting or critical revision, approval of the final version, and accountability for the work. Advisory members and community partners who meet those criteria are authors, and others are acknowledged. Agreeing on this at the start is much easier than at the end.

**Sarah:** Budgets come next. What makes a good budget line?

**Kiffer:** Every line is a quantity multiplied by a unit cost, and the budget justification explains both the calculation and the need. Take transcription. Four focus groups of about ninety minutes produce about three hundred and sixty minutes of audio. At two dollars a minute, a professional transcriber costs seven hundred and twenty dollars. The interviewers transcribe the twenty-four interviews themselves from speech recognition drafts, so that work sits under personnel, and the justification says clean verbatim transcripts are needed for the coding in the analysis plan.

**Sarah:** And reviewers check those numbers against the protocol.

**Kiffer:** They do. If the protocol says twenty-four interviews, the budget should pay for twenty-four honoraria and enough staff time to transcribe twenty-four recordings. The illustrative Cedar Valley total is one hundred and fifty-three thousand, four hundred and eighty dollars, with the data access costs added once the provider's quote arrives. The justification also lists in-kind contributions, like the university hosting the survey platform.

**Sarah:** Then there is file organization. This is the part I suspect students will skip.

**Kiffer:** I hope they do not. A two-year project produces thousands of files. The Cedar Valley folders are numbered in the order the work happens. Raw data are kept read-only, so the original export can always be recovered, and cleaned data are written only by scripts. The linking key, signed consent forms and audio sit in a separate restricted folder that only two named people can open.

**Sarah:** And the file names follow a rule.

**Kiffer:** The name starts with the date, written as year, month and day, followed by the project code, the content and a two-digit version number, with no spaces. The international date format makes files sort in date order, and two-digit versions keep version ten after version nine. The lesson also advises against the word final in file names, because documents are so often revised again after being called final.

**Sarah:** How is version control different from that?

**Kiffer:** Version control is any system that records changes over time. Documents use named versions and a change log. Shared institutional storage keeps an automatic history, which is a useful safety net but does not say what changed or why. And analysis code is best tracked with Git, which records each change with a short message. One firm rule is that data never go into a code repository.

**Sarah:** Finally, there is the data management plan. Students met this in the ethics lesson.

**Kiffer:** In Lesson five it appeared mainly as a privacy safeguard. Here it becomes an operating document. A data management plan describes how the data will be collected, documented, stored, protected, shared and preserved. The lesson walks through seven headings, from data collection through to responsibilities and resources, and fills each one in for Cedar Valley.

**Sarah:** What guides the plan?

**Kiffer:** Two sets of principles guide it. The FAIR principles ask that data be findable, accessible, interoperable and reusable. The CARE principles for Indigenous data governance, which are collective benefit, authority to control, responsibility and ethics, add attention to the people and purposes behind the data. In Canada they sit alongside the First Nations principles of ownership, control, access and possession that Lesson four turned into questions for research agreements.

**Sarah:** Let's close by putting it together. What does the Cedar Valley plan look like in one view?

**Kiffer:** The design is explanatory sequential, with equal priority for both strands and a cohort design for the emergency department question. The Gantt chart covers twenty-four months, and because interviewees are chosen from survey respondents, the critical path runs through both strands.

**Sarah:** And the rest of the plan?

**Kiffer:** A responsibility matrix names one accountable person for each task, the budget shows a calculation for every line, and the folders, file names and version control keep every file findable. The data management plan ties these together, and the team reviews it at each milestone and whenever the protocol changes.

**Sarah:** That is a concrete list. What comes next week?

**Kiffer:** Lesson seven covers sampling, recruitment and measurement: who your study population is, how you reach them, and how you turn concepts into measures you can actually collect.

**Sarah:** Thank you, Kiffer.

**Kiffer:** Thank you, Sarah. I will see everyone next week.
