First Steps in Analysis
Research Methods in Health Sciences
Learning objectives for this lesson:
- Organize survey data in a rectangular layout and protect the raw file with double data entry, validation rules and a read-only copy.
- Write data dictionary entries that record each variable's name, label, type, values, missing codes and derivation.
- Use the data dictionary of a real survey, the Canadian Social Connection Survey, to look up variables and identify missing codes.
- Calculate and interpret frequency tables, valid percentages, means, medians, standard deviations and cross-tabulations, and explain why missing codes are converted to missing values before analysis.
- Build and interpret a descriptive Table 1 for a defined analytic sample, with column percentages and missing values reported.
- Build a starter qualitative codebook and code an interview transcript with deductive and inductive codes and analytic memos.
- Describe seven qualitative analytic approaches and match each to the research questions and data collection methods it suits.
This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University. It is the applied research methods course of the Public Health Assessment and Analysis series.
Data Entry, Codebooks and Data Dictionaries
Learning Objectives for this section
- Organize research data in a rectangular layout with one row per participant, one column per variable and one value per cell.
- Describe how double data entry, validation rules and a read-only raw file protect data quality.
- Write a data dictionary entry that gives a variable's name, label, type, permitted values, missing codes and source.
- Explain why numeric missing codes must be converted before analysis, and distinguish the main reasons a value can be missing.
- Use the metadata of the Canadian Social Connection Survey to look up what a variable means and to identify its missing codes.
Introduction
Analysis begins long before the first table is produced. The fictional Cedar Valley Social Connection Study, which this course follows, has finished collecting its regional survey of adults aged 65 and older: 1,600 people completed it, and 392 of them (24.5 percent) scored 6 or higher on the three-item UCLA Loneliness Scale. Before the team can report those numbers with confidence, someone has to make sure that the file holding them is complete, correctly laid out and fully documented. This section covers the three habits that prevent most of them: entering data in a consistent layout, writing a data dictionary, and handling missing values deliberately. It then applies these ideas to a real Canadian dataset that Section 2 analyzes.
The graduate research assistant on the Cedar Valley team exports the survey from REDCap as a comma-separated values (CSV) file. Of the 1,600 surveys, 560 were returned on paper and were typed by hand into the same REDCap form, with a random tenth re-entered as a check (Lesson 8). Dr. Maya Hart asks for three things before anyone calculates a percentage: an untouched copy of the raw export, a data dictionary that explains every column, and a written rule for each kind of missing answer.
1.1 Data Entry and the Shape of a Dataset
One row, one column, one value
Almost every statistical program expects data in a rectangular layout. Each row holds one unit of observation, which in a survey is usually one participant. Each column holds one variable, which is a characteristic that can take different values, such as age or loneliness score. Each cell holds one value. A unique participant identifier in the first column links each row to the consent record and, later, to linked administrative data (Lesson 9), without storing the person's name in the analysis file. Broman and Woo (2018) give practical rules for spreadsheets that follow from this layout: use one header row with short names, put nothing but data in the data cells, never use colour or bold type to carry information, keep codes consistent (always "Woman", never a mix of "W", "woman" and "F"), write dates in the year-month-day form (2026-03-14), and avoid merged cells and blank rows.
| participant_id | age | gender | community | ucla_1 | ucla_2 | ucla_3 |
|---|---|---|---|---|---|---|
| CV-00012 | 72 | 1 | 1 | 2 | 2 | 3 |
| CV-01733 | 68 | 2 | 2 | 1 | 1 | 1 |
| CV-03106 | 81 | 1 | 4 | 3 | 2 | 2 |
| CV-04580 | 75 | 2 | 3 | 1 | 2 | -88 |
The table above shows four illustrative rows of the Cedar Valley raw file. The numbers in the gender, community and UCLA columns are codes, and a reader cannot interpret them without the data dictionary in Section 1.2. The value -88 for participant CV-04580 is a missing code meaning that the person chose not to answer the third loneliness item.
Entering paper forms
Typing paper forms into a file introduces errors such as transposed digits, skipped lines and values entered in the wrong column. The standard protection is double data entry: two people enter the same forms independently, a program compares the two files cell by cell, and every disagreement is checked against the paper original. A second protection is a set of validation rules (also called range checks) that refuse impossible values at the moment of entry, such as an age of 6 in a study of people aged 65 and older or a UCLA item score of 4 on a scale that runs from 1 to 3. Electronic capture in REDCap or Qualtrics applies these rules as participants answer (Lesson 8), although exported files still need checking because a rule only catches the errors someone anticipated.
Protecting the raw file
The raw export is the study's primary record. The team saves it once, marks it read-only, and never edits it. Every change, from fixing a typing error to creating a new variable, is made by code in a script that reads the raw file and writes a separate analysis file. Anyone can then rerun the script to reproduce the analysis file, and a cleaning mistake can be found and reversed, in keeping with the file organization habits of Lesson 6.
1.2 Codebooks and Data Dictionaries
A data dictionary is a table that describes every variable in a dataset. For each variable it records the variable name used in the file, a variable label that gives the full question or a plain description, the type of data, the permitted values and what each code means (the value labels), the missing codes, the units, and the source or derivation. The word codebook is often used for the same document, and some codebooks also add the frequency of each value. Section 3 uses "codebook" in a different sense, for the list of qualitative codes applied to interview text, so this lesson keeps the two terms apart by calling the quantitative document a data dictionary. Lesson 3 began a variable list from the Cedar Valley DAG, and Lesson 8 added the survey items; the data dictionary is where that list becomes complete.
| Name | Label | Type | Values and codes | Missing codes | Source or derivation |
|---|---|---|---|---|---|
| participant_id | Study identification number | Text | CV-00001 to CV-05250 | None permitted | Assigned when the sample was drawn (Lesson 7) |
| age | Age in years | Numeric (ratio) | 65 to 110 | -99 not answered | Survey question 1 |
| gender | Gender | Categorical (nominal) | 1 Woman; 2 Man; 3 Another gender | -88 prefer not to answer; -99 not answered | Survey question 2 |
| community | Community of residence | Categorical (nominal) | 1 Cedar City; 2 Riverside; 3 North Bench; 4 Kestrel Lake; 5 Other | -99 not answered | Survey question 4 |
| ucla_1 | How often do you feel that you lack companionship? | Categorical (ordinal) | 1 Hardly ever; 2 Some of the time; 3 Often | -88; -99 | UCLA item 1 (Hughes et al., 2004) |
| ucla_total | Three-item UCLA Loneliness Scale score | Numeric, derived | 3 to 9 | Left empty if any item is missing | ucla_1 + ucla_2 + ucla_3 |
| lonely | Lonely by the study cut-off | Categorical (binary), derived | 1 score of 6 or higher; 0 score of 3 to 5 | Left empty if ucla_total is empty | Recoded from ucla_total |
Two entries deserve attention. The derived variables ucla_total and lonely do not appear on the survey; they are created in the cleaning script, and the dictionary states the rule that creates them so that anyone can check it. The three UCLA items are each scored from 1 to 3 (Hughes et al., 2004), so the total runs from 3 to 9, and the Cedar Valley team decided in its protocol that a score of 6 or higher would count as lonely.
HSCI 410 Lesson 1 Section 2 builds a full study data dictionary on these principles, and HSCI 410 Lesson 2 Section 1 uses the data dictionary to set valid ranges for data cleaning.
1.3 Missing Values and Missing Codes
A value can be missing for different reasons, and the reasons matter for analysis. A question may be not applicable because branching logic skipped it (a question about a spouse's health is skipped for someone who is not married). A participant may decline to answer or answer that they do not know. A participant may stop before reaching the question, or may never be shown it. Many datasets record these reasons with numeric missing codes such as -77, -88 and -99, which keep the reasons distinct in the raw file.
Numeric missing codes become dangerous when they are treated as real numbers. Suppose four Cedar Valley participants report ages of 72, 68, 81 and 75, and a fifth has the code -99. The mean of the four real ages is 74.0 years. If the -99 is left in, the computed mean of the five values is 39.4 years, which looks like a plausible adult age and could pass unnoticed. Before any analysis, the cleaning script must convert each missing code to the value that the software recognizes as missing. Many programs display that value as NA (not available). Summaries such as the mean are then calculated from the recorded values only, and the number missing is reported beside them, as Section 2 shows.
Rule of thumb for missing data
Keep the reason for missingness in the raw file, convert every missing code to NA in the analysis file, and report how many values are missing for each variable in your tables. HSCI 341 Lesson 3 Section 4 applies missing-value codes to questionnaire coding, and HSCI 410 Lesson 2 Section 2 extends them to range checks, verification and scripted cleaning.
1.4 A Real Dataset: the Canadian Social Connection Survey
The Cedar Valley data are fictional, so the series practises on a real dataset. The Canadian Social Connection Survey (CSCS) is a national online survey of social connection, loneliness and health among people living in Canada, and a de-identified public version is available on GitHub. The public file combines several survey years, and this lesson uses the 2021 wave. The worked examples below show what the file contains and how its data dictionary and missing values look in practice.
The public file holds three components: the survey responses, a metadata table that serves as the data dictionary, and a summary of the scales. The responses form a rectangular file with 13,219 rows and 3,247 columns. Keeping only the rows whose survey year is 2021 leaves 4,045 rows, one for each respondent in the 2021 wave, each with the same 3,247 columns.
Reading the names and the dictionary
With 3,247 columns, nobody memorizes the variables. The CSCS names begin with a prefix for each part of the survey, and the metadata record the question text for each variable in each wave.
A search of the column names for those beginning with LONELY_ucla finds eight variables. They follow the structure of a good dictionary: three items stored as labelled categories, three numeric versions (_num), the derived total (_score) and the derived yes or no category (_score_y_n).
| Variable name | Contents |
|---|---|
| LONELY_ucla_loneliness_scale_companionship | Companionship item, labelled categories |
| LONELY_ucla_loneliness_scale_left_out | Left out item, labelled categories |
| LONELY_ucla_loneliness_scale_isolated | Isolated item, labelled categories |
| LONELY_ucla_loneliness_scale_companionship_num | Companionship item, numeric |
| LONELY_ucla_loneliness_scale_left_out_num | Left out item, numeric |
| LONELY_ucla_loneliness_scale_isolated_num | Isolated item, numeric |
| LONELY_ucla_loneliness_scale_score | Derived total score, 3 to 9 |
| LONELY_ucla_loneliness_scale_score_y_n | Derived category: No (3 to 5) or Yes (6 to 9) |
The 2021 dictionary has 387 entries, far fewer than 3,247 columns, because the file also holds derived variables and questions from other years. Looking up LONELY_ucla_loneliness_scale_left_out in the 2021 dictionary returns its question text: "Indicate how often each of the statements below is descriptive of you. - How often do you feel left out?" Looking up GEO_housing_household_size returns nothing, meaning that the 2021 metadata have no entry for it. Later waves list a question about how many people the respondent lives with, but an analyst would need to find out how the 2021 values were produced before using them. This is what an undocumented variable looks like in practice.
Inspecting a variable and finding its missing codes
Age (DEMO_age) is stored as a whole number. The 2021 respondents were aged 16 to 100, with a first quartile of 28, a median of 34, a mean of 39.8 and a third quartile of 52. Self-rated physical health (WELLNESS_self_rated_physical_health) is stored as a categorical variable with six categories in order: Poor, Fair, Good, Very good, Excellent, and "Presented but no response". The last category is a missing code stored as a category: the question was shown and left blank. The frequency table below for the "left out" item shows both kinds of missingness in the CSCS.
| How often do you feel left out? | Number of respondents |
|---|---|
| Hardly Ever | 1,312 |
| Some of the time | 1,657 |
| Often | 631 |
| Presented but no response | 24 |
No recorded answer (NA) | 421 |
| Total | 4,045 |
Twenty-four respondents saw the item and skipped it, and 421 have no recorded answer (NA, not available), for example because they stopped before reaching the question. Section 2 converts the "Presented but no response" category to NA before computing percentages.
Write a data dictionary entry for each of three CSCS variables, with the six columns used in the Cedar Valley table above. The first is DEMO_relationship_status, whose categories are Single and not dating, Single and dating, In a relationship, and "Presented but no response". The second is WELLNESS_self_rated_physical_health, whose categories are listed in the worked example above. The third is LONELY_ucla_loneliness_scale_score_y_n; state its derivation in words.
Reflection
A volunteer at a seniors' centre has typed 40 paper surveys from the centre's own survey of its members into a spreadsheet. The Age column contains 72, 68, seventy and 999. The Gender column contains W, woman, F and Man. The three UCLA loneliness items (each meant to be coded 1 for Hardly ever, 2 for Some of the time and 3 for Often) contain some blank cells and some cells with -88, which the volunteer says means "prefer not to answer". The header spans two merged rows, and rows highlighted in yellow mean "call this person back". Identify the problems in this file, explain what you would do about each one without editing the raw file, and write a data dictionary entry for the first UCLA item with a name, label, type, values and codes, missing codes, and source.
The file breaks several layout rules. The merged two-row header should become one header row with short names such as age, gender and ucla_1. The yellow highlighting stores information in colour, so I would add a column callback coded 1 or 0. In Age, the text seventy mixes words with numbers, and 999 looks like an undocumented missing code; I would check both against the paper forms and record any remaining unknown value as -99, documented in the dictionary. Gender uses four spellings for two categories, so the cleaning script should recode W, woman and F to one code for Woman after confirming on the forms what F meant. The blank UCLA cells and the -88 values mean different things, so I would code blanks as -99 (not answered) and keep -88 (prefer not to answer) in the raw file, then convert both to NA in the analysis file. All changes go in a script that reads the raw file and writes a new one. Because these forms were typed by hand, I would also ask a second person to enter them independently and compare the files.
Dictionary entry: name ucla_1; label "How often do you feel that you lack companionship?"; type categorical (ordinal); values 1 Hardly ever, 2 Some of the time, 3 Often; missing codes -88 prefer not to answer, -99 not answered; source survey question 10, item 1 of the three-item UCLA Loneliness Scale (Hughes et al., 2004).
Minimum 20 characters required.
Question 1: In a rectangular dataset prepared for analysis, what does each row usually represent?
Question 2: Five Cedar Valley participants report ages of 72, 68, 81 and 75, and a fifth has the missing code -99. With the code left in, the mean is 39.4 years. What should the cleaning script do?
NA) in the analysis file, and the number missing is reported. Filling in the mean (option a) invents a value, and the raw file is never edited (option b).Question 3: Which part of a data dictionary entry states that ucla_total is the sum of ucla_1, ucla_2 and ucla_3?
Question 4: In the 2021 CSCS, the item on feeling left out shows 24 respondents as "Presented but no response" and 421 as NA. What is the best interpretation?
NA means no answer was recorded, for example because the respondent stopped before reaching the question. Option b reverses the two meanings.Frequency Tables, Descriptive Statistics and a Table 1
Learning Objectives for this section
- Choose an appropriate summary (counts and percentages, or a mean, standard deviation and median) from a variable's level of measurement.
- Produce a frequency table with percentages, and explain the difference between a percentage of all participants and a valid percentage.
- Compute and interpret the mean, median and standard deviation, and explain how missing values are handled when these summaries are calculated.
- Build a cross-tabulation and distinguish row percentages from column percentages.
- Define an analytic sample and assemble a descriptive Table 1, with the missing values reported.
Introduction
Descriptive statistics summarize who is in a dataset and what their answers look like. They come before any comparison or model, for two reasons. First, they let the team check the data: an age of 340 or a loneliness score of 12 shows up immediately in a frequency table. Second, readers need to know who was studied before they can judge whether the findings apply to anyone else. In most health research articles this description appears as the first table, which is why it is called Table 1. This section works through frequency tables, numeric summaries and cross-tabulations, and then builds a Table 1 from the CSCS 2021 data for respondents aged 65 and older, the age group that the fictional Cedar Valley study recruits.
2.1 Matching the Summary to the Variable
Lesson 7 introduced levels of measurement. They decide which summary is appropriate. Categorical variables, whether nominal (categories with no order, such as community of residence) or ordinal (ordered categories, such as self-rated health from poor to excellent), are summarized with counts and percentages. Numeric variables measured on an interval or ratio scale, such as age or the number of emergency department visits, are summarized with a measure of centre (the mean or median) and a measure of spread (the standard deviation or the interquartile range).
| Level of measurement | Cedar Valley example | Usual summary |
|---|---|---|
| Nominal | Community of residence | n (%) |
| Ordinal | Self-rated physical health | n (%) for each level |
| Binary | Lonely (UCLA score 6 or higher) | n (%) |
| Interval or ratio, roughly symmetric | Age | Mean (standard deviation) |
| Ratio, skewed | Emergency department visits in a year | Median (interquartile range) |
HSCI 410 Lesson 2 Section 3 extends these summaries to measures of shape, including skew, and to transformations of skewed variables.
2.2 Frequency Tables
A frequency table lists each value of a variable with the number of participants who gave it. Adding percentages makes groups of different sizes comparable. The Cedar Valley team reports that 392 of 1,600 respondents, or 24.5 percent, scored 6 or higher on the UCLA scale. A percentage always has a denominator, and the analyst must say what it is. A valid percentage uses only the participants with a recorded answer as the denominator, while a percentage of all participants includes those with missing values. The difference is small when little is missing and large when much is missing.
The CSCS analysis starts by keeping respondents aged 65 and older, so that the practice data resemble the Cedar Valley population.
Keeping the 2021 respondents aged 65 or more leaves 486 people. The frequency table below counts each possible UCLA total and adds a count of missing values.
| UCLA total | Number of respondents |
|---|---|
| 3 | 100 |
| 4 | 51 |
| 5 | 58 |
| 6 | 81 |
| 7 | 40 |
| 8 | 39 |
| 9 | 64 |
| Missing (no score) | 53 |
One hundred had the lowest possible score of 3, 64 had the highest possible score of 9, and 53 have no score because at least one of the three items is missing. Reading a frequency table like this one is also a data check: every value lies within the permitted range of 3 to 9.
The derived variable LONELY_ucla_loneliness_scale_score_y_n groups the scores at the same cut-off the Cedar Valley team uses: 209 respondents scored 3 to 5 and 224 scored 6 to 9. Leaving out the missing values gives valid percentages with a denominator of 433 (209 plus 224), so 48.3 percent were not classed as lonely and 51.7 percent of older respondents with a score were classed as lonely. Had the 53 respondents without a score been kept in the denominator, the figure would have been 224 of 486, or 46.1 percent, which understates the proportion among those who answered. The CSCS figure is much higher than the 24.5 percent in the fictional Cedar Valley survey. The CSCS recruited its participants online without drawing them at random from a sampling frame, so its percentage describes these respondents and should not be read as the prevalence of loneliness among older Canadians (Lesson 7).
2.3 Describing Numeric Variables
The mean is the sum of the values divided by their number. The median is the middle value when the values are sorted, so half of the participants lie below it and half above. The standard deviation (SD) describes how far values typically lie from the mean: a small SD means that most values are close to the mean, and a large SD means that they are spread out. The interquartile range runs from the first quartile (the value below which a quarter of participants lie) to the third quartile (the value below which three quarters lie), and it contains the middle half of the data.
The mean and the median agree when values are spread symmetrically, and they part ways when a few values are extreme. Suppose nine Cedar Valley participants had 0, 0, 0, 0, 1, 1, 2, 3 and 11 emergency department visits in a year. The total is 18, so the mean is 2.0 visits, while the median is 1 visit. One person with 11 visits pulls the mean upward, and the median better describes the typical participant. Counts of health service use are often skewed in this way, which is why the table in Section 2.1 recommends the median for them.
Of the 486 respondents aged 65 and older, 433 have a UCLA score and 53 do not. A mean cannot be calculated from a set of values that includes unknowns, so each summary below uses the 433 recorded scores, and the number missing is reported beside them.
| Summary of the UCLA score | Value |
|---|---|
| Respondents with a score | 433 |
| Missing | 53 |
| Mean | 5.65 |
| Standard deviation | 2.09 |
| Minimum | 3 |
| First quartile | 4 |
| Median | 6 |
| Third quartile | 7 |
| Maximum | 9 |
Rounded to one decimal place, as reports usually give them, the mean and SD are 5.7 and 2.1. The median is 6, and the interquartile range runs from 4 to 7.
2.4 Cross-Tabulations
A cross-tabulation (or two-way table) counts participants in every combination of two categorical variables. It is the first step toward asking whether two characteristics go together. Before tabulating, any missing code stored as a category must be converted to NA, or it will appear as a real group in every table.
Among the 486 respondents aged 65 and older, relationship status has four categories, one of which is the missing code "Presented but no response". The cleaning step replaces that category with NA and removes the now-empty category. The ten respondents are still counted, now as missing.
| Relationship status | Before conversion | After conversion |
|---|---|---|
| Single and not dating | 187 | 187 |
| Single and dating | 22 | 22 |
| In a relationship | 267 | 267 |
| Presented but no response | 10 | Category removed |
Missing (NA) | 0 | 10 |
The cross-tabulation below places gender in the rows and the loneliness group in the columns for the 433 respondents with a UCLA score. Column percentages divide each count by its column total, so each column adds to 100 percent. Row percentages divide each count by its row total, so each row adds to 100 percent.
| Gender | n, score 3 to 5 | n, score 6 to 9 | Column %, 3 to 5 | Column %, 6 to 9 | Row %, 3 to 5 | Row %, 6 to 9 |
|---|---|---|---|---|---|---|
| Man | 72 | 60 | 34.4 | 26.8 | 54.5 | 45.5 |
| Woman | 136 | 157 | 65.1 | 70.1 | 46.4 | 53.6 |
| Non-binary | 1 | 7 | 0.5 | 3.1 | 12.5 | 87.5 |
Column percentages answer the question "what are the people in each loneliness group like?" Among respondents classed as lonely, 70.1 percent were women, 26.8 percent were men and 3.1 percent were non-binary. Among those not classed as lonely, 65.1 percent were women. Column percentages are the form used in a Table 1, where each column describes one group.
Row percentages answer the question "how common is loneliness in each gender group?" Among men, 45.5 percent were classed as lonely; among women, 53.6 percent were. Row percentages suit a comparison of the outcome across groups. The row for non-binary respondents (87.5 percent) rests on 8 people, so a single respondent changes it by 12.5 percentage points.
A cell with 1 person, such as the non-binary respondent who was not classed as lonely, gives an unstable percentage and can make a person identifiable in a small community. Many data custodians set a minimum cell size for published tables, and a study that uses linked administrative data must follow the rules in its data access agreement (Lesson 9). Common solutions are to combine small categories or to suppress the count.
HSCI 410 Lesson 2 Section 3 returns to frequency tables and cross-tabulations and adds the odds ratio to the cross-tabulation.
2.5 Building a Table 1
A Table 1 describes the analytic sample, the participants included in the analysis, usually overall and separately for the groups that the study compares. The STROBE reporting guideline for observational studies asks authors to give the characteristics of study participants and the number with missing data for each variable (von Elm et al., 2007), and Lesson 12 returns to STROBE and to table design. Building the table involves five decisions, which the cards below summarize.
The first decision defines the analytic sample. Keeping the respondents whose loneliness group is recorded leaves 433 of the 486 respondents aged 65 and older: 209 with scores of 3 to 5 and 224 with scores of 6 to 9. Age is roughly symmetric, so it is summarized as a mean (SD): 71.3 years (5.3) overall, 71.3 (4.9) in the group with scores of 3 to 5 and 71.3 (5.8) in the group with scores of 6 to 9.
Each categorical row comes from a cross-tabulation of the characteristic by loneliness group, with an overall column that adds the two groups together, and the counts are converted to column percentages. For relationship status, 7 respondents have no recorded status (1 in the group with scores of 3 to 5 and 6 in the group with scores of 6 to 9). They are reported on a separate row, and the percentages are calculated without them, so the overall denominator is 426. Education and self-rated physical health, after its missing-code category is converted to NA, are handled in the same way. The table below brings the rows together.
| Characteristic | Overall (n = 433) | UCLA score 3 to 5 (n = 209) | UCLA score 6 to 9 (n = 224) |
|---|---|---|---|
| Age in years, mean (SD) | 71.3 (5.3) | 71.3 (4.9) | 71.3 (5.8) |
| Gender, n (%) | |||
| Woman | 293 (67.7) | 136 (65.1) | 157 (70.1) |
| Man | 132 (30.5) | 72 (34.4) | 60 (26.8) |
| Non-binary | 8 (1.8) | 1 (0.5) | 7 (3.1) |
| Relationship status, n (%) | |||
| In a relationship | 238 (55.9) | 136 (65.4) | 102 (46.8) |
| Single and dating | 20 (4.7) | 5 (2.4) | 15 (6.9) |
| Single and not dating | 168 (39.4) | 67 (32.2) | 101 (46.3) |
| Missing | 7 | 1 | 6 |
| Education, n (%) | |||
| Bachelor's degree or higher | 149 (34.4) | 81 (38.8) | 68 (30.4) |
| Less than a bachelor's degree | 284 (65.6) | 128 (61.2) | 156 (69.6) |
| Self-rated physical health, n (%) | |||
| Excellent | 38 (9.7) | 26 (14.0) | 12 (5.9) |
| Very good | 110 (28.2) | 60 (32.3) | 50 (24.5) |
| Good | 132 (33.8) | 66 (35.5) | 66 (32.4) |
| Fair | 87 (22.3) | 32 (17.2) | 55 (27.0) |
| Poor | 23 (5.9) | 2 (1.1) | 21 (10.3) |
| Missing | 43 | 23 | 20 |
Table 1. Characteristics of CSCS 2021 respondents aged 65 and older, overall and by score on the three-item UCLA Loneliness Scale. Percentages are column percentages among respondents with a recorded value. Of 486 respondents aged 65 and older, 53 without a UCLA score were excluded. Loneliness scores range from 3 to 9; a score of 6 or higher was classed as lonely.
Reading a Table 1
The table describes the analytic sample. Mean age is the same in both groups. Respondents classed as lonely were more often single and not dating (46.3 percent compared with 32.2 percent) and more often rated their physical health as poor (10.3 percent compared with 1.1 percent) or fair (27.0 percent compared with 17.2 percent). These are descriptions of differences in this sample. The table contains no statistical tests, and on its own it cannot show whether the differences would hold in other samples or whether loneliness and health influence one another. Statistical tests are taught in HSCI 341 Lesson 6 Section 4 and HSCI 410 Lesson 3 Section 1, regression models in HSCI 410 Lesson 3, and the causal questions require the reasoning of HSCI 341. The table also shows the analyst's choices openly: the row order runs from the most to the least favourable health rating, the overall column comes first, and the footnote gives the denominator and the exclusions.
Using Table 1, write two sentences that describe the education rows. Then write two sentences that describe the self-rated physical health rows, and state the denominator of the percentages in each column, given that 43 respondents (23 with scores of 3 to 5 and 20 with scores of 6 to 9) have no recorded rating.
Reflection
The table below describes 433 respondents aged 65 and older in the Canadian Social Connection Survey (CSCS) 2021, a survey whose participants were recruited online without a sampling frame. Respondents are grouped by score on the three-item UCLA Loneliness Scale (3 to 5 not lonely; 6 to 9 lonely). Relationship status, n (%), by group: in a relationship, overall 238 (55.9), not lonely 136 (65.4), lonely 102 (46.8); single and dating, 20 (4.7), 5 (2.4), 15 (6.9); single and not dating, 168 (39.4), 67 (32.2), 101 (46.3); missing 7, 1, 6. Gender: non-binary, overall 8 (1.8), not lonely 1 (0.5), lonely 7 (3.1). Write a short results paragraph describing the relationship status rows. State the denominator used for the percentages, comment on the non-binary row, and say what the table cannot tell a reader.
Among the 433 respondents aged 65 and older with a UCLA score, 55.9 percent were in a relationship, 39.4 percent were single and not dating, and 4.7 percent were single and dating. Respondents classed as lonely were less often in a relationship than those not classed as lonely (46.8 percent compared with 65.4 percent) and more often single and not dating (46.3 percent compared with 32.2 percent). The percentages are column percentages among respondents with a recorded relationship status, so the overall denominator is 426 (433 minus 7 missing), the lonely group denominator is 218 and the not-lonely denominator is 208; the 7 missing values are reported separately.
The non-binary row rests on 8 people, with a single person in the not-lonely column, so its percentages are unstable and the cell could identify someone. In a report I would either combine it with another category, if that did not hide something important to the study, or suppress the small count in line with the data custodian's rules.
The table describes this sample only. It contains no statistical tests, it cannot show whether relationship status affects loneliness or the reverse, and because the CSCS was recruited online without a sampling frame, its percentages should not be read as the prevalence of these characteristics among older Canadians.
Minimum 20 characters required.
Question 1: Which summary best suits the number of emergency department visits in a year, when most people have 0 or 1 visit and a few have many?
Question 2: Of 486 CSCS respondents aged 65 and older, 224 were classed as lonely, 209 were not, and 53 had no UCLA score. What is the valid percentage classed as lonely?
Question 3: In a cross-tabulation with gender in the rows and loneliness group in the columns, 70.1 percent of respondents classed as lonely were women, and 53.6 percent of women were classed as lonely. Which statement is correct?
Building a First Codebook and Coding a Transcript
Learning Objectives for this section
- Define a qualitative code, a coded segment and a memo, and distinguish first-cycle codes from the categories and themes built from them later.
- Distinguish deductive codes, drawn from the research question, interview guide or theory, from inductive codes that arise from the data, including in vivo codes.
- Build a starter codebook in which each code has a name, a definition, rules for when to use and not use it, and an example.
- Code a short interview transcript step by step, applying deductive codes, adding inductive codes and writing memos.
- Describe how a team revises a codebook, compares coding and chooses tools for coding.
Introduction
Part 2 of this lesson turns from numbers to words. The qualitative strand of the fictional Cedar Valley Social Connection Study includes 24 semi-structured interviews with older adults living alone and four focus groups: two with older adults, one with family caregivers, and one with clinic staff and community connectors. Lesson 10 prepared a clean verbatim transcript of one interview. A single interview of 54 minutes can produce fifteen or more pages of text, and 24 of them produce several hundred. Coding is the first systematic step in making that volume of text manageable without losing what participants said. This section builds a starter codebook for the Cedar Valley interviews and uses it to code the excerpt from Lesson 10. It gives a working method for a first coding pass, and HSCI 841 Lessons 5 to 12 develop qualitative analysis in depth.
3.1 What a Code Is
A code is a short label, usually a word or phrase, that a researcher attaches to a passage of text to record what the passage is about or what it means for the research question (Saldaña, 2021). The passage is a coded segment; it may be a phrase, a sentence or a few sentences that express one idea. One segment can carry more than one code, and the same code is applied to every segment, in every transcript, that fits its definition. Coding lets the researcher gather every segment with the code "transportation" from all 24 interviews and read them side by side.
Saldaña (2021) distinguishes first-cycle coding, the initial labelling of segments, from second-cycle coding, in which codes are grouped, compared and organized into larger categories and themes. A theme is a pattern of meaning that runs across many segments and participants and says something about the research question. This lesson stops at first-cycle coding. Section 4 shows how different analytic approaches carry codes forward. HSCI 841 Lesson 5 Section 1.4 teaches the move from codes to themes, and HSCI 841 Lesson 6 turns codes into conceptual models.
Coding goes together with writing memos. A memo is a short dated note in which the researcher records an idea, a question, a possible connection between codes or a reaction to the data. Analytic memos written during coding continue the reflexive memos of Lesson 10 and often hold the first drafts of later themes.
3.2 Deductive and Inductive Codes
Codes come from two directions. Deductive codes (also called a priori codes) are written before coding begins. They come from the research question, the topics of the interview guide, the causal web drawn in Lesson 3, or an existing theory. For example, Weiss (1973) distinguished emotional loneliness, the absence of a close attachment figure, from social loneliness, the absence of a wider network of friends and acquaintances, and a team could write a code for each. Inductive codes are created during coding, when a passage expresses something important that no existing code captures. One useful kind is the in vivo code, which uses the participant's own words as the label, such as "a room full of strangers". Most applied health studies combine the two, starting with a short deductive list and adding inductive codes as the data require. Fereday and Muir-Cochrane (2006) describe this hybrid approach and the record-keeping it needs.
| Deductive codes | Inductive codes | |
|---|---|---|
| Source | Research question, interview guide, theory, earlier studies | The data themselves |
| When written | Before coding begins | During coding |
| Strength | They keep coding tied to the question and make team coding consistent from the start. | They capture ideas the team did not anticipate, in participants' terms. |
| Risk | Coders may force text into codes that do not fit and overlook what is new. | The list can grow very long, with overlapping codes, unless it is reviewed. |
| Cedar Valley example | Transportation and mobility, taken from the causal web | Concealing loneliness, added after reading interview 14 |
3.3 Building a Starter Codebook
A qualitative codebook is the list of codes with the rules for applying them. MacQueen and colleagues (1998), writing about team-based analysis, recommended that each entry give a short code name, a brief definition, a full definition, guidance on when to use the code, guidance on when not to use it, and an example. The "when not to use" rule is the part beginners most often leave out, and it is the part that keeps two similar codes apart. A starter codebook holds the deductive codes written before coding begins; it is version 1.0, and it will change.
The Cedar Valley team wrote seven deductive codes from its SPIDER question (how older adults living alone experience social connection after a move to a smaller town), the interview guide and the causal web.
| Code | Definition | Use when | Do not use when |
|---|---|---|---|
| MOVING | Changes in social life that the participant links to moving home | The participant compares life before and after a move. | A change is linked to something else, such as bereavement. |
| FAMILY | Contact with relatives and its frequency, form and quality | Calls, visits, help or tension with family are described. | The person is a friend or neighbour; use COMMUNITY. |
| COMMUNITY | Contact with friends, neighbours, groups, programs and community places | Libraries, churches, centres, groups or neighbours are described. | Contact is only with paid health care providers; use HEALTH. |
| TRANSPORT | The ability to travel, including driving, transit and rides | Getting places, or being unable to, is described. | Travel is mentioned only as a setting with no effect on contact. |
| HEALTH | Health conditions, function and health care that affect social contact | Illness, surgery, mobility or appointments shape contact. | Health is mentioned with no link to social life. |
| TECHNOLOGY | Phones, video calls and the internet used to stay in touch | A device or connection helps or hinders contact. | A phone call is mentioned with no comment on the medium; code the relationship only. |
| LONELINESS | The participant's own descriptions of feeling lonely, isolated or alone | A feeling of disconnection is described or expressed. | The passage describes the amount of contact without a feeling. |
The table shows four of the six columns that MacQueen and colleagues recommended. The full codebook also has a brief definition for quick reference and an example segment for each code, which the team adds as soon as it finds a clear example in the data. HSCI 841 Lesson 5 Section 2.4 extends the six columns to seven elements by splitting the example into positive and negative examples.
3.4 Coding a Transcript Step by Step
Read the transcript once from start to finish without coding, ideally while listening to the recording. Write a memo of a few sentences on your first impressions: what seemed most important to the participant, what surprised you, and what you want to look for on the second reading.
On the second reading, mark each passage that expresses one idea relevant to the research question. Keep enough text that the segment makes sense when it is read alone, away from the transcript. Passages that do not bear on the question, such as small talk at the start, can be left uncoded.
For each segment, check the codebook and apply every code whose definition and "use when" rule fit. Read the "do not use when" rule each time, especially for similar codes.
When a segment expresses something important that no code captures, write a new code, with a definition, in a list of proposed codes. Do not stretch an existing code to cover it.
Record connections between codes, questions to ask in later interviews, and your own reactions. Date each memo and note the segment that prompted it.
After the first two or three transcripts, the team reviews the proposed codes, adds the useful ones with full definitions, merges or splits codes that overlap, and saves the result as a new version (1.1, 1.2 and so on) with the date and the reason for each change. Earlier transcripts are then rechecked against the revised codes.
The excerpt below is invented for teaching; no real person was interviewed. It comes from minute 18 to minute 21 of a 54-minute interview with P14, given the pseudonym "Ruth", a woman aged 78 who lives alone and moved from Cedar City to Kestrel Lake about two years before the interview. Lesson 10 transcribed it in clean verbatim form. Read it in full before looking at the coding that follows it.
I: You mentioned that you moved to Kestrel Lake about two years ago. Can you tell me what the first few months were like?
P14: It was harder than I expected. In Cedar City I knew everybody on my street, and I never had to plan to see people. I'd go to the bank or the pharmacy and I'd run into someone I knew. Here, nobody knew me from Adam. I'd walk to the store and come home and realize I hadn't spoken to anyone except the cashier.
I: What was that like for you?
P14: [pause] Quiet. Very quiet. My daughter, [daughter's name], phones every Sunday, and that helps, but it isn't the same as somebody dropping by. I didn't tell her how I felt, because she was the one who found me this place, and she worries enough already.
I: You said you didn't tell her. Can you say a bit more about that?
P14: I didn't want to be a burden. When you get to my age, people start deciding things for you. If I'd said I was lonely, she'd have had me in a care home by Christmas. [laughs] So I said to her, "I'm fine, Mom's fine."
I: How did things change, if they did?
P14: The first winter was the worst. I'd stopped driving after my cataract surgery, and there are no buses out here, so when the roads were bad I could go a week, a whole week, without leaving the house. What turned it around was the library. [Librarian's name], the lady who runs it, started a coffee morning on Tuesdays, and she phoned me herself to ask me to come. I wouldn't have gone if she hadn't called. Now I go every week, and a fellow from the coffee group drives me to my eye appointments in Cedar City.
I: It sounds like that phone call mattered.
P14: It did. It really did. I think people assume older folks will just show up if you put a notice on the board. I'd seen the notice. I just couldn't walk into a room full of strangers on my own. Not at my age.
I: Is there anything that still makes it hard to stay in touch with people?
P14: The internet out here is poor, so the video calls with my grandson freeze up. He thinks I'm pulling faces. And in a small town everybody knows your business. I'm careful what I say at coffee, because it goes around.
A first coding pass
The table shows one coder's first pass. Deductive codes are in capitals; proposed inductive codes are in italics, with in vivo codes in quotation marks.
| Segment | Codes | Memo |
|---|---|---|
| "In Cedar City I knew everybody on my street ... I'd run into someone I knew." | MOVING; Chance encounters | Contact in the old neighbourhood happened without planning. The move removed it. |
| "I'd walk to the store and come home and realize I hadn't spoken to anyone except the cashier." | MOVING; Chance encounters | The same errands no longer produce contact. Ask other participants about everyday errands. |
| "Quiet. Very quiet." | LONELINESS; "Very quiet" | She answers a feeling question with a description of silence after a pause. |
| "My daughter ... phones every Sunday, and that helps, but it isn't the same as somebody dropping by." | FAMILY | Regular calls are valued but are described as different from in-person visits. |
| "I didn't want to be a burden ... she'd have had me in a care home by Christmas." | FAMILY; LONELINESS; Concealing loneliness; "a burden" | Hiding loneliness to protect independence. Check whether other participants describe this. |
| "I'd stopped driving after my cataract surgery, and there are no buses out here ... a whole week without leaving the house." | TRANSPORT; HEALTH | Health, rural transit and winter weather combine. This links to the transportation node in the causal web. |
| "She phoned me herself to ask me to come. I wouldn't have gone if she hadn't called." | COMMUNITY; Personal invitation | The program existed, and the personal call is what made her attend. |
| "A fellow from the coffee group drives me to my eye appointments in Cedar City." | COMMUNITY; TRANSPORT; HEALTH | A social contact became practical help with health care access. |
| "I just couldn't walk into a room full of strangers on my own." | Personal invitation; "a room full of strangers" | A notice was not enough. Possible implication for the health authority's programs. |
| "The internet out here is poor, so the video calls with my grandson freeze up." | TECHNOLOGY; FAMILY | Rural internet quality limits contact with family. |
| "In a small town everybody knows your business. I'm careful what I say at coffee." | COMMUNITY; Small-town visibility | A new connection also brings exposure. This may limit what she shares. |
The pass proposes four inductive codes: Chance encounters, Concealing loneliness, Personal invitation and Small-town visibility. Each now needs a definition and "use when" and "do not use when" rules before it enters version 1.1 of the codebook. For example, Personal invitation could be defined as "a direct, individual request from a person to join an activity, described as affecting whether the participant attended", to be used when the participant links attendance to being asked, and not used for general publicity such as notices. One interview cannot show whether these codes recur; the memos flag them so the coders can watch for them in the remaining 23 interviews.
3.5 Coding as a Team, and Tools for Coding
When more than one person codes, the team agrees on the codebook first. Two coders then code the same two or three transcripts independently, compare their coding segment by segment, and discuss each disagreement. Disagreements usually reveal an unclear definition, which the team rewrites. Some teams then calculate a measure of intercoder agreement such as Cohen's kappa, while others, particularly those using reflexive approaches, treat discussion itself as the check. HSCI 841 Lesson 5 covers agreement measures and when they suit an approach.
Coding can be done with highlighters on paper, comments in a word processor, or a spreadsheet with one row per segment, as in the table above. Qualitative data analysis software stores the transcripts, codes and memos together and retrieves every segment with a given code in one step. Taguette is a free, open-source option that HSCI 841 uses; NVivo, ATLAS.ti, MAXQDA and Dedoose are commercial programs. Whatever the tool, transcripts must be de-identified before coding, and a cloud-based program may only be used if the consent form and the data management plan permit it (Lessons 5 and 10).
Copy the excerpt into a document or spreadsheet. Using the seven deductive codes and the four proposed inductive codes, code it without looking at the table above, then compare. Write a full codebook entry (name, definition, use when, do not use when, example) for Concealing loneliness, and propose one further inductive code that you think the first pass missed, with a one-sentence memo explaining why.
Reflection
The following excerpt is invented for teaching. P09, a fictional man aged 81 who lives alone in Riverside, says: "Since my wife died I don't really go to the Legion anymore. It was her that kept the calendar. I've got the dog, and I talk to the neighbour over the fence most mornings, but evenings are long. My son wants me to get one of those tablets for video calls, but I can't see the screen well enough." The starter codebook has seven codes: MOVING (changes in social life linked to moving home; do not use when the change is linked to something else, such as bereavement); FAMILY (contact with relatives); COMMUNITY (contact with friends, neighbours, groups and community places); TRANSPORT (ability to travel); HEALTH (health conditions or function that affect social contact); TECHNOLOGY (devices or connections that help or hinder contact); and LONELINESS (the participant's own descriptions of feeling lonely, isolated or alone). Divide the excerpt into segments, apply the codes, propose at least one inductive code with a definition and a "use when" rule, and write a two-sentence memo.
Segment 1, "Since my wife died I don't really go to the Legion anymore. It was her that kept the calendar": COMMUNITY. MOVING does not apply, because the change is linked to bereavement. I propose an inductive code, Loss of a social organizer: a partner or other person who arranged social activities has died or left, and the participant describes reduced participation as a result; use when the participant links less contact to losing the person who planned it.
Segment 2, "I've got the dog, and I talk to the neighbour over the fence most mornings": COMMUNITY for the neighbour, with a second proposed code, Animal companionship, for the dog.
Segment 3, "but evenings are long": LONELINESS, with an in vivo code, "evenings are long".
Segment 4, "My son wants me to get one of those tablets for video calls, but I can't see the screen well enough": FAMILY, TECHNOLOGY and HEALTH, because a vision problem limits a device that could support contact with his son.
Memo: P09 describes contact that is regular in the morning and absent in the evening, which suggests looking at the timing of loneliness across interviews. His account of his wife as the person who kept the calendar is similar to Ruth's reliance on the librarian's invitation, so both may point to the role of a person who initiates contact.
Minimum 20 characters required.
Question 1: A coder labels Ruth's phrase "a room full of strangers" with exactly those words. What kind of code is this?
Question 2: Which code in the Cedar Valley coding is deductive?
Question 3: Why does each codebook entry include a "do not use when" rule?
Question 4: After coding three transcripts, a coder finds an important idea that no existing code captures. What should happen next?
Qualitative Analytic Approaches and the Data They Suit
Learning Objectives for this section
- Describe the aim, typical data and main steps of thematic analysis, qualitative content analysis, framework analysis, rapid qualitative analysis, grounded theory, interpretative phenomenological analysis and narrative analysis.
- Match each approach to the data collection methods it suits, including in-depth interviews, key informant interviews, focus groups, open-ended survey responses and documents.
- Use the research question, the data, the timeline and the team to choose an approach for a study.
- Explain how the first-cycle coding of Section 3 feeds into each approach, and locate where HSCI 841 teaches each one.
Introduction
Coding is common to most qualitative work, but what happens to the codes afterwards depends on the analytic approach. An approach is a recognized way of moving from data to findings, with its own aims, procedures and standards of quality. Thematic analysis looks for patterns across participants; interpretative phenomenological analysis stays close to each person's experience; narrative analysis keeps stories whole. Ideally the approach is chosen when the study is designed and written into the protocol (Lesson 6), because it affects how many participants are needed, how interviews are run and how transcripts are prepared. This section introduces seven approaches used in health research, maps them to the data collection methods of Lesson 10, and shows how the fictional Cedar Valley team chose among them. It is an orientation. HSCI 841 Lesson 1 Section 4 relates the quality criteria in Section 4.3 to systematic, transparent and replicable analysis, and Section 4.4 lists the HSCI 841 lesson for each approach.
4.1 Seven Analytic Approaches
Each card below gives the aim of the approach, the data it usually works with, its main steps and a foundational source.
The approaches share some steps. Each begins with close reading, and most involve coding. They differ in what they produce: a set of themes, a set of categories with frequencies, a comparison matrix, a rapid summary for decision-makers, an explanatory theory, a detailed account of a few people's experience, or an analysis of stories. They also differ in how they treat a codebook. Framework analysis, rapid analysis, directed content analysis and codebook forms of thematic analysis use a structured codebook like the one built in Section 3. Reflexive thematic analysis, grounded theory and interpretative phenomenological analysis develop codes more freely and treat a fixed codebook with caution.
What each approach would do with Ruth's interview
The coded excerpt from Section 3 shows how the same first-cycle codes lead in different directions. The table describes what an analyst using each approach would do next with interview CV-INT-14.
| Approach | Next step with the coded excerpt |
|---|---|
| Thematic analysis | Compare Personal invitation and COMMUNITY segments across all 24 interviews to see whether a theme about being drawn in by someone known is developing. |
| Qualitative content analysis | Count how many participants describe a transportation barrier, and report the categories of barrier with their frequencies. |
| Framework analysis | Write a short summary of Ruth's coded data under each code in the matrix row for P14, among the Kestrel Lake participants. |
| Rapid qualitative analysis | Fill in a one-page template with Ruth's main points under each topic of the interview guide, ready for the interim brief. |
| Grounded theory | Compare Concealing loneliness with similar incidents in other interviews, and recruit further participants who moved recently to test an emerging idea about how connection is rebuilt after a move. |
| Interpretative phenomenological analysis | Read the whole interview closely to interpret how Ruth makes sense of hiding her loneliness from her daughter, before moving to other cases. |
| Narrative analysis | Keep the account whole and examine its structure: the move as a setting, the first winter as the complication, the librarian's call as the turning point and the weekly coffee morning as the resolution. |
4.2 Matching Approaches to Data Collection Methods
Lesson 10 introduced in-depth interviews, key informant interviews and focus groups. Two other common sources of qualitative data are open-ended survey questions and documents such as policies, meeting minutes or program records. The table maps the seven approaches to these sources. "Strong fit" means that the approach was designed for that kind of data or is routinely used with it; "possible" means that it is used with that data with some adaptation; a blank cell means that the combination is uncommon.
| Approach | In-depth interviews | Key informant interviews | Focus groups | Open-ended survey responses | Documents |
|---|---|---|---|---|---|
| Thematic analysis | Strong fit | Strong fit | Strong fit | Strong fit | Possible |
| Qualitative content analysis | Possible | Possible | Possible | Strong fit | Strong fit |
| Framework analysis | Strong fit | Strong fit | Strong fit | Possible | Possible |
| Rapid qualitative analysis | Possible | Strong fit | Strong fit | Possible | |
| Grounded theory | Strong fit | Possible | Possible | Possible | |
| Interpretative phenomenological analysis | Strong fit | ||||
| Narrative analysis | Strong fit | Possible |
Two patterns stand out. Interpretative phenomenological analysis and narrative analysis need long, personal accounts from individuals, so they depend on one-to-one interviews (or, for narrative analysis, diaries and written life stories), and focus groups suit them poorly because the group conversation interrupts each person's account. Qualitative content analysis and thematic analysis can handle short answers from many people, which makes them the usual choices for open-ended survey questions. The table is a guide drawn from common practice; methodologists debate some cells, and HSCI 841 examines those debates.
4.3 Choosing an Approach
Four questions narrow the choice. What does the research question ask for: patterns, counts, comparisons, an explanation of a process, the depth of individual experience, or stories? What data will there be, and how much? When are findings needed, and by whom? Who will do the analysis, and how many people? The diagram pairs typical answers with the approaches that suit them.
The Cedar Valley decision
The Cedar Valley team considered several options for its 24 interviews and four focus groups. The tabs summarize its reasoning, which the team recorded in a memo and then in the protocol.
For the 24 interviews, the team chose framework analysis. The SPIDER question asks about experiences across participants, the health authority wants to know whether experiences differ between Cedar City and the smaller communities, and four people will share the coding. After first-cycle coding with the codebook from Section 3, each interview will be summarized in a matrix with one row per participant and one column per code, and the rows will be grouped by community so that patterns such as the role of transportation in Kestrel Lake can be compared with those in Cedar City.
The health authority asked for early findings within six weeks of the last focus group to inform planning for its community programs. For the focus group with clinic staff and community connectors, the team chose rapid qualitative analysis: the moderator and note-taker will complete a summary template organized by the topics of the focus group guide within two days of the session, and the summaries will feed a short brief. The full transcript will later be added to the framework matrix.
The team discussed narrative analysis of participants' accounts of moving, since several interviews, including Ruth's, tell the move as a story with a hard first winter and a turning point. It also discussed interpretative phenomenological analysis for a small group of recently widowed participants. It recorded both as possible follow-up studies, because each would need longer, less structured interviews than its guide provides and analytic training that the current team does not have.
Quality in qualitative analysis
Whatever the approach, readers judge qualitative findings by how carefully they were produced. Lincoln and Guba (1985) proposed criteria of credibility, transferability, dependability and confirmability, which researchers address with practices such as an audit trail of codebook versions and memos, reflexive memos about the researcher's own position, and a clear description of the setting. The reporting guidelines COREQ and SRQR, which Lesson 12 introduces, list what a report should include. HSCI 841 Lesson 1 Section 4 relates these criteria to systematic, transparent and replicable analysis.
4.4 Where HSCI 841 Picks Up
HSCI 841, the series' graduate course in qualitative research methods and analysis, assumes the skills of this lesson: preparing transcripts, building a codebook and completing a first coding pass. The table shows where each approach is developed.
| Approach or skill | HSCI 841 lesson |
|---|---|
| Finding themes, the phases of thematic analysis, inductive and deductive coding, codebooks and intercoder agreement | Lesson 5, Themes and Codebooks, Sections 1 to 3 |
| Framework analysis and matrices | Lesson 7, Comparing Variables and Grounded Theory, Section 1 |
| Rapid qualitative analysis | Lesson 7, Comparing Variables and Grounded Theory, Section 1.7 |
| Conceptual models | Lesson 6, Analysis Frameworks and Conceptual Models, Section 3 |
| Grounded theory and constant comparison | Lesson 7, Comparing Variables and Grounded Theory |
| Qualitative and quantitative content analysis | Lesson 8, Content Analysis |
| Narrative analysis | Lesson 9, Schema and Narrative Analysis |
| Interviewing for interpretative phenomenological analysis; phenomenology and interpretative phenomenological analysis as a methodology | Lesson 4, Qualitative Data Collection, Section 3.1 (interviewing); Lesson 2, Research Questions, Theory and Literature, Section 2.7 (methodology) |
| Discourse analysis, analytic induction and computational text analysis | Lessons 10 to 12 |
Reflection
Primary care clinics in a fictional health region want to understand how clinic staff decide whether to refer an older patient to a community connector. The research team will conduct 12 key informant interviews of about 30 minutes with physicians, nurses and receptionists. The clinics need findings for a planning meeting five weeks after the last interview, and two analysts will share the work. The seven approaches available are thematic analysis, qualitative content analysis, framework analysis, rapid qualitative analysis, grounded theory, interpretative phenomenological analysis and narrative analysis. Choose an approach, justify it using four considerations (what the question asks for, the data, the timeline and the team), describe the first two analytic steps, and name one approach you would not use, with the reason.
I would use rapid qualitative analysis. The question asks for a practical description of how staff make referral decisions, which suits a structured summary of what each informant reports. The data are 12 short key informant interviews organized around a guide, so a summary template with one heading per guide topic (for example, how patients are identified, what information staff use, barriers to referral and suggestions) fits them well. The timeline of five weeks leaves little room for full transcription and line-by-line coding of every interview, and rapid analysis was developed for decisions on this kind of schedule. Two analysts can share the work if both complete templates the same way, which is checked by having both summarize the first two interviews and compare.
Step one is to finalize the template and complete it for each interview within a few days, from the recording and notes. Step two is to transfer the summaries into a matrix with one row per informant and one column per topic, grouped by role, so that patterns, such as differences between physicians and receptionists, can be read across rows.
I would not use interpretative phenomenological analysis, because it is designed for in-depth interviews about how a few people make sense of a significant personal experience, and these short professional interviews about a work process do not provide that kind of account. Framework analysis would be a reasonable alternative if the timeline were longer.
Minimum 20 characters required.
Question 1: A team must give a health authority findings from eight key informant interviews within four weeks. Which approach suits this best?
Question 2: Which approach keeps each participant's story whole and examines its structure, such as the setting, complication, turning point and resolution?
Question 3: Why do focus groups fit poorly with interpretative phenomenological analysis?
Question 4: Which approach places coded data in a matrix with one row per participant and one column per code, so that cases and groups can be compared?
Final Assessment
Bringing It All Together
This lesson took the first analytic steps with both kinds of data that the fictional Cedar Valley Social Connection Study collects. Part 1 prepared survey data for analysis: a rectangular layout, a read-only raw file, a data dictionary and deliberate handling of missing codes. It then used the 2021 wave of the Canadian Social Connection Survey to work through frequency tables, descriptive statistics, cross-tabulations and a Table 1 for respondents aged 65 and older.
Part 2 prepared interview data for analysis. It defined codes, segments and memos, distinguished deductive from inductive codes, built a starter codebook with rules for when to use and not use each code, and coded an excerpt from the interview with Ruth. It then described seven analytic approaches, matched them to the data collection methods they suit, and showed how the Cedar Valley team chose framework analysis for its interviews and rapid analysis for an interim brief.
The two parts follow the same logic. In each, written rules (a data dictionary or a codebook) make the work repeatable, missing or unclear material is handled openly, and the first summary (a Table 1 or a coding table) prepares the ground for analyses taught in HSCI 410 and HSCI 841.
Key Takeaways from this lesson
- A dataset ready for analysis has one row per participant, one column per variable and one value per cell, with a unique identifier in place of names.
- The raw file is kept read-only, and every change is made in a script that writes a separate analysis file.
- A data dictionary records each variable's name, label, type, values, missing codes and derivation, and it is best written before cleaning begins.
- Numeric missing codes must be converted to NA before analysis, because a code such as -99 treated as a number distorts every summary.
- Categorical variables are summarized with counts and percentages, and numeric variables with a mean and standard deviation or a median and interquartile range.
- Every percentage has a denominator, and a table should state whether it uses valid or total denominators and report the number missing.
- Column percentages describe the make-up of each group in a Table 1, and row percentages compare the outcome across groups.
- A starter codebook holds deductive codes with definitions and rules for use, and it grows through dated versions as inductive codes are added.
- First-cycle coding labels segments of text; categories and themes are built from the codes in later analysis.
- The analytic approach is chosen from the research question, the data, the timeline and the team, and Section 4.4 lists the HSCI 841 lesson for each approach.
Core Concepts Reviewed
Section 1: rectangular data, double data entry, validation rules, the read-only raw file, data dictionaries, derived variables, missing codes, and the documentation of a real survey dataset.
Section 2: levels of measurement, frequency tables, valid percentages, mean, median, standard deviation, interquartile range, cross-tabulations, row and column percentages, the analytic sample and Table 1.
Section 3: codes, coded segments, memos, first-cycle and second-cycle coding, deductive, inductive and in vivo codes, the starter codebook, and team coding.
Section 4: thematic analysis, qualitative content analysis, framework analysis, rapid qualitative analysis, grounded theory, interpretative phenomenological analysis and narrative analysis, matched to data collection methods.
The final reflection asks you to compare the quantitative and qualitative steps of this lesson and to plan the first analytic steps for a new dataset and two new transcripts.
Reflection
Part 1 of this lesson prepared and described survey data, and Part 2 prepared and coded interview data. Consider these pairs: a data dictionary and a qualitative codebook; a missing code converted to NA and a passage left uncoded or given a proposed new code; the definition of an analytic sample and the selection of segments for coding; a Table 1 and a coding table or framework matrix. Choose two of these pairs and explain what each element of the pair does and how the two are similar and different. Then describe, in three or four sentences, how an analyst would take a new survey dataset and two new interview transcripts through the first analytic steps of this lesson.
A data dictionary and a qualitative codebook both make an analysis repeatable by writing down rules that would otherwise live in one person's head. The dictionary states what each column means, which codes are permitted and how derived variables such as a UCLA total are calculated; the codebook states what each code means and when to use or not use it. They differ in when they are settled. A data dictionary is ideally fixed before cleaning, while a codebook grows through dated versions as inductive codes are added.
A Table 1 and a framework matrix both arrange data so that groups can be compared. Table 1 summarizes each characteristic with counts, percentages or means for each loneliness group. A framework matrix places a written summary of each participant's coded data under each code, grouped, for example, by community. Table 1 reduces people to numbers, whereas the matrix keeps each person's words and context.
With a new survey dataset, the analyst would define the analytic sample, convert missing codes to NA and build a Table 1 with a footnote on denominators and exclusions. With two new transcripts, the analyst would read both in full, write five to eight deductive codes with all six codebook columns, code the first transcript with memos, revise the codebook to version 1.1, and then code the second.
Minimum 30 characters required.
Final Knowledge Assessment
Question 1: Which practice best protects the raw survey export?
Question 2: The Cedar Valley data dictionary says that ucla_total is left empty if any item is missing. A participant answered 2 and 3 on the first two items and declined the third (-88). What is ucla_total in the analysis file?
Question 3: Which summary belongs in a Table 1 row for age in a sample whose ages are roughly symmetric?
Question 4: A Table 1 reports 168 (39.4) respondents as single and not dating in the overall column of 433, with 7 missing. What denominator gives 39.4 percent?
Question 5: Why does the Table 1 in this lesson contain no p-values?
Question 6: Among CSCS respondents aged 65 and older, 51.7 percent were classed as lonely, compared with 24.5 percent in the fictional Cedar Valley survey. What is the most defensible reading of the CSCS figure?
Question 7: Which description best fits an analytic memo written during coding?
Question 8: Which statement about deductive and inductive codes is accurate?
Question 9: Which set of codes did the first coding pass apply to Ruth's sentence "A fellow from the coffee group drives me to my eye appointments in Cedar City"?
Question 10: A survey adds the open-ended question "What makes it hard to stay in touch with people?", answered in a sentence or two by 1,600 people. Which approaches are the usual fit?
Question 11: Which approach builds a theory of a social process through constant comparison and theoretical sampling, alternating data collection and analysis?
Question 12: Why did the Cedar Valley team choose framework analysis for its 24 interviews?
Question 13: Which pairing of a data problem and a step from this lesson is correct?
NA and removes the empty category. An undocumented variable needs investigation, a small cell may need combining or suppression, and passages that fit no code may need a new inductive code.Glossary: Key Terms, People & Frameworks
📚 Reference page, available throughout the lesson
Search the terms, tools and people introduced in this lesson.