HSCI 207 · Lesson 11

First Steps in Analysis

Research Methods in Health Sciences

Learning objectives for this lesson:

  • Organize survey data in a rectangular layout and protect the raw file with double data entry, validation rules and a read-only copy.
  • Write data dictionary entries that record each variable's name, label, type, values, missing codes and derivation.
  • Use the data dictionary of a real survey, the Canadian Social Connection Survey, to look up variables and identify missing codes.
  • Calculate and interpret frequency tables, valid percentages, means, medians, standard deviations and cross-tabulations, and explain why missing codes are converted to missing values before analysis.
  • Build and interpret a descriptive Table 1 for a defined analytic sample, with column percentages and missing values reported.
  • Build a starter qualitative codebook and code an interview transcript with deductive and inductive codes and analytic memos.
  • Describe seven qualitative analytic approaches and match each to the research questions and data collection methods it suits.

This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University. It is the applied research methods course of the Public Health Assessment and Analysis series.

Lesson 11 · HSCI 207

First Steps in Analysis

A short guided walkthrough before you work through the lesson at your own pace.

Research Methods in Health Sciences
Two parts, four sections

Numbers first, then words

Part 1: Quantitative

Section 1 prepares and documents survey data and introduces a real Canadian survey. Section 2 produces frequency tables, descriptive statistics and a Table 1.

Part 2: Qualitative

Section 3 builds a starter codebook and codes a transcript. Section 4 maps seven analytic approaches to the data they suit.

Running case

The Cedar Valley Social Connection Study (fictional)

1,600completed surveys of adults aged 65 and older
24.5%scored 6 or higher on the UCLA Loneliness Scale
24 + 4interviews and focus groups

Because the study is fictional, the worked examples use real data from the Canadian Social Connection Survey.

How to work through it

Reflections, checks and worked examples

Each section ends with a reflection and a short knowledge check. The final page has a thirteen-question assessment.

Worked examples

The lesson builds a Table 1 from a real dataset and a first-cycle coding of an interview transcript with a starter codebook.

Section 1 of 5

Data Entry, Codebooks and Data Dictionaries

⏱ Estimated reading time: 40 minutes
Section 1 of 5

Data Entry, Codebooks and Data Dictionaries

Preparing and documenting a survey file before analysis.

The shape of a dataset

One row, one column, one value

One row per participant

Each row is a unit of observation, identified by a study number in place of a name.

One column per variable

Each column has one short name, such as ucla_1.

One value per cell

Colour, merged cells and notes never carry data.

Broman and Woo (2018) set out practical rules for organizing data in spreadsheets.

Protecting data quality

Three habits that prevent most errors

Double data entry

Two people type the same paper forms, and the files are compared cell by cell.

Validation rules

Impossible values, such as an age of 6 in a study of older adults, are refused at entry.

A read-only raw file

The export is never edited; a script makes every change and writes a new file.

Documentation

What a data dictionary records

Variable nameVariable labelTypeValues and value labelsMissing codesSource or derivation

For a derived variable, the dictionary states the rule: ucla_total is the sum of the three items and runs from 3 to 9.

Missing codes

A code treated as a number distorts every summary

74.0mean age of 72, 68, 81 and 75 years
39.4mean age with the missing code -99 left in

Convert every missing code to NA in the analysis file and report how many values are missing.

Real data

The Canadian Social Connection Survey, 2021 wave

4,045respondents in 2021
3,247columns across all waves
387entries in the 2021 data dictionary

A blank answer is recorded as “Presented but no response”; a question never answered is NA.

Carry forward

From a documented file to a first table

You now have a dataset laid out for analysis, a dictionary to explain it, and a rule for every missing value.

Section 2 uses frequency tables, percentages, means and standard deviations to describe respondents aged 65 and older and assembles a Table 1.

Learning Objectives for this section

  • Organize research data in a rectangular layout with one row per participant, one column per variable and one value per cell.
  • Describe how double data entry, validation rules and a read-only raw file protect data quality.
  • Write a data dictionary entry that gives a variable's name, label, type, permitted values, missing codes and source.
  • Explain why numeric missing codes must be converted before analysis, and distinguish the main reasons a value can be missing.
  • Use the metadata of the Canadian Social Connection Survey to look up what a variable means and to identify its missing codes.

Introduction

Analysis begins long before the first table is produced. The fictional Cedar Valley Social Connection Study, which this course follows, has finished collecting its regional survey of adults aged 65 and older: 1,600 people completed it, and 392 of them (24.5 percent) scored 6 or higher on the three-item UCLA Loneliness Scale. Before the team can report those numbers with confidence, someone has to make sure that the file holding them is complete, correctly laid out and fully documented. This section covers the three habits that prevent most of them: entering data in a consistent layout, writing a data dictionary, and handling missing values deliberately. It then applies these ideas to a real Canadian dataset that Section 2 analyzes.

Case: the Cedar Valley survey export

The graduate research assistant on the Cedar Valley team exports the survey from REDCap as a comma-separated values (CSV) file. Of the 1,600 surveys, 560 were returned on paper and were typed by hand into the same REDCap form, with a random tenth re-entered as a check (Lesson 8). Dr. Maya Hart asks for three things before anyone calculates a percentage: an untouched copy of the raw export, a data dictionary that explains every column, and a written rule for each kind of missing answer.

1.1 Data Entry and the Shape of a Dataset

One row, one column, one value

Almost every statistical program expects data in a rectangular layout. Each row holds one unit of observation, which in a survey is usually one participant. Each column holds one variable, which is a characteristic that can take different values, such as age or loneliness score. Each cell holds one value. A unique participant identifier in the first column links each row to the consent record and, later, to linked administrative data (Lesson 9), without storing the person's name in the analysis file. Broman and Woo (2018) give practical rules for spreadsheets that follow from this layout: use one header row with short names, put nothing but data in the data cells, never use colour or bold type to carry information, keep codes consistent (always "Woman", never a mix of "W", "woman" and "F"), write dates in the year-month-day form (2026-03-14), and avoid merged cells and blank rows.

participant_idagegendercommunityucla_1ucla_2ucla_3
CV-000127211223
CV-017336822111
CV-031068114322
CV-04580752312-88

The table above shows four illustrative rows of the Cedar Valley raw file. The numbers in the gender, community and UCLA columns are codes, and a reader cannot interpret them without the data dictionary in Section 1.2. The value -88 for participant CV-04580 is a missing code meaning that the person chose not to answer the third loneliness item.

Entering paper forms

Typing paper forms into a file introduces errors such as transposed digits, skipped lines and values entered in the wrong column. The standard protection is double data entry: two people enter the same forms independently, a program compares the two files cell by cell, and every disagreement is checked against the paper original. A second protection is a set of validation rules (also called range checks) that refuse impossible values at the moment of entry, such as an age of 6 in a study of people aged 65 and older or a UCLA item score of 4 on a scale that runs from 1 to 3. Electronic capture in REDCap or Qualtrics applies these rules as participants answer (Lesson 8), although exported files still need checking because a rule only catches the errors someone anticipated.

Protecting the raw file

The raw export is the study's primary record. The team saves it once, marks it read-only, and never edits it. Every change, from fixing a typing error to creating a new variable, is made by code in a script that reads the raw file and writes a separate analysis file. Anyone can then rerun the script to reproduce the analysis file, and a cleaning mistake can be found and reversed, in keeping with the file organization habits of Lesson 6.

Raw export read-only Cleaning script every change in code Analysis file derived variables Table 1 first results Data dictionary names, labels, codes, missing values and derivations
The raw export is never edited. A script produces the analysis file, and the data dictionary documents every variable along the way.

1.2 Codebooks and Data Dictionaries

A data dictionary is a table that describes every variable in a dataset. For each variable it records the variable name used in the file, a variable label that gives the full question or a plain description, the type of data, the permitted values and what each code means (the value labels), the missing codes, the units, and the source or derivation. The word codebook is often used for the same document, and some codebooks also add the frequency of each value. Section 3 uses "codebook" in a different sense, for the list of qualitative codes applied to interview text, so this lesson keeps the two terms apart by calling the quantitative document a data dictionary. Lesson 3 began a variable list from the Cedar Valley DAG, and Lesson 8 added the survey items; the data dictionary is where that list becomes complete.

NameLabelTypeValues and codesMissing codesSource or derivation
participant_idStudy identification numberTextCV-00001 to CV-05250None permittedAssigned when the sample was drawn (Lesson 7)
ageAge in yearsNumeric (ratio)65 to 110-99 not answeredSurvey question 1
genderGenderCategorical (nominal)1 Woman; 2 Man; 3 Another gender-88 prefer not to answer; -99 not answeredSurvey question 2
communityCommunity of residenceCategorical (nominal)1 Cedar City; 2 Riverside; 3 North Bench; 4 Kestrel Lake; 5 Other-99 not answeredSurvey question 4
ucla_1How often do you feel that you lack companionship?Categorical (ordinal)1 Hardly ever; 2 Some of the time; 3 Often-88; -99UCLA item 1 (Hughes et al., 2004)
ucla_totalThree-item UCLA Loneliness Scale scoreNumeric, derived3 to 9Left empty if any item is missingucla_1 + ucla_2 + ucla_3
lonelyLonely by the study cut-offCategorical (binary), derived1 score of 6 or higher; 0 score of 3 to 5Left empty if ucla_total is emptyRecoded from ucla_total

Two entries deserve attention. The derived variables ucla_total and lonely do not appear on the survey; they are created in the cleaning script, and the dictionary states the rule that creates them so that anyone can check it. The three UCLA items are each scored from 1 to 3 (Hughes et al., 2004), so the total runs from 3 to 9, and the Cedar Valley team decided in its protocol that a score of 6 or higher would count as lonely.

Variable namesClick to explore
Variable labelsClick to explore
Value labelsClick to explore
DerivationsClick to explore

HSCI 410 Lesson 1 Section 2 builds a full study data dictionary on these principles, and HSCI 410 Lesson 2 Section 1 uses the data dictionary to set valid ranges for data cleaning.

1.3 Missing Values and Missing Codes

A value can be missing for different reasons, and the reasons matter for analysis. A question may be not applicable because branching logic skipped it (a question about a spouse's health is skipped for someone who is not married). A participant may decline to answer or answer that they do not know. A participant may stop before reaching the question, or may never be shown it. Many datasets record these reasons with numeric missing codes such as -77, -88 and -99, which keep the reasons distinct in the raw file.

Numeric missing codes become dangerous when they are treated as real numbers. Suppose four Cedar Valley participants report ages of 72, 68, 81 and 75, and a fifth has the code -99. The mean of the four real ages is 74.0 years. If the -99 is left in, the computed mean of the five values is 39.4 years, which looks like a plausible adult age and could pass unnoticed. Before any analysis, the cleaning script must convert each missing code to the value that the software recognizes as missing. Many programs display that value as NA (not available). Summaries such as the mean are then calculated from the recorded values only, and the number missing is reported beside them, as Section 2 shows.

Rule of thumb for missing data

Keep the reason for missingness in the raw file, convert every missing code to NA in the analysis file, and report how many values are missing for each variable in your tables. HSCI 341 Lesson 3 Section 4 applies missing-value codes to questionnaire coding, and HSCI 410 Lesson 2 Section 2 extends them to range checks, verification and scripted cleaning.

1.4 A Real Dataset: the Canadian Social Connection Survey

The Cedar Valley data are fictional, so the series practises on a real dataset. The Canadian Social Connection Survey (CSCS) is a national online survey of social connection, loneliness and health among people living in Canada, and a de-identified public version is available on GitHub. The public file combines several survey years, and this lesson uses the 2021 wave. The worked examples below show what the file contains and how its data dictionary and missing values look in practice.

Worked example: the shape of the CSCS file

The public file holds three components: the survey responses, a metadata table that serves as the data dictionary, and a summary of the scales. The responses form a rectangular file with 13,219 rows and 3,247 columns. Keeping only the rows whose survey year is 2021 leaves 4,045 rows, one for each respondent in the 2021 wave, each with the same 3,247 columns.

Reading the names and the dictionary

With 3,247 columns, nobody memorizes the variables. The CSCS names begin with a prefix for each part of the survey, and the metadata record the question text for each variable in each wave.

Worked example: the UCLA loneliness variables

A search of the column names for those beginning with LONELY_ucla finds eight variables. They follow the structure of a good dictionary: three items stored as labelled categories, three numeric versions (_num), the derived total (_score) and the derived yes or no category (_score_y_n).

Variable nameContents
LONELY_ucla_loneliness_scale_companionshipCompanionship item, labelled categories
LONELY_ucla_loneliness_scale_left_outLeft out item, labelled categories
LONELY_ucla_loneliness_scale_isolatedIsolated item, labelled categories
LONELY_ucla_loneliness_scale_companionship_numCompanionship item, numeric
LONELY_ucla_loneliness_scale_left_out_numLeft out item, numeric
LONELY_ucla_loneliness_scale_isolated_numIsolated item, numeric
LONELY_ucla_loneliness_scale_scoreDerived total score, 3 to 9
LONELY_ucla_loneliness_scale_score_y_nDerived category: No (3 to 5) or Yes (6 to 9)

The 2021 dictionary has 387 entries, far fewer than 3,247 columns, because the file also holds derived variables and questions from other years. Looking up LONELY_ucla_loneliness_scale_left_out in the 2021 dictionary returns its question text: "Indicate how often each of the statements below is descriptive of you. - How often do you feel left out?" Looking up GEO_housing_household_size returns nothing, meaning that the 2021 metadata have no entry for it. Later waves list a question about how many people the respondent lives with, but an analyst would need to find out how the 2021 values were produced before using them. This is what an undocumented variable looks like in practice.

Inspecting a variable and finding its missing codes

Worked example: types, summaries and missing codes

Age (DEMO_age) is stored as a whole number. The 2021 respondents were aged 16 to 100, with a first quartile of 28, a median of 34, a mean of 39.8 and a third quartile of 52. Self-rated physical health (WELLNESS_self_rated_physical_health) is stored as a categorical variable with six categories in order: Poor, Fair, Good, Very good, Excellent, and "Presented but no response". The last category is a missing code stored as a category: the question was shown and left blank. The frequency table below for the "left out" item shows both kinds of missingness in the CSCS.

How often do you feel left out?Number of respondents
Hardly Ever1,312
Some of the time1,657
Often631
Presented but no response24
No recorded answer (NA)421
Total4,045

Twenty-four respondents saw the item and skipped it, and 421 have no recorded answer (NA, not available), for example because they stopped before reaching the question. Section 2 converts the "Presented but no response" category to NA before computing percentages.

Try it: write three dictionary entries

Write a data dictionary entry for each of three CSCS variables, with the six columns used in the Cedar Valley table above. The first is DEMO_relationship_status, whose categories are Single and not dating, Single and dating, In a relationship, and "Presented but no response". The second is WELLNESS_self_rated_physical_health, whose categories are listed in the worked example above. The third is LONELY_ucla_loneliness_scale_score_y_n; state its derivation in words.

Reflection

A volunteer at a seniors' centre has typed 40 paper surveys from the centre's own survey of its members into a spreadsheet. The Age column contains 72, 68, seventy and 999. The Gender column contains W, woman, F and Man. The three UCLA loneliness items (each meant to be coded 1 for Hardly ever, 2 for Some of the time and 3 for Often) contain some blank cells and some cells with -88, which the volunteer says means "prefer not to answer". The header spans two merged rows, and rows highlighted in yellow mean "call this person back". Identify the problems in this file, explain what you would do about each one without editing the raw file, and write a data dictionary entry for the first UCLA item with a name, label, type, values and codes, missing codes, and source.

Model answer

The file breaks several layout rules. The merged two-row header should become one header row with short names such as age, gender and ucla_1. The yellow highlighting stores information in colour, so I would add a column callback coded 1 or 0. In Age, the text seventy mixes words with numbers, and 999 looks like an undocumented missing code; I would check both against the paper forms and record any remaining unknown value as -99, documented in the dictionary. Gender uses four spellings for two categories, so the cleaning script should recode W, woman and F to one code for Woman after confirming on the forms what F meant. The blank UCLA cells and the -88 values mean different things, so I would code blanks as -99 (not answered) and keep -88 (prefer not to answer) in the raw file, then convert both to NA in the analysis file. All changes go in a script that reads the raw file and writes a new one. Because these forms were typed by hand, I would also ask a second person to enter them independently and compare the files.

Dictionary entry: name ucla_1; label "How often do you feel that you lack companionship?"; type categorical (ordinal); values 1 Hardly ever, 2 Some of the time, 3 Often; missing codes -88 prefer not to answer, -99 not answered; source survey question 10, item 1 of the three-item UCLA Loneliness Scale (Hughes et al., 2004).

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: In a rectangular dataset prepared for analysis, what does each row usually represent?

In the rectangular layout each row is one unit of observation (usually one participant), each column is one variable and each cell holds one value. Option a describes a column.

Question 2: Five Cedar Valley participants report ages of 72, 68, 81 and 75, and a fifth has the missing code -99. With the code left in, the mean is 39.4 years. What should the cleaning script do?

Missing codes must be converted to the value the software treats as missing (NA) in the analysis file, and the number missing is reported. Filling in the mean (option a) invents a value, and the raw file is never edited (option b).

Question 3: Which part of a data dictionary entry states that ucla_total is the sum of ucla_1, ucla_2 and ucla_3?

The source or derivation column records where a variable comes from, including the rule that creates a derived variable. Value labels give the meaning of each code, and the variable label gives the question or description.

Question 4: In the 2021 CSCS, the item on feeling left out shows 24 respondents as "Presented but no response" and 421 as NA. What is the best interpretation?

"Presented but no response" means the question was shown and left blank. NA means no answer was recorded, for example because the respondent stopped before reaching the question. Option b reverses the two meanings.
Section 2 of 5

Frequency Tables, Descriptive Statistics and a Table 1

⏱ Estimated reading time: 40 minutes
Section 2 of 5

Frequency Tables, Descriptive Statistics and a Table 1

Describing who is in the data before comparing anyone.

Choosing a summary

The level of measurement decides

Categorical

Nominal, ordinal and binary variables are summarized as n (%).

Numeric

Symmetric values take a mean (SD); skewed values take a median (IQR).

Counts of health service use are usually skewed, so the median describes the typical participant better.

Frequency tables

Every percentage has a denominator

51.7%224 of 433 with a score (valid percentage)
46.1%224 of all 486, counting the 53 missing

An online sample without a sampling frame describes its respondents and does not estimate prevalence.

Missing values in summaries

Summaries use the recorded values

Recorded values only

The mean and standard deviation are calculated from respondents with a score, and the number missing is reported.

Missing-code categories

After a missing-code category becomes NA, the empty category is removed.

The mean loneliness score is 5.65 and the median is 6.

Cross-tabulations

Row and column percentages answer different questions

Column: 70.1%

Of respondents classed as lonely, 70.1% were women.

Row: 53.6%

Of women, 53.6% were classed as lonely.

The non-binary row rests on 8 people, so its percentages are unstable and small cells may need to be combined or suppressed.

Building a Table 1

Five decisions

Define the analytic sampleChoose the columnsChoose the rowsChoose the summariesReport missing data

The analytic sample is 433 of 486 respondents; 53 without a loneliness score are excluded and the exclusion is reported.

Reading the table

Describing differences without testing them

Single and not dating

46.3% of the lonely group compared with 32.2% of the group not classed as lonely.

Poor physical health

10.3% compared with 1.1%; mean age was 71.3 years in both groups.

Statistical tests are taught in HSCI 341 Lesson 6 Section 4 and HSCI 410 Lesson 3 Section 1.

Carry forward

From numbers to words

Part 1 produced a documented file and a descriptive table. Part 2 does the same work for interview data.

Section 3 builds a qualitative codebook, which plays the role for transcripts that the data dictionary plays for survey columns.

Learning Objectives for this section

  • Choose an appropriate summary (counts and percentages, or a mean, standard deviation and median) from a variable's level of measurement.
  • Produce a frequency table with percentages, and explain the difference between a percentage of all participants and a valid percentage.
  • Compute and interpret the mean, median and standard deviation, and explain how missing values are handled when these summaries are calculated.
  • Build a cross-tabulation and distinguish row percentages from column percentages.
  • Define an analytic sample and assemble a descriptive Table 1, with the missing values reported.

Introduction

Descriptive statistics summarize who is in a dataset and what their answers look like. They come before any comparison or model, for two reasons. First, they let the team check the data: an age of 340 or a loneliness score of 12 shows up immediately in a frequency table. Second, readers need to know who was studied before they can judge whether the findings apply to anyone else. In most health research articles this description appears as the first table, which is why it is called Table 1. This section works through frequency tables, numeric summaries and cross-tabulations, and then builds a Table 1 from the CSCS 2021 data for respondents aged 65 and older, the age group that the fictional Cedar Valley study recruits.

2.1 Matching the Summary to the Variable

Lesson 7 introduced levels of measurement. They decide which summary is appropriate. Categorical variables, whether nominal (categories with no order, such as community of residence) or ordinal (ordered categories, such as self-rated health from poor to excellent), are summarized with counts and percentages. Numeric variables measured on an interval or ratio scale, such as age or the number of emergency department visits, are summarized with a measure of centre (the mean or median) and a measure of spread (the standard deviation or the interquartile range).

Level of measurementCedar Valley exampleUsual summary
NominalCommunity of residencen (%)
OrdinalSelf-rated physical healthn (%) for each level
BinaryLonely (UCLA score 6 or higher)n (%)
Interval or ratio, roughly symmetricAgeMean (standard deviation)
Ratio, skewedEmergency department visits in a yearMedian (interquartile range)

HSCI 410 Lesson 2 Section 3 extends these summaries to measures of shape, including skew, and to transformations of skewed variables.

2.2 Frequency Tables

A frequency table lists each value of a variable with the number of participants who gave it. Adding percentages makes groups of different sizes comparable. The Cedar Valley team reports that 392 of 1,600 respondents, or 24.5 percent, scored 6 or higher on the UCLA scale. A percentage always has a denominator, and the analyst must say what it is. A valid percentage uses only the participants with a recorded answer as the denominator, while a percentage of all participants includes those with missing values. The difference is small when little is missing and large when much is missing.

The CSCS analysis starts by keeping respondents aged 65 and older, so that the practice data resemble the Cedar Valley population.

Worked example: loneliness among CSCS respondents aged 65 and older

Keeping the 2021 respondents aged 65 or more leaves 486 people. The frequency table below counts each possible UCLA total and adds a count of missing values.

UCLA totalNumber of respondents
3100
451
558
681
740
839
964
Missing (no score)53

One hundred had the lowest possible score of 3, 64 had the highest possible score of 9, and 53 have no score because at least one of the three items is missing. Reading a frequency table like this one is also a data check: every value lies within the permitted range of 3 to 9.

The derived variable LONELY_ucla_loneliness_scale_score_y_n groups the scores at the same cut-off the Cedar Valley team uses: 209 respondents scored 3 to 5 and 224 scored 6 to 9. Leaving out the missing values gives valid percentages with a denominator of 433 (209 plus 224), so 48.3 percent were not classed as lonely and 51.7 percent of older respondents with a score were classed as lonely. Had the 53 respondents without a score been kept in the denominator, the figure would have been 224 of 486, or 46.1 percent, which understates the proportion among those who answered. The CSCS figure is much higher than the 24.5 percent in the fictional Cedar Valley survey. The CSCS recruited its participants online without drawing them at random from a sampling frame, so its percentage describes these respondents and should not be read as the prevalence of loneliness among older Canadians (Lesson 7).

2.3 Describing Numeric Variables

The mean is the sum of the values divided by their number. The median is the middle value when the values are sorted, so half of the participants lie below it and half above. The standard deviation (SD) describes how far values typically lie from the mean: a small SD means that most values are close to the mean, and a large SD means that they are spread out. The interquartile range runs from the first quartile (the value below which a quarter of participants lie) to the third quartile (the value below which three quarters lie), and it contains the middle half of the data.

The mean and the median agree when values are spread symmetrically, and they part ways when a few values are extreme. Suppose nine Cedar Valley participants had 0, 0, 0, 0, 1, 1, 2, 3 and 11 emergency department visits in a year. The total is 18, so the mean is 2.0 visits, while the median is 1 visit. One person with 11 visits pulls the mean upward, and the median better describes the typical participant. Counts of health service use are often skewed in this way, which is why the table in Section 2.1 recommends the median for them.

Worked example: describing the UCLA score

Of the 486 respondents aged 65 and older, 433 have a UCLA score and 53 do not. A mean cannot be calculated from a set of values that includes unknowns, so each summary below uses the 433 recorded scores, and the number missing is reported beside them.

Summary of the UCLA scoreValue
Respondents with a score433
Missing53
Mean5.65
Standard deviation2.09
Minimum3
First quartile4
Median6
Third quartile7
Maximum9

Rounded to one decimal place, as reports usually give them, the mean and SD are 5.7 and 2.1. The median is 6, and the interquartile range runs from 4 to 7.

2.4 Cross-Tabulations

A cross-tabulation (or two-way table) counts participants in every combination of two categorical variables. It is the first step toward asking whether two characteristics go together. Before tabulating, any missing code stored as a category must be converted to NA, or it will appear as a real group in every table.

Worked example: converting a missing-code category

Among the 486 respondents aged 65 and older, relationship status has four categories, one of which is the missing code "Presented but no response". The cleaning step replaces that category with NA and removes the now-empty category. The ten respondents are still counted, now as missing.

Relationship statusBefore conversionAfter conversion
Single and not dating187187
Single and dating2222
In a relationship267267
Presented but no response10Category removed
Missing (NA)010
Worked example: gender by loneliness, with column and row percentages

The cross-tabulation below places gender in the rows and the loneliness group in the columns for the 433 respondents with a UCLA score. Column percentages divide each count by its column total, so each column adds to 100 percent. Row percentages divide each count by its row total, so each row adds to 100 percent.

Gendern, score 3 to 5n, score 6 to 9Column %, 3 to 5Column %, 6 to 9Row %, 3 to 5Row %, 6 to 9
Man726034.426.854.545.5
Woman13615765.170.146.453.6
Non-binary170.53.112.587.5

Column percentages answer the question "what are the people in each loneliness group like?" Among respondents classed as lonely, 70.1 percent were women, 26.8 percent were men and 3.1 percent were non-binary. Among those not classed as lonely, 65.1 percent were women. Column percentages are the form used in a Table 1, where each column describes one group.

Row percentages answer the question "how common is loneliness in each gender group?" Among men, 45.5 percent were classed as lonely; among women, 53.6 percent were. Row percentages suit a comparison of the outcome across groups. The row for non-binary respondents (87.5 percent) rests on 8 people, so a single respondent changes it by 12.5 percentage points.

A cell with 1 person, such as the non-binary respondent who was not classed as lonely, gives an unstable percentage and can make a person identifiable in a small community. Many data custodians set a minimum cell size for published tables, and a study that uses linked administrative data must follow the rules in its data access agreement (Lesson 9). Common solutions are to combine small categories or to suppress the count.

HSCI 410 Lesson 2 Section 3 returns to frequency tables and cross-tabulations and adds the odds ratio to the cross-tabulation.

2.5 Building a Table 1

A Table 1 describes the analytic sample, the participants included in the analysis, usually overall and separately for the groups that the study compares. The STROBE reporting guideline for observational studies asks authors to give the characteristics of study participants and the number with missing data for each variable (von Elm et al., 2007), and Lesson 12 returns to STROBE and to table design. Building the table involves five decisions, which the cards below summarize.

1. Define the analytic sampleClick to explore
2. Choose the columnsClick to explore
3. Choose the rowsClick to explore
4. Choose the summariesClick to explore
5. Report missing dataClick to explore
Worked example: assembling the CSCS Table 1

The first decision defines the analytic sample. Keeping the respondents whose loneliness group is recorded leaves 433 of the 486 respondents aged 65 and older: 209 with scores of 3 to 5 and 224 with scores of 6 to 9. Age is roughly symmetric, so it is summarized as a mean (SD): 71.3 years (5.3) overall, 71.3 (4.9) in the group with scores of 3 to 5 and 71.3 (5.8) in the group with scores of 6 to 9.

Each categorical row comes from a cross-tabulation of the characteristic by loneliness group, with an overall column that adds the two groups together, and the counts are converted to column percentages. For relationship status, 7 respondents have no recorded status (1 in the group with scores of 3 to 5 and 6 in the group with scores of 6 to 9). They are reported on a separate row, and the percentages are calculated without them, so the overall denominator is 426. Education and self-rated physical health, after its missing-code category is converted to NA, are handled in the same way. The table below brings the rows together.

CharacteristicOverall (n = 433)UCLA score 3 to 5 (n = 209)UCLA score 6 to 9 (n = 224)
Age in years, mean (SD)71.3 (5.3)71.3 (4.9)71.3 (5.8)
Gender, n (%)
  Woman293 (67.7)136 (65.1)157 (70.1)
  Man132 (30.5)72 (34.4)60 (26.8)
  Non-binary8 (1.8)1 (0.5)7 (3.1)
Relationship status, n (%)
  In a relationship238 (55.9)136 (65.4)102 (46.8)
  Single and dating20 (4.7)5 (2.4)15 (6.9)
  Single and not dating168 (39.4)67 (32.2)101 (46.3)
  Missing716
Education, n (%)
  Bachelor's degree or higher149 (34.4)81 (38.8)68 (30.4)
  Less than a bachelor's degree284 (65.6)128 (61.2)156 (69.6)
Self-rated physical health, n (%)
  Excellent38 (9.7)26 (14.0)12 (5.9)
  Very good110 (28.2)60 (32.3)50 (24.5)
  Good132 (33.8)66 (35.5)66 (32.4)
  Fair87 (22.3)32 (17.2)55 (27.0)
  Poor23 (5.9)2 (1.1)21 (10.3)
  Missing432320

Table 1. Characteristics of CSCS 2021 respondents aged 65 and older, overall and by score on the three-item UCLA Loneliness Scale. Percentages are column percentages among respondents with a recorded value. Of 486 respondents aged 65 and older, 53 without a UCLA score were excluded. Loneliness scores range from 3 to 9; a score of 6 or higher was classed as lonely.

Reading a Table 1

The table describes the analytic sample. Mean age is the same in both groups. Respondents classed as lonely were more often single and not dating (46.3 percent compared with 32.2 percent) and more often rated their physical health as poor (10.3 percent compared with 1.1 percent) or fair (27.0 percent compared with 17.2 percent). These are descriptions of differences in this sample. The table contains no statistical tests, and on its own it cannot show whether the differences would hold in other samples or whether loneliness and health influence one another. Statistical tests are taught in HSCI 341 Lesson 6 Section 4 and HSCI 410 Lesson 3 Section 1, regression models in HSCI 410 Lesson 3, and the causal questions require the reasoning of HSCI 341. The table also shows the analyst's choices openly: the row order runs from the most to the least favourable health rating, the overall column comes first, and the footnote gives the denominator and the exclusions.

Try it: describe two rows of Table 1

Using Table 1, write two sentences that describe the education rows. Then write two sentences that describe the self-rated physical health rows, and state the denominator of the percentages in each column, given that 43 respondents (23 with scores of 3 to 5 and 20 with scores of 6 to 9) have no recorded rating.

Reflection

The table below describes 433 respondents aged 65 and older in the Canadian Social Connection Survey (CSCS) 2021, a survey whose participants were recruited online without a sampling frame. Respondents are grouped by score on the three-item UCLA Loneliness Scale (3 to 5 not lonely; 6 to 9 lonely). Relationship status, n (%), by group: in a relationship, overall 238 (55.9), not lonely 136 (65.4), lonely 102 (46.8); single and dating, 20 (4.7), 5 (2.4), 15 (6.9); single and not dating, 168 (39.4), 67 (32.2), 101 (46.3); missing 7, 1, 6. Gender: non-binary, overall 8 (1.8), not lonely 1 (0.5), lonely 7 (3.1). Write a short results paragraph describing the relationship status rows. State the denominator used for the percentages, comment on the non-binary row, and say what the table cannot tell a reader.

Model answer

Among the 433 respondents aged 65 and older with a UCLA score, 55.9 percent were in a relationship, 39.4 percent were single and not dating, and 4.7 percent were single and dating. Respondents classed as lonely were less often in a relationship than those not classed as lonely (46.8 percent compared with 65.4 percent) and more often single and not dating (46.3 percent compared with 32.2 percent). The percentages are column percentages among respondents with a recorded relationship status, so the overall denominator is 426 (433 minus 7 missing), the lonely group denominator is 218 and the not-lonely denominator is 208; the 7 missing values are reported separately.

The non-binary row rests on 8 people, with a single person in the not-lonely column, so its percentages are unstable and the cell could identify someone. In a report I would either combine it with another category, if that did not hide something important to the study, or suppress the small count in line with the data custodian's rules.

The table describes this sample only. It contains no statistical tests, it cannot show whether relationship status affects loneliness or the reverse, and because the CSCS was recruited online without a sampling frame, its percentages should not be read as the prevalence of these characteristics among older Canadians.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: Which summary best suits the number of emergency department visits in a year, when most people have 0 or 1 visit and a few have many?

Counts of service use are usually skewed, and a few high values pull the mean upward. The median and interquartile range describe the typical participant and the middle half of the data. The mean and SD (option a) suit roughly symmetric values such as age.

Question 2: Of 486 CSCS respondents aged 65 and older, 224 were classed as lonely, 209 were not, and 53 had no UCLA score. What is the valid percentage classed as lonely?

A valid percentage uses only respondents with a recorded answer (209 + 224 = 433) as the denominator, giving 51.7 percent. Option b keeps the 53 missing respondents in the denominator and understates the proportion among those who answered.

Question 3: In a cross-tabulation with gender in the rows and loneliness group in the columns, 70.1 percent of respondents classed as lonely were women, and 53.6 percent of women were classed as lonely. Which statement is correct?

A column percentage describes the make-up of one column (the lonely group), so 70.1 percent is a column percentage. A row percentage describes one row (women), so 53.6 percent is a row percentage.
Section 3 of 5

Building a First Codebook and Coding a Transcript

⏱ Estimated reading time: 35 minutes
Section 3 of 5

Building a First Codebook and Coding a Transcript

Codes, memos and a first coding pass on a Cedar Valley interview.

Key terms

Codes, segments and memos

Code

A short label recording what a passage is about.

Coded segment

The passage that carries the code; it can carry several codes.

Memo

A dated note of an idea, question or connection.

First-cycle coding labels segments; categories and themes are built later (Saldaña, 2021).

Where codes come from

Deductive and inductive codes

Deductive

Written before coding from the question, the interview guide, the causal web or theory.

Inductive

Created during coding when a passage says something no code captures.

I just couldn’t walk into a room full of strangers on my own.P14, “Ruth”, invented excerpt
Starter codebook

Seven deductive codes, each with rules

MOVINGFAMILYCOMMUNITYTRANSPORTHEALTHTECHNOLOGYLONELINESS

Each entry gives a name, a short and a full definition, when to use the code, when not to use it, and an example (MacQueen et al., 1998).

The process

Coding a transcript in six steps

1. Read it all

Read without coding and write a first memo.

2. Segment

Mark passages that each express one idea.

3. Apply codes

Check each definition and its exclusion rule.

4. Add codes

Propose inductive codes where nothing fits.

5. Write memos

Record connections and questions as you go.

6. Revise

Save a new, dated codebook version and recheck.

A first pass on Ruth’s interview

Deductive codes plus four proposed inductive codes

I wouldn’t have gone if she hadn’t called.Coded COMMUNITY and Personal invitation
Chance encountersConcealing lonelinessPersonal invitationSmall-town visibility
Coding with others

Agreeing on the codebook, comparing the coding

Team coding

Two coders code the same transcripts independently, compare, and rewrite unclear definitions.

Tools

Paper, comments, a spreadsheet, or software such as Taguette, NVivo or ATLAS.ti.

Transcripts are de-identified before coding, and cloud tools are used only if consent and the data management plan allow it.

Carry forward

From codes to an analytic approach

The same coded excerpt can lead in different directions.

Section 4 shows what seven analytic approaches would do next with Ruth’s codes, and which kinds of data each approach suits.

Learning Objectives for this section

  • Define a qualitative code, a coded segment and a memo, and distinguish first-cycle codes from the categories and themes built from them later.
  • Distinguish deductive codes, drawn from the research question, interview guide or theory, from inductive codes that arise from the data, including in vivo codes.
  • Build a starter codebook in which each code has a name, a definition, rules for when to use and not use it, and an example.
  • Code a short interview transcript step by step, applying deductive codes, adding inductive codes and writing memos.
  • Describe how a team revises a codebook, compares coding and chooses tools for coding.

Introduction

Part 2 of this lesson turns from numbers to words. The qualitative strand of the fictional Cedar Valley Social Connection Study includes 24 semi-structured interviews with older adults living alone and four focus groups: two with older adults, one with family caregivers, and one with clinic staff and community connectors. Lesson 10 prepared a clean verbatim transcript of one interview. A single interview of 54 minutes can produce fifteen or more pages of text, and 24 of them produce several hundred. Coding is the first systematic step in making that volume of text manageable without losing what participants said. This section builds a starter codebook for the Cedar Valley interviews and uses it to code the excerpt from Lesson 10. It gives a working method for a first coding pass, and HSCI 841 Lessons 5 to 12 develop qualitative analysis in depth.

3.1 What a Code Is

A code is a short label, usually a word or phrase, that a researcher attaches to a passage of text to record what the passage is about or what it means for the research question (Saldaña, 2021). The passage is a coded segment; it may be a phrase, a sentence or a few sentences that express one idea. One segment can carry more than one code, and the same code is applied to every segment, in every transcript, that fits its definition. Coding lets the researcher gather every segment with the code "transportation" from all 24 interviews and read them side by side.

Saldaña (2021) distinguishes first-cycle coding, the initial labelling of segments, from second-cycle coding, in which codes are grouped, compared and organized into larger categories and themes. A theme is a pattern of meaning that runs across many segments and participants and says something about the research question. This lesson stops at first-cycle coding. Section 4 shows how different analytic approaches carry codes forward. HSCI 841 Lesson 5 Section 1.4 teaches the move from codes to themes, and HSCI 841 Lesson 6 turns codes into conceptual models.

Coding goes together with writing memos. A memo is a short dated note in which the researcher records an idea, a question, a possible connection between codes or a reaction to the data. Analytic memos written during coding continue the reflexive memos of Lesson 10 and often hold the first drafts of later themes.

Segments passages of text First-cycle codes labels on segments Categories groups of codes Themes patterns of meaning HSCI 207 Lesson 11 HSCI 841 Lessons 5 to 12 Memos record ideas at every step
Coding moves from passages of text to labels, then to groups of labels and to themes. This lesson covers the first two steps.

3.2 Deductive and Inductive Codes

Codes come from two directions. Deductive codes (also called a priori codes) are written before coding begins. They come from the research question, the topics of the interview guide, the causal web drawn in Lesson 3, or an existing theory. For example, Weiss (1973) distinguished emotional loneliness, the absence of a close attachment figure, from social loneliness, the absence of a wider network of friends and acquaintances, and a team could write a code for each. Inductive codes are created during coding, when a passage expresses something important that no existing code captures. One useful kind is the in vivo code, which uses the participant's own words as the label, such as "a room full of strangers". Most applied health studies combine the two, starting with a short deductive list and adding inductive codes as the data require. Fereday and Muir-Cochrane (2006) describe this hybrid approach and the record-keeping it needs.

Deductive codesInductive codes
SourceResearch question, interview guide, theory, earlier studiesThe data themselves
When writtenBefore coding beginsDuring coding
StrengthThey keep coding tied to the question and make team coding consistent from the start.They capture ideas the team did not anticipate, in participants' terms.
RiskCoders may force text into codes that do not fit and overlook what is new.The list can grow very long, with overlapping codes, unless it is reviewed.
Cedar Valley exampleTransportation and mobility, taken from the causal webConcealing loneliness, added after reading interview 14
Descriptive codesClick to explore
In vivo codesClick to explore
Process codesClick to explore
Emotion codesClick to explore

3.3 Building a Starter Codebook

A qualitative codebook is the list of codes with the rules for applying them. MacQueen and colleagues (1998), writing about team-based analysis, recommended that each entry give a short code name, a brief definition, a full definition, guidance on when to use the code, guidance on when not to use it, and an example. The "when not to use" rule is the part beginners most often leave out, and it is the part that keeps two similar codes apart. A starter codebook holds the deductive codes written before coding begins; it is version 1.0, and it will change.

The Cedar Valley team wrote seven deductive codes from its SPIDER question (how older adults living alone experience social connection after a move to a smaller town), the interview guide and the causal web.

CodeDefinitionUse whenDo not use when
MOVINGChanges in social life that the participant links to moving homeThe participant compares life before and after a move.A change is linked to something else, such as bereavement.
FAMILYContact with relatives and its frequency, form and qualityCalls, visits, help or tension with family are described.The person is a friend or neighbour; use COMMUNITY.
COMMUNITYContact with friends, neighbours, groups, programs and community placesLibraries, churches, centres, groups or neighbours are described.Contact is only with paid health care providers; use HEALTH.
TRANSPORTThe ability to travel, including driving, transit and ridesGetting places, or being unable to, is described.Travel is mentioned only as a setting with no effect on contact.
HEALTHHealth conditions, function and health care that affect social contactIllness, surgery, mobility or appointments shape contact.Health is mentioned with no link to social life.
TECHNOLOGYPhones, video calls and the internet used to stay in touchA device or connection helps or hinders contact.A phone call is mentioned with no comment on the medium; code the relationship only.
LONELINESSThe participant's own descriptions of feeling lonely, isolated or aloneA feeling of disconnection is described or expressed.The passage describes the amount of contact without a feeling.

The table shows four of the six columns that MacQueen and colleagues recommended. The full codebook also has a brief definition for quick reference and an example segment for each code, which the team adds as soon as it finds a clear example in the data. HSCI 841 Lesson 5 Section 2.4 extends the six columns to seven elements by splitting the example into positive and negative examples.

3.4 Coding a Transcript Step by Step

Step 1: Read the whole transcript before codingv

Read the transcript once from start to finish without coding, ideally while listening to the recording. Write a memo of a few sentences on your first impressions: what seemed most important to the participant, what surprised you, and what you want to look for on the second reading.

Step 2: Divide the text into segmentsv

On the second reading, mark each passage that expresses one idea relevant to the research question. Keep enough text that the segment makes sense when it is read alone, away from the transcript. Passages that do not bear on the question, such as small talk at the start, can be left uncoded.

Step 3: Apply the deductive codesv

For each segment, check the codebook and apply every code whose definition and "use when" rule fit. Read the "do not use when" rule each time, especially for similar codes.

Step 4: Add inductive codes where nothing fitsv

When a segment expresses something important that no code captures, write a new code, with a definition, in a list of proposed codes. Do not stretch an existing code to cover it.

Step 5: Write memos as you gov

Record connections between codes, questions to ask in later interviews, and your own reactions. Date each memo and note the segment that prompted it.

Step 6: Revise the codebook and record the changev

After the first two or three transcripts, the team reviews the proposed codes, adds the useful ones with full definitions, merges or splits codes that overlap, and saves the result as a new version (1.1, 1.2 and so on) with the date and the reason for each change. Earlier transcripts are then rechecked against the revised codes.

Case: interview CV-INT-14 with "Ruth"

The excerpt below is invented for teaching; no real person was interviewed. It comes from minute 18 to minute 21 of a 54-minute interview with P14, given the pseudonym "Ruth", a woman aged 78 who lives alone and moved from Cedar City to Kestrel Lake about two years before the interview. Lesson 10 transcribed it in clean verbatim form. Read it in full before looking at the coding that follows it.

I: You mentioned that you moved to Kestrel Lake about two years ago. Can you tell me what the first few months were like?

P14: It was harder than I expected. In Cedar City I knew everybody on my street, and I never had to plan to see people. I'd go to the bank or the pharmacy and I'd run into someone I knew. Here, nobody knew me from Adam. I'd walk to the store and come home and realize I hadn't spoken to anyone except the cashier.

I: What was that like for you?

P14: [pause] Quiet. Very quiet. My daughter, [daughter's name], phones every Sunday, and that helps, but it isn't the same as somebody dropping by. I didn't tell her how I felt, because she was the one who found me this place, and she worries enough already.

I: You said you didn't tell her. Can you say a bit more about that?

P14: I didn't want to be a burden. When you get to my age, people start deciding things for you. If I'd said I was lonely, she'd have had me in a care home by Christmas. [laughs] So I said to her, "I'm fine, Mom's fine."

I: How did things change, if they did?

P14: The first winter was the worst. I'd stopped driving after my cataract surgery, and there are no buses out here, so when the roads were bad I could go a week, a whole week, without leaving the house. What turned it around was the library. [Librarian's name], the lady who runs it, started a coffee morning on Tuesdays, and she phoned me herself to ask me to come. I wouldn't have gone if she hadn't called. Now I go every week, and a fellow from the coffee group drives me to my eye appointments in Cedar City.

I: It sounds like that phone call mattered.

P14: It did. It really did. I think people assume older folks will just show up if you put a notice on the board. I'd seen the notice. I just couldn't walk into a room full of strangers on my own. Not at my age.

I: Is there anything that still makes it hard to stay in touch with people?

P14: The internet out here is poor, so the video calls with my grandson freeze up. He thinks I'm pulling faces. And in a small town everybody knows your business. I'm careful what I say at coffee, because it goes around.

A first coding pass

The table shows one coder's first pass. Deductive codes are in capitals; proposed inductive codes are in italics, with in vivo codes in quotation marks.

SegmentCodesMemo
"In Cedar City I knew everybody on my street ... I'd run into someone I knew."MOVING; Chance encountersContact in the old neighbourhood happened without planning. The move removed it.
"I'd walk to the store and come home and realize I hadn't spoken to anyone except the cashier."MOVING; Chance encountersThe same errands no longer produce contact. Ask other participants about everyday errands.
"Quiet. Very quiet."LONELINESS; "Very quiet"She answers a feeling question with a description of silence after a pause.
"My daughter ... phones every Sunday, and that helps, but it isn't the same as somebody dropping by."FAMILYRegular calls are valued but are described as different from in-person visits.
"I didn't want to be a burden ... she'd have had me in a care home by Christmas."FAMILY; LONELINESS; Concealing loneliness; "a burden"Hiding loneliness to protect independence. Check whether other participants describe this.
"I'd stopped driving after my cataract surgery, and there are no buses out here ... a whole week without leaving the house."TRANSPORT; HEALTHHealth, rural transit and winter weather combine. This links to the transportation node in the causal web.
"She phoned me herself to ask me to come. I wouldn't have gone if she hadn't called."COMMUNITY; Personal invitationThe program existed, and the personal call is what made her attend.
"A fellow from the coffee group drives me to my eye appointments in Cedar City."COMMUNITY; TRANSPORT; HEALTHA social contact became practical help with health care access.
"I just couldn't walk into a room full of strangers on my own."Personal invitation; "a room full of strangers"A notice was not enough. Possible implication for the health authority's programs.
"The internet out here is poor, so the video calls with my grandson freeze up."TECHNOLOGY; FAMILYRural internet quality limits contact with family.
"In a small town everybody knows your business. I'm careful what I say at coffee."COMMUNITY; Small-town visibilityA new connection also brings exposure. This may limit what she shares.

The pass proposes four inductive codes: Chance encounters, Concealing loneliness, Personal invitation and Small-town visibility. Each now needs a definition and "use when" and "do not use when" rules before it enters version 1.1 of the codebook. For example, Personal invitation could be defined as "a direct, individual request from a person to join an activity, described as affecting whether the participant attended", to be used when the participant links attendance to being asked, and not used for general publicity such as notices. One interview cannot show whether these codes recur; the memos flag them so the coders can watch for them in the remaining 23 interviews.

3.5 Coding as a Team, and Tools for Coding

When more than one person codes, the team agrees on the codebook first. Two coders then code the same two or three transcripts independently, compare their coding segment by segment, and discuss each disagreement. Disagreements usually reveal an unclear definition, which the team rewrites. Some teams then calculate a measure of intercoder agreement such as Cohen's kappa, while others, particularly those using reflexive approaches, treat discussion itself as the check. HSCI 841 Lesson 5 covers agreement measures and when they suit an approach.

Coding can be done with highlighters on paper, comments in a word processor, or a spreadsheet with one row per segment, as in the table above. Qualitative data analysis software stores the transcripts, codes and memos together and retrieves every segment with a given code in one step. Taguette is a free, open-source option that HSCI 841 uses; NVivo, ATLAS.ti, MAXQDA and Dedoose are commercial programs. Whatever the tool, transcripts must be de-identified before coding, and a cloud-based program may only be used if the consent form and the data management plan permit it (Lessons 5 and 10).

Try it: code the excerpt yourself

Copy the excerpt into a document or spreadsheet. Using the seven deductive codes and the four proposed inductive codes, code it without looking at the table above, then compare. Write a full codebook entry (name, definition, use when, do not use when, example) for Concealing loneliness, and propose one further inductive code that you think the first pass missed, with a one-sentence memo explaining why.

Reflection

The following excerpt is invented for teaching. P09, a fictional man aged 81 who lives alone in Riverside, says: "Since my wife died I don't really go to the Legion anymore. It was her that kept the calendar. I've got the dog, and I talk to the neighbour over the fence most mornings, but evenings are long. My son wants me to get one of those tablets for video calls, but I can't see the screen well enough." The starter codebook has seven codes: MOVING (changes in social life linked to moving home; do not use when the change is linked to something else, such as bereavement); FAMILY (contact with relatives); COMMUNITY (contact with friends, neighbours, groups and community places); TRANSPORT (ability to travel); HEALTH (health conditions or function that affect social contact); TECHNOLOGY (devices or connections that help or hinder contact); and LONELINESS (the participant's own descriptions of feeling lonely, isolated or alone). Divide the excerpt into segments, apply the codes, propose at least one inductive code with a definition and a "use when" rule, and write a two-sentence memo.

Model answer

Segment 1, "Since my wife died I don't really go to the Legion anymore. It was her that kept the calendar": COMMUNITY. MOVING does not apply, because the change is linked to bereavement. I propose an inductive code, Loss of a social organizer: a partner or other person who arranged social activities has died or left, and the participant describes reduced participation as a result; use when the participant links less contact to losing the person who planned it.

Segment 2, "I've got the dog, and I talk to the neighbour over the fence most mornings": COMMUNITY for the neighbour, with a second proposed code, Animal companionship, for the dog.

Segment 3, "but evenings are long": LONELINESS, with an in vivo code, "evenings are long".

Segment 4, "My son wants me to get one of those tablets for video calls, but I can't see the screen well enough": FAMILY, TECHNOLOGY and HEALTH, because a vision problem limits a device that could support contact with his son.

Memo: P09 describes contact that is regular in the morning and absent in the evening, which suggests looking at the timing of loneliness across interviews. His account of his wife as the person who kept the calendar is similar to Ruth's reliance on the librarian's invitation, so both may point to the role of a person who initiates contact.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: A coder labels Ruth's phrase "a room full of strangers" with exactly those words. What kind of code is this?

An in vivo code uses the participant's own words as the label. It is inductive, because it comes from the data, and it is a first-cycle code; categories are built later by grouping codes.

Question 2: Which code in the Cedar Valley coding is deductive?

Deductive codes are written before coding from the question, the guide, the causal web or theory. The three italicized codes were proposed during coding and are therefore inductive.

Question 3: Why does each codebook entry include a "do not use when" rule?

MacQueen and colleagues (1998) included exclusion guidance because it separates codes that could overlap, such as FAMILY and COMMUNITY, which makes coding consistent across coders and transcripts.

Question 4: After coding three transcripts, a coder finds an important idea that no existing code captures. What should happen next?

Codebooks are expected to change. The coder proposes an inductive code with a definition, the team reviews it, and the revised codebook is saved as a new version; earlier transcripts are rechecked. Stretching an existing code (option b) hides what is new.
Section 4 of 5

Qualitative Analytic Approaches and the Data They Suit

⏱ Estimated reading time: 30 minutes
Section 4 of 5

Qualitative Analytic Approaches

Seven approaches, the data they suit, and how to choose among them.

Patterns, categories and comparison

Four approaches that work across many participants

Thematic analysis

Patterns of meaning across a dataset, in six phases (Braun & Clarke, 2006).

Qualitative content analysis

Systematic categories, sometimes counted (Hsieh & Shannon, 2005).

Framework analysis

A matrix of cases by codes for comparison (Ritchie & Spencer, 1994).

Rapid qualitative analysis

Structured summaries and matrices for timely findings.

Process, experience and story

Three approaches that go deep

Grounded theory

A theory of a process, built by constant comparison and theoretical sampling.

Interpretative phenomenological analysis

How a few people make sense of a major experience, case by case.

Narrative analysis

Whole stories, their structure and their telling (Riessman, 2008).

Matching approach to data

What each kind of data suits

In-depth interviews

Suit every approach, and are needed for the phenomenological and narrative approaches.

Focus groups and key informants

Suit thematic, framework and rapid analysis.

Open-ended survey answers

Suit content analysis and thematic analysis.

Choosing

Four questions, and the Cedar Valley decision

Ask

What does the question need? What data are there? When are findings due? Who will analyze?

Cedar Valley

Framework analysis for the 24 interviews, and rapid analysis of the staff focus group for an interim brief.

Before the final assessment

Where this lesson leads

HSCI 841

Lesson 5 themes and codebooks; Lesson 6 conceptual models; Lesson 7 framework, rapid and grounded theory analysis; Lesson 8 content analysis; Lesson 9 narrative analysis.

What HSCI 841 assumes

Prepared transcripts, a starter codebook and a completed first coding pass, as in this lesson.

Complete the reflection and the knowledge check below, then go on to the final assessment.

Learning Objectives for this section

  • Describe the aim, typical data and main steps of thematic analysis, qualitative content analysis, framework analysis, rapid qualitative analysis, grounded theory, interpretative phenomenological analysis and narrative analysis.
  • Match each approach to the data collection methods it suits, including in-depth interviews, key informant interviews, focus groups, open-ended survey responses and documents.
  • Use the research question, the data, the timeline and the team to choose an approach for a study.
  • Explain how the first-cycle coding of Section 3 feeds into each approach, and locate where HSCI 841 teaches each one.

Introduction

Coding is common to most qualitative work, but what happens to the codes afterwards depends on the analytic approach. An approach is a recognized way of moving from data to findings, with its own aims, procedures and standards of quality. Thematic analysis looks for patterns across participants; interpretative phenomenological analysis stays close to each person's experience; narrative analysis keeps stories whole. Ideally the approach is chosen when the study is designed and written into the protocol (Lesson 6), because it affects how many participants are needed, how interviews are run and how transcripts are prepared. This section introduces seven approaches used in health research, maps them to the data collection methods of Lesson 10, and shows how the fictional Cedar Valley team chose among them. It is an orientation. HSCI 841 Lesson 1 Section 4 relates the quality criteria in Section 4.3 to systematic, transparent and replicable analysis, and Section 4.4 lists the HSCI 841 lesson for each approach.

4.1 Seven Analytic Approaches

Each card below gives the aim of the approach, the data it usually works with, its main steps and a foundational source.

Thematic analysisClick to explore
Qualitative content analysisClick to explore
Framework analysisClick to explore
Rapid qualitative analysisClick to explore
Grounded theoryClick to explore
Interpretative phenomenological analysisClick to explore
Narrative analysisClick to explore

The approaches share some steps. Each begins with close reading, and most involve coding. They differ in what they produce: a set of themes, a set of categories with frequencies, a comparison matrix, a rapid summary for decision-makers, an explanatory theory, a detailed account of a few people's experience, or an analysis of stories. They also differ in how they treat a codebook. Framework analysis, rapid analysis, directed content analysis and codebook forms of thematic analysis use a structured codebook like the one built in Section 3. Reflexive thematic analysis, grounded theory and interpretative phenomenological analysis develop codes more freely and treat a fixed codebook with caution.

What each approach would do with Ruth's interview

The coded excerpt from Section 3 shows how the same first-cycle codes lead in different directions. The table describes what an analyst using each approach would do next with interview CV-INT-14.

ApproachNext step with the coded excerpt
Thematic analysisCompare Personal invitation and COMMUNITY segments across all 24 interviews to see whether a theme about being drawn in by someone known is developing.
Qualitative content analysisCount how many participants describe a transportation barrier, and report the categories of barrier with their frequencies.
Framework analysisWrite a short summary of Ruth's coded data under each code in the matrix row for P14, among the Kestrel Lake participants.
Rapid qualitative analysisFill in a one-page template with Ruth's main points under each topic of the interview guide, ready for the interim brief.
Grounded theoryCompare Concealing loneliness with similar incidents in other interviews, and recruit further participants who moved recently to test an emerging idea about how connection is rebuilt after a move.
Interpretative phenomenological analysisRead the whole interview closely to interpret how Ruth makes sense of hiding her loneliness from her daughter, before moving to other cases.
Narrative analysisKeep the account whole and examine its structure: the move as a setting, the first winter as the complication, the librarian's call as the turning point and the weekly coffee morning as the resolution.

4.2 Matching Approaches to Data Collection Methods

Lesson 10 introduced in-depth interviews, key informant interviews and focus groups. Two other common sources of qualitative data are open-ended survey questions and documents such as policies, meeting minutes or program records. The table maps the seven approaches to these sources. "Strong fit" means that the approach was designed for that kind of data or is routinely used with it; "possible" means that it is used with that data with some adaptation; a blank cell means that the combination is uncommon.

ApproachIn-depth interviewsKey informant interviewsFocus groupsOpen-ended survey responsesDocuments
Thematic analysisStrong fitStrong fitStrong fitStrong fitPossible
Qualitative content analysisPossiblePossiblePossibleStrong fitStrong fit
Framework analysisStrong fitStrong fitStrong fitPossiblePossible
Rapid qualitative analysisPossibleStrong fitStrong fitPossible
Grounded theoryStrong fitPossiblePossiblePossible
Interpretative phenomenological analysisStrong fit
Narrative analysisStrong fitPossible

Two patterns stand out. Interpretative phenomenological analysis and narrative analysis need long, personal accounts from individuals, so they depend on one-to-one interviews (or, for narrative analysis, diaries and written life stories), and focus groups suit them poorly because the group conversation interrupts each person's account. Qualitative content analysis and thematic analysis can handle short answers from many people, which makes them the usual choices for open-ended survey questions. The table is a guide drawn from common practice; methodologists debate some cells, and HSCI 841 examines those debates.

4.3 Choosing an Approach

Four questions narrow the choice. What does the research question ask for: patterns, counts, comparisons, an explanation of a process, the depth of individual experience, or stories? What data will there be, and how much? When are findings needed, and by whom? Who will do the analysis, and how many people? The diagram pairs typical answers with the approaches that suit them.

If the study needs... consider patterns of meaning across many participants Thematic analysis categories and counts from a lot of short text Content analysis comparison across cases or groups by a team Framework analysis findings for a decision within weeks Rapid qualitative analysis an explanation of how a process unfolds Grounded theory depth on how a few people make sense of an experience Interpretative phenomenological analysis how people tell the story of a change in their lives Narrative analysis A study can combine approaches for different strands or audiences.
A simplified decision map. Real choices also weigh the team's training, the number of participants and the expectations of the intended audience.

The Cedar Valley decision

The Cedar Valley team considered several options for its 24 interviews and four focus groups. The tabs summarize its reasoning, which the team recorded in a memo and then in the protocol.

For the 24 interviews, the team chose framework analysis. The SPIDER question asks about experiences across participants, the health authority wants to know whether experiences differ between Cedar City and the smaller communities, and four people will share the coding. After first-cycle coding with the codebook from Section 3, each interview will be summarized in a matrix with one row per participant and one column per code, and the rows will be grouped by community so that patterns such as the role of transportation in Kestrel Lake can be compared with those in Cedar City.

The health authority asked for early findings within six weeks of the last focus group to inform planning for its community programs. For the focus group with clinic staff and community connectors, the team chose rapid qualitative analysis: the moderator and note-taker will complete a summary template organized by the topics of the focus group guide within two days of the session, and the summaries will feed a short brief. The full transcript will later be added to the framework matrix.

The team discussed narrative analysis of participants' accounts of moving, since several interviews, including Ruth's, tell the move as a story with a hard first winter and a turning point. It also discussed interpretative phenomenological analysis for a small group of recently widowed participants. It recorded both as possible follow-up studies, because each would need longer, less structured interviews than its guide provides and analytic training that the current team does not have.

Quality in qualitative analysis

Whatever the approach, readers judge qualitative findings by how carefully they were produced. Lincoln and Guba (1985) proposed criteria of credibility, transferability, dependability and confirmability, which researchers address with practices such as an audit trail of codebook versions and memos, reflexive memos about the researcher's own position, and a clear description of the setting. The reporting guidelines COREQ and SRQR, which Lesson 12 introduces, list what a report should include. HSCI 841 Lesson 1 Section 4 relates these criteria to systematic, transparent and replicable analysis.

4.4 Where HSCI 841 Picks Up

HSCI 841, the series' graduate course in qualitative research methods and analysis, assumes the skills of this lesson: preparing transcripts, building a codebook and completing a first coding pass. The table shows where each approach is developed.

Approach or skillHSCI 841 lesson
Finding themes, the phases of thematic analysis, inductive and deductive coding, codebooks and intercoder agreementLesson 5, Themes and Codebooks, Sections 1 to 3
Framework analysis and matricesLesson 7, Comparing Variables and Grounded Theory, Section 1
Rapid qualitative analysisLesson 7, Comparing Variables and Grounded Theory, Section 1.7
Conceptual modelsLesson 6, Analysis Frameworks and Conceptual Models, Section 3
Grounded theory and constant comparisonLesson 7, Comparing Variables and Grounded Theory
Qualitative and quantitative content analysisLesson 8, Content Analysis
Narrative analysisLesson 9, Schema and Narrative Analysis
Interviewing for interpretative phenomenological analysis; phenomenology and interpretative phenomenological analysis as a methodologyLesson 4, Qualitative Data Collection, Section 3.1 (interviewing); Lesson 2, Research Questions, Theory and Literature, Section 2.7 (methodology)
Discourse analysis, analytic induction and computational text analysisLessons 10 to 12

Reflection

Primary care clinics in a fictional health region want to understand how clinic staff decide whether to refer an older patient to a community connector. The research team will conduct 12 key informant interviews of about 30 minutes with physicians, nurses and receptionists. The clinics need findings for a planning meeting five weeks after the last interview, and two analysts will share the work. The seven approaches available are thematic analysis, qualitative content analysis, framework analysis, rapid qualitative analysis, grounded theory, interpretative phenomenological analysis and narrative analysis. Choose an approach, justify it using four considerations (what the question asks for, the data, the timeline and the team), describe the first two analytic steps, and name one approach you would not use, with the reason.

Model answer

I would use rapid qualitative analysis. The question asks for a practical description of how staff make referral decisions, which suits a structured summary of what each informant reports. The data are 12 short key informant interviews organized around a guide, so a summary template with one heading per guide topic (for example, how patients are identified, what information staff use, barriers to referral and suggestions) fits them well. The timeline of five weeks leaves little room for full transcription and line-by-line coding of every interview, and rapid analysis was developed for decisions on this kind of schedule. Two analysts can share the work if both complete templates the same way, which is checked by having both summarize the first two interviews and compare.

Step one is to finalize the template and complete it for each interview within a few days, from the recording and notes. Step two is to transfer the summaries into a matrix with one row per informant and one column per topic, grouped by role, so that patterns, such as differences between physicians and receptionists, can be read across rows.

I would not use interpretative phenomenological analysis, because it is designed for in-depth interviews about how a few people make sense of a significant personal experience, and these short professional interviews about a work process do not provide that kind of account. Framework analysis would be a reasonable alternative if the timeline were longer.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: A team must give a health authority findings from eight key informant interviews within four weeks. Which approach suits this best?

Rapid qualitative analysis summarizes each interview in a structured template and combines the summaries in a matrix, which suits key informant interviews and short timelines. Grounded theory needs repeated rounds of data collection and analysis.

Question 2: Which approach keeps each participant's story whole and examines its structure, such as the setting, complication, turning point and resolution?

Narrative analysis treats the story as the unit of analysis (Riessman, 2008). The other approaches divide accounts into coded segments and compare them across participants.

Question 3: Why do focus groups fit poorly with interpretative phenomenological analysis?

Interpretative phenomenological analysis examines in detail how particular individuals make sense of an experience, case by case, so it relies on one-to-one in-depth interviews. Word counts (option c) belong to summative content analysis.

Question 4: Which approach places coded data in a matrix with one row per participant and one column per code, so that cases and groups can be compared?

Framework analysis (Ritchie & Spencer, 1994; Gale et al., 2013) summarizes each participant's coded data in a case-by-code matrix. The Cedar Valley team chose it to compare communities.
Section 5 of 5

Final Assessment

⏱ Estimated time: 25 minutes

Bringing It All Together

This lesson took the first analytic steps with both kinds of data that the fictional Cedar Valley Social Connection Study collects. Part 1 prepared survey data for analysis: a rectangular layout, a read-only raw file, a data dictionary and deliberate handling of missing codes. It then used the 2021 wave of the Canadian Social Connection Survey to work through frequency tables, descriptive statistics, cross-tabulations and a Table 1 for respondents aged 65 and older.

Part 2 prepared interview data for analysis. It defined codes, segments and memos, distinguished deductive from inductive codes, built a starter codebook with rules for when to use and not use each code, and coded an excerpt from the interview with Ruth. It then described seven analytic approaches, matched them to the data collection methods they suit, and showed how the Cedar Valley team chose framework analysis for its interviews and rapid analysis for an interim brief.

The two parts follow the same logic. In each, written rules (a data dictionary or a codebook) make the work repeatable, missing or unclear material is handled openly, and the first summary (a Table 1 or a coding table) prepares the ground for analyses taught in HSCI 410 and HSCI 841.

Key Takeaways from this lesson

  • A dataset ready for analysis has one row per participant, one column per variable and one value per cell, with a unique identifier in place of names.
  • The raw file is kept read-only, and every change is made in a script that writes a separate analysis file.
  • A data dictionary records each variable's name, label, type, values, missing codes and derivation, and it is best written before cleaning begins.
  • Numeric missing codes must be converted to NA before analysis, because a code such as -99 treated as a number distorts every summary.
  • Categorical variables are summarized with counts and percentages, and numeric variables with a mean and standard deviation or a median and interquartile range.
  • Every percentage has a denominator, and a table should state whether it uses valid or total denominators and report the number missing.
  • Column percentages describe the make-up of each group in a Table 1, and row percentages compare the outcome across groups.
  • A starter codebook holds deductive codes with definitions and rules for use, and it grows through dated versions as inductive codes are added.
  • First-cycle coding labels segments of text; categories and themes are built from the codes in later analysis.
  • The analytic approach is chosen from the research question, the data, the timeline and the team, and Section 4.4 lists the HSCI 841 lesson for each approach.

Core Concepts Reviewed

Section 1: rectangular data, double data entry, validation rules, the read-only raw file, data dictionaries, derived variables, missing codes, and the documentation of a real survey dataset.

Section 2: levels of measurement, frequency tables, valid percentages, mean, median, standard deviation, interquartile range, cross-tabulations, row and column percentages, the analytic sample and Table 1.

Section 3: codes, coded segments, memos, first-cycle and second-cycle coding, deductive, inductive and in vivo codes, the starter codebook, and team coding.

Section 4: thematic analysis, qualitative content analysis, framework analysis, rapid qualitative analysis, grounded theory, interpretative phenomenological analysis and narrative analysis, matched to data collection methods.

The final reflection asks you to compare the quantitative and qualitative steps of this lesson and to plan the first analytic steps for a new dataset and two new transcripts.

Reflection

Part 1 of this lesson prepared and described survey data, and Part 2 prepared and coded interview data. Consider these pairs: a data dictionary and a qualitative codebook; a missing code converted to NA and a passage left uncoded or given a proposed new code; the definition of an analytic sample and the selection of segments for coding; a Table 1 and a coding table or framework matrix. Choose two of these pairs and explain what each element of the pair does and how the two are similar and different. Then describe, in three or four sentences, how an analyst would take a new survey dataset and two new interview transcripts through the first analytic steps of this lesson.

Model answer

A data dictionary and a qualitative codebook both make an analysis repeatable by writing down rules that would otherwise live in one person's head. The dictionary states what each column means, which codes are permitted and how derived variables such as a UCLA total are calculated; the codebook states what each code means and when to use or not use it. They differ in when they are settled. A data dictionary is ideally fixed before cleaning, while a codebook grows through dated versions as inductive codes are added.

A Table 1 and a framework matrix both arrange data so that groups can be compared. Table 1 summarizes each characteristic with counts, percentages or means for each loneliness group. A framework matrix places a written summary of each participant's coded data under each code, grouped, for example, by community. Table 1 reduces people to numbers, whereas the matrix keeps each person's words and context.

With a new survey dataset, the analyst would define the analytic sample, convert missing codes to NA and build a Table 1 with a footnote on denominators and exclusions. With two new transcripts, the analyst would read both in full, write five to eight deductive codes with all six codebook columns, code the first transcript with memos, revise the codebook to version 1.1, and then code the second.

Minimum 30 characters required.

✓ Reflection saved

Final Knowledge Assessment

Final Assessment, this lesson: First Steps in Analysis (13 Questions)

Question 1: Which practice best protects the raw survey export?

The raw export is the primary record. Keeping it unchanged and making all changes in a script means the analysis file can be reproduced and any cleaning mistake can be traced and reversed.

Question 2: The Cedar Valley data dictionary says that ucla_total is left empty if any item is missing. A participant answered 2 and 3 on the first two items and declined the third (-88). What is ucla_total in the analysis file?

The derivation rule in the dictionary says the total is missing if any item is missing. Option c shows what happens when a missing code is treated as a number, and option d invents a value.

Question 3: Which summary belongs in a Table 1 row for age in a sample whose ages are roughly symmetric?

For a numeric variable with a roughly symmetric distribution, the mean and standard deviation summarize centre and spread. The CSCS Table 1 reports age as 71.3 (5.3).

Question 4: A Table 1 reports 168 (39.4) respondents as single and not dating in the overall column of 433, with 7 missing. What denominator gives 39.4 percent?

The percentages are valid column percentages: 168 divided by 426 (433 minus 7 missing) is 39.4 percent. Dividing by 433 would give 38.8 percent.

Question 5: Why does the Table 1 in this lesson contain no p-values?

A descriptive Table 1 shows who was studied. Statistical tests are taught in HSCI 341 Lesson 6 Section 4 and HSCI 410 Lesson 3 Section 1. The groups are 209 and 224, so option b is also false.

Question 6: Among CSCS respondents aged 65 and older, 51.7 percent were classed as lonely, compared with 24.5 percent in the fictional Cedar Valley survey. What is the most defensible reading of the CSCS figure?

The CSCS sample was not drawn at random from a sampling frame, so its percentage describes its respondents. Comparing it with another survey's figure as if both were prevalence estimates (option b) ignores how each sample was obtained.

Question 7: Which description best fits an analytic memo written during coding?

Memos record the analyst's thinking as coding proceeds. Option d describes a codebook, and the memos often contain the first drafts of what later become themes.

Question 8: Which statement about deductive and inductive codes is accurate?

Deductive codes are written in advance from the question, guide, causal web or theory; inductive codes, including in vivo codes, are created from the data. Many applied studies combine them, an approach that Fereday and Muir-Cochrane (2006) describe as a hybrid of inductive and deductive coding.

Question 9: Which set of codes did the first coding pass apply to Ruth's sentence "A fellow from the coffee group drives me to my eye appointments in Cedar City"?

The fellow is a community contact (COMMUNITY), he provides rides (TRANSPORT), and the rides are to health care appointments (HEALTH). One segment can carry several codes.

Question 10: A survey adds the open-ended question "What makes it hard to stay in touch with people?", answered in a sentence or two by 1,600 people. Which approaches are the usual fit?

Short answers from many people suit approaches that categorize or find patterns across a large dataset. Interpretative phenomenological and narrative analysis need long individual accounts, and theoretical sampling is not possible with a completed survey.

Question 11: Which approach builds a theory of a social process through constant comparison and theoretical sampling, alternating data collection and analysis?

Grounded theory (Glaser & Strauss, 1967; Charmaz, 2014) alternates collection and analysis and chooses later participants to test emerging ideas. HSCI 841 Lesson 7 teaches it.

Question 12: Why did the Cedar Valley team choose framework analysis for its 24 interviews?

The health authority wanted to know whether experiences differ between Cedar City and smaller communities, and four people share the coding. A case-by-code matrix grouped by community supports both. Option d describes the team's rapid analysis of one focus group.

Question 13: Which pairing of a data problem and a step from this lesson is correct?

Section 2 converts the missing-code category to NA and removes the empty category. An undocumented variable needs investigation, a small cell may need combining or suppression, and passages that fit no code may need a new inductive code.
✦ Complete the final reflection above before submitting

Congratulations!

You have successfully completed this lesson: First Steps in Analysis.

You can now prepare and document a survey dataset, inspect a real dataset and its data dictionary, produce frequency tables, descriptive statistics, cross-tabulations and a Table 1, build a starter codebook, code an interview transcript with memos, and choose a qualitative analytic approach that fits a study's question and data.

Lesson 12, Writing Up Research, uses these results. It shows how to build the background, methods, results and discussion sections of a short report, how to design tables for categorical, continuous and qualitative data, including a Table 1 and a themes-with-quotes table, and how to cite sources in APA 7 style.

Reference

Glossary: Key Terms, People & Frameworks

📚 Reference page, available throughout the lesson

Search the terms, tools and people introduced in this lesson.

Core Concepts: Preparing and Describing Quantitative Data
Rectangular data A layout with one row per unit of observation (usually a participant), one column per variable and one value per cell.
Double data entry Two people enter the same paper forms independently, and the two files are compared so that every disagreement can be checked against the original.
Validation rule A check that refuses an impossible value at the moment of entry, such as an age outside the study's range.
Data dictionary A table describing every variable in a dataset: its name, label, type, values and value labels, missing codes, units and source or derivation.
Derived variable A variable created from other variables by a stated rule, such as a scale total calculated from its items.
Missing code A value, such as -88 or -99, used in a raw file to record why an answer is missing; it must be converted to a missing value before analysis.
Frequency table A table listing each value of a variable with the number of participants who have it, often with percentages.
Valid percentage A percentage whose denominator includes only participants with a recorded answer.
Median The middle value when values are sorted, with half of the participants below it and half above.
Standard deviation A measure of how far values typically lie from the mean; larger values mean more spread.
Interquartile range The range from the first to the third quartile, containing the middle half of the values.
Cross-tabulation A table counting participants in every combination of the categories of two variables.
Row and column percentages Percentages that make each row, or each column, of a cross-tabulation add to 100; column percentages describe each group in a Table 1.
Table 1 The descriptive table, usually first in a health research article, that summarizes the characteristics of the analytic sample overall and by comparison group.
Analytic sample The participants included in an analysis after stated exclusions, such as those missing the main outcome.
Core Concepts: Coding Qualitative Data
Code A short label attached to a segment of text to record what the segment is about or what it means for the research question.
Memo A short dated note in which a researcher records ideas, questions and connections that arise during data collection or coding.
First-cycle coding The initial labelling of segments of data, before codes are grouped into categories and themes in second-cycle coding.
Deductive code A code written before coding begins, from the research question, interview guide, causal web or theory.
Inductive code A code created during coding because a passage expresses something important that no existing code captures.
In vivo code An inductive code that uses a participant's own words as its label, such as "a room full of strangers".
Qualitative codebook The list of codes for a qualitative study, each with a name, definition, rules for when to use and not use it, and an example; it is revised in dated versions.
Frameworks & Tools
Canadian Social Connection Survey (CSCS) A national online survey of social connection, loneliness and health among people living in Canada, with a de-identified public dataset used in the series.
Thematic analysis An approach that identifies patterns of meaning (themes) across a dataset, set out in six phases by Braun and Clarke (2006).
Qualitative content analysis An approach that sorts text systematically into categories and may count them, in conventional, directed or summative forms (Hsieh & Shannon, 2005).
Framework analysis An approach that summarizes coded data in a matrix of cases by codes or themes so that cases and groups can be compared (Ritchie & Spencer, 1994).
Rapid qualitative analysis An approach that summarizes each interview in a structured template by topic and combines the summaries in a matrix to provide timely findings.
Grounded theory An approach that builds a theory of a social process from data through constant comparison and theoretical sampling.
Interpretative phenomenological analysis An approach that examines in detail how particular people make sense of a significant life experience, case by case (Smith et al., 2009).
Narrative analysis An approach that treats the story as the unit of analysis and examines its content, structure and performance (Riessman, 2008).
Key People
Virginia Braun and Victoria Clarke Psychologists who set out the six phases of thematic analysis (2006) and later described its reflexive form.
Johnny Saldaña Author of The Coding Manual for Qualitative Researchers, which describes first-cycle and second-cycle coding methods.
Kathleen M. MacQueen Lead author of a 1998 guide to codebook development for team-based qualitative analysis.
Jane Ritchie and Liz Spencer Social researchers who developed framework analysis for applied policy research (1994).
Barney Glaser and Anselm Strauss Sociologists who introduced grounded theory in The Discovery of Grounded Theory (1967).
Kathy Charmaz Sociologist who developed constructivist grounded theory.
Jonathan A. Smith Psychologist who developed interpretative phenomenological analysis.
Catherine Kohler Riessman Sociologist whose work on narrative methods describes how researchers analyze the content, structure and performance of stories.
No matching entries. Try a different search term.