HSCI 841 – Lesson 12

Computational Text Analysis, Cultural Domain Analysis & LLM-Assisted Coding

The Final Lesson

Learning objectives for this lesson:

  • Apply keyword-in-context (KWIC), word-frequency, keyness, and collocation analysis to interview corpora
  • Distinguish TF-IDF, keyness (chi-squared/log-likelihood), and raw frequency, and select the right measure for the analytic question
  • Conduct cultural domain analysis, free listing (with Smith's salience), pile sorts (with MDS), triad tests, and the Romney-Weller-Batchelder consensus model
  • Build a term co-occurrence network, compute centrality and community structure, and visualize it
  • Articulate the opportunities and risks of LLM-assisted qualitative coding, speed, scale, reproducibility, hallucination, prompt drift, bias amplification
  • Design a calibration-and-validation workflow for LLM coding using a hand-coded reference set and Krippendorff's alpha
  • Write a prompt that applies a study codebook reliably to a transcript, and audit the resulting codings
  • Disclose the use of LLM tools defensibly in a methods section
  • Check that consent, terms of use, and privacy law permit a model to process interview transcripts before any are uploaded

This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University based on Bernard, H. R., Wutich, A., & Ryan, G. W. (2017). Analyzing Qualitative Data: Systematic Approaches (2nd ed.). SAGE. This lesson covers Chapters 17, 18, and 19, and extends to LLM-assisted coding, a methodology that postdates Bernard, Wutich, and Ryan’s 2017 edition and is the most rapidly changing topic in the field.

Section 1 of 5

KWIC, Word Counts, and Keyness Analysis

⏱ Estimated reading time: 35 minutes
Lesson 12

Computational Text Analysis, Cultural Domain Analysis & LLM-Assisted Coding

When close reading hits its limits: three computational families, plus the large language model as coding assistant.

Section 1 of 4

KWIC, Word Counts, and Keyness Analysis

The workhorse statistics of corpus linguistics applied to your interview data.

The oldest technique

Keyword-in-context

...the corner where she used to sit empty now since March... ...pulled the kitchen closer to the window each morning... ...his reading still faces the television I never watch... chairchairchair Target word column-aligned; surrounding context printed left and right

The keyword-in-context list makes the distribution of senses inspectable. Forty-seven occurrences of chair is a frequency count; reading the lines is analysis.

Counting words

Frequency and lexical diversity

Word frequency

After stopword removal, rank words by count. The top 50 surface candidate terms for keyword-in-context analysis. The surprising entries are the analytically productive ones.

Type-token ratio

Types = distinct words. Tokens = all word instances. The ratio measures lexical variety. Length-corrected versions (moving-average type-token ratio) allow fair cross-document comparison.

In health research, lower lexical diversity in transcripts has been studied as a proxy for cognitive constriction in dementia and for emotional constriction in depression-related samples.

Document-distinctive vocabulary

Term frequency × inverse document frequency

Term Frequency (TF) How often a word appears in this document × Inverse Document Frequency Rarity of the word across the whole corpus

Term frequency inverse document frequency is highest for words common in this document and rare in others. Words like loneliness score low; wahda scores high for P15.

Comparing subcorpora

Keyness: older versus younger participants

Older participants (P05, P11, P17, P20)

deadquietvisitfadingradiowalker

Younger participants (P01, P10, P12, P19)

onlinescreenscrollingdiscordfollowers

Keyness with small subcorpora (4 documents per group) is statistically fragile. Treat individual word p-values as pointers, not confirmed tests. Read every keyword in its keyword-in-context context before drawing any interpretive conclusion.

Words that travel together

Collocations and n-grams

Collocation

Word pairs co-occurring more often than chance. empty chair, quiet apartment, nobody there.

Metaphor surface

Recurring phrases reveal shared symbolic vocabulary: fading at the edges, shrinks around you.

N-gram

Any sequence of n consecutive words. Bigrams and trigrams extend the feature space to multi-word concepts.

Carry forward

What to take into the next section

  • Keyword-in-context disambiguates senses; frequency statistics cannot.
  • Term frequency inverse document frequency surfaces case-distinctive vocabulary; keyness surfaces group-distinctive vocabulary.
  • Computational statistics are front-loaders for close reading, not substitutes for it.

Introduction and Overview

For the first eleven lessons of this course you have read transcripts line by line, hand-coded passages in Taguette, built codebooks, compared subgroups, written analytic memos, and worked your way through grounded theory, content analysis, schema analysis, narrative analysis, discourse analysis, and analytic induction. All of that work is what Bernard, Wutich, and Ryan call close reading. It is irreplaceable. But it does not scale. A 20-transcript loneliness corpus is at the upper edge of what a single analyst can read three or four times during a single graduate term. A 200-transcript corpus is not analyzable in that way. A 20,000-document corpus, the kind that increasingly arrives from social-media scraping, electronic health-record narrative fields, or large-scale open-text survey items, is unanalyzable in that way for any human team.

Computational text analysis is the family of techniques that lets you ask quantitative questions of a text corpus without first reducing the texts to codes (Grimmer, Roberts, & Stewart, 2022; Wikipedia contributors, n.d.-a). The questions are simple: which words appear most often? which words appear distinctively in subgroup A compared to subgroup B? which words travel together? where in the corpus does a given word appear, and in what immediate context? The answers are computed in seconds for corpora the size of yours and in minutes for corpora a thousand times the size. Chapter 17 of Bernard, Wutich, and Ryan covers the foundational techniques: keyword-in-context (KWIC), word frequencies, type-token ratios, TF-IDF, keyness, collocations, and n-grams. This first section of this lesson walks you through each of them, applied to the loneliness corpus.

One framing point before we begin. Computational text analysis is sometimes presented as an alternative to close reading. It is not. Bernard, Wutich, and Ryan's stance, and the stance of this course, is that computational techniques are front-loaders for close reading. They surface candidates for the analyst to read, in context, and decide whether the pattern is real. A keyness analysis that identifies tired as significantly more frequent in older participants than in younger participants is not a finding. It is a pointer toward passages worth reading closely. The finding is what you say after you have read those passages and decided whether the word is doing analytically interesting work or simply registering a generational vocabulary tic.

Learning Objectives for this section

  • Apply keyword-in-context (KWIC) analysis to the loneliness corpus.
  • Compute and interpret word frequencies, type-token ratios, and lexical diversity.
  • Distinguish raw frequency, TF-IDF, and keyness, and select the right measure for an analytic question.
  • Interpret a keyness analysis comparing older and younger participants in the loneliness corpus.
  • Identify collocations (words that co-occur more often than chance).
  • Understand n-grams as multi-word features and when to use them.

1.1 KWIC: The Oldest Computational Text-Analysis Technique

Keyword in Context. The oldest computational text-analysis technique, dating to medieval concordances and computerized in the 1950s. Generates lists of every occurrence of a target word with surrounding context. Used to disambiguate polysemy and to see how a term is actually deployed in a corpus before interpreting frequency counts.

Term Frequency × Inverse Document Frequency. A weighting that boosts terms frequent in a particular document and penalizes terms common across the whole corpus. Surfaces what is distinctive about each document. The standard input to document classification and retrieval systems.

Keyness. Statistical comparison of word frequencies in a target corpus vs a reference corpus. 'Which words are unusually frequent in 2026 health policy documents compared to 2010?' Log-likelihood and chi-squared tests are common. Output: a ranked list of distinguishing terms.

Collocations: words that co-occur within a defined window more often than chance would predict ('mental + health', 'public + health'). N-grams: consecutive word sequences (2-grams, 3-grams). Both reveal multi-word concepts that single-word analysis misses.

The keyword-in-context concordance is the oldest digital text-analysis technique, predating personal computers. The original KWIC concordances were produced on mainframes in the 1950s and 1960s, most famously for biblical scholarship and for early lexicography (Bernard, Wutich & Ryan, 2017, Ch. 17). The idea is simple: pick a target word; for every occurrence of that word in the corpus, print the word plus a fixed number of words on either side. The output is a vertical list of one-line excerpts with the target word column-aligned in the middle. A human can scan a 200-line KWIC of a single word in two or three minutes and develop a sense of how the word is being used that no other technique offers as cheaply.

The reason KWIC remains analytically valuable in 2026 is that it does something that quantitative text statistics do not: it preserves local context. A word frequency table tells you that chair appears 47 times in the loneliness corpus; it does not tell you that 32 of those 47 are references to a chair where an absent person used to sit, that 11 are references to the participant's own seated immobility, and that 4 are incidental mentions of furniture. The KWIC reveals the distribution of senses in seconds. The numerical-frequency analyst who skipped the KWIC step might mistakenly write “chairs are an important physical motif in the loneliness corpus,” which is true but misses the more specific finding that empty chairs are an important physical motif.

The technical mechanics of KWIC are trivial. The interpretive demands are not. A good KWIC reading is one where you have looked at every line of the concordance, classified each occurrence into a small number of senses, and decided which sense is the analytically productive one. You should treat KWIC as a five-to-ten-minute exercise per target word, not as a one-line command whose output you skim.

To try KWIC on the loneliness corpus, open the 20 transcripts in the course data bundle, HSCI_841_loneliness_data.zip, search each one for chair, and copy every occurrence with about six words on either side into a single list. Read the list line by line and classify each occurrence into a sense (the empty chair of an absent person, the participant’s own chair as a marker of immobility, an incidental mention of furniture); the distribution of senses is itself a finding. Repeat the exercise for alone and lonely. Bernard, Wutich, and Ryan (and your transcripts) treat these as distinct, and a KWIC contrast surfaces the distinction concretely.

1.2 Word Frequencies, Type-Token Ratios, and Lexical Diversity

The simplest computational text statistic is the word frequency. After removing stopwords (the, and, of, but, was, are…), you compute how many times each remaining word appears in the corpus and rank them. The top 50 or 100 words are usually a mix of the obvious (loneliness, alone, people, feel) and the surprising. The surprising ones are the analytically valuable findings. In the loneliness corpus, the top 50 content words include chair, quiet, radio, wednesday, and fading; each of those repays a KWIC read and yields a candidate theme.

Word frequency analysis is also the foundation for two further statistics: the type-token ratio (TTR) and lexical diversity. A token is any occurrence of a word; a type is a distinct word. The phrase “the chair, the empty chair” contains 5 tokens and 3 types (the, chair, empty). The ratio of types to tokens is one measure of how lexically varied a text is. A high TTR means the speaker uses many different words; a low TTR means they recycle a small vocabulary. TTR is sensitive to document length (longer documents inevitably have lower TTRs because common words repeat), so for cross-document comparison you typically use a length-corrected measure such as the Moving-Average Type-Token Ratio (MATTR) or Mean Segmental TTR (MSTTR).

In health research, lexical diversity has been used as a coarse proxy for cognitive function (lower diversity in dementia transcripts) and for emotional constriction (lower diversity in some depression measures). In your loneliness corpus, a comparison of lexical diversity across participants is a defensible analytic move, it would let you ask, for example, whether the participants with the most circumscribed social worlds also speak with the most circumscribed vocabularies. Bernard, Wutich, and Ryan would emphasize that any such finding requires close-reading confirmation; a numerical difference is a pointer, not a conclusion.

What to look for: the top-50 list of content words surfaces the candidate words for KWIC analysis. A ranking of participants by MATTR identifies the participants with the most and least varied vocabularies. Read those transcripts again in light of their MATTR rank, and decide whether the numerical difference registers something analytically meaningful (a genuine constriction of social or emotional vocabulary) or a stylistic or dialectal difference.

1.3 TF-IDF: Term Frequency × Inverse Document Frequency

Raw word frequency tells you which words appear most often in the corpus. It does not tell you which words are distinctive of a particular document. The word loneliness appears many times in every transcript; it is not informative about which transcript is which. The word wahda appears in only one transcript (P15, Amira) and there several times; it is highly informative about that transcript.

Term Frequency × Inverse Document Frequency (TF-IDF) is the classical measure designed to surface this kind of document-distinctive vocabulary (Wikipedia contributors, n.d.-b). The idea is to weight each word's frequency in a document by how rarely it appears in the rest of the corpus. The term-frequency component rewards words that are common in the focal document; the inverse-document-frequency component penalizes words that are common in many other documents. The product is highest for words that are common in this document and rare elsewhere. TF-IDF was developed for information retrieval (which documents are most relevant to a search query?) in the 1970s and has become the workhorse statistic for document similarity, search ranking, and feature selection in supervised text classification.

In qualitative analysis, TF-IDF is most useful for surfacing what a single transcript is distinctively about. A TF-IDF ranking of P15 (Amira, recent refugee from Syria) is likely to surface wahda, Aleppo, family, before, and other terms that capture what is distinctive about Amira's account of loneliness compared to the other nineteen participants. Bernard, Wutich, and Ryan treat this as a tool for case characterization: which words best describe what makes this case different from the rest of the corpus?

Worked example: vocabulary fingerprints

Ranking each transcript’s words by TF-IDF gives a vocabulary fingerprint for each participant. In the loneliness corpus, P15 (Amira) surfaces wahda and Syria-related terms; P11 (Helen) surfaces fading, radio, marie, and wednesday; and P05 (Linda) surfaces terms tied to her widowhood and her terrier. These fingerprints are useful both as descriptive case sketches in your findings section and as cross-checks on the candidate themes that emerged in earlier coding.

1.4 Keyness: Comparing Two Subcorpora

Keyness analysis is the comparative cousin of TF-IDF. Instead of asking which words distinguish a single document from the corpus, it asks which words distinguish a group of documents (a subcorpus) from another group of documents. The standard implementation uses either the chi-squared test or the log-likelihood ratio test on the 2×2 frequency table of (word X / not-word X) by (group A / group B), one word at a time. The output is a ranked list of words with their test statistic, p-value, and frequency in each subcorpus. The words at the top are the words that are statistically more distinctive of one subcorpus than the other.

Keyness is widely used in corpus linguistics (e.g., comparing British versus American English corpora) and increasingly in health research (comparing patient-experience text by diagnosis, by gender, by treatment arm in a trial). For your loneliness corpus, the most natural comparison is across age. The corpus contains four participants in their seventies and eighties (P05 Linda, P11 Helen, P17 Jacob, P20 Frank) and four in their twenties (P01 Maya, P10 Daniel, P12 Tyler, P19 Rose). A keyness analysis comparing these two subcorpora surfaces the vocabulary of late-life loneliness against the vocabulary of early-adult loneliness, and the contrast is consequential. The older subcorpus has more keywords like dead, quiet, visit, fading, radio, walker; the younger has more keywords like online, screen, followers, group chat, discord, scrolling.

The interpretive point is not that the words are themselves the finding. The interpretive point is that the structural worlds in which the two cohorts experience loneliness are different, and the keyness list is a vocabulary-level signature of that structural difference. The finding the keyness analysis points you toward, that the technology mediation of loneliness in young adulthood and the embodied immobility of loneliness in old age are different empirical objects deserving different policy responses, is the analytic claim that goes in the write-up. The keyness list is the evidence trail.

Diverging horizontal bar chart of keyness values. Older participants' distinctive words (walker, fading, radio, visit, quiet, dead) extend to the left; younger participants' distinctive words (scrolling, discord, followers, group chat, screen, online) extend to the right. Bar length is the log-likelihood keyness statistic.
Illustrative keyness output for the loneliness corpus. Words distinctive of the older subcorpus describe embodied immobility and physical absence; words distinctive of the younger subcorpus describe screen-mediated contact. The two vocabularies are signatures of structurally different experiences of loneliness.

Two interpretive cautions apply to the figure above. First, keyness with small subcorpora (four documents per group) is statistically fragile, so individual word p-values are best treated as pointers to passages, with confirmation left to the close read. Second, read every keyword in its KWIC context before drawing any interpretive conclusion. The numerical contrast is the front end; the close read is the analysis.

1.5 Collocations: Words That Travel Together

A collocation is a pair (or larger sequence) of words that co-occurs more often than chance would predict. The standard test is again log-likelihood or chi-squared on the 2×2 contingency table of (word A present / absent) by (word B present / absent) within a sliding window of n words. High-scoring collocations tell you which word pairs are linguistically bonded in the corpus. The output for a loneliness corpus might include empty chair, quiet apartment, group chat, tired all the time, nobody there, last conversation, and similar pairings.

Collocations are useful for two analytic purposes. First, they surface candidate multi-word concepts that single-word frequency analysis misses. Empty chair as a collocation captures something neither empty nor chair alone does; it is the unit of meaning. Second, they surface candidate conventional metaphors, recurring word combinations that participants are drawing on a shared symbolic vocabulary to use. In the loneliness corpus, recurring phrases like fading at the edges (Helen), shrinks around you (also Helen), and empty space (multiple participants) are conventional metaphors in the Lakoff-Johnson sense; collocation analysis surfaces them efficiently.

Worked example: the neighborhoods of alone and lonely

Comparing the words that fall within a few words of alone with those that fall near lonely makes the contrast visible. The alone neighborhood tends to surface neutral or even positive context words (home, quiet, peace, comfortable), because participants describe being alone in many ways and only some of them are lonely. The lonely neighborhood tends to surface affectively heavier neighbors (tired, hollow, fading, empty, dead, miss). The contrast operationalizes Bernard, Wutich, and Ryan’s claim, and your participants’ direct statements, that alone and lonely are distinct concepts.

1.6 N-grams

An n-gram is a sequence of n consecutive words. Unigrams (n=1) are single words; bigrams (n=2) are pairs; trigrams (n=3) are triples. The collocations above were a special case of n-grams: bigrams and trigrams whose components co-occur statistically more than chance. Plain n-gram frequency analysis, without the statistical-collocation filter, is also useful, especially when you want to find conventional fixed phrases (at the end of the day, most of the time, I don't know) or when you are setting up a more elaborate downstream model that consumes n-gram features (a topic model, a supervised classifier, a semantic network on bigrams).

Technically, an n-gram feature set is built by joining each run of n consecutive words into a single feature (for example, empty_chair) and then counting those features like any other words. The cost is that the resulting feature space is much larger than the unigram space (typically 10–20× more features), which slows downstream computation; the benefit is that n-grams capture lexical structure that unigrams cannot.

1.7 What These Techniques Do and Do Not Tell You

Free listing & free recallClick to explore
Code suggestionClick to explore
Inter-coder benchmarkingClick to explore
Critical cautionsClick to explore

Computational text statistics are useful diagnostics. They are not, on their own, qualitative analysis. Bernard, Wutich, and Ryan are explicit on this: a frequency table, a keyness list, a TF-IDF ranking, and a collocation report are pointers. They surface candidates for close reading. The analytic move, deciding what the patterns mean, whether they support a theme, whether they tell you something about the phenomenon or merely about the vocabulary your participants happen to share, remains the analyst's work. Treat this section's tools as the first hour of a longer day. They tell you where to look. The looking is still your job.

Reflection

Of the techniques in this section, KWIC, frequency, TTR/MATTR, TF-IDF, keyness, collocations, which one would be most analytically useful for a specific research question about the loneliness interviews? Why? Be specific about what it would surface and how that would feed into a close-reading move.

Model answerA strong response names a specific technique, names what it would surface, and names the close-reading move that follows. Examples: “My question compares the loneliness of recent immigrants with that of long-time residents. Keyness analysis is the right tool because it operationalizes the question directly, which words are distinctively used by the immigrant subcorpus, and the resulting keyword list (e.g., before, family, language, belong, possibly an Arabic loan-word like wahda) will point me to passages where the comparison is most concretely visible.” Or: “My question is about the embodied, sensory dimension of late-life loneliness. KWIC analysis of bodily-sense words (tired, cold, quiet, still, shrinks) is the right tool because it lets me sample every occurrence in context and decide which are genuine embodiment claims versus filler usage.” A weak answer names a technique without specifying what the surfaced output would be or what you would do with it.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: Which of the following is the analytically most defensible use of a keyness analysis comparing older and younger participants in the loneliness corpus?

Keyness is a pointer for close reading, not a finding by itself. With small subcorpora, individual p-values are fragile, but the ranked list of candidate words is still useful as a starting point for close-reading interpretation.

Question 2: Why is raw word frequency a poor measure for surfacing the words most distinctive of a single transcript?

TF-IDF weights term frequency by inverse document frequency, so words common across the whole corpus are de-emphasized and document-distinctive vocabulary surfaces. Raw frequency does not do this and is therefore dominated by common words.

Question 3: KWIC analysis is most accurately described as:

KWIC is a concordance technique, for every occurrence of a target word, it shows the word in its immediate context. Its analytic value comes from making local context inspectable cheaply.
Section 2 of 5

Cultural Domain Analysis: Free Lists, Pile Sorts, Triads, Consensus

⏱ Estimated reading time: 35 minutes
Section 2 of 4

Cultural Domain Analysis

Free listing, pile sorts, triad tests, and consensus analysis. Measuring shared mental models.

Free listing

Free listing and Smith's salience

Smith's Salience per item family (high freq, early rank) pets language-class (rare, late)

The formula: for each item, sum (L − Rp + 1) / L across all participants, then divide by total participants. Full credit to first-mentioned; diminishing credit down the list.

Pile sorts

Pile sorts and multidimensional scaling

mother partner best friend group chat pet Intimate cluster Mediated cluster Floating

The dimensions are unlabeled by the algorithm. Part of the analytic work is proposing a label (e.g., dim 1 = intimacy, dim 2 = embodiment) and arguing for it.

Triad tests

Triad tests

Present three items. Ask: which is most different from the other two?

Repeated across many triads, this accumulates pairwise-similarity data at finer resolution than a pile sort, but the number of triads scales as C(n, 3) = n! / (3! × (n−3)!).

For n = 15 items: 455 triads, too many for one participant. Balanced incomplete block designs reduce the load while preserving coverage.

When to use: Small domains (n < 10), populations where pile sorts are cognitively demanding (children, cognitive impairment), or when the additional discriminatory precision is worth the design complexity.

Consensus analysis

Romney-Weller-Batchelder consensus model

What it estimates

Per-participant competence scores and a competence-weighted cultural answer for each domain item.

The eigenvalue ratio test

First / second eigenvalue ≥ 3: supports single-culture model.
Ratio < 2: multiple sub-cultures present; partition the sample.

Published applications include lay theories of HIV transmission risk, postpartum mood, acceptable harm-reduction supplies, and concepts of a good death.

Carry forward

What to take into the next section

  • Free listing surfaces the cultural core (high Smith's salience) versus the periphery of a domain.
  • Pile sorts and multidimensional scaling reveal implicit category structure without imposing it.
  • The eigenvalue ratio test is the diagnostic for consensus: ratio ≥ 3 supports a shared model; below 2 means multiple sub-models.

Introduction and Overview

Cultural domain analysis is a family of techniques developed in cognitive anthropology in the 1950s through 1980s to study how people in a cultural group mentally organize a delimited domain of knowledge: kinds of illness, kinds of edible plants, kinds of kin, kinds of risk. The starting premise is that a cultural domain is a shared mental model. The techniques in this section, free listing, pile sorts, triad tests, and consensus analysis, are designed to measure the structure of that shared model and to estimate each participant's degree of competence: how closely their understanding of the domain tracks the group consensus.

For public-health audiences, cultural domain analysis is consequential because most health categories (kinds of risk, kinds of symptom, kinds of treatment, kinds of support) are domain-organized in exactly this way. A public-health communication campaign that treats “the public” as a single audience with a single mental model often fails because the model is in fact heterogeneous across groups. Cultural domain analysis measures the heterogeneity precisely and tells you which sub-populations share a model and which do not.

The intellectual lineage is short and important. The technique was systematized by Romney, Weller, and Batchelder in their 1986 paper Culture as Consensus: A Theory of Culture and Informant Accuracy, published in American Anthropologist. The Romney-Weller-Batchelder (RWB) consensus model treats culture statistically: there is a true cultural answer for each item in the domain, and participants approximate it with varying competence. The model estimates competence from inter-informant agreement (using a factor-analytic decomposition) and produces, for each item, a best-estimate cultural answer that is the competence-weighted average of all participants' answers. The model has been applied to hundreds of public-health questions, from kinds of risk for HIV transmission to the cultural domain of acceptable cooking fuels in low-income households.

Learning Objectives for this section

  • Conduct a free-listing exercise and compute Smith's salience.
  • Distinguish single-sort and successive-sort pile sorts, and analyze pile-sort data with multidimensional scaling (MDS).
  • Conduct a triad test and read its output.
  • State the Romney-Weller-Batchelder consensus model assumptions and interpret the eigenvalue ratio test for cultural agreement.
  • Apply cultural domain methods to a public-health question: which dimensions of social support, kinds of risk, or kinds of treatment do your participants share a model of?

2.1 Free Listing: The Foundational Elicitation Technique

Free listing, the foundationv

Ask 30+ informants to list all the items they can think of in a domain. Aggregate the lists. Two key outputs: (1) salience scores per item (frequency + rank), (2) the empirical map of the domain’s vocabulary. Most cultural-domain studies start here.

Pile sorts, revealing structurev

Give informants 20-40 items (typically from free lists) and ask them to sort into piles. Aggregate proximity matrices. Multidimensional scaling and hierarchical clustering reveal the implicit category structure. The result is a map of how the domain hangs together.

Triad tests, precision toolv

Present three items and ask which one doesn’t belong. Repeated across many triads, this elicits the underlying dimensions along which informants distinguish items. More cognitively demanding but more discriminating than pile sorts; useful for refining a known structure.

Consensus analysis (Romney-Weller-Batchelder)v

A statistical model that uses the pattern of agreement across informants to (a) estimate each informant’s competence in the domain, (b) estimate the ‘answer key’ that the group implicitly shares. The defensible alternative to majority voting in cultural surveys. Particularly important when cultural consensus itself is the research question.

Free listing is the simplest cultural domain technique and almost always the one to start with. You ask a participant a simple prompt: “List all the kinds of X you can think of.” You let them list freely, in their own order, for as long as they wish to continue. You record both the items and the order. You collect free lists from a sample of n participants (typically 20–40 is sufficient; less than 15 is usually too few for reliable salience estimates). You then analyze the combined list to identify (a) which items are mentioned by the most participants (frequency), (b) which items tend to be mentioned early (rank), and (c) which items are both common and early (Smith's salience).

Smith's salience is a single number per item, calculated as:

S = ( ∑ ((L − Rp + 1) / L) ) / N

where, for each item, you sum across the N participants the quantity (L − Rp + 1) / L, with L being the length of participant p's list and Rp being the rank of the item in that list (R = 1 for first-mentioned). The numerator gives full credit (= 1) to an item mentioned first on a participant's list, half credit to an item halfway down the list, and so on. The sum is divided by N (the total number of participants). The resulting score ranges from 0 to 1; higher is more salient.

Items with the highest Smith's salience are the core items of the domain, the ones a randomly chosen member of the group is most likely to think of first when asked about the domain. Items with low salience are peripheral, idiosyncratic to one or two participants, or sub-domain-specific. In a free-listing exercise on “kinds of social support,” the high-salience items might be family, friends, partner; the low-salience items might be online community, religious community, therapist. The contrast tells you what the participants' default model of social support contains, and what they leave out.

🔎 Free-listing exercise (worked example using the loneliness corpus)

The loneliness interviews are semi-structured rather than free-list, but several questions in the interview guide elicit list-like responses (Domain 2, Q7: “Are there situations or times that reliably bring it on?”; Domain 4, Q11: “What helps?”). For the worked example below, treat each participant's response to Q11 (“What helps?”) as their free list of coping strategies. Extract the items mentioned and their order. Compute Smith's salience across the 20 participants. The high-salience items are the cultural core of coping with loneliness in this sample; the low-salience items are idiosyncratic.

A real free-listing study would use the prompt “List all the things that help you when you are lonely” and would record the listing directly. The retrospective extraction from semi-structured interviews is a defensible second-best.

Smith’s S can be computed by hand or in a spreadsheet. For each participant’s list, give each item the score (L − R + 1) / L, add each item’s scores across participants, and divide by the number of participants, so that participants who did not list an item contribute 0.

Worked example: Smith’s salience for “What helps?”

The course data bundle, HSCI_841_loneliness_data.zip, includes the free lists extracted from the 20 transcripts (freelist_what_helps.csv), with one row per participant, item, and rank. The 20 lists name 31 different items, with an average list length of 4.25. The five most salient items are shown below.

ItemParticipants listing it (n / 20)Smith’s S
peer-group100.43
therapy90.25
friends60.22
television50.14
cooking30.11

Peer groups stand out as the cultural core of coping in this sample: half the participants name one, usually near the top of their list. Therapy is listed almost as often but further down the lists, which lowers its salience. Most of the remaining items are named by one to four participants and sit in the periphery. For a public-health intervention designer, the cultural core tells you what most people will already know about; the periphery tells you what an information campaign would need to introduce.

2.2 Pile Sorts: Eliciting Implicit Structure

A pile sort is conducted as follows. You write each item of the domain on a card (or display them on a screen). You hand the deck to a participant and say: “Sort these into piles so that the things in each pile are similar to each other. Make as many or as few piles as you like.” You record the resulting partition: which items the participant placed together. You repeat across n participants and aggregate the pile sorts into a single similarity matrix: cell (i, j) of the matrix is the number of participants who put items i and j in the same pile.

The aggregate similarity matrix is then analyzed with multidimensional scaling (MDS) or hierarchical clustering to recover the implicit structure of the domain. MDS produces a 2-D map where items are points and the distances between them reflect their dissimilarity. Items that participants reliably grouped together appear close on the map; items that were almost never grouped together appear far apart. The resulting map is the visualization of the group's shared mental structure of the domain.

A successive pile sort is the same exercise repeated: after the first sort, you ask the participant to combine piles (“Which of these piles could be combined into a larger pile? Which combinations would still feel similar?”) until they reach a small number of super-piles, and you record the hierarchy. The result is a hierarchical clustering for each participant, which can also be aggregated. The successive pile sort is more informative than the single sort but slower; the single sort is the standard for most studies.

For a worked example consistent with the loneliness study: imagine 30 cards, each naming a kind of social relationship (mother, father, sibling, partner, best friend, ex-partner, coworker, neighbour, pet, online friend, religious-community member, therapist, GP, phone-pal, group-chat acquaintance, …). A pile-sort study with 25 BC residents would produce an MDS map of how these relationships cluster in our shared mental model. The clustering would tell you which relationships participants treat as functionally equivalent for the purpose of countering loneliness, which they treat as substitutes, and which they treat as categorically different. Such a study would be policy-actionable: if “pet” clusters with “close friend” in the shared model, then a pet-based intervention is closer to the cultural category of friendship than to the cultural category of distraction, with consequences for how the intervention is framed and evaluated.

Worked example: pile sorts of ten kinds of relationship

The course data bundle includes a small pile-sort dataset (pilesort_relationships.csv) in which 20 participants sorted ten kinds of social relationship into piles. Counting, for every pair of items, how many participants placed the two in the same pile gives the similarity matrix, and classical MDS of that matrix (available in most statistical packages) gives a two-dimensional map. In this dataset the three kin items (mother, partner, adult child) sit together, the three close non-kin items (best friend, neighbour, support group) form a second cluster, the three mediated ties (coworker, group chat, online friend) form a third, and pet sits on its own, which is the interesting result.

The algorithm does not label the dimensions of the map. Part of your analytic work is to inspect the map and propose a labelling (for example, dimension 1 = “intimacy”, dimension 2 = “embodiment”). Proposed dimensional labels are an interpretive claim that needs argument, and they come from the analyst.

2.3 Triad Tests

The triad test addresses a methodological problem with pile sorts: pile sorts assume participants can hold and partition an entire deck of items at once, which is cognitively heavy for large decks. The triad test breaks the problem into manageable pieces. You present three items at a time and ask: “Which of these three is most different from the other two?” (Some variants instead ask which two are most similar; the information is equivalent.) Across many triads, you accumulate enough pairwise-similarity data to reconstruct the same kind of similarity matrix that the pile sort produces, but with finer resolution.

The cost is the number of triads. For a domain of n items, there are C(n, 3) = n!/(3!(n−3)!) possible triads; for n = 15, that is 455 triads, which is too many for any participant to complete. Triad-test designers use balanced incomplete block designs (Borgatti's UCINET implementation has a wizard for this) that present each participant with a stratified subset of triads such that every pair of items appears in approximately equal numbers of triads across the sample. The standard analysis is again MDS on the aggregated similarity matrix.

Triad tests have largely been displaced in contemporary practice by pile sorts (which are faster) and by ratings (which are easier to administer remotely). They remain the gold standard for very small domains (n < 10) and for studies where the cognitive load of holding the full deck would compromise data quality (children, participants with cognitive impairment).

2.4 Consensus Analysis: The Romney-Weller-Batchelder Model

The pile sort and the triad test produce an aggregate similarity matrix for the group. The Romney-Weller-Batchelder consensus analysis goes one step further: it treats the participants themselves as items to be analyzed, and asks how much do they agree with each other? The output is two-fold: a per-participant competence score (how closely the participant's answers track the group consensus) and an estimated cultural answer for each item (the competence-weighted average of all participants' answers).

The model has three formal assumptions: (1) there is a single shared cultural truth for each item in the domain (one knowledge system, not several); (2) participants' answers are independent samples from their own competence-driven approximation to the truth; (3) each participant's competence is approximately constant across items. The first assumption is the most consequential and the most testable. If the data violate assumption 1, if there are two or more distinct cultural sub-models in your sample, the consensus analysis will tell you so.

The diagnostic is the eigenvalue ratio test. Consensus analysis runs a factor analysis on the participant-by-participant agreement matrix. If there is a single shared culture, the first eigenvalue will be much larger than the second, the rule of thumb is that the first should be at least three times the second. If the first/second ratio is closer to 1, there is no single consensus; the sample contains multiple sub-cultures. In that case you should not report a single “cultural answer”; you should report separate analyses for each subgroup (and explain in your discussion why the consensus broke down).

In contemporary public health, consensus analysis has been used to study, among other things, lay theories of HIV transmission risk in different communities (where the eigenvalue ratio test repeatedly identifies sub-cultures), kinds of postpartum mood, kinds of acceptable harm-reduction supply, and the cultural domain of “a good death.” The technique's strength is that it gives a formal statistical test for whether your participants share a cultural model; its weakness is the strong single-culture assumption and the modest sensitivity of the eigenvalue test in small samples.

The course data bundle includes a consensus dataset (consensus_loneliness.csv) in which the 20 participants answer 12 yes/no statements about loneliness. The computation follows the model directly. The agreement matrix records, for each pair of participants, the proportion of statements on which they give the same answer. Its eigen-decomposition gives the eigenvalue ratio and each participant’s competence (the participant’s loading on the first factor), and the cultural answer for each statement is a competence-weighted vote. If the ratio of the first to the second eigenvalue is at least 3, the sample supports a single-culture model and the cultural answers can be read as the group’s best estimate of the truth for each item. If the ratio is between 2 and 3, the support is marginal, so treat the cultural answers with caution and explore subgroups. If the ratio is below 2, there is no single consensus; partition the sample (for example, by age or by community), run the analysis separately on each subgroup, and compare.

2.5 When to Reach for Cultural Domain Analysis in a Health Study

Key insight - Augment, don't replace

The 2026 state of qualitative analysis: LLMs and computational text tools can augment the careful researcher but cannot replace them. They speed up scaffolding tasks (KWIC, code suggestion, summarization), allow new kinds of analysis (corpus-scale keyness comparison, semantic clustering), and lower the cost of methodological triangulation (human + computational coding). What they cannot do is warrant interpretive claims, that work still belongs to a human analyst who can be held accountable for it. The defensible 2026 workflow combines computational and manual methods explicitly, names each method’s contribution, and reports the human judgment that arbitrated between them.

Cultural domain analysis is the right tool when your research question is about a shared mental model: a structured set of categories that the participants treat as having a known list of members, a known set of relations among the members, and a known set of acceptable answers. It is the wrong tool when the phenomenon is fundamentally idiographic (each participant's experience is its own object) or when the “domain” is too open-ended to have a closed list of items.

For the loneliness interviews, cultural domain methods are not the dominant analytic approach (the interviews are too open-ended), but they could plausibly feature in a sub-analysis. The most natural application would be a free-listing study of kinds of social support or kinds of coping, possibly with an MDS on a derived pile-sort matrix. The result would be a structured map of the shared cultural model of how loneliness is countered, which would complement the close-reading account that anchors the rest of the paper.

2.6 Applications in Health Research

ACTIVITY Try it - Prompt an LLM as a code-suggestion engine

Take 3-5 short passages from the loneliness interview dataset that you have already coded by hand; the transcripts are synthetic, so they can be sent to a third-party model, whereas real interview data need the governance checks in Section 4.4 first. Draft a prompt:

  • Role: 'You are a qualitative health research assistant helping me code interview transcripts about [topic]'.
  • Codebook: Provide your existing codes with one-line definitions.
  • Examples: Provide 2-3 already-coded passages.
  • Task: Provide one new passage and ask: 'Which codes apply, and to which phrase? Provide a one-sentence rationale per code.'

Then compare LLM output to your own coding. Where does it agree? Where does it diverge? Why?

This exercise reveals both LLM strengths (consistency, speed) and weaknesses (over-literal reading, missed irony, stereotyping). Both are useful diagnostic findings.

ApplicationTechniqueWhat it surfaces
Lay theories of HIV transmissionFree listing + consensusThe cultural core of perceived risks; identification of sub-cultures with discrepant models
Kinds of postpartum moodFree listing + pile sort + MDSThe lay nosology of mood states that screening tools must map onto
Acceptable harm-reduction suppliesPile sort + consensusWhich supplies participants treat as functionally equivalent for provincial policy
A “good death”Free listing + consensus by cohortGenerational and cultural variation in end-of-life values
Kinds of social supportFree listing + pile sort + MDSWhich relationships participants treat as substitutes for one another
Indigenous concepts of wellnessFree listing + community-based consensusConcepts not captured by Western biomedical frameworks

Reflection

Imagine a study of loneliness has space for one cultural-domain-analysis sub-study. Which of the four techniques (free listing, pile sort, triad test, consensus analysis) would you propose, and what would the prompt be? What would a positive result look like? What would a negative result (e.g., low eigenvalue ratio) tell you?

Model answerA good answer names a domain, a technique, a prompt, and an interpretation rule. Example: “I would conduct a free-listing study with the prompt ‘List all the things that help when you are lonely’ across 30 participants stratified by age. Smith's salience would surface the cultural core of coping (likely family, friends, pets, activities). I would then run consensus analysis on a follow-up acceptability rating of each item across participants. A high eigenvalue ratio (≥ 3) would license the conclusion that participants share a single cultural model of coping; a low ratio (< 2) would license the more interesting conclusion that kinds of coping is not a unified domain in this population but instead splits across age strata, which would have direct implications for designing age-tailored loneliness interventions.” A weak answer names a technique without specifying the prompt or the interpretation criterion.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: Smith's salience is calculated such that an item is most salient when it is:

Smith's S combines frequency (how many participants mention the item) and rank (how early they mention it). The most salient items are common AND mentioned early; the least salient are mentioned rarely AND late.

Question 2: The Romney-Weller-Batchelder consensus model's eigenvalue ratio test is used to assess:

The first/second eigenvalue ratio of the participant-by-participant agreement matrix is the diagnostic for the single-culture assumption. A ratio ≥ 3 supports single-culture; a low ratio indicates multiple sub-cultures or no consensus.

Question 3: A pile-sort study presents 25 cards to each of 30 participants and asks them to sort the cards into similarity piles. The aggregated similarity matrix is then analyzed with multidimensional scaling. The resulting 2-D map will best show:

MDS produces an unlabeled 2-D map where items participants grouped together are close and items they rarely grouped are far. The analyst must propose interpretations of the dimensions; the algorithm does not label them. The result is a visualization of the implicit (shared) structure.
Section 3 of 5

Semantic Network Analysis: Text as Network

⏱ Estimated reading time: 35 minutes
Section 3 of 4

Semantic Network Analysis

Text as a co-occurrence network: centrality, community detection, and visualization.

Building the network

Feature co-occurrence matrix

Context unit: sliding window (n = 5 tokens) ...the chair was empty after Marie died... co-occurrence edge Sentence-level and sliding-window are the standard choices for interview corpora
Centrality measures

Four ways a word can matter in a network

Degree

Count of neighbors. Breadth of connection.

Betweenness

On many shortest paths. The conceptual bridge.

Closeness

Short average path to all others. Near the centre.

Eigenvector

Connected to well-connected nodes. Conceptual prestige.

Community detection

Louvain method: theme clusters emerge from the data

Embodied tired, body quiet, fading Social family, friends group, work Coping dog, walk radio, music Candidate communities from Louvain; labelling is the analyst's interpretive work
Carry forward

What to take into the next section

  • Louvain communities are candidate themes from the corpus vocabulary, not final codes.
  • Agreement between network communities and hand-coded themes is computational corroboration; divergence is a finding to investigate.
  • A defensible semantic network figure colors nodes by community, sizes them by degree or eigenvector centrality, and explains the context-unit choice in the caption.

Introduction and Overview

Chapter 19 of Bernard, Wutich, and Ryan treats text as network. The idea is straightforward: take a corpus, identify the content words, and treat each word as a node. Connect two words with an edge whenever they co-occur within a defined window, in the same sentence, the same paragraph, or some sliding span of n tokens. The result is a weighted graph: nodes are words, edges are co-occurrences, edge weights are co-occurrence counts. Once you have the graph, the entire toolkit of network analysis becomes available: centrality measures, community detection, density, path length, visualizations.

Semantic network analysis is one of the most interpretively rich computational text techniques in Bernard, Wutich, and Ryan’s discussion because the network is a visualization of the corpus's conceptual structure. Words that participants talked about in the same breath cluster together in the network. Words that mediate between clusters, high-betweenness words, are the conceptual bridges of the corpus. Word communities, dense sub-graphs identified by modularity-based clustering algorithms, are candidate themes that no human coder identified explicitly but that the participants' language organized implicitly.

The technique has been used productively in studies of patient experience (semantic networks of pain descriptions), public communication (semantic networks of vaccine-hesitant tweets), policy text (semantic networks of climate-adaptation reports), and qualitative methods textbooks themselves (semantic networks of methods-section vocabularies across decades). For your loneliness corpus, a semantic network of the top 30–50 content words will visualize the conceptual scaffolding of how the participants collectively articulate the experience of loneliness, with the major thematic clusters (embodiment / coping / loss / connection) emerging as graph communities rather than as analyst-imposed codes.

Learning Objectives for this section

  • Build a term co-occurrence matrix with sentence, paragraph, or sliding-window definitions.
  • Convert the matrix to a network and visualize it.
  • Compute degree, betweenness, closeness, and eigenvector centrality, and interpret each.
  • Run community detection (Louvain method) and read the resulting partitions as candidate themes.
  • Compute network-level statistics: density, mean degree, clustering coefficient.
  • Recognize the limits of semantic network analysis, what it surfaces and what it misses.

3.1 Building a Co-Occurrence Matrix

The starting point of a semantic network is the feature co-occurrence matrix (FCM). For a vocabulary of v words, the FCM is a v × v symmetric matrix; entry (i, j) is the number of times words i and j appeared in the same context unit. The context unit can be any of several things, and the choice matters:

  • Sentence-level co-occurrence: two words co-occur if they are in the same sentence. Tightest, most likely to reflect direct conceptual association. Produces sparser networks.
  • Paragraph-level co-occurrence: two words co-occur if they are in the same paragraph. Looser; reflects topical co-occurrence over a few sentences of conversation.
  • Document-level co-occurrence: two words co-occur if they appear anywhere in the same transcript. Loosest; suitable for small corpora where finer-grained co-occurrence would be too sparse.
  • Sliding-window co-occurrence: two words co-occur if they are within n tokens of each other. The sliding window is the standard in computational linguistics; n = 5 to 10 is typical for content-word co-occurrence.

For interview transcripts of the length typical of your loneliness corpus (2,000–4,000 tokens per transcript, 30–90 minute interviews), a sentence-level or 5-token-window FCM is the right starting point. Document-level co-occurrence is too coarse: every word would co-occur with every other word in a single transcript, and the resulting network would be uninformative.

What to expect: a network of the 50 most frequent content words in the loneliness corpus (the transcripts are in HSCI_841_loneliness_data.zip), with edges whose thickness reflects co-occurrence weight, places the most central words (loneliness, alone, feel, people, time) in the middle and the more specific words (chair, radio, wednesday, fading, group, screen) on the periphery in clusters. That first picture is a quick look; the centrality measures and community detection in the following sections are the analysis.

3.2 From Co-Occurrence Matrix to Network: Centrality

The co-occurrence matrix is already a network in matrix form: each word is a node, and each non-zero cell is an edge weighted by the co-occurrence count. Network-analysis software reads the matrix directly as a weighted, undirected graph, which gives access to dozens of network statistics, several community-detection algorithms, and layout tools for visualization. The first statistics to compute are four centrality measures, each of which describes a different role a word can play in the network.

  • Degree: a measure of breadth, how many different words this one connects to.
  • Betweenness: a measure of bridging; this word sits on many shortest paths between other words and is the conceptual hinge of the network.
  • Closeness: a measure of accessibility; this word is on average close to everything else and sits in the conceptual middle.
  • Eigenvector: a measure of prestige; this word is connected to other well-connected words and is at the heart of the conceptual structure.

In the loneliness corpus, expect loneliness, alone, feel, and people to score high on most measures, and expect words like tired, quiet, radio, group, and family to score high on degree but lower on betweenness, because they are nodes of specific thematic clusters and do little bridging between clusters.

3.3 Community Detection: Louvain and Beyond

A network community is a sub-set of nodes that are more densely connected to each other than to the rest of the graph. Community detection algorithms partition a graph into communities by optimizing some criterion, most commonly modularity, which measures how much denser the within-community edges are compared to what would be expected if edges were placed randomly. High modularity means a strong community structure; low modularity means the graph is more like a single dense blob.

The Louvain method (Blondel et al., 2008) is the most widely used modularity-optimization algorithm. It is fast (runs in seconds on networks of tens of thousands of nodes), deterministic up to node ordering, and produces interpretable communities of varying granularity. The output is a partition: each node is assigned to one community. Alternative algorithms (Walktrap, Infomap, label propagation) can be applied as cross-checks; in qualitative work, if the same broad communities emerge under two different algorithms, you can be more confident that the partition is robust.

For your loneliness semantic network, the Louvain method will likely identify three to five communities, which often map onto interpretable theme clusters: an “embodied / sensory” cluster (tired, body, sleep, quiet, fading), a “social / structural” cluster (family, friends, group, chat, work), a “coping / activity” cluster (dog, walk, music, radio, church), and perhaps a “meaning / time” cluster (life, years, dead, gone, missing). The communities are candidates for thematic labels, not labels themselves; the labelling is your interpretive work.

What success looks like: modularity above about 0.3 indicates a meaningful community structure, and above 0.5 is strong. Three to five communities is typical for a 30–50 node semantic network from a 20-transcript corpus. When a second algorithm such as Walktrap is run as a cross-check, a Rand index near 1 means the two partitions agree closely; a value near 0 means they disagree, and neither partition should be treated as the authoritative grouping.

3.4 Visualizing the Semantic Network

A semantic-network figure is one of the most striking visualizations available for a qualitative paper. It looks scientific to a quantitative reader, it is genuinely informative to a qualitative reader, and it gives a journal reviewer something concrete to look at. The standard layout algorithms for force-directed graph drawing (Fruchterman-Reingold, Kamada-Kawai, ForceAtlas2) all aim to position nodes so that connected nodes are close and disconnected nodes are far, while avoiding overlap. Choose a layout, color nodes by their community, size nodes by their degree or eigenvector centrality, and label nodes with their words.

What the figure communicates: each colored cluster is a candidate theme. Reading the cluster as a list of words and then returning to the transcripts to find the passages where those words actually co-occurred is how you turn a graph community into an analytic theme. The figure belongs in the findings section of a report.

3.5 Network-Level Statistics

Beyond per-node centralities and per-community partitions, a few network-level statistics characterize the corpus as a whole. Density is the fraction of all possible edges that are actually present; high density means the vocabulary is tightly interconnected. Average path length is the mean shortest-path distance between any two nodes; low average path length means concepts are conceptually close even when they are not directly linked. The clustering coefficient measures the extent to which neighbors of a node are also neighbors of each other; high clustering means the corpus has small densely-connected sub-graphs (theme clusters), low clustering means it is more like a single distributed web.

What a comparison with chance tells you: if your loneliness network has higher clustering than a random graph of the same size and density, with a similar mean distance, it has the “small-world” property of being locally dense and globally short. This is the expected shape for a semantic network of natural language, because the underlying conceptual structure of human discourse is small-world. If your network lacks this property, something has probably gone wrong; often the co-occurrence window is too coarse and the network is fully connected.

3.6 What Semantic Network Analysis Shows You and What It Misses

Semantic networks visualize conceptual co-occurrence. They do not measure what participants meant by the co-occurring words. The word chair connected to the word empty in the network is a finding only if the chair-empty pairs in the corpus are doing the same analytic work; the network statistic does not tell you that. The interpretive work, reading the actual passages in which the words co-occur, remains the analyst's task, exactly as for the KWIC and keyness techniques in an earlier section.

What semantic networks add over those simpler techniques is a synoptic view of the corpus's conceptual structure. The Louvain communities are a kind of unsupervised theme proposal: clusters of words that the corpus itself groups together. When those communities track the themes you developed through hand-coding, you have computational corroboration of your analytic structure. When they diverge from your hand-coded themes, the divergence is itself a finding: either your hand coding missed something the corpus is doing, or the corpus is doing something at the level of vocabulary co-occurrence that does not survive when you read it closely. A related class of fully probabilistic theme-discovery methods, latent Dirichlet allocation (Blei, Ng, & Jordan, 2003; Wikipedia contributors, n.d.-c) and its structural-topic-model extension (Roberts, Stewart, Tingley, et al., 2014), treats each document as a mixture of latent topics and is widely used in the broader computational-text-as-data tradition (Wikipedia contributors, n.d.-d); the co-occurrence-network approach in this lesson is the more interpretively transparent cousin. In plain words, a topic model sorts a corpus's vocabulary into a handful of recurring word-bundles and treats each transcript as a mixture of them, whereas the co-occurrence network here lays the same kind of structure out as a graph you can read node by node.

Reflection

Sketch (in words) what you predict the Louvain communities of your loneliness corpus's semantic network would look like. How many communities would there be? What kinds of words would cluster together? Then think about which of those predicted clusters would correspond to a theme you have already identified through close reading, and which would represent something the close reading missed.

Model answerA defensible prediction names three to five communities, each with a few representative words. Example: “I would expect (a) an embodied / sensory cluster: tired, body, sleep, quiet, fading, ache, slow; (b) a social-structural cluster: family, friends, group, kids, work, partner; (c) a coping / activity cluster: walk, dog, music, radio, church, book; (d) a temporal / loss cluster: years, dead, gone, miss, used-to, before. The first two cleanly map onto themes I have already coded. The fourth probably also does. The third may surface a sub-structure I had treated as a single theme, the difference between embodied solo coping (walk, dog, music) and mediated solo coping (radio, screen, online) might emerge as two sub-clusters within what I had hand-coded as a single 'coping' category. That would be a useful refinement of the codebook.” The point is not to be right; the point is to make a specific, testable prediction that the data can then confirm or refine.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: In a semantic network built from interview transcripts, an edge between two words represents:

Edges in a co-occurrence network represent the simple fact that the two words appeared in the same context unit. The context unit (sentence / paragraph / sliding window) is a design choice that affects how the network looks. Word embeddings and dependency parses are different network constructions, not the standard for the technique in Bernard, Wutich & Ryan, Ch 19.

Question 2: A word with high betweenness centrality in a semantic network is best interpreted as:

Betweenness measures how often a node is on the shortest path between pairs of other nodes. High-betweenness words are conceptual hinges: removing them would disconnect parts of the network. This is distinct from high-frequency, which is captured by degree or strength.

Question 3: The Louvain community-detection algorithm partitions a network by optimizing:

Louvain (and most modern community-detection algorithms) optimize modularity. High modularity means strong community structure; modularity above ~0.3 is meaningful, above 0.5 is strong.
Section 4 of 5

LLM-Assisted Qualitative Coding: Opportunities, Risks, Discipline

⏱ Estimated reading time: 45 minutes
Section 4 of 4

LLM-Assisted Qualitative Coding

Three opportunities, five risks, and the disciplined workflow that makes it a defensible method.

What an LLM is

A plausibility engine

Transcript + Codebook + Prompt Large Language Model Coded output JSON array of (passage, code)

The working model: a plausibility engine that produces outputs looking as fluent whether correct or wrong.

Three opportunities

Three genuine advantages

Speed

Weeks of coding in hours. Enables codebook iteration with rapid turnaround.

Scale

Corpora of thousands, not dozens. New research designs become feasible.

Reproducibility

Fixed prompt + fixed model = approximately stable output. Also reproduces errors exactly.

Five risks

Five serious risks

Hallucination

Fluent, confident output not supported by the source. Indistinguishable in style from correct output.

Prompt drift

Nearly identical prompts can produce divergent outputs. Requires version control.

Opaque reasoning

The model's explanation of its coding is another generated text, not a transparent audit trail.

Bias amplification

Training-data biases mean the model under-recognizes non-Western and underrepresented cultural concepts.

Methodological invisibility

Using a large language model without disclosure. A research-integrity issue that has led to retractions.

The disciplined workflow

The disciplined workflow

Before any upload: check consent and terms of use; de-identify Hand-codecalibration set Write promptrole+codebook+format Apply tocalibration set ComputeKrippendorff α α < 0.6 → iterate prompt α ≥ 0.7 → lock prompt Apply to fullcorpus Audit 30random pairs Disclosein methods
Carry forward

What the large language model cannot do

  • Develop the codebook, that is earlier modules.
  • Interpret findings or write the discussion, interpretive claims are the analyst's.
  • Detect the absence of an expected code, a fundamental limit, not a passing one.
  • Write the positionality statement, only you can.

The three non-negotiable conditions for any large language model use: disclose (model, version, date, prompt); validate (Krippendorff's alpha against hand-coded reference); audit (verbatim-quote hallucination check).

Introduction and Overview

Bernard, Wutich, and Ryan's second edition was published in 2017. ChatGPT was released to the public in late 2022. Claude was released in 2023. Llama-class open-weight models capable of running locally on a laptop arrived in 2023–2024. The most consequential recent technical development in qualitative methods since Bernard, Wutich, and Ryan’s second edition went to press, the arrival of large language models capable of reading a transcript and applying a codebook to it in seconds, is not covered in Bernard, Wutich, and Ryan (2017). This section is your introduction to it.

The methodological literature on LLM-assisted qualitative coding has grown rapidly since 2023. Gilardi, Alizadeh, & Kubli (2023) showed that ChatGPT outperformed crowdworkers on several annotation tasks at a fraction of the cost. Ziems et al. (2024) systematically benchmarked LLMs on computational social science tasks and concluded that they perform credibly on many classification and extraction tasks but unevenly on tasks requiring deep cultural knowledge. Bail (2024) argues that generative AI has real potential across social-science methods while flagging problems that accuracy benchmarks alone do not surface: bias inherited from training data, and further challenges of ethics, replication, environmental cost, and a proliferation of low-quality research. Non-disclosure of AI assistance, which this section calls methodological invisibility, is a related research-integrity problem that emerging journal policies increasingly address. Bender, Gebru, McMillan-Major, & Shmitchell (2021) provide the field's most-cited articulation of the ethical and epistemic risks of treating large language models as authoritative knowledge sources.

This section's purpose is to give you a disciplined working stance: LLMs are a tool, like Taguette or NVivo; they require validation, disclosure, and human oversight; they are not a replacement for analytic judgment. We walk through what LLMs can and cannot do, what the risks are, what the methodological discipline looks like, and how to actually use an LLM to apply a codebook to a transcript with the validation that any defensible qualitative method requires.

Learning Objectives for this section

  • Define what an LLM is, in plain terms, at the level needed to use it responsibly as a research tool.
  • Identify the three main opportunities of LLM-assisted coding (speed, scale, reproducibility) and the five main risks (hallucination, prompt drift, opaque reasoning, bias amplification, methodological invisibility).
  • Specify a calibration-and-validation workflow for LLM coding using a hand-coded reference set and an agreement statistic.
  • Write a prompt that applies a codebook to a transcript reliably.
  • Audit LLM outputs against the source transcripts and detect hallucination.
  • Disclose LLM use defensibly in a methods section.
  • Confirm that consent, terms of use, and privacy law permit LLM processing of transcripts, and de-identify them before upload.

4.1 What an LLM Is, in Plain Terms

Background

This box summarizes the mechanism at the level a researcher needs. HSCI 241 Lesson 6, Sections 1.2 to 1.4, give a fuller account of the mechanism and of retrieval-augmented generation (optional reading).

A large language model (LLM) works with tokens, short pieces of text such as a word, part of a word, or a punctuation mark. In training, it reads a very large body of text and learns to predict the next token from the tokens before it; developers then tune it on instructions and human ratings so that it follows requests. When it answers, it repeats one step many times: it assigns a probability to every possible next token, selects one, adds it to the text, and calculates again. A setting called temperature controls how much randomness enters that selection, which is one reason the same prompt can return different outputs. What the model has learned stops at its knowledge cutoff, the date after which its training text ends.

The working model to hold is that an LLM is a plausibility engine: it produces text that is a plausible continuation of the prompt. When a task is well specified, plausible output is often correct. When it is not, the model can produce fluent, confident text that the source does not support, such as an invented quotation attributed to a participant. This failure is called hallucination. It is the central risk of LLM-assisted coding, and Section 4.3 returns to it.

For qualitative coding, the relevant capability is that contemporary LLMs (Claude 3.5/4, GPT-4/5, Llama 3+ and similar) can read a several-thousand-token transcript and a codebook of a few dozen codes, and produce a coded version of the transcript, either as a list of (passage, code) pairs or as an annotated copy of the original text. They do this in seconds. A 20-transcript codebook application that would take a careful human coder two weeks of focused work runs in under an hour.

A note on terminology

“Generative AI”, “LLM”, “foundation model”, and “chatbot” are sometimes used interchangeably in the methodological literature but they mean different things. Generative AI is the broadest term, including image and audio generation as well as text. LLM is specifically a text-generation model trained on natural language. Foundation model is the most general term for a large pre-trained model adaptable to multiple downstream tasks. Chatbot is a user-facing application built on top of an LLM (ChatGPT, Claude.ai, Gemini). In a methods section you should be specific: name the model and version (e.g., “Claude 3.5 Sonnet, accessed via the Anthropic API in November 2025”), not the umbrella category.

4.2 The Three Opportunities

The case for LLM-assisted coding is built on three claims. None are absolute; each requires the disciplinary work that the rest of this section is about.

Speed. A codebook of 30 codes applied to a 20-transcript corpus by a human team takes weeks. The same application by an LLM takes minutes to hours, depending on whether you process one transcript at a time or run them in parallel. For a research program where the bottleneck has been the cost of human coder time, this is transformative. It enables analyses of corpora that were previously infeasible, and it enables iteration on the codebook with feedback turnaround in minutes rather than weeks.

Scale. Closely related but distinct. With human-only coding, the size of the corpus an academic team can analyze is bounded by the team's labor budget. A doctoral student plus a research assistant might analyze 30–60 transcripts in a year. The same team using an LLM as a coding assistant can scale to thousands of transcripts, with the caveats below. This unlocks new research designs (longitudinal corpora, multi-site studies, social-media analyses) that simply did not have a human-coding answer.

Reproducibility. An LLM applied with a fixed prompt to a fixed text, at fixed inference parameters (temperature, top-p, seed), produces approximately the same output every time. Two researchers running the same prompt against the same model and the same data should get nearly identical codings, nearly, not exactly, because some randomness remains in commercial APIs and because the underlying model is occasionally re-versioned. The reproducibility is, in this narrow sense, better than the reproducibility of two human coders, who never produce identical codings even with the best calibration. This is a real methodological advantage, but it cuts both ways: the model also reproduces its mistakes exactly. A systematically wrong human coder can be caught by a co-coder noticing the disagreement; a systematically wrong LLM has no internal co-coder. The reproducibility claim holds only when the model version, the date of access, the prompt, and the inference parameters are recorded, as in the AI-assisted search log of HSCI 241 Lesson 6, Section 3.4 (optional reading). Consumer chat interfaces give the analyst less control over those parameters than an API does, so the claim is weaker for work done through a chat window.

4.3 The Five Risks

The five risks below are not arguments against using LLMs. They are arguments against using them carelessly. Each risk has a known mitigation; the disciplined LLM-assisted workflow in Section 4.4 builds on those mitigations.

Risk 1: Hallucination. The model produces fluent, confident output that is not supported by the source text. For coding tasks, this most often manifests as the LLM inventing a quote attributed to a participant, or applying a code to a passage that does not contain the content the code names. Hallucination is the central methodological risk. A coding output where 8 of 100 codes are applied to invented or misread passages will look as professional as a coding output where all 100 codes are correct, because LLMs produce uniformly fluent prose either way. The only mitigation is auditing: read a random sample of the LLM's codes back against the source transcripts.

Risk 2: Prompt drift. Two prompts that look almost identical can produce divergent outputs. The difference between “Apply the codebook below to the transcript” and “Identify which codes from the codebook apply to each passage in the transcript” is small to a human reader and material to the LLM. Prompt drift becomes a methodological risk when the prompt is being iteratively refined during a study without version control: the codings produced under prompt v2 are not directly comparable to the codings produced under prompt v1. The mitigation is to fix the prompt before doing the production run and to record both the prompt and the model version in the methods section.

Risk 3: Opaque reasoning. When a human coder applies a code, you can ask them why and they can give a defensible answer. When an LLM applies a code, the “reasoning” it can articulate is itself another generated text and may or may not reflect the actual computation that produced the code assignment. Asking the model “why did you apply code X here?” produces a plausible-sounding rationalization that you cannot independently verify. This is a real epistemic limit and you should be honest about it in the methods section: the LLM is a black box that produces outputs; the audit, not the model's self-explanation, is what licenses you to trust an output.

Risk 4: Bias amplification. LLMs reflect the biases in their training data. For most qualitative coding tasks, the relevant biases are the cultural and linguistic biases of contemporary English-language web text. The model may under-recognize concepts from non-Western or non-English knowledge traditions; it may over-recognize concepts that are heavily represented in the training distribution; it may apply codes with subtle valence shifts that map onto the dominant culture's framing of the phenomenon. For studies with participants whose worldviews are not well-represented in the training data (Indigenous communities, recent immigrants, queer / trans participants, religious minorities), this risk is particularly acute. The mitigation is a high-quality hand-coded calibration set drawn from the actual study population, not a benchmark from another study.

Risk 5: Methodological invisibility. A researcher uses an LLM to do some or all of the coding, does not disclose it, and presents the resulting analysis as if it were hand-coded. This is a research-integrity issue, not a technical one, but it is widespread. Several recently retracted papers used LLMs without disclosure. The mitigation is straightforward and is required by emerging journal policies: disclose every use of an LLM in the methods section, including the model name and version, the date of access, the prompt or prompts, and the validation procedure. Treat LLM use the way you treat any other analytic tool: name it, validate it, and report the validation.

This course's stance, stated plainly

In this course's view, LLMs are admissible in qualitative analysis with three non-negotiable conditions, which apply once the data-governance check that opens the workflow in Section 4.4 has been passed. (1) You disclose every LLM use in the methods section, with model name, version, date, and prompt. (2) You validate every LLM-coded analysis against a hand-coded reference set drawn from your own data, and you report the validation statistic (Krippendorff's alpha or similar). (3) You audit a random sample of LLM outputs against the source transcripts for hallucination, and you report the audit. An analysis that does not meet all three conditions is not admissible. An analysis that meets all three is treated like any other defensible qualitative-methods choice.

The course's deeper position is that LLMs are neither a savior nor a threat. They are a tool that handles a specific kind of work (large-scale, well-specified, repetitive coding) better than humans, and a different kind of work (interpretive judgment, theoretical synthesis, ethical reasoning about cases) substantially worse than humans. Treating them as the right tool for the right tasks, with validation, is what disciplined practice looks like.

4.4 The Disciplined Workflow

The following workflow is the methodological default for LLM-assisted coding in this course. It is not the only defensible workflow, but it is the simplest one that satisfies the three non-negotiable conditions above.

  1. Confirm that data governance permits the processing. Check that the participants' consent and the dataset's terms of use allow the transcripts to be processed by a third-party model. Apply TCPS 2 and British Columbia privacy law, including the rules on where personal information may be stored and processed, de-identify the transcripts before upload, and prefer an institutionally approved or locally run model. HSCI 207 Lesson 10, Section 4.4, and HSCI 241 Lesson 6, Section 3.5, cover the same questions for automated transcription and for evidence work (optional reading).
  2. Hand-code a calibration set. Take a stratified sample of 3–5 transcripts from your corpus and code them by hand, the way you have done in earlier modules. This is your reference set. The codebook you use must be the same codebook you will apply with the LLM; the reference codings are what you will validate the LLM against.
  3. Write the prompt. The prompt has three components: (a) a role / task statement, (b) the full codebook with definitions, (c) instructions for the output format. See Section 4.5 for a worked example.
  4. Apply the LLM to the calibration set. Run the prompt against each transcript in the calibration set. Save the outputs.
  5. Compute agreement against the hand-coded reference. For each passage in the calibration set, compare the human code to the LLM code. Compute Krippendorff's alpha or Cohen's kappa across the set. If agreement is low (alpha < 0.6), iterate on the prompt. If agreement is acceptable (alpha ≥ 0.7), proceed. The thresholds are debated but 0.7 is the conventional acceptable-agreement floor for nominal data.
  6. Document the prompt-validation history. Record every prompt iteration in your audit trail along with the resulting agreement statistic. The final accepted prompt is the one you report.
  7. Apply the accepted prompt to the full corpus. Run the LLM on the remaining transcripts.
  8. Audit a random sample for hallucination. Take 30 randomly sampled (passage, code) pairs from the LLM outputs. For each, look up the passage in the source transcript and confirm that (a) the passage exists as quoted, (b) the assigned code is plausibly supported. Report the audit results in the methods section.
  9. Disclose. The methods section names the model and version, the prompt, the calibration sample size, the agreement statistic, the audit result, and the governance steps taken (Section 4.9). An appendix carries the full prompt verbatim and the codebook.

4.5 A Sample Prompt That Applies a Codebook

The prompt below is a template that has been tested across several qualitative studies. Adapt the codebook section to your specific codebook; the surrounding scaffolding is what makes the output auditable.

PROMPTSample prompt: applying a loneliness codebook to a transcript

This is the prompt you would paste into Claude, ChatGPT, or any equivalent LLM interface. The pattern is platform-agnostic; you can also send the same prompt programmatically through the provider’s API.

# --- ROLE AND TASK ---
You are a qualitative-research coding assistant. You will receive a codebook
and a transcript of an interview about loneliness. Your task is to apply the
codebook to the transcript, identifying every passage that matches one or more
codes.

# --- CODEBOOK ---
The codebook has 12 codes. Apply them strictly according to the definitions
below. Do not invent new codes. Do not modify the codes.

1. somatic-absence  -- Reference to a physical place or object whose meaning
   is the absence of a specific person (e.g., the empty chair, the unused
   side of the bed, the silent kitchen).

2. embodied-loneliness  -- Reference to a bodily sensation of loneliness:
   tiredness, ache, cold, weight, hollow, fading. Bodily as opposed to
   affective.

3. affective-loneliness  -- Reference to a feeling of loneliness in affective
   terms: sad, empty, low, isolated, longing.

4. social-shrinkage  -- Reference to the participant's social world becoming
   smaller over time, either through death, mobility loss, or other
   structural change.

5. mediated-connection  -- Reference to maintaining connection through
   technology, phone, or other mediation rather than in person (group chat,
   WhatsApp, phone-pal, video call, social media).

6. embodied-coping  -- Reference to a coping strategy that involves the body
   or an activity: walking, the dog, exercise, music, gardening.

7. relational-coping  -- Reference to a coping strategy involving another
   person or relationship: friends, family, partner, neighbor, professional.

8. loneliness-as-cost-of-love  -- Statement that loneliness is the price of
   having loved (typically older participants reflecting on widowhood or
   long-loss).

9. choosing-solitude  -- Statement that the participant has chosen to be
   alone and that the alone-ness is desired rather than imposed.

10. structural-cause  -- Attribution of loneliness to a structural cause:
    immigration, housing, work conditions, the pandemic, ageism.

11. internal-cause  -- Attribution of loneliness to an internal cause:
    personality, mental-health history, neurodivergence, self-criticism.

12. existential-meaning  -- Statement that the loneliness is meaning-making
    or meaning-giving: solitude as part of the human condition, loneliness
    as teacher.

# --- OUTPUT FORMAT ---
Return a JSON array. Each element of the array is an object with three keys:
  "passage":     the exact verbatim passage from the transcript, no edits
  "code":        the code from the codebook above, exact name
  "rationale":   one short sentence explaining why this code applies

Rules:
- Use only the codes above. If a passage matches no code, do not include it.
- A passage may receive more than one code. Return one JSON object per code
  per passage (so a doubly-coded passage produces two objects with the same
  "passage" but different "code" values).
- Do not summarize, paraphrase, or shorten passages. Quote verbatim.
- Do not invent passages. If you are uncertain whether a passage is verbatim,
  do not include it.
- Return only the JSON array. Do not include any other text.

# --- TRANSCRIPT ---
[Paste the transcript here, including the participant ID and metadata header]

Why this scaffolding is necessary: The verbatim-quote requirement and the JSON output format make the LLM's output mechanically auditable. You can take any output passage, search for it in the source transcript, and verify it exists. If it does not exist, the LLM has hallucinated, and you have caught it. Without these structural constraints, the audit step becomes much harder.

4.6 The Worked Example: 5 Loneliness Transcripts, Hand vs LLM

The following walkthrough runs the workflow end-to-end on the loneliness interview dataset. The numbers come from the hand and LLM coding files supplied with the dataset and are illustrative of what the workflow produces.

StageWhat you doTypical result
1. Hand code 5 transcripts P01 Maya, P05 Linda, P11 Helen, P15 Amira, P20 Frank. Apply the 12-code loneliness codebook in Taguette. 130 passages coded in the shipped file wk12_hand_codings.csv, one code each; a passage that matched two codes was resolved to the more specific one.
2. Apply the LLM prompt to the same 5 transcripts Run the prompt in Section 4.5, copying each transcript in turn. Save the JSON outputs. 144 codings in wk12_llm_codings.csv: 126 of the hand-coded passages, plus 18 the hand coder did not mark.
3. Align passages between hand and LLM For each LLM passage, find the closest hand-coded passage (overlapping text). Build a paired table. 88% of LLM codings match a hand-coded passage; the other 12% are LLM-only (some legitimate, some hallucinated).
4. Compute Krippendorff's alpha Treat the human and LLM as two coders; compute alpha on the matched passages. First prompt: alpha = 0.62. Below threshold; iterate.
5. Iterate the prompt Reading the disagreement matrix: LLM over-applies structural-cause; under-applies somatic-absence. Tighten the definitions in the prompt; add 2 worked examples for each problematic code. Second prompt: alpha = 0.81 on the shipped files. Acceptable. Lock the prompt.
6. Apply the locked prompt to the remaining 15 transcripts Run on P02–P04, P06–P10, P12–P14, P16–P19. ~430 (passage, code) pairs across the 15 transcripts.
7. Audit 30 random pairs for hallucination For each, verify the passage exists verbatim in the source. Typical: 27–30 are verbatim. The shipped LLM file contains three paraphrases, so the audit here returns 27 of 30 (0.90). Report the rate.
8. Methods section paragraph Disclose: model + version + access date; prompt in appendix; alpha; audit result. One transparent paragraph that a reviewer can evaluate.

The hand and LLM codings used in the table, outputs/wk12_hand_codings.csv (130 hand-coded passages) and outputs/wk12_llm_codings.csv (144 LLM codings), are in the course data bundle, HSCI_841_loneliness_data.zip. Once the matched passages are laid out as two columns of codes, one for the hand coder and one for the LLM, any reliability calculator returns Krippendorff’s alpha for nominal codes, and Cohen’s kappa on the complete pairs gives a direct comparison. Cross-tabulating the two columns gives a confusion matrix.

What to do with the confusion matrix: the diagonal counts are agreements and the off-diagonal counts are disagreements. Reading the rows and columns of the largest off-diagonal counts identifies the code pairs the LLM systematically confuses (for example, embodied-loneliness and affective-loneliness, or structural-cause and internal-cause). Those confusions are what to address in the next prompt iteration: tighten the definitions, add worked examples, and re-run the agreement check.

4.7 Auditing for Hallucination

The single most useful audit you can run on LLM-coded data is the verbatim-quote check. The prompt in Section 4.5 instructed the model to return passages verbatim; the audit verifies that the model obeyed. The procedure is simple: take the LLM output, randomly sample 30 (passage, code) pairs, and for each one, search the source transcript for the passage. If the passage exists as quoted, the audit passes for that pair. If it does not exist, the model has hallucinated.

The audit needs no special software. Draw 30 rows at random from the LLM output, open each source transcript, search for the quoted passage, and record whether it appears verbatim. What success looks like: a verbatim rate above 95% is typical for well-prompted contemporary LLMs on transcripts of the length and style you are working with. A rate below 80% means the prompt is failing to enforce the verbatim constraint and should be re-engineered, and you should also consider whether the model version you are using is appropriate for the task. Report the verbatim rate in your methods section as the headline hallucination-audit statistic, and keep the audit log for the appendix.

4.8 What an LLM Cannot Do

Closing this section honestly requires naming what the LLM workflow does not replace. The LLM does not develop the codebook (you do, in earlier modules). It does not interpret the findings (you do, in the discussion). It does not write the positionality statement (only you can). It does not decide which themes are central to the paper's argument (you decide). It is not a co-author. It is a coding assistant whose output you validate, audit, and take responsibility for. The methods section reports its role; the paper's interpretive claims are yours.

One specific limitation that is often missed: the LLM is poor at recognizing the absence of a code. Hand coders are trained to notice when a participant says something that conspicuously does not deploy a code that the rest of the corpus deploys frequently. That absent-pattern detection is part of why qualitative analysis is interpretive. LLMs do not currently do it well; they apply codes to passages but rarely flag passages-where-an-expected-code-would-go. This is a fundamental rather than passing limitation, and the appropriate response is to do the absence-checking yourself, by hand, after the LLM has done the routine application work.

4.9 The Disclosure Paragraph in a Methods Section

Here is a template disclosure paragraph that satisfies the three non-negotiable conditions of Section 4.3. It also reports the fields that HSCI 241 Lesson 6, Section 3.6, takes from the 2025 joint position statement that endorses the RAISE recommendations on AI use in evidence synthesis (both optional reading): the purpose and the stages affected, the justification for the tool, the judgements not delegated to it, related interests, and limitations. To these it adds the fields specific to LLM-assisted coding (model, version, access date, prompt, calibration sample, agreement statistic, and audit result) and a governance sentence stating where the model ran, what de-identification preceded upload, and that consent and the dataset's terms of use permitted the processing. Adapt the numbers, the bracketed details, and the model name to the actual use.

Methods section disclosure (template)

“Codebook application across the 20-transcript corpus was assisted by Claude 3.5 Sonnet (Anthropic), accessed via the API in November 2025, at the coding stage only, to apply the codebook reported in Appendix A to the transcripts. I chose the tool because it returns verbatim passages in a structured format that can be audited against the source transcripts. The transcripts were de-identified before upload, the model ran on [the provider's servers in country, or an institutionally approved or local installation], and the participants' consent and the dataset's terms of use permitted processing by a third-party model. The full prompt is provided in Appendix B. The prompt was calibrated against a hand-coded reference set of five transcripts (P01, P05, P11, P15, P20), which I coded in Taguette using the codebook reported in Appendix A. After two prompt-iteration cycles, Krippendorff's alpha between the human reference codings and the LLM codings on the calibration set reached 0.78 (n = 142 passages; nominal codes); both the iterations and the final agreement are documented in the audit trail (Appendix C). The locked prompt was then applied to the remaining 15 transcripts. A random sample of 30 LLM-coded (passage, code) pairs was audited against the source transcripts to detect hallucination; 29 of 30 passages were verbatim, yielding a verbatim rate of 96.7%. The one non-verbatim case is documented and discussed in Appendix C. The LLM did not develop the codebook, identify themes, or make interpretive decisions; all interpretive claims in the findings and discussion are mine, and the LLM functioned as a coding assistant only. I have no financial or other relationship with the model's provider. Limitations of this use include the model's weaker recognition of concepts that are underrepresented in its training data and the variation of commercial models over time, which means the coding cannot be reproduced exactly.”

Reflection

Suppose you are writing the methods section of a qualitative study of the loneliness interviews and must disclose your use (or non-use) of an LLM. Sketch the disclosure paragraph. If you used an LLM, what was the prompt-iteration history, what was the validation agreement statistic, and what was the hallucination audit result? If you did not use one, why not, was it a methodological choice, an access choice, or something else? Either answer is defensible if it is reflective.

Model answerA defensible used-LLM answer specifies the model and version, names a prompt-iteration history with at least two cycles, names the final agreement (e.g., Krippendorff's alpha around 0.75–0.80), and names the audit result (e.g., 95% verbatim). It also acknowledges what the LLM did not do: it did not develop the codebook, it did not interpret the themes, it did not write the discussion. A defensible did-not-use-LLM answer names a real reason: the small corpus size made hand coding feasible and analytically richer; the population (e.g., refugee participants whose worldview is poorly represented in LLM training data) made bias amplification an unacceptable risk; the time required to validate the LLM exceeded the time saved on a 20-transcript corpus. Both stances are admissible. What is not admissible is a stance that does not engage the question, quietly using an LLM and not disclosing, or refusing to use one without a stated reason. A methods section needs the analyst on the record.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: The central methodological risk of LLM-assisted qualitative coding is:

Hallucination is the central risk because LLM outputs are uniformly fluent regardless of correctness. An 8%-hallucination output looks as professional as a 0%-hallucination output. The mitigation is auditing, reading a random sample of outputs against the source transcripts.

Question 2: This course's stance on LLM use specifies three non-negotiable conditions:

The three conditions are disclosure, validation against a hand-coded reference, and a hallucination audit. The course is neither against LLM use nor uncritical of it; the position is that LLMs are a tool that requires the same kind of methodological discipline as any other.

Question 3: In the disciplined LLM-assisted workflow, the calibration step compares LLM codings to:

The reference set must come from the same study (the same codebook applied to the same kind of transcripts) because it is the only basis for testing whether the LLM has captured the codebook's specific operationalizations. Benchmarks from other studies cannot substitute.
Section 5 of 5

Final Assessment & Course Completion

⏱ Estimated time: 30 minutes

Bringing It All Together

This lesson is the final lesson of this course. The five sections of this module, KWIC and keyness (an earlier section), cultural domain analysis (an earlier section), semantic network analysis (an earlier section), LLM-assisted coding (an earlier section), and this final assessment (this section), complete the methodological arc that began in an earlier lesson with the operational definition of qualitative data analysis. The arc has carried you through theory and literature (an earlier lesson), sampling (an earlier lesson), data collection (an earlier lesson), themes and codebooks (an earlier lesson), analytic frameworks (an earlier lesson), constant-comparative grounded theory (an earlier lesson), content analysis (an earlier lesson), schema and narrative analysis (an earlier lesson), discourse analysis (an earlier lesson), analytic induction and QCA (an earlier lesson), and now into the computational and machine-assisted methods of this last lesson.

A qualitative study report is where these methods come together. A defensible report takes a qualitative dataset, asks a defensible research question, chooses appropriate methods, applies them transparently, produces findings that are evidenced and interpretive, and is written up in journal-article format with its codebook, audit trail, and positionality statement. If, two years from now, you are submitting a qualitative public-health paper to a journal, the analytic moves in that paper can be the ones first practised in this course on the loneliness interview dataset.

Key Takeaways from this lesson

  • Computational text analysis is a front-loader for close reading, not a replacement for it. KWIC, frequency, TF-IDF, keyness, collocations, and n-grams surface candidates; the close reading turns the candidates into findings.
  • KWIC is the oldest computational text technique and remains analytically valuable because it preserves local context, enabling rapid sense-disambiguation that pure numerical statistics cannot deliver.
  • TF-IDF surfaces document-distinctive vocabulary; keyness surfaces subcorpus-distinctive vocabulary. Both are versions of comparing focal vocabulary against a background.
  • Cultural domain analysis (Bernard, Wutich & Ryan Ch 18) measures shared mental models with free listing (Smith's salience), pile sorts (with MDS), triad tests, and consensus analysis (Romney-Weller-Batchelder).
  • The eigenvalue ratio test is the diagnostic for the single-culture assumption of consensus analysis: a ratio ≥ 3 supports a shared model; a low ratio indicates multiple sub-cultures.
  • Semantic network analysis treats text as a co-occurrence graph; centrality measures (degree, betweenness, closeness, eigenvector) and community detection (Louvain) reveal implicit conceptual structure.
  • LLM-assisted coding is admissible under three non-negotiable conditions, disclosure, validation against a hand-coded reference set, and a hallucination audit, once a data-governance check has confirmed that consent, terms of use, and privacy law permit the processing.
  • The central risk of LLM coding is hallucination; the mitigation is a verbatim-quote constraint in the prompt plus a random-sample audit against the source transcripts.
  • A qualitative study report in journal-article format brings the course together: a methods section that meets the systematic-transparent-replicable standard of an earlier lesson, with the codebook, audit trail, and positionality statement in its appendices.

Core Concepts Reviewed

An earlier section: KWIC concordances; word frequencies; type-token ratio and lexical diversity (MATTR, MSTTR); TF-IDF; keyness (chi-squared and log-likelihood); collocations; n-grams; the principle that computational text statistics are pointers for close reading.

An earlier section: Free listing and Smith's salience; pile sorts and successive pile sorts; multidimensional scaling (MDS); triad tests; the Romney-Weller-Batchelder consensus model; competence scores; the eigenvalue ratio test for the single-culture assumption; applications in public-health domains.

An earlier section: Term co-occurrence matrices (FCM); context units (sentence, paragraph, window); the four centralities (degree, betweenness, closeness, eigenvector); community detection (Louvain, Walktrap, modularity); network-level statistics (density, mean path length, clustering coefficient); the small-world property of semantic networks.

An earlier section: LLMs as plausibility engines; the three opportunities (speed, scale, reproducibility); the five risks (hallucination, prompt drift, opaque reasoning, bias amplification, methodological invisibility); data governance before upload; the disciplined workflow (calibrate, prompt, iterate, lock, apply, audit, disclose); the prompt template; agreement statistics for human-LLM comparison; the hallucination audit; the disclosure paragraph.

The final reflection below is the last reflection of the course. It is forward-looking. It asks what you carry from this course into your next piece of qualitative research, the one not assigned by a syllabus.

Reflection

Looking back across the course, name one method or stance from it that you can imagine actually using in your next piece of work, the one not assigned by a course. What is the method or stance, what kind of question or data would it be answering, and what do you need to be able to do it well that you did not know how to do before this course?

Model answerA strong closing reflection names something specific and connects it back to the trained capability. Examples: “Constant-comparative grounded theory (an earlier lesson) is the method I expect to actually use, on a planned study of why nurses leave acute-care wards. What I now know how to do that I did not before is develop a codebook in dialogue with the data rather than imposing one and writing memos that themselves become part of the analytic record.” Or: “The disciplined LLM-coding workflow (this lesson) is the method I expect to use most, because most of my future studies will involve corpora too large to hand-code. What I now know how to do is validate an LLM against a hand-coded reference set and audit for hallucination; what I did not know before this course is that doing this is methodologically required.” Or: “The stance, transparency about positionality, is the durable thing. I now know how to write a methods section that declares who I am to the analysis, and that is something I will carry into every paper, qualitative or quantitative, for the rest of my career.” The point is to leave the course with a specific, articulable carry-forward. Vague answers (“I learned to be more reflective”) miss the opportunity that this last reflection is offering.

Minimum 30 characters required.

✓ Reflection saved

Final Knowledge Assessment

Final Assessment for this lesson: Computational Text Analysis, Cultural Domain Analysis & LLM-Assisted Coding (15 Questions)

Question 1: Which of the following best characterizes the analytic role of KWIC, frequency, TF-IDF, keyness, and collocations in a qualitative study?

Computational text statistics are pointers for close reading. The numerical output is the front edge of the analysis, not the finding. Bernard, Wutich, and Ryan are explicit on this and the course adopts the same stance.

Question 2: A keyness analysis comparing older and younger participants in the loneliness corpus identifies radio, fading, visit, and walker as words distinctive of the older subcorpus. Which of the following is the most defensible analytic next step?

Keyness words are pointers. The next step is to read the words in their KWIC contexts and develop interpretive claims supported by the passages, not by the test statistic alone.

Question 3: Smith's salience for an item in a free-listing study combines two pieces of information:

Smith's S = (∑((L − R + 1) / L)) / N, combining how many participants mention the item with how early they mention it. The highest-salience items are common AND early.

Question 4: The Romney-Weller-Batchelder consensus model's eigenvalue ratio test diagnoses:

The first/second eigenvalue ratio of the participant agreement matrix tests the single-culture assumption. A ratio ≥ 3 supports one shared model; lower ratios suggest sub-cultures, requiring stratified analysis.

Question 5: A pile-sort study produces an aggregate similarity matrix that is then analyzed with multidimensional scaling. The 2-D MDS map produces:

MDS produces an unlabeled map. The interpretation of the dimensions is the analyst's interpretive contribution and must be argued, not assumed from the algorithm.

Question 6: In a semantic network built from interview transcripts, an edge between two words represents:

Edges in a co-occurrence semantic network are simple co-occurrence within a context unit. The context unit is a design choice with real consequences for how the network looks.

Question 7: A word with high betweenness centrality in a semantic network is best interpreted as:

Betweenness measures how often a node lies on the shortest paths between pairs of other nodes. High betweenness identifies bridges between communities, not raw frequency.

Question 8: The Louvain community-detection algorithm partitions a network by optimizing:

Louvain optimizes modularity. Modularity > 0.3 indicates meaningful community structure; > 0.5 is strong.

Question 9: The central methodological risk of LLM-assisted qualitative coding is:

Hallucination is the central risk because LLM output is uniformly fluent regardless of correctness. The mitigation is a verbatim-quote constraint in the prompt and a random-sample audit against the source transcripts.

Question 10: This course's stance on LLM-assisted coding specifies three non-negotiable conditions:

The three conditions are disclosure, validation, and audit. The course's position is that LLMs are admissible with discipline; they are not admissible without it.

Question 11: In the disciplined LLM-coding workflow, the calibration reference set should consist of:

The reference set must come from the same study because it tests whether the LLM has captured the codebook's specific operationalizations. Benchmarks from elsewhere cannot substitute.

Question 12: The Krippendorff's alpha (or Cohen's kappa) threshold conventionally accepted as the floor for adequate human-LLM agreement in qualitative coding is approximately:

0.7 is the conventional floor for acceptable agreement on nominal data; 0.8 is preferred. Below 0.6 is generally not adequate and requires further prompt iteration or methodological revision.

Question 13: A verbatim-quote audit on 30 randomly sampled LLM-coded passages from a 20-transcript corpus shows that 29 of 30 passages are verbatim. The most defensible report in the methods section is:

Disclosure is unconditional. Always report the audit numbers (even, especially, when they are good) and locate the failure cases in an appendix so reviewers can see exactly what failed.

Question 14: Which of the following is NOT one of the five risks of LLM-assisted coding identified in this lesson?

The five risks are hallucination, prompt drift, opaque reasoning, bias amplification, and methodological invisibility. Computational cost is a practical consideration but not one of the methodological risks identified.

Question 15: According to this lesson, a qualitative study report in journal-article format includes:

A journal-article report of a qualitative study has the standard sections plus appendices that make the analysis auditable: the codebook, an audit-trail summary, and a positionality statement.
✦ Complete the final reflection above before submitting

Course complete

You have finished this lesson: Computational Text Analysis, Cultural Domain Analysis & LLM-Assisted Coding, the final lesson of the course.

Across twelve lessons you have built up an analytic vocabulary that runs from the operational definition of QDA (an earlier lesson), through theory and literature (an earlier lesson), sampling and data collection (earlier lessons), themes and codebooks (an earlier lesson), conceptual models (an earlier lesson), grounded theory (an earlier lesson), content analysis (an earlier lesson), schema and narrative analysis (an earlier lesson), discourse analysis (an earlier lesson), analytic induction and QCA (an earlier lesson), and the computational and LLM-assisted methods of this final lesson. Each of those toolkits was applied to the loneliness interview dataset. You now have an analytic kit you can carry into your next study.

If you use an LLM at any stage of a future study, the governance step in Section 4.4 comes first, and the disclosure paragraph in Section 4.9 is a template for reporting it.

Treat the course as the beginning of a longer engagement with qualitative research, not the end. Bernard, Wutich, and Ryan is a reference text; keep your copy. Re-read it as you encounter new methodological problems in your own work. The reference text is denser than any single course can make legible; you will get more out of it on the second pass than on the first.

Reference

Glossary: Key Terms, People & Concepts

📚 Reference page: available throughout the lesson

This glossary collects the key concepts, people, and computational tools introduced in this lesson. Use it as a reference while you work through the material, or as a review before the final assessment. Type in the search box to filter entries.

Computational Text Analysis
KWIC (Key Word In Context) A concordance technique that displays each occurrence of a target word with a window of surrounding context (typically 5–10 words on each side). KWIC makes local context inspectable cheaply and is the standard bridge between computational pattern-finding and close reading.
Word Frequency The raw count of how often each word (or stem) appears in a corpus, usually after tokenization, lowercasing, and stop-word removal. Useful as a starting descriptive but dominated by globally common words; rarely the final analytic output.
TF-IDF Term-Frequency × Inverse-Document-Frequency. A weighting that down-weights words that appear in many documents (so “the,” “and,” and topic-specific common words like “loneliness” get demoted) and elevates words that are distinctive to specific documents. Useful for surfacing document-distinctive vocabulary.
Keyness A statistic (typically log-likelihood or chi-squared) comparing word frequencies between a target subcorpus and a reference subcorpus. Words with high keyness are over-represented in the target. Keyness is a pointer to passages worth close reading, not a finding in itself.
Collocations Word pairs that appear together more often than chance would predict, identified with statistics like log-likelihood, MI (mutual information), or t-score. Collocations surface multi-word units (e.g., “social isolation,” “mental health”) that single-word frequency misses.
N-grams Contiguous sequences of n tokens. Bigrams (n=2) and trigrams (n=3) are the most common in qualitative analysis; they capture short phrases that single-word tokenization destroys.
Corpus / Subcorpus A corpus is the full body of texts under analysis (e.g., the 20 loneliness interview transcripts). A subcorpus is a defined subset (e.g., transcripts from participants aged 65+) used for comparative analyses such as keyness.
Document-Feature Matrix (DFM) A sparse matrix with one row per document and one column per feature (typically a word or n-gram), cells containing counts. All downstream computations (frequencies, TF-IDF, keyness, network construction) are built from the DFM.
Cultural Domain Analysis
Cultural Domain Analysis (CDA) A family of structured-elicitation techniques (free listing, pile sorts, triad tests, rankings) for mapping the shared mental categories a group uses to organize a topic. Developed in cognitive anthropology and adopted in public health for codebook development and intervention design.
Free Listing & Smith's S A free list asks participants to list all the items in a domain (e.g., “all the things that make someone feel lonely”). Smith's S combines frequency (how many participants list an item) with rank (how early they list it) to produce a salience score: items that are common AND early have the highest salience.
Pile Sort Participants are given a deck of cards (one item per card) and asked to sort them into piles of similar items. Aggregated across participants, the co-pile frequencies form a similarity matrix that can be analyzed with MDS or clustering to reveal the implicit category structure of the domain.
Triad Tests Participants are shown sets of three items and asked which one is the “odd one out.” Repeated across triads, the choices generate an item-by-item similarity matrix usable for the same downstream analyses as pile sorts but in higher resolution.
Consensus Analysis A statistical procedure (Romney, Weller & Batchelder, 1986) that uses participant-by-participant agreement to estimate (a) whether the participants share a single cultural model and (b) each participant's cultural competence, how closely their answers match the group's shared model.
Cultural Competence In the consensus model, the proportion of the shared cultural answer key each participant has command of. Competence scores range from 0 to 1; high-competence participants are the de facto cultural experts on the domain.
Eigenvalue Ratio Test In consensus analysis, the ratio of the first to the second eigenvalue of the participant-by-participant agreement matrix. A ratio ≥ 3 supports the single-culture assumption (one shared model); a lower ratio indicates sub-cultures and the analysis should be stratified.
Multidimensional Scaling (MDS) A dimensionality-reduction technique that takes a similarity matrix (e.g., from a pile sort) and produces a low-dimensional spatial map, items that are similar are close in the map, dissimilar items are far apart. The analyst must propose substantive interpretations of the dimensions; the algorithm does not label them.
Semantic Network Analysis
Semantic Network A graph in which nodes are words (or concepts) and edges represent some form of association, most commonly co-occurrence within a context unit (sentence, paragraph, or sliding window) across a corpus. Used in Bernard, Wutich & Ryan, Ch 19 as a tool for visualizing the conceptual structure of a body of text.
Co-occurrence Matrix A symmetric word-by-word matrix in which cell (i, j) records how often word i and word j appear together within the chosen context unit. The co-occurrence matrix is the input to semantic-network construction.
Degree / Strength Centrality Degree is the number of edges a node has; strength is the sum of edge weights for a node. In a semantic network, high-degree words are the “hubs” that connect to many other words.
Betweenness Centrality A centrality measure capturing how often a node sits on the shortest path between other nodes. High-betweenness words are conceptual bridges: removing them would disconnect parts of the network. Distinct from raw frequency.
Community Detection Algorithms (Louvain, Leiden, walktrap, etc.) that partition a network into densely-connected clusters (communities). In semantic networks, communities surface thematic clusters of co-occurring words.
Louvain Algorithm A widely-used community-detection algorithm that optimizes modularity through a fast, iterative, agglomerative procedure.
Modularity A quality score for a network partition: the extent to which within-community edges are denser than would be expected if edges were placed at random. Modularity above ~0.3 indicates meaningful community structure; above 0.5 is strong.
LLM-Assisted Coding
Large Language Model (LLM) A neural network (transformer architecture) trained on very large text corpora to predict the next token, then fine-tuned for instruction-following. In qualitative analysis, LLMs (GPT-4, Claude, Gemini, open-source models) can be prompted to apply codebooks, summarize transcripts, or surface candidate themes. Admissible with discipline; not admissible without.
Hallucination An LLM output that is fluent, confident, and wrong, invented quotes, misattributed claims, codes applied to passages that don't support them. The central methodological risk in LLM-assisted coding because hallucinated output is stylistically indistinguishable from correct output. Mitigated by verbatim-quote prompt constraints and random-sample audits.
Prompt Drift The phenomenon that LLM outputs change as the prompt is refined across iterations, often without the analyst noticing the cumulative effect. Mitigated by version-controlling prompts and reporting the final prompt verbatim in the methods section or appendix.
Opaque Reasoning LLMs do not produce inspectable rationales for their classifications. Even “chain-of-thought” outputs are post-hoc narrations, not the actual computation. Implication: an LLM coding cannot be defended on the grounds that the model “reasoned” correctly; it must be defended by validation against a hand-coded reference set.
Bias Amplification LLMs are trained on internet-scale text and inherit its demographic, linguistic, and ideological skews. When applied to qualitative data they may systematically under-code or over-code particular kinds of speech (e.g., AAVE, non-native English, working-class registers). Mitigated by stratified validation across participant subgroups.
Methodological Invisibility The temptation to use LLM assistance without disclosing it, a violation of the transparency commitment. The course's position is that all LLM use must be disclosed in the methods section, with the prompt, the model and version, the date of use, and the validation results.
Calibration / Reference Subset A hand-coded set of transcripts (typically 10–20% of the corpus, drawn from the same study) used to validate LLM coding output. The LLM is run on the same transcripts and the two codings are compared with an intercoder-agreement statistic. The reference set must come from the same study; benchmarks from other studies cannot substitute.
Intercoder Reliability (LLM) An agreement statistic (Cohen's κ, Krippendorff's α) computed between LLM coding and a hand-coded reference. Conventional floors: 0.7 acceptable, 0.8 preferred. Below 0.6 the LLM should not be used as-is; prompt or codebook must be revised.
LLM Disclosure A paragraph in the methods section that states (a) which model and version was used, (b) on what date(s), (c) the exact prompt, (d) the validation results against a hand-coded reference, (e) the hallucination audit results. Unconditional, required even when results are good.
Verbatim-Quote Constraint A prompt instruction requiring the LLM to return only verbatim text from the source transcript when quoting, and to flag any inability to do so. Reduces (but does not eliminate) hallucination of fabricated quotes; the audit checks whether the constraint operated.
Key People
A. Kimball Romney (1925–2016) Cognitive anthropologist; lead author of the foundational cultural-consensus paper (Romney, Weller & Batchelder, 1986, American Anthropologist). His work established that cultural shared-ness could be measured as a statistical property of participant agreement.
Susan C. Weller Medical anthropologist and co-author of the cultural-consensus model. Her later textbook, Systematic Data Collection (with A. Kimball Romney, 1988), is the standard reference for free listing, pile sorts, and triads.
William H. Batchelder (1940–2018) Mathematical psychologist who, with Romney and Weller, provided the formal statistical machinery of the consensus model. The Romney-Weller-Batchelder (RWB) procedure is named for the three.
Klaus Krippendorff (1932–2022) Communication scholar; author of Content Analysis: An Introduction to Its Methodology (multiple editions). Developed Krippendorff's α, the most general intercoder-agreement statistic and the recommended choice for LLM-vs-human comparisons.
James W. Pennebaker Social psychologist; developer of LIWC (Linguistic Inquiry and Word Count), a dictionary-based text-analysis tool that scores texts on linguistic and psychological dimensions. A canonical example of category-based computational text analysis and a useful contrast with embedding-based approaches.
No matching entries. Try a different search term.