Finding Themes & Building Codebooks
Qualitative Research Methods & Analysis in Public Health
Learning objectives for this lesson:
- Distinguish themes from codes, categories, and concepts, the terminology that introductory texts use loosely and that Bernard, Wutich, and Ryan tighten up
- Apply the twelve Ryan & Bernard (2003) techniques for finding themes to a small corpus of loneliness transcripts
- Describe the six phases of thematic analysis and distinguish its coding-reliability, codebook, and reflexive variants
- Differentiate inductive, deductive, and hybrid coding strategies and justify which is appropriate for a given research question
- Build a structured codebook with code names, brief definitions, full definitions, inclusion criteria, exclusion criteria, and positive/negative exemplars
- Explain coding mechanics: hierarchical codes, multiple codes per passage, axial coding, and the use of in vivo codes
- Compute and interpret percent agreement, Cohen's kappa, and Krippendorff's alpha, and identify when intercoder reliability is the wrong measure
- Operate the Taguette workflow: upload, code, export, and analyze coded extracts
- Produce a preliminary codebook tested on 3–5 transcripts, with a one-page memo on what coding revealed
This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University based on Bernard, H. R., Wutich, A., & Ryan, G. W. (2017). Analyzing Qualitative Data: Systematic Approaches (2nd ed.). SAGE. This lesson covers Chapters 5 and 6 (pp. 101–160).
What Themes Are, and Twelve Techniques for Finding Them
Introduction and Overview
Earlier lessons gave you the upstream apparatus of a qualitative project: an operational definition of QDA, a research question, a sampling logic, and a data-collection procedure. You now arrive at the central act of the analysis. You have transcripts. You have read them. You sense that something is going on across them. You suspect there are patterns. The question of this lesson is: how do you find those patterns systematically, and how do you record what you find in a form another analyst could follow?
Two activities sit at the heart of analytic work on text: finding themes and building a codebook. These twin moves are the connective tissue of every major qualitative analytic tradition, from thematic analysis (Braun & Clarke, 2006; Braun & Clarke, 2019) through qualitative content analysis (Hsieh & Shannon, 2005) to the systematic coding manuals used in applied health research (Saldaña, 2021). They are tightly coupled but not the same. Finding themes is the discovery phase: the search for what recurs, what surprises, what is missing, what coheres. Building a codebook is the codification phase: turning what you found into operational rules that you (and other analysts) can apply consistently across the rest of the corpus. This lesson works through both, drawing on Chapters 5 and 6 of Bernard, Wutich, and Ryan (2017).
Learning Objectives for this section
- Define theme, code, category, and concept and explain how the terms relate.
- Recognize that themes are found by an analyst, not discovered in nature.
- Identify all twelve Ryan and Bernard (2003) techniques for finding themes and recognize when each is most useful.
- Apply at least four of the twelve techniques to passages from the loneliness dataset.
- Distinguish first-cycle coding, which labels segments, from second-cycle coding, which builds categories and themes from codes.
- Describe Braun and Clarke's six phases of thematic analysis, the phases in which the twelve techniques are used, and the coding-reliability, codebook, and reflexive variants.
1.1 Themes, Codes, Categories, Concepts: A Precise Vocabulary
A theme is a recurrent meaning or pattern across a body of qualitative data. Themes are interpretive; they exist at a higher level of abstraction than what is literally in the text. 'Stigma as barrier to disclosure' is a theme; 'said the word stigma' is not.
A code is a label applied to a segment of data. Codes can be descriptive (close to what was said) or interpretive (analyst's reading). Codes are the granular units that, when aggregated, support claims about themes.
A category is a higher-level grouping of related codes. The hierarchy is typically: data segment → code → category → theme. Some traditions blur category and theme; the distinction is most useful in framework analysis and applied policy work.
A concept is a portable theoretical idea that can travel beyond the original study. 'Allostatic load,' 'cultural safety,' 'social capital' are concepts. Theme is local; concept is general. Grounded theory aims at concepts; thematic analysis is content to stop at themes.
The terms theme, code, category, and concept get used interchangeably in much of the published qualitative literature, including in papers that are otherwise methodologically careful. Bernard, Wutich, and Ryan (2017, Ch. 5) are precise where most writers are not. Adopting their vocabulary will save you grief when you write your methods section and will save the reader confusion about what you actually did.
A theme is a recurring abstract idea you identify in the data. It is the analyst's product. “Loneliness as the cost of love” is a theme. “Spatial metaphor for absence” is a theme. Themes are typically expressed as short phrases rather than single words. They sit at a higher level of abstraction than the individual statements that supply evidence for them.
A code is the operational label you attach to a segment of text to mark what that passage is about. Codes are the working units of analysis: they are what you actually mark up in Taguette or NVivo. A single theme may be supported by several codes; a single passage may receive multiple codes. The code is the marker; the theme is what the markers, when assembled, are about.
A category is a grouping of related codes. In a hierarchical codebook, categories are the parent nodes and codes are the children. “Coping strategies” is a category that might contain codes for “phoning a confidant,” “watching comfort television,” “going to a coffee shop to be around people,” and “cooking food from home.” The category organizes; the codes do the marking.
A concept is the most abstract of the four. A concept is a theoretically meaningful idea that may organize many themes. Liminality is a concept. Embodiment is a concept. Structural exclusion is a concept. Concepts are typically borrowed from theoretical literatures and used to organize themes into something a discipline can argue about.
| Term | Level of abstraction | Example from the loneliness dataset |
|---|---|---|
| Code | Lowest: operational marker | chair-absent-spouse (applied to Linda's “Bill's chair” passage and similar) |
| Theme | Mid: recurring idea | Spatial objects standing in for absent people |
| Category | Mid: organizing bucket | Material traces of relationship loss (contains codes for chair, side of bed, photograph, kitchen, mobility aids) |
| Concept | Highest: theoretical | Embodied memory; material culture of grief |
The hierarchy also describes an order of work. Saldaña (2021) divides coding into two cycles. First-cycle coding is the initial labelling of data segments: you read a transcript and attach codes to passages, usually staying close to what was said. Second-cycle coding works on those codes rather than on the raw text: you group, compare, and reorganize them into categories and then into themes. Codes therefore come first, and themes are built from them. A code such as chair-absent-spouse is applied in the first cycle; the theme about objects standing in for absent people is assembled in the second, once several codes and many passages point the same way. Section 1.4 shows how thematic analysis organizes the two cycles into phases, and Section 3.3 returns to them when it introduces axial coding.
Themes are found, not discovered
Bernard, Wutich, and Ryan are emphatic that themes do not emerge from data the way fossils emerge from rock, a point Braun and Clarke (2019) make forcefully in their reflexive reframing of thematic analysis. The analyst notices them, names them, decides which to keep, and decides where the boundaries are. The phrasing “themes emerged from the data,” ubiquitous in published qualitative papers, obscures the analytic work and is one of Bernard, Wutich, and Ryan's pet peeves. A more honest phrasing is “we identified the following themes through inductive coding” or “the themes below were developed iteratively as we read across the 20 transcripts.”
1.2 Ryan and Bernard's Twelve Techniques for Finding Themes
Repetitions, linguistic connectors, word lists/KWIC. Look for words and phrases that recur, that link causal claims ('because', 'so that'), or that cluster around your concepts. Computationally tractable; useful as a first pass before manual coding.
Emic categories, metaphors and analogies, theory-related material. Look for the participant's own vocabulary ('feeling left behind'), figurative language ('drowning in the system'), and segments that engage existing theory. These techniques find themes that matter to participants or to the field.
Transitions, similarities and differences, co-occurrence. Look at moments of change in narrative, contrasts across cases, and codes that appear together. These produce the relational themes, ones describing how concepts move, contrast, or combine.
Missing data (what people don't say), cutting-and-sorting (pile sort), metacoding. Notice what is absent (silences are themes too), physically rearrange coded segments to discover groupings, and code the codes themselves to surface higher-order patterns. The most labour-intensive techniques but often the most revealing.
In a now-classic 2003 paper, Gery Ryan and H. Russell Bernard (2003) catalogued twelve techniques qualitative researchers use to find themes. The paper became foundational because it disaggregated what had previously been described as “immersion” or “reading deeply” into a set of specifiable operations. Bernard, Wutich, and Ryan (Ch. 5) reproduce and update the list. Each technique answers a slightly different question about the data, and most projects use several in combination. Below we work through all twelve, with examples drawn from the loneliness dataset.
Technique 1: Repetitions
Take a short paragraph from one of your transcripts (or sample text). Code it twice:
- Descriptive coding: one code per phrase, close to the surface meaning. Aim for 5-8 codes in a paragraph.
- Interpretive coding: one or two codes per paragraph, capturing what you think this passage is doing. Aim for 1-3 higher-level codes.
Compare the two coding passes. The descriptive layer organizes the corpus; the interpretive layer carries the argument. Most defensible thematic analyses move iteratively between them.
The most obvious technique. What words, phrases, ideas recur across transcripts? If multiple participants reach for the same word to describe something, that word is doing analytic work. In the loneliness corpus, the word chair appears in at least eight transcripts as a stand-in for an absent person: Linda's “Bill's chair,” Frank's chair, Helen's mention of the chair her brother used to sit in when he visited, and others. The word tired recurs across caregiver and bereaved participants in a way that goes beyond ordinary fatigue: it appears to mark a specific exhaustion of grieving-as-work. Repetition is what nearly all theme-finding starts with, and what every other technique builds on.
Technique 2: Indigenous Typologies and Categories (Emic Terms)
What categories do participants themselves use? When a participant reaches for a non-English word, a slang term, or a phrase that operates as a category in their world, the analyst should pay attention. In the loneliness corpus, Amira uses the Arabic word wahda to name something that the English category of “loneliness” cannot fully hold, a loneliness specific to having been the sole survivor of a particular life. Aarav uses ekantam and ekakitatvam as paired Sanskrit-Hindi terms that distinguish chosen solitude from involuntary aloneness. Marcus speaks of “code-switching” loneliness, an experience he names that has no neat one-word English equivalent. These emic categories are gifts; they often become themes that organize an entire section of your eventual paper.
Technique 3: Metaphors and Analogies
People describe abstract experiences (especially feelings) by reaching for concrete images that map onto them. Cataloguing the metaphors a corpus uses is a powerful theme-finding move. The loneliness corpus is dense with spatial metaphors of absence and erosion: Maya feels “hollow” in her chest; Sarah describes loneliness as her “witness-less hours”; Helen describes it as “fading at the edges”; Frank uses imagery of a slow disappearance; Maya again talks about feeling she could “disappear and nobody would notice”; Linda talks about “walking around with that absence.” Read together, these metaphors converge: loneliness is repeatedly figured as a thinning or vanishing of the self. That convergence is a theme that you would never have seen if you had read the metaphors one at a time.
Technique 4: Transitions
What does the participant move to right after they say what they say? Transitions are turn-taking shifts and topic changes. They tell you what feels related in the participant's mind. In Linda's transcript, every passage about Bill's chair is followed by a passage about Rufus the dog: “I haven't moved it... the dog is what keeps me up in the morning.” The transition from the empty chair to the dog suggests an analytic linkage you might otherwise have missed: the dog is functioning as the affective replacement for the chair's absence. In Maya's transcript, the loneliness topic transitions repeatedly to her phone, then food, then TV, signalling that her coping repertoire is digitally mediated.
Technique 5: Similarities and Differences (Constant Comparison)
Read two passages side by side. What is the same? What is different? This is the constant-comparative move that Glaser and Strauss made foundational to grounded theory and that you will meet again in a later lesson. As a theme-finding technique, it is most useful when you have already identified candidate themes and want to test whether they hold up across subgroups. In the loneliness corpus, the bereaved-spouse loneliness of Linda (age 67, widow of three years) and Frank (age 81, widower of one year) share most features but differ on duration of mourning and the role of children: Linda's adult sons are present (David in Toronto, Michael in Calgary), while Frank's are estranged. The comparison sharpens the theme rather than dissolving it.
Technique 6: Linguistic Connectors
Words like because, since, as a result, therefore, that's why, and so are causal connectives. They mark places where the participant is explaining a relationship between events or states. Searching for them is a fast way to find passages where causal accounts of loneliness appear. Maya's transcript contains: “I just, I felt like everyone's life is still happening and I'm just here. Like I left and life closed up where I used to be.” The connective like here is not strictly causal but it is doing relational work. Linda's transcript: “If you love deeply for a long time, you will, eventually, be lonely deeply for a long time. The two are connected.” That explicit linkage of love and loneliness is a participant's own causal model, surfaced by attention to a connective.
Technique 7: Missing Data, What People Don't Say
What the corpus does not contain is as analytically informative as what it does. If you asked every participant about coping and three avoided answering, that pattern of avoidance is itself data. In the loneliness corpus, several participants conspicuously do not use the word “lonely” about themselves until very late in the interview, even though the interview is explicitly about loneliness. Marcus repeatedly substitutes “disconnected” or “invisible.” Older men in the corpus avoid the word in a way younger women do not. That gendered pattern of avoidance is a finding. Missing data is also useful at the level of what the interview guide asked but participants deflected: questions about professional help, in particular, are routinely deflected.
Technique 8: Theory-Related Material
Bernard, Wutich, and Ryan call this “looking through a theoretical lens.” If you bring a specific theory to the data, such as Cacioppo and Patrick's loneliness-versus-aloneness distinction, Weiss's social-emotional loneliness typology, structural/situational/existential models, you can deliberately scan the corpus for passages that confirm, complicate, or contradict the theory. This is more deductive than the previous seven techniques (which start from the data) and we will return to it when we distinguish inductive from deductive coding in a later section. As an example: Cacioppo and Patrick (2008) argue that loneliness and being alone are dissociable states. The loneliness corpus is full of explicit articulations of that distinction (Maya: “loneliness is different than being alone”; Helen: “I have lived alone all my adult life… the loneliness I have now is different from the solitude I had at 50”). The theoretical lens both helps you see the pattern and gives you a way to write about it.
Technique 9: Cutting and Sorting (The Pile-Sort)
A physical, tactile method. Print out striking quotes from your corpus, one per index card. Spread them out on a table. Group cards that feel related. Move cards around as your sense of groupings evolves. Give each group a name. This technique externalizes the analytic work and uses spatial cognition to find patterns that screen-based reading misses. It is especially useful when you have 20+ candidate themes and need to consolidate them into a manageable set. Bernard, Wutich, and Ryan recommend actual cutting-and-sorting (literal scissors) at least once per project; the tactile experience is what makes it work. We will do a digital version of this in a later section workflow.
Technique 10: Word Lists and KWIC (Keyword-in-Context)
Generate a frequency list of all words in the corpus and inspect the high-frequency content words (after removing stopwords like “the” and “and”). For words that catch your eye, generate a KWIC concordance: every occurrence of the word with a few words of context on either side. KWIC concordances are how you check whether participants are using a word the same way. If “chair” appears 23 times across the corpus, are all 23 instances chairs-as-stand-ins-for-people, or are some just literal pieces of furniture? The KWIC concordance answers that quickly. This is also the natural bridge to computational text analysis (a later module); a later section of this lesson works through a small word-list and KWIC exploration of the loneliness transcripts.
Technique 11: Co-occurrence
Which codes (or words) appear together more often than chance would predict? Co-occurrence is the first move toward axial coding (covered in a later section) and toward concept-mapping. If your code chair-absent-spouse co-occurs in nearly every transcript with codes for pet-as-companion or volunteer-coping, that co-occurrence is a finding. Co-occurrence is also how you build the network displays you will meet later in the course. Co-occurrence is straightforward to tabulate once your coded extracts are in a long-format table, with one row per passage and code.
Technique 12: Metacoding (Codes About Codes)
Key insight - A theme is a claim, not a topic
Beginners label themes with single nouns: Stigma. Access. Identity. These are topics. A theme is a defensible analytic claim about a pattern. Better theme labels are short statements: 'Stigma is managed strategically, not passively endured' or 'Access is described less in terms of geography than in terms of trust'. Themes-as-claims are easier to evidence, easier to dispute, and easier to write up. Themes-as-topics are easier to fall in love with and harder to defend.
Once you have a working set of codes, you can step up a level and label the codes themselves. Are some of your codes affective (about feelings) and others behavioural (about coping actions)? Are some emic (using participants' own words) and others etic (using your analyst-imposed terminology)? Sorting your codes by type is metacoding, and it often reveals structural patterns in your codebook that you would not have seen by staring at the codes one at a time. In the loneliness corpus, a useful metacoding move is to sort all codes into four buckets: experiential (what loneliness feels like), causal (what triggers it), responsive (what people do about it), and interpretive (what people think it means). The four-bucket frame becomes a candidate structure for the findings section of your eventual paper.
You do not need to use all twelve
The twelve techniques are a toolbox, not a checklist. A defensible project will use three or four of them, deliberately, and report which ones it used. Repetitions + indigenous categories + metaphors + missing data is a typical opening combination. The point of having twelve in your awareness is that when one technique is not turning anything up, you have eleven others to try. Your eventual methods section should name which techniques you used and why.
1.3 Working an Example Through Four Techniques
Take three passages from the loneliness corpus and run them through four of the techniques to see how theme-finding actually works.
“It's, okay, this is going to sound dramatic, but it feels like being hungry. Like, in my chest. It's a physical thing. I get this feeling, especially at night, where my chest just feels, hollow isn't the right word, it's like an ache.”
“I don't know that I'd say it has a physical feeling exactly. It's more like a weight. A weight that I carry around. I sleep on one side of the bed still, the right side, my side, and the left side is undisturbed for three years. So it's that. It's like half of my life is just not there anymore, and I'm walking around with that absence.”
“It feels like, fading. Like fading at the edges. When you do not speak for days, you become less real to yourself. Your voice sounds strange when you do speak.”
Technique 1 (Repetitions): All three passages reach for embodied descriptions of loneliness. The words physical, chest, weight, body, voice, real recur. Loneliness is a bodily experience for these participants, as much as a cognitive one.
Technique 3 (Metaphors): Three different metaphors, all converging on the same image. Maya: hunger, hollow, ache (interior absence). Linda: weight, half a life not there, walking around with absence (carrying something missing). Helen: fading, less real, voice strange (thinning of the self). The metaphors are not the same, but the analytic abstraction over them is: loneliness as the somatic register of an absence.
Technique 5 (Similarities and Differences): The three passages share embodiment but differ on what is absent. For Maya it is a future not yet built; for Linda it is a specific deceased person; for Helen it is the simple act of social contact. A theme that holds across the differences: the body registers absence as presence (you feel something missing as something there).
Technique 7 (Missing Data): Notice what none of the three say. None invokes psychiatric vocabulary. None mentions a therapist or a medication. None reaches for the word “depression” even though the descriptions are compatible with depressive symptomatology. The absence of clinical framing is itself a finding: these participants are describing loneliness as a normal embodied condition, not as a disorder. That has implications for how an intervention should be framed.
Four techniques applied to three passages have already moved from codes on single passages toward a candidate theme: loneliness as embodied absence, narrated outside clinical vocabularies. The codes that carry it, such as somatic-absence, go into the codebook, and the candidate theme is tested as more transcripts are coded. Section 1.4 sets out the procedure for that testing.
1.4 Thematic Analysis: From Codes to Themes
The vocabulary of Section 1.1 and the techniques of Sections 1.2 and 1.3 come together in thematic analysis, the most widely used procedure for moving from coded passages to reported themes. Braun and Clarke (2006) set out the procedure as six phases, and their later writing renamed several of them to stress that the analyst generates themes from codes (Braun & Clarke, 2019, 2021). The phases are recursive. You move back and forth between them, and a theme reviewed in Phase 4 can send you back to the coded data of Phase 2.
| Phase | What you do | Ryan and Bernard techniques that fit |
|---|---|---|
| 1. Familiarization | Read and reread every transcript, listen to the audio if you have it, and keep brief notes on first impressions. | Repetitions; metaphors and analogies; indigenous typologies; word lists and KWIC as a first scan |
| 2. Coding | Work systematically through the dataset and label every segment that bears on the research question. This is first-cycle coding. | Repetitions; indigenous typologies (in vivo codes); metaphors; transitions; linguistic connectors; theory-related material; missing data |
| 3. Generating initial themes | Group codes that seem to share a central idea into candidate themes, and collate the coded extracts for each. Second-cycle coding begins here. | Similarities and differences; co-occurrence; cutting and sorting; metacoding |
| 4. Developing and reviewing themes | Check each candidate theme against its coded extracts and then against the whole dataset; split, merge, or discard themes that do not hold. | Similarities and differences (constant comparison); missing data and negative cases; KWIC to check how far a pattern reaches |
| 5. Refining, defining, and naming themes | Write a short definition of each theme that states its central organizing concept, its scope, and its boundaries, and give it a concise name. | Indigenous typologies and in vivo phrases for names; metacoding to settle how themes relate |
| 6. Writing up | Weave the analytic narrative and the data extracts into a report that answers the research question. | Choice of exemplar quotations; a statement of which techniques were used |
Three variants of thematic analysis
Braun and Clarke (2019, 2021) argue that thematic analysis is a family of methods rather than a single method. They distinguish three variants that differ in how codes and themes are produced and in what counts as quality.
- Coding-reliability thematic analysis uses a structured codebook, often developed early, which several coders apply independently, and it measures their agreement with statistics such as Cohen's kappa (Section 3.7). Themes are often summaries of what participants said about a topic and may be identified before coding begins. Boyatzis (1998) is a classic example. The approach treats coding as something that can be more or less accurate, an assumption close to postpositivism.
- Codebook thematic analysis also uses a structured codebook to organize and chart the data, but it does not measure agreement between coders. Framework analysis (Ritchie & Spencer, 1994) and template analysis (King, 2012) belong to this group. Themes are often partly developed in advance, and the approach suits applied team projects with fixed information needs.
- Reflexive thematic analysis treats coding as an open and organic process carried out by a researcher whose subjectivity is a resource for the analysis. It uses no fixed codebook and no measure of coder agreement. Themes are developed late, from codes, as patterns of shared meaning united by a central organizing concept, and quality rests on reflexivity, depth of engagement, and the coherence of the account.
The hybrid, codebook-based workflow taught in Sections 2 to 4 of this lesson belongs to the coding-reliability family, because it ends with an intercoder agreement statistic. Section 3.9 explains when that statistic is the wrong measure, and reflexive thematic analysis is one of those cases. A qualitative paper that reports themes and sub-themes should name the procedure that produced them, whether one of these variants or a named alternative such as framework analysis or grounded theory, and its quality claims should match that choice.
Worked example: two loneliness transcripts
Phase 1, familiarization. Reading both transcripts in full, you note that Maya keeps returning to her body and to the sense that life went on without her, and that Linda keeps returning to objects in her home and to her late husband, Bill.
Phase 2, coding. First-cycle codes on Maya's transcript include somatic-absence (“it's like an ache”), life-moved-on (“Like I left and life closed up where I used to be”), and loneliness-vs-aloneness (“loneliness is different than being alone”). On Linda's transcript they include somatic-absence (“A weight that I carry around”), chair-absent-spouse (Bill's chair), loneliness-as-cost-of-love (“The two are connected”), and coping-pet (Rufus the dog). Repetitions, metaphors, and linguistic connectors do most of the work in this phase.
Phase 3, generating initial themes. Laying the coded extracts side by side (similarities and differences, co-occurrence), you see that somatic-absence and chair-absent-spouse share a central idea: something missing is felt as a presence, in the body or in the house. A candidate theme follows: loneliness is carried as the felt trace of an absence. The coping codes form a useful category, but “coping with loneliness” on its own is a topic summary, because it names an area without making a claim about it.
Phase 4, developing and reviewing themes. Two transcripts cannot support a theme. You check the candidate against its coded extracts and then against the other eighteen transcripts, looking for negative cases such as participants who describe loneliness with no bodily or material reference, and you narrow or split the theme if such cases are common.
Phase 5, refining, defining, and naming themes. You write a definition that states the theme's central organizing concept and its boundary with neighbouring themes such as loneliness-as-cost-of-love, and you consider an in vivo name drawn from Linda: “walking around with that absence.”
Phase 6, writing up. The theme is reported with its definition, exemplar quotations from more than one participant identified by participant code, and an analytic commentary on what the pattern means for the research question.
Check your understanding
1. In which phase of thematic analysis do cutting and sorting and co-occurrence do most of their work, and why there?
Phase 3, generating initial themes. Both techniques operate on codes and coded extracts rather than on raw text, so they belong to second-cycle work, in which codes are grouped and compared to build candidate themes.
2. A team writes a fixed codebook before coding, has two analysts apply it independently, and reports Krippendorff's alpha. Which variant of thematic analysis is this, and why would a reflexive thematic analysis report no such statistic?
Coding-reliability thematic analysis. Reflexive thematic analysis treats the researcher's interpretation as part of how meaning is produced, so it has no fixed codebook for coders to agree on, and it judges quality by reflexivity and the coherence of the account rather than by agreement between coders.
Reflection
Pick any one of the twelve techniques you found most or least intuitive. Briefly explain why, and describe a passage from a transcript you have read (or could read) where the technique would or would not work well. The point is not to defend the technique; it is to articulate, for yourself, when it would and would not earn its keep.
Minimum 20 characters required.
Question 1: According to Bernard, Wutich, and Ryan, what is the difference between a theme and a code?
Question 2: Which technique would best help you discover that older men in the loneliness corpus systematically avoid the word “lonely” itself?
Question 3: Amira describes her loneliness using the Arabic word wahda, which she says holds a meaning that the English “loneliness” cannot. Which technique is most directly engaged?
Inductive, Deductive & Hybrid Coding, and Codebook Architecture
Introduction and Overview
Section 1 ended with candidate themes, and the order in which they arose matters. Codes come first: first-cycle coding labels segments of text, and themes are built later, in second-cycle work, by grouping and comparing those codes (Sections 1.1 and 1.4). Ryan and Bernard's twelve techniques serve both moments. Repetitions, metaphors, and linguistic connectors guide first-cycle reading, while similarities and differences, co-occurrence, cutting and sorting, and metacoding guide the grouping of codes into themes. This section turns to the codes themselves and to the codebook that writes them down as operational rules applied consistently across the corpus, the bridge between recognition and analysis that Boyatzis (1998) and Braun and Clarke (2006) describe as central to rigorous thematic analysis. The two big design questions for coding are: where do the codes come from (inductive, deductive, or hybrid), and what does a defensible codebook look like (the architecture). This section addresses both.
Learning Objectives for this section
- Distinguish inductive (data-up), deductive (theory-down), and hybrid coding strategies.
- Match each strategy to the kind of research question it best serves.
- Build a codebook entry containing all seven required elements (name, brief definition, full definition, inclusion criteria, exclusion criteria, positive example, negative example).
- Recognize that the codebook is a living document, revisable with audit-trail documentation.
Background: Deductive, inductive, and hybrid coding
Codes come from two directions. Deductive (a priori) codes are written before coding begins, from the research question, the topics of the interview guide, an existing theory, or earlier studies. They keep coding tied to the question and make team coding consistent from the start, but they can push text into codes that do not fit and overlook what is new. Inductive codes are created during coding, when a passage expresses something important that no existing code captures; in vivo codes, which use a participant's own words as the label, are one kind. They capture what the researcher did not anticipate, in participants' terms, but the list can grow long and overlapping unless it is reviewed. Hybrid coding combines the two: a short deductive list, the provisional codebook, is applied to the first transcripts, and inductive codes are added as the data require (Fereday & Muir-Cochrane, 2006). A hybrid design needs a dated record of when each code was added or revised.
HSCI 207 Lesson 11, Section 3.2 (First Steps in Analysis) teaches these definitions with a worked example and is optional reading. The subsections below assume them and concentrate on how each direction plays out with the loneliness dataset.
2.1 Inductive Coding (Data-Up)
Inductive coding in practice follows a recognizable arc. After three or four transcripts, you have a working set of perhaps 30–50 provisional codes; you consolidate, merge, and rename until you have a coherent codebook of 8–15 codes you can apply across the remaining transcripts.
Inductive coding suits exploratory studies, under-described phenomena, and projects in the grounded-theory tradition (Lesson 7). Its risk is that, without theoretical anchoring, it can produce codebooks that are descriptive but not analytically interesting, what Charmaz (2014) warns against as “coding too close to the data.”
For the loneliness dataset, an inductive start is appropriate. The dataset is rich, the phenomenon is contested, and the participants speak in distinct vocabularies. Starting from their language is the right move.
2.2 Deductive Coding (Theory-Down)
Frameworks used for deductive coding in health research include the WHO determinants of health, the CFIR implementation framework, and the Cacioppo and Patrick loneliness model, and an interest holder may bring a framework of its own. When you code against one of them, codes that are not present in the data are recorded as absent, which is itself a result.
Deductive coding is the right move when you are testing or extending an existing framework, when you are working in a confirmatory mode, or when your study is part of a multi-site collaboration that needs a shared coding scheme. Its virtue is that it produces results comparable across studies. Its risk is that you may miss things the data are saying that the framework was not designed to see.
An applied example: if you were coding the loneliness transcripts deductively using Weiss's (1973) social-emotional loneliness typology, you would have two pre-specified codes, namely social loneliness (deficits in a social network) and emotional loneliness (absence of a close attachment figure), and you would tag each passage as one, the other, both, or neither. You would learn the distribution of the two types across the corpus and would have framework-comparable results. You would also miss the embodiment theme we developed in an earlier section, because Weiss's framework does not contain it.
2.3 Hybrid Coding (the Practical Default)
Most contemporary qualitative health research uses a hybrid approach. The deductive starting list is called a provisional codebook, and the codebook then grows from the bottom up while keeping its theoretical anchor.
The Fereday and Muir-Cochrane (2006) hybrid framework is a widely cited operationalization of this approach, and Hsieh and Shannon's (2005) "directed" qualitative content analysis is a closely related variant. Bernard, Wutich, and Ryan (Ch. 6) endorse hybrid coding as the practical default for applied health research because it preserves the inductive openness that makes qualitative work valuable while keeping the deductive anchoring that makes it interpretable by quantitatively trained reviewers.
The worked examples in this lesson use a hybrid strategy: you start with two or three theoretically motivated codes (perhaps experiential loneliness, causal accounts, and coping strategies, derived from the interview guide structure) and let the rest develop inductively from the first three to five transcripts you code.
| Strategy | Codes come from | Best fit | Risk |
|---|---|---|---|
| Inductive | The data | Exploration; under-described phenomena; grounded theory | Descriptive without theoretical purchase; long codebooks |
| Deductive | Theory or framework | Confirmatory work; multi-site studies; testing established models | Misses what the framework was not built to see |
| Hybrid | Both, in sequence | Most applied health research; this course's default | Requires explicit documentation of when codes were added or revised |
2.4 The Anatomy of a Codebook Entry
A codebook is a structured document. Each entry describes one code with enough specificity that another analyst could apply it consistently (MacQueen, McLellan, Kay, & Milstein, 1998; DeCuir-Gunby, Marshall, & McCulloch, 2011). Bernard, Wutich, and Ryan (Ch. 6) recommend seven elements per entry. All seven matter; cutting any of them is the most common reason why intercoder reliability later turns out to be low.
- Code name. Short, mnemonic, unique. Use hyphens or underscores, not spaces. Example:
chair-absent-spouse. - Brief definition. One sentence. Example: “A physical object (chair, side of bed, photograph) that the participant marks as standing in for the absence of a deceased or departed partner.”
- Full definition. A paragraph. When to apply the code; what range of cases it covers; how it relates to neighbouring codes. The full definition is what an analyst reads when in doubt.
- Inclusion criteria. Bullet-pointed. What features must be present for the code to apply. Example: the participant explicitly names a physical object; the object is associated with an absent person; the participant gives the object affective weight.
- Exclusion criteria. Bullet-pointed. What looks similar but does not count. Example: mention of a physical object without affective weight (“the chair in the corner”) does not count; mention of an absent person without an associated object does not count.
- Positive example. A direct quote from the corpus that clearly fits. Example: Linda P05: “Bill sat in that chair every evening for thirty-some years. And it's still there. I haven't moved it. I haven't sat in it. I haven't given it away. It's just there. And every evening I look at it and it's empty.”
- Negative example. A near-miss quote that does not fit. Example: Helen P11: “I have this walker because my hip gave out.” A physical object is named, but it is not associated with an absent person.
The eighth element: the memo
Bernard, Wutich, and Ryan recommend a seven-element structure. Many experienced researchers add an eighth: a memo space attached to each code, where the analyst records how the code evolved, why edge cases were resolved a particular way, and what the code's relationship to neighbouring codes turned out to be. Memos are the audit trail for the codebook itself. We strongly recommend including a memo column in any codebook you build, because it is the record a methods section is later written from.
Crosswalk: two layouts for the same codebook
If you took HSCI 207, you met a six-column codebook in Lesson 11, Section 3.3 (First Steps in Analysis), also based on MacQueen and colleagues (1998). The two layouts describe the same artefact with different labels.
| HSCI 207 codebook column | Element in this lesson |
|---|---|
| Code name | 1. Code name |
| Brief definition | 2. Brief definition |
| Full definition | 3. Full definition |
| Use when | 4. Inclusion criteria |
| Do not use when | 5. Exclusion criteria |
| Example | 6. Positive example and 7. Negative example |
The seven-element entry adds no new kind of information. It splits the example into a passage that clearly fits and a near miss, because the near miss shows a second coder where the boundary of the code lies.
2.5 A Worked Codebook Entry
Here is a complete codebook entry for a code drawn from the loneliness corpus. We will return to this entry in a later section when we walk through the Taguette workflow.
somatic-absenceBrief definition: Participant describes loneliness as a bodily sensation that registers the absence of someone or something.
Full definition: This code applies to passages where the participant locates loneliness in the body (chest, weight, fatigue, voice, hunger, ache, hollowness, fading) and the embodied sensation is described as a registering of an absence rather than as a free-standing physical symptom. The code is distinct from fatigue-grief (which is exhaustion specifically tied to grief work) and from illness-talk (which is the description of medical symptoms). When in doubt, apply somatic-absence if the participant uses a bodily metaphor and explicitly or implicitly links it to the absence of a person, role, or life-phase.
Inclusion criteria:
- The passage contains a bodily reference (chest, weight, hunger, ache, fading, voice, body, physical, real).
- The bodily reference is figurative or interpretive (not a literal medical complaint).
- The bodily reference is tied (explicitly or by clear inference) to the absence of a person, role, or part of life.
Exclusion criteria:
- Literal medical symptoms or complaints (apply
illness-talkinstead). - Mental-state descriptions without a body reference (apply
affective-loneliness). - Bodily references not tied to absence (e.g., “I was tired from work”).
Positive example: Linda P05: “It's more like a weight. A weight that I carry around… I'm walking around with that absence.”
Negative example: Helen P11: “My hip gave out two years ago.” (Literal medical complaint, so apply illness-talk.)
Memo: Added on iteration 2 after noticing the convergence of Maya's “hollow”/“ache,” Linda's “weight,” and Helen's “fading.” Distinguished from affective-loneliness after a near-miss in Sarah's transcript (“witness-less hours” was tagged both ways; resolved by requiring a body reference for somatic-absence).
2.6 The Codebook as a Living Document
Bernard, Wutich, and Ryan are clear that a codebook is not built once and frozen. You will revise it. The standard expectation is that revisions are documented in an audit trail that records: (a) the date of revision, (b) which codes changed, (c) what they changed to, and (d) the justification. When a codebook is revised midway through a project, the earlier transcripts must be re-coded under the new scheme, not the old one, or the codebook becomes inconsistent across the corpus.
The practical implication is that you should not commit to a final codebook until you have read most of the transcripts you intend to code. A preliminary codebook developed on three to five transcripts is a first draft, and the codebook reported at the end of a study is that draft revised at least twice on the basis of what the rest of the corpus reveals.
Reflection
Think about how you would code the loneliness dataset. Would you adopt a primarily inductive, deductive, or hybrid coding strategy? Name two of the codes you anticipate ending up with, and indicate whether each one comes from theory (deductive) or from the data (inductive). There is no wrong answer; the point is to commit to a strategy you can defend in a methods section.
loneliness-vs-aloneness, drawn from Cacioppo and Patrick's distinction, and structural-versus-existential, drawn from a public-health loneliness typology in the published literature. From there I will let the rest develop inductively from the first three transcripts I code. I anticipate ending up with codes like somatic-absence (inductive, from the convergence of Maya's, Linda's, and Helen's embodied metaphors) and coping-pet (inductive, from Linda's Rufus and Maya's neighbour's cat).” The point is to articulate a strategy that you can defend, and to recognize that 'pure inductive' and 'pure deductive' are rare in applied work; most defensible projects are hybrid.Minimum 20 characters required.
Question 1: Which coding strategy is the practical default in applied health research?
Question 2: Which of the following is NOT one of the seven required elements of a codebook entry?
Question 3: A codebook is revised midway through coding. What do Bernard, Wutich, and Ryan require for the analysis to remain methodologically defensible?
Coding Mechanics & Intercoder Reliability
Introduction and Overview
Earlier sections covered the conceptual side of coding: what themes are, where codes come from, what a codebook looks like. This section addresses two operational matters that determine whether your coding is methodologically defensible. The first is coding mechanics: how passages get marked up, how codes relate to each other, how the analyst handles complications. The second is intercoder reliability: when a second analyst applies your codebook to the same passages, how much agreement should you expect, how do you measure it, and what does the resulting number mean?
Learning Objectives for this section
- Apply the four basic coding mechanics: hierarchical codes, multiple codes per passage, axial coding, and in vivo codes.
- Compute percent agreement, Cohen's kappa, and Krippendorff's alpha for a pair of coders.
- Interpret the magnitude of kappa using Landis and Koch (1977) thresholds.
- Identify when intercoder reliability is the wrong measure to seek.
3.1 Hierarchical Codes (Parent and Child)
A codebook is rarely flat. Codes nest under broader codes, which nest under categories, which sit under the codebook root. The hierarchy is what makes a large codebook navigable and what allows you to aggregate findings at different levels. In the loneliness corpus, a plausible partial hierarchy looks like this:
LONELINESS_EXPERIENCE/ somatic-absence affective-loneliness social-invisibility temporal-fading CAUSAL_ACCOUNTS/ bereavement-onset migration-onset life-stage-transition structural-isolation COPING_STRATEGIES/ coping-pet coping-volunteer coping-phone-confidant coping-comfort-media coping-cooking-home-food INTERPRETIVE_FRAMES/ loneliness-as-cost-of-love loneliness-as-failure loneliness-as-rebuilding loneliness-as-invisible-to-society
The four parent categories at the top are the metacoding bins from an earlier section (Technique 12). Each child code can be applied independently to passages. When you report your findings, you can report at the category level (“coping strategies appeared in all 20 transcripts”) or at the code level (“coping-pet appeared in 3 of 20 transcripts”), depending on what your argument needs.
3.2 Multiple Codes Per Passage
A single passage may instantiate more than one code. This is normal and expected. Linda's chair passage simultaneously instantiates chair-absent-spouse (a specific code), somatic-absence (the absence as carried), and loneliness-as-cost-of-love (the interpretive frame she later articulates). All three codes are applied to overlapping or identical text spans. Taguette and the major QDA packages handle multi-coding natively.
The implication for analysis is that when you later count code occurrences, you are counting passage-code pairs, not unique passages. A transcript with 80 unique coded passages might have 130 code applications because many passages got two or three codes. Both numbers are meaningful; report whichever supports your argument and be clear about which you are reporting.
3.3 Axial Coding (Relationships Between Codes)
The term axial coding comes from Strauss and Corbin's (1990) grounded-theory tradition; for an exhaustive catalog of coding methods including in vivo and axial styles, see Saldaña (2021). It refers to a second pass over the data, after initial coding, in which the analyst attends to the relationships between codes rather than to the codes themselves. Axial-coding questions include: which codes co-occur? Which codes appear in sequence (and in which order)? Which codes seem to be causes of which others, in participants' own accounts? Which codes are mutually exclusive in practice? In Saldaña's terms (Section 1.1), axial coding is a second-cycle method, because it works on codes already applied in the first cycle rather than on the raw text.
Axial coding is what turns a flat codebook into a model. In the loneliness corpus, axial coding might reveal that bereavement-onset codes are nearly always followed in the same transcript by coping-pet codes; that migration-onset codes co-occur with cooking-home-food codes; that loneliness-as-cost-of-love appears only in transcripts that also contain somatic-absence. These relationships are not in the codes themselves; they are in the pattern of co-occurrence. A later section shows how to tabulate co-occurrence from your coded extracts.
3.4 In Vivo Codes (Using Participants' Exact Words)
An in vivo code is a code whose name is taken verbatim from a participant's speech. The convention is to set in vivo codes in quotation marks in the codebook to mark their origin. In the loneliness corpus, defensible in vivo codes include: “wahda” (Amira's word for a refugee-specific loneliness), “witness-less hours” (Sarah's phrase for the loneliness of being unobserved), “the cost of love” (Linda's interpretation), “code-switching loneliness” (Marcus's articulation of the loneliness of moving between cultural registers), and “fading at the edges” (Helen's metaphor).
In vivo codes do two analytic things at once. First, they keep the participant's voice in the codebook, which protects against analyst over-abstraction. Second, they give the eventual paper memorable language: reviewers and readers remember “wahda” in a way they do not remember refugee-specific-loneliness. Most well-written qualitative findings sections have at least three or four section headers built from in vivo codes. We will use Amira's “wahda” as a section header in the worked findings exercise of a later lesson.
3.5 Intercoder Reliability: Why It Matters
Once you have a codebook, the central question is: can someone else apply it the way you intended? Intercoder reliability is the operational answer to that question. Two (or more) analysts independently code the same passages using the same codebook; you compute the agreement; you decide whether the agreement is good enough.
Bernard, Wutich, and Ryan (Ch. 6) frame intercoder reliability as the most common operational standard for the third of the three methodological commitments (replicability, from an earlier lesson). It is not the only test of replicability, but it is the one most often expected by quantitatively trained reviewers in public-health journals. Methodologically, it serves three purposes: it forces the codebook to be specific enough to be applied consistently; it reveals which codes are unclear and need refinement; and it gives you a defensible number to report in your methods section.
3.6 Percent Agreement (Simple but Flawed)
The most intuitive measure of agreement: count the passages on which two coders agree, divide by the total number of passages coded, multiply by 100. If coders agree on 85 of 100 passages, percent agreement is 85%.
The problem with percent agreement is that it does not adjust for the agreement you would expect by chance. If two coders are using a codebook with only two codes (apply / do not apply), and one of the codes is used 90% of the time, two random coders would agree about 82% of the time by accident. An 85% percent-agreement score in that setting reflects only a few percentage points of real agreement above chance. The chance-adjusted measures below are designed to fix this problem.
3.7 Cohen's Kappa (Two Coders, Nominal Categories)
Cohen's kappa (Cohen, 1960) is the chance-corrected agreement measure used most often in qualitative health research. The formula is:
κ = (po − pe) / (1 − pe)
Where po is the observed proportion of agreement (the percent agreement, expressed as a decimal) and pe is the proportion of agreement expected by chance, given the marginal distributions of each coder's codings (that is, how often each coder used each code overall). Kappa ranges from −1 (perfect disagreement) through 0 (chance-level agreement) to +1 (perfect agreement). In plain terms, kappa asks how much of the coders' agreement is real rather than lucky: it takes the agreement they actually reached, subtracts the agreement two people would fall into by chance, and rescales what remains so that flawless agreement scores 1 and chance-level agreement scores 0.
Landis and Koch (1977) proposed interpretive thresholds for kappa that have become the field standard. They are guidance, not law; Bernard, Wutich, and Ryan are clear that the appropriate threshold depends on the stakes of the coding and the nature of the categories.
| Kappa value | Landis & Koch (1977) interpretation |
|---|---|
| < 0.00 | Poor (worse than chance) |
| 0.00–0.20 | Slight |
| 0.21–0.40 | Fair |
| 0.41–0.60 | Moderate |
| 0.61–0.80 | Substantial |
| 0.81–1.00 | Almost perfect |
The convention in applied health research is that kappa ≥ 0.60 (substantial) is a defensible threshold for publication, and kappa ≥ 0.80 (almost perfect) is excellent. If your kappa is below 0.60 on a given code, the standard response is to refine the codebook entry for that code (usually by tightening the inclusion or exclusion criteria) and re-code the disputed passages.
3.8 Krippendorff's Alpha (the Preferred Measure)
Cohen's kappa has three limitations: it handles only two coders, it does not gracefully handle missing data, and it assumes the codes are nominal categories (mutually exclusive and unordered). For projects with more than two coders, with intermittent missingness, or with codes that are ordered (e.g., severity levels), Cohen's kappa is the wrong tool.
Krippendorff's alpha (Krippendorff, 2018) generalizes the kappa idea to handle all three limitations. It accommodates any number of coders, any pattern of missing codings, and any level of measurement: nominal (unordered labels), ordinal (ranked levels such as low, medium, high), or interval and ratio (numeric scales). It is the measure most contemporary methodologists recommend, including Bernard, Wutich, and Ryan.
The formula is more complex than kappa's, since alpha is built on the difference between observed and expected disagreement computed from a coincidence matrix. For a small two-coder table it can still be worked through by hand, and any intercoder reliability calculator will compute it for larger tables. A later section works through a two-coder example. Interpretively, Krippendorff's own guidance is that α ≥ 0.80 is the usual standard for drawing firm conclusions, while 0.667 ≤ α < 0.80 supports only tentative conclusions. Below 0.667, the codebook is not yet reliable and needs revision.
Which measure should you use?
With two coders, nominal codes, and no missing data, Cohen's kappa is acceptable and is what most reviewers expect. If your codes are ordered (e.g., severity of loneliness on a 1–3 scale), or if you have three coders, or if there is missingness, use Krippendorff's alpha. Bernard, Wutich, and Ryan recommend defaulting to alpha because it generalizes: if you can compute alpha, you can report it for any future project, though kappa remains the most common reported statistic in published health qualitative work.
3.9 When Intercoder Reliability Is the Wrong Measure
Not all qualitative work asks for intercoder reliability. Bernard, Wutich, and Ryan are explicit about this and so are many methodologists writing in the interpretivist tradition. In two situations, the reliability framing is actually misleading.
The first is interpretivist or constructivist work. In Charmaz-style constructivist grounded theory (Charmaz, 2014), the analyst's interpretation is understood to be partly constitutive of what the data mean. Two competent analysts may legitimately reach different but defensible interpretations of the same passages. The standard of evaluation is not identity (the kappa standard) but coherence, that is, whether each interpretation is internally consistent, evidence-based, and methodologically transparent. For this kind of work, the equivalent quality check is investigator triangulation (multiple analysts compare interpretations and document where they differ and why) rather than a kappa score.
The second is narrative and discourse-analytic work (later lessons). Narrative analysis attends to the structure of a single telling; discourse analysis attends to the rhetorical work an utterance performs. Neither tradition typically treats codings as nominal categories applied independently to passages, so a kappa is not the right object. The methods sections of credible papers in these traditions usually report on member checking (Lesson 6, Section 2.8), on theoretical sensitivity, and on the writing trail rather than on kappa.
Reflexive thematic analysis (Section 1.4) belongs with these traditions, because it treats the researcher's interpretation as a resource and uses no fixed codebook to agree on. For a study that follows Bernard, Wutich, and Ryan's systematic approach and uses inductive-deductive hybrid coding, as the worked examples in this lesson do, intercoder reliability is the right measure, and Section 4 computes it. You should still know the defensible qualitative methodologies for which it is the wrong measure, so that you read the absence of a kappa in an interpretivist paper as a methodological choice.
3.10 The QDA Software Landscape
Before we move into the workflow in a later section, a quick orientation to the software landscape. There are four commercial options and a growing free-and-open-source ecosystem. None of them does anything you cannot do by hand on a small corpus; what they buy you is speed, consistency, and the ability to scale.
| Tool | Type | Strengths | Limitations |
|---|---|---|---|
| NVivo (Lumivero) | Commercial, desktop | Industry standard; excellent multi-coder workflow; rich visualization | Expensive licence; closed format; steep learning curve |
| ATLAS.ti | Commercial, desktop and cloud | Strong network views; good multimedia support | Expensive licence; some workflow quirks |
| MAXQDA | Commercial, desktop | Mixed-methods friendly; good visualizations | Expensive; smaller user base than NVivo |
| Dedoose | Commercial, browser-based, subscription | Cloud collaboration; lower monthly cost | Subscription model means access ends when you stop paying |
| Taguette (this course's pick) | Free, open-source, browser or local | Free; open data format (SQLite + CSV/HTML export); transferable beyond the course; runs locally | Fewer visualizations than commercial tools; smaller feature set |
We chose Taguette for this course because it is free (every student in the world can use it), it is open-source (your coded data are not held hostage by a licence), and its export format is standard (CSV and HTML, which any analysis tool can read). The features it lacks, namely advanced visualizations and sophisticated network views, can be produced from its CSV exports with a spreadsheet or a free diagramming tool, or drawn by hand. The combination is a complete, transferable workflow.
Reflection
Imagine you compute Cohen's kappa for a codebook you are developing with a second coder, and the result is κ = 0.52 (moderate) overall, with one specific code (somatic-absence) scoring κ = 0.31 (fair). What is your next step? Be concrete: what would you do to bring the kappa up?
somatic-absence; (2) read them side by side and articulate why you each coded as you did; (3) revise the inclusion and exclusion criteria so the disagreements would have been resolvable from the codebook text alone (e.g., specify that the bodily reference must be figurative not literal, or that a clear absence-link must be present); (4) re-code the disputed passages under the revised entry; (5) re-compute kappa on the same passages and on a fresh set. The overall kappa of 0.52 is also addressable: typically two or three codes are dragging the average down, and fixing them brings the overall up. Document all revisions in the audit trail. Do not silently change codings to inflate kappa; that defeats the entire purpose of intercoder reliability.Minimum 20 characters required.
Question 1: Why is percent agreement an inadequate measure of intercoder reliability on its own?
Question 2: Using Landis and Koch (1977) thresholds, a Cohen's kappa of 0.72 is interpreted as:
Question 3: Which is the primary advantage of Krippendorff's alpha over Cohen's kappa?
The Taguette Workflow on the Loneliness Dataset
Introduction and Overview
The first three sections of this lesson laid out the conceptual apparatus: themes versus codes, the twelve theme-finding techniques, inductive versus deductive coding, codebook architecture, coding mechanics, intercoder reliability. This section turns operational. You will see, end to end, how to find candidate themes in the loneliness corpus with word lists and keyword searches, how to build a codebook for them, how to apply the codebook in Taguette, how to summarize the exported coded extracts, and how to compute Krippendorff's alpha on a small two-coder reliability check. The section ends with what a first coding pass produces.
Learning Objectives for this section
- Use word frequencies and keyword-in-context (KWIC) concordances as theme-finding aids.
- Run a complete Taguette project: upload, code, export.
- Summarize coded extracts as code-frequency, code-by-participant, and co-occurrence tables.
- Compute percent agreement, Cohen's kappa, and Krippendorff's alpha from a two-coder agreement table.
- Describe the working documents that a first coding pass produces.
4.1 Step 1: Gather the Transcripts
The loneliness transcripts are in HSCI_841_loneliness_data.zip (the 20 interview transcripts, participant tables and codebooks). After you unzip it, the transcripts folder holds the 20 plain-text files (P01_Maya.txt through P20_Frank.txt). Each transcript opens with a metadata header (participant ID, age, gender, occupation, etc.) followed by the interview proper. Together the 20 files are the corpus you will work with in the steps below.
4.2 Step 2: Repetitions and KWIC as Theme-Finding Aids
Theme-finding technique 1 (Repetitions) and technique 10 (Word lists and KWIC) are easy to support with simple tools. A free word-frequency or concordance tool will count recurring words across the corpus and pull keyword-in-context concordances, and the search function of any text editor will find each occurrence of a word that catches your eye. The result is not the end of theme-finding; it is the front edge of it. The patterns the computer surfaces are then read closely by you, in their original transcripts, to decide whether they support a theme.
- Build a word list for the corpus: count how often each content word appears across the 20 transcripts, setting aside stopwords such as the, and, and of.
- Search the transcripts for chair and record each occurrence with about five words of context on either side.
- Repeat the search for words that emerged as candidate themes in an earlier section: tired, hollow, and fading.
- Search for the linguistic connectors because and since (Technique 6) and collect the causal accounts they introduce.
What to look for: The word list will surface obvious words (loneliness, alone, people, feel) and a few less obvious ones that turn into candidate themes. The chair search shows every chair-mention with its context, letting you check whether the chair-as-stand-in-for-absent-person reading is supported across transcripts beyond Linda's. In this corpus all four mentions of chair come from Linda's interview, so the reading belongs to her account and would need other evidence before it could be treated as a shared theme. The connector search pulls the causal accounts out of the corpus so they can be read together.
4.3 Step 3: Set Up the Taguette Project
Word lists and concordances surface patterns; Taguette is where you mark passages with codes. The two work together: the searches in Step 2 suggest what to look for, and Taguette records where you found it. The workflow below assumes you set up Taguette in Lesson 1 (Section 4.5); if not, return to those setup steps first.
The transcripts are in HSCI_841_loneliness_data.zip (the 20 interview transcripts, participant tables and codebooks).
- Open your Loneliness Interviews Taguette project (or create it if you have not).
- Upload the 3–5 transcripts you intend to code first. Recommended starter set: P01 Maya, P05 Linda, P11 Helen, P15 Amira, P20 Frank. The variation across these five is deliberate (age, gender, life-stage, immigration, life-circumstance) and gives you the widest analytic surface for codebook development.
- Create the parent codebook structure as a small set of broad codes:
experiential,causal,coping,interpretive(the four metacoding bins from an earlier section, Technique 12). - As you read transcript 1, highlight passages and tag them with provisional codes, drawing on the candidate themes you identified in Step 2 and on whatever else strikes you. Expect to create 20–40 provisional codes on this first transcript.
- After transcript 1, consolidate. Merge codes that turned out to be the same thing; rename codes whose names did not turn out to be right; sort codes under the four parent categories.
- Repeat for transcripts 2–5, expecting the codebook to grow more slowly each time (the curve flattens; this is theoretical saturation in miniature, which we will revisit in a later lesson).
- Once you have coded all 3–5 transcripts, export the codebook (Project → Codebook → Export) and the coded extracts (Project → Highlights → Export as CSV).
The Taguette export is the input to Step 4: the CSV contains one row per highlighted passage with columns for document, code, and the passage text.
4.4 Step 4: Summarize the Coded Extracts
Taguette's CSV export gives you a long-format table: one row per (passage, code) pair. Multi-coded passages appear in multiple rows. This is the format you want for almost any quantitative analysis of your qualitative coding: code frequencies, co-occurrence, code-by-participant matrices, comparison across subgroups.
The file taguette_export_week5.csv in HSCI_841_loneliness_data.zip is a worked Taguette export, and codebook_week5.csv holds its 17 codes. Open the export in a spreadsheet; each row gives the transcript (document), the code (tag), and the passage (content). Then build three tables, using a pivot table or by sorting and counting.
- Code frequencies: count how many times each code was applied across the corpus.
- Code by participant: count each code within each transcript, with one row per transcript and one column per code.
- Co-occurrence: for each pair of codes, count the transcripts in which both codes appear. This is transcript-level co-occurrence; passage-level co-occurrence can be counted the same way.
What the output gives you: The code-frequency table tells you which codes carried the most weight. The code-by-participant table is the basis of any subgroup comparison you might do in a later lesson. The co-occurrence table is the starting point for axial coding and for the network displays of a later module.
4.5 Step 5: Compute Krippendorff's Alpha for a Two-Coder Reliability Check
For a two-coder reliability check, you and a second coder independently code one shared transcript using the same provisional codebook. You then compute Krippendorff's alpha (or Cohen's kappa) on the resulting codings, by hand for a small table or with any intercoder reliability calculator. The worked example below uses a wide-format table where rows are passages and columns are coders, with the cell value being the code each coder assigned.
Two coders applied four codes to the same 12 passages: 1 = somatic-absence, 2 = coping-pet, 3 = affective-loneliness, and 4 = loneliness-as-cost-of-love. The codings are simulated for illustration; in practice the table comes from the two coders' Taguette exports.
| Passage | Coder A | Coder B | Agree? |
|---|---|---|---|
| 1 | 1 | 1 | Yes |
| 2 | 1 | 1 | Yes |
| 3 | 2 | 2 | Yes |
| 4 | 1 | 3 | No |
| 5 | 3 | 3 | Yes |
| 6 | 4 | 4 | Yes |
| 7 | 2 | 2 | Yes |
| 8 | 1 | 1 | Yes |
| 9 | 3 | 1 | No |
| 10 | 2 | 2 | Yes |
| 11 | 4 | 4 | Yes |
| 12 | 1 | 1 | Yes |
Percent agreement: the coders agree on 10 of 12 passages (83%). They disagree on Passage 4 (A: somatic-absence, B: affective-loneliness) and Passage 9 (A: affective-loneliness, B: somatic-absence).
Cohen's kappa: each coder used code 1 five times, code 2 three times, code 3 twice, and code 4 twice, so the agreement expected by chance is (5 × 5 + 3 × 3 + 2 × 2 + 2 × 2) / (12 × 12) = 42/144 = 0.29. Kappa is (0.83 − 0.29) / (1 − 0.29) = 0.76.
Krippendorff's alpha (nominal): pooling both coders gives 24 codings (code 1 ten times, code 2 six times, code 3 four times, and code 4 four times). Each of the two disagreements contributes two mismatched pairings, so observed disagreement is 4/24 = 0.17. Expected disagreement is 1 − (10 × 9 + 6 × 5 + 4 × 3 + 4 × 3) / (24 × 23) = 1 − 144/552 = 0.74. Alpha is 1 − 0.17/0.74 = 0.77.
Percent agreement (83%) overstates agreement, while Cohen's kappa (0.76) and Krippendorff's alpha (0.77) are lower because they correct for chance. Report alpha (or kappa) in a methods section, and report percent agreement only as a descriptive supplement.
4.6 The Iterative Cycle
The five steps above describe one pass through the workflow. In practice you will iterate. After computing reliability on a first transcript, you will revise the codebook for codes that scored poorly, re-code the disputed passages, and only then move on to the remaining transcripts. The audit trail (Section 2.6) records each iteration. A full analysis of the twenty transcripts usually goes through three or four iterations of this cycle, each documented, each producing a better codebook than the one before.
4.7 What a First Coding Pass Produces
A first pass through this workflow on three to five transcripts produces three working documents: a preliminary codebook of eight to twelve codes, each with all seven required elements; a Taguette project with those transcripts coded; and a one-page methods memo recording what the coding revealed, how the codebook changed, and what to revise next. In the terms of Section 1.4, these documents belong to the coding phase. Candidate themes are generated from them once more of the dataset has been coded, and the earlier design work on the research question, the sampling rationale, and the data-collection procedure now becomes analysis of transcripts.
Reflection
Imagine you have just finished coding three transcripts using a 10-code hybrid codebook. You compute Krippendorff's alpha with a second coder and you get α = 0.71 overall. What do you do next, and why? Be specific.
Minimum 20 characters required.
Question 1: A search that returns every occurrence of a target word with a window of surrounding words, so you can check whether participants use the word consistently, is a computational implementation of which theme-finding technique?
Final Assessment
Bringing It All Together
This lesson is the heaviest analytic lesson in the first half of this course. It is the lesson that converts the conceptual apparatus of the first four lessons (definitions, design, sampling, collection) into operational coding moves. The twelve theme-finding techniques (an earlier section), the six phases of thematic analysis (an earlier section), the inductive/deductive/hybrid distinction (an earlier section), the seven-element codebook entry (an earlier section), the four coding mechanics (an earlier section), and the intercoder reliability statistics (an earlier section) are tools you will use, in some combination, in every qualitative project for the rest of your career.
A first coding pass is the first piece of real analytic work on the loneliness dataset. The preliminary codebook it produces is revised in Lesson 6 when a conceptual model is added, expanded in Lesson 7 when constant-comparative analysis runs across subgroups, and stress-tested in Lessons 8 to 11, which run content, narrative, discourse, and analytic-induction passes on the same coded extracts. The codebook is the spine of any coding-based qualitative analysis, and this lesson is where it is built.
Key Takeaways from this lesson
- Themes are found, not discovered. The analyst notices, names, and bounds them. Replace “themes emerged” with “we identified the following themes” in your writing.
- Theme, code, category, and concept are different levels of abstraction. Code is the operational marker; theme is the recurring idea; category is the organizing bucket; concept is the theoretical idea. Codes come first (first-cycle coding), and themes are built from them (second-cycle coding).
- Thematic analysis moves through six recursive phases: familiarization, coding, generating initial themes, developing and reviewing themes, refining, defining, and naming themes, and writing up. Its coding-reliability, codebook, and reflexive variants differ in whether they use a codebook and whether they measure coder agreement.
- Ryan and Bernard (2003) catalogue twelve techniques for finding themes: repetitions, indigenous typologies, metaphors, transitions, similarities/differences, linguistic connectors, missing data, theory-related material, cutting and sorting, word lists and KWIC, co-occurrence, and metacoding.
- Coding strategies come in three flavours: inductive (data-up), deductive (theory-down), hybrid (the practical default in applied health research).
- A codebook entry has seven required elements: name, brief definition, full definition, inclusion criteria, exclusion criteria, positive example, negative example. A memo column is a strongly recommended eighth.
- Coding mechanics include hierarchical codes, multiple codes per passage, axial coding (relationships between codes), and in vivo codes (using participants' exact words).
- Intercoder reliability: percent agreement is intuitive but uncorrected for chance; Cohen's kappa is the field-standard chance-corrected measure for two coders with nominal codes; Krippendorff's alpha generalizes to any number of coders, any level of measurement, and missing data.
- Reliability is the wrong measure for interpretivist and narrative work, where the standard is coherence among defensible interpretations rather than identity of codings.
- The Taguette workflow is end-to-end transferable: word lists and concordances for theme discovery, Taguette for coding, and simple tables built from the exports for code-frequency and reliability analysis.
Core Concepts Reviewed
An earlier section: Theme vs code vs category vs concept; first-cycle and second-cycle coding; the six phases of thematic analysis and its three variants; the twelve Ryan and Bernard theme-finding techniques (repetitions, indigenous typologies, metaphors, transitions, similarities/differences, linguistic connectors, missing data, theory-related material, cutting and sorting, word lists and KWIC, co-occurrence, metacoding); worked examples from the loneliness corpus (chair, wahda, fading, somatic-absence).
An earlier section: Inductive vs deductive vs hybrid coding; the seven-element codebook entry (name, brief def, full def, inclusion, exclusion, +example, −example); the codebook as a living document with audit trail.
An earlier section: Hierarchical codes; multiple codes per passage; axial coding; in vivo codes; percent agreement; Cohen's kappa with Landis and Koch thresholds; Krippendorff's alpha; when reliability is the wrong measure; the QDA software landscape (NVivo, ATLAS.ti, MAXQDA, Dedoose, Taguette).
An earlier section: The end-to-end Taguette workflow: theme discovery with word frequencies and KWIC, Taguette coding, export and summary of coded extracts, Krippendorff's alpha from a two-coder table. What a first coding pass produces.
The final reflection below asks you to take stock of one specific decision you would make in your own coding work. Concrete, specific, and defensible: that is the standard any coding plan should meet.
Reflection
In one paragraph, set out a provisional coding plan for a first pass through the loneliness dataset. Name the transcripts you would code first, the coding strategy you would adopt (inductive, deductive, hybrid), three theme-finding techniques you would use, two anchor codes you anticipate, and how you would handle intercoder reliability with a second coder. Treat it as a working plan you could defend in a methods memo.
loneliness-vs-aloneness from Cacioppo and Patrick, and a structural-versus-existential typology code from the public-health loneliness literature), and let the rest develop inductively. The three theme-finding techniques I will use are repetitions (because a word list surfaces these quickly), metaphors (because the loneliness corpus is metaphor-dense), and missing data (to surface what participants avoid). My two anticipated inductive codes are somatic-absence (from Maya/Linda/Helen embodied metaphors) and coping-pet (from Linda's Rufus and Maya's neighbour's cat). For intercoder reliability, a second coder and I will independently code P11 (Helen) using my provisional codebook; I will compute Krippendorff's alpha from our two-coder table and aim for alpha ≥ 0.70 on this first pass, revising any code that scores below 0.60 before coding the remaining transcripts.” The point is to commit to a plan that is concrete enough to execute and to defend, with the recognition that it will be revised.Minimum 30 characters required.
Final Knowledge Assessment
Question 1: Bernard, Wutich, and Ryan's preferred phrasing in writing up qualitative findings is:
Question 2: How many techniques for finding themes do Ryan and Bernard (2003) catalogue?
Question 3: Amira's use of the Arabic word wahda to name a loneliness specific to refugee experience is an example of which theme-finding technique?
Question 4: What is the difference between a code and a category?
coping-pet, coping-phone-confidant, and coping-volunteer are codes within it.Question 5: A researcher begins with two codes drawn from a published theoretical framework, applies them to the first transcripts, and lets new codes emerge inductively as they read. Which coding strategy is this?
Question 6: Which of the following is NOT one of the seven required elements of a codebook entry?
Question 7: An in vivo code is:
“wahda” (Amira) and “witness-less hours” (Sarah) are in vivo codes. They keep the participant's voice in the codebook and give the eventual paper memorable language.Question 8: What does axial coding attend to?
Question 9: Why is simple percent agreement an inadequate measure of intercoder reliability on its own?
Question 10: Using Landis and Koch (1977) thresholds, a Cohen's kappa of 0.45 is interpreted as:
Question 11: Krippendorff's alpha has at least three advantages over Cohen's kappa. Which of the following is NOT one of them?
Question 12: Intercoder reliability is the wrong measure for which kind of qualitative work?
Question 13: Why did this course choose Taguette over NVivo or ATLAS.ti?
Question 14: Which practice belongs to coding-reliability thematic analysis rather than to reflexive thematic analysis?
Glossary: Themes, Codes, Codebooks & Reliability
📚 Reference page, available throughout the lesson
This glossary collects the key concepts, people, and methodological terms introduced in this lesson. Use it as a reference while you work through the material, or as a review before the final assessment. Type in the search box to filter entries.
coping-pet, coping-phone-confidant, and so on.
“wahda” (Amira), “witness-less hours” (Sarah), “fading at the edges” (Helen).