methods
Gottschalk-Gleser Content Analysis: Coding Affect in Speech
How Gottschalk-Gleser content analysis scores affect in speech, clause by clause—the seven depression subscales, the weighted coding, and its NLP legacy.

Long before a neural network scored a tweet for sentiment, a psychiatrist sat with a typed transcript and a pencil, tagging one clause at a time. Gottschalk-Gleser content analysis is that method: a system from the 1960s for measuring affect from what people actually say, clause by clause, with weighted scores. It reads less like a questionnaire and more like a labeling scheme—every scorable clause is a span, tagged with a category and a weight. In hindsight, that makes it one of the earliest serious attempts at the computational linguistics of affect, and a close cousin of the way an annotation tool tags text today.
The design choice that still matters is the weighting. A clause where the speaker blames themselves counts more than one blaming others; an active suicidal wish counts more than either. Those weights turn a transcript into numbers, and the numbers were reliable enough that trained coders were held to an inter-scorer agreement of 0.80 or better (Gottschalk & Gleser, 1969).
What is Gottschalk-Gleser content analysis?
Gottschalk-Gleser content analysis is a method for quantifying psychological states—anxiety, hostility, hopelessness, and more—from the content of a person’s speech, developed by psychiatrist Louis Gottschalk and psychologist Goldine Gleser (Gottschalk & Gleser, 1969). Instead of asking a patient to rate themselves, it takes a short, standardized speech sample and scores what the person says.
The standard elicitation is disarmingly simple: the subject is asked to talk for about five minutes about “any interesting or dramatic personal life experience” (Gottschalk & Gleser, 1969). That ambiguous prompt is deliberate—it invites projection, so the content reveals affect the person might not report on a checklist. The recording is transcribed, and the transcript becomes the data.
What separates the method from a word-counting dictionary is its unit of analysis. Scoring happens clause by clause, so the coder can capture the relationship inside a sentence—who did or felt what, about whom—rather than tallying emotionally charged words in isolation. That relational reading is the whole point, and it’s why the same word can score differently depending on the clause it sits in.
How does Gottschalk-Gleser content analysis scoring work?
Gottschalk-Gleser scoring assigns each scorable clause to a thematic category, multiplies by a weight that reflects how strongly the clause signals the state, corrects for how much the person spoke, and sums the result into a magnitude score (Gottschalk & Gleser, 1969). The correction for verbal output matters: without it, a talkative subject would score higher on everything simply for producing more clauses.
For the depression scale, the per-subscale score is computed as sqrt((weighted_sum + 0.5) × 100 / word_count), and the seven subscale scores are summed into a total depression score (Gottschalk & Hoigaard-Martin, 1986). You don’t need to memorize the formula; the shape is what counts. Weighted evidence, normalized by word count, turned into a number that’s comparable across people and across sessions for the same person.
Here is the seven-subscale structure of the depression scale, with the weighting logic for each. Flat-weight subscales score a present clause as 1; the perspective-graded subscales scale the weight by whose experience the clause describes.
| Subscale (abbr.) | What the clause references | Weighting |
|---|---|---|
| Hopelessness (HOP) | Despair, futility, lack of hope or support | Flat 1 |
| Self-accusation (SAC) | Guilt, shame, hostility turned inward | 1 denial/impersonal → 2 others → 3 self → 4 active suicidal |
| Psychomotor retardation (PMR) | Slowing in thinking, feeling, or action | Flat 1 |
| Somatic concerns (SOM) | Bodily complaints; sleep, appetite, energy | Flat 1 |
| Death & mutilation (DAM) | Death, dying, injury, physical damage | 1 inanimate → 2 others → 3 self |
| Separation (SEP) | Abandonment, loss, ostracism | 1 inanimate → 2 others → 3 self |
| Hostility outward (HOS) | Aggression and criticism toward others | 1 covert/impersonal → 2 others → 3 overt from self |
Why does the method weight self-references more heavily?
The weights encode psychological proximity: the closer a piece of content is to the self, the more strongly it counts. On subscales like self-accusation, death and mutilation, and separation, a clause about the speaker is weighted 3, a clause about other people is weighted 2, and an impersonal or denied reference is weighted 1 (Gottschalk & Hoigaard-Martin, 1986).
The clearest case is self-accusation, where an active suicidal wish reaches a weight of 4—the highest in the depression scale. That ordering is a clinical claim rendered as arithmetic: “I’d be better off dead” is treated as a stronger depressive signal than “I feel guilty,” which in turn outranks “people can be too hard on themselves.” The method makes the coder decide whose experience a clause is about, then lets that decision move the score.
This is exactly the judgment that a flat keyword count throws away. “Death” in “my grandfather’s death” and “death” in “I wish I were dead” are the same token but very different signals, and the weighted, perspective-graded design is what keeps them apart.
The seven subscales of the depression scale
The depression scale decomposes low mood into seven distinct thematic streams, so two people with the same total can still have very different profiles (Gottschalk & Hoigaard-Martin, 1986). Hopelessness captures despair and futility; self-accusation captures guilt, shame, and hostility turned inward, up to active suicidal content; psychomotor retardation captures a sense of slowing in thought or action.
The remaining four cover the somatic and interpersonal texture of depression. Somatic concerns picks up bodily complaints and problems with sleep, appetite, or energy, while death and mutilation covers references to dying, injury, and physical damage.
Then separation covers abandonment, loss, and ostracism, and hostility outward covers aggression and criticism aimed at others—included because externalized anger is part of many depressive presentations, not separate from them.
Reading the profile rather than the single number is the practical skill here. A transcript loaded on separation and hopelessness reads differently from one loaded on somatic concerns and psychomotor retardation, even at an identical total—the same reason a PHQ-9 total of 14 can describe two clinically distinct people.
What does the Hope Scale add?
The Hope Scale is the optimistic mirror image of the depression work: a companion Gottschalk-Gleser scale that scores a verbal sample for expressed hope of favorable outcomes (Gottschalk, 1974). It matters because affect isn’t only about counting deficits—the presence of hope carries information a symptom checklist tends to miss.
Its validation is what makes hope scale content analysis worth knowing. In the original study, Hope scores correlated negatively with Hamilton and BPRS depression ratings and, more strikingly, predicted favorable clinical outcome and longer survival time in patients with terminal cancer (Gottschalk, 1974). A positively-keyed content-analysis measure that tracks outcome is a genuinely different lens from a deficit count, and it rounds out the picture the depression subscales give you.
Is Gottschalk-Gleser content analysis reliable?
Yes—with a real caveat about the cost of getting there. Gottschalk and Gleser held trained coders to an inter-scorer reliability of 0.80 or better, and the scales accumulated construct-validation evidence against psychological, physiological, pharmacological, and biochemical criteria over decades (Gottschalk & Gleser, 1969). The method also validated across languages and cultures, which is unusual for an instrument this tied to natural speech (Gottschalk & Lolas, 1989).
That 0.80 bar is the honest part of the story. Reaching it takes practice against previously scored samples and ongoing monitoring, which is why the method never became routine clinical practice despite its research track record. If you plan to score verbal content analysis of depression with more than one coder, treat agreement as something to measure, not assume—the same discipline covered in our guide to Cohen’s kappa and inter-rater reliability. A weighted coding scheme, where two coders can agree a clause is self-accusation but split on whether it’s weight 3 or 4, is precisely where a weighted agreement coefficient earns its keep.
The original computational linguistics of affect
Here’s the part most summaries skip: the Gottschalk-Gleser scales were being machine-scored in 1975. Collaborating with computer scientists, Gottschalk used Woods’ Augmented Transition Network parser on a PDP-10 mainframe to score the Hostility Outward scale from typescripts, and the computer’s scores correlated about 0.80 with expert human coders—matching the lower bound of acceptable human agreement (Gottschalk, Hausmann & Brown, 1975).
What’s remarkable is why it was hard, because the same problem defines NLP today. Earlier automated systems, like Philip Stone’s General Inquirer, tagged content word by word—and, as the Gottschalk group put it, that approach throws away who did what to whom, misreads idioms like “kick the bucket,” and can’t tell “I’ll get you” from “get” as a neutral verb. Their fix was to parse each clause, identify the actor and the recipient of the action verb, and score the relationship. That is clause-level, relationship-aware semantics—the exact thing token-classification models spent the 2010s rediscovering.
| Method | Unit of analysis | Relationship & perspective aware? | Weighted? | Handles idiom (e.g. “kick the bucket”)? |
|---|---|---|---|---|
| Gottschalk-Gleser | Grammatical clause | Yes—scores who did or felt what about whom | Yes—1–4 by psychological proximity | Yes—the clause is read for meaning, not matched token by token |
| General Inquirer (Stone et al., 1966) | Single word | No—a dictionary tag per word | No—category counts | No—matches words in isolation |
| LIWC (Tausczik & Pennebaker, 2010) | Single word | No—bag-of-words category counts | No—percentage of words per category | No—dictionary match, no clause parse |
Frame it that way and the lineage is clear. Psycholinguistic content analysis of the Gottschalk-Gleser kind is a direct ancestor of modern affective computing: a hand-built, weighted, span-level annotation scheme, validated against outcomes, running on the premise that meaning lives in structured clauses rather than isolated words. The tools changed; the core representation did not.
Coding a verbal sample with an AI agent today
The method’s design is an unusually good match for a modern annotation workspace, because it was always a span-tagging task. Consider a short synthetic sample—the kind of thing the five-minute prompt produces:
Prompt: Tell me about something that’s been on your mind lately.
Subject: I keep letting everyone down at work. I’ve barely slept in weeks. And honestly, some days I think they’d all be better off without me around.
Three clauses, three different codes. “I keep letting everyone down” is self-accusation directed at the self—weight 3. “I’ve barely slept in weeks” is a somatic concern—a flat weight of 1. And “they’d all be better off without me” is self-accusation crossing into active suicidal content—weight 4, the clause you never want a total to hide.
| Tagged clause (synthetic) | Subscale | Weight |
|---|---|---|
| “I keep letting everyone down” | Self-accusation (SAC)—directed at self | 3 |
| “I've barely slept in weeks” | Somatic concerns (SOM)—flat | 1 |
| “they'd all be better off without me” | Self-accusation (SAC)—active suicidal | 4 |
An agent that tags each clause with a subscale, a weight, and the exact span of text behind it produces an auditable score: a reviewer can check every weighted decision against the words that justified it.
That auditability is the reason weighted span annotation belongs in software rather than on paper. The coder—human or agent—attaches a rationale and a span to each clause, disagreements surface at the clause level where reliability is actually measured, and the arithmetic is done for you. It’s the 1969 workflow, minus the pencil and the post-processing.
Limitations and honest cautions
The method’s strengths are also its constraints. Because it scores what a person chooses to say in five minutes, it’s sensitive to how talkative and how guarded someone is on the day, and the word-count correction only partly offsets that. It depends on a good transcript, so transcription errors and disfluencies propagate into the score. And the training burden behind that 0.80 reliability is real—this is a research instrument, not a bedside quick-screen (Gottschalk & Gleser, 1969).
It is also, firmly, not a diagnostic test. A magnitude score locates a verbal sample relative to normative data; it does not confirm a disorder, and turning content-analysis output into clinical decisions is a job for a clinician with the full picture. The value is in measurement—tracking change over time, comparing conditions, quantifying affect that a self-report scale can’t reach—not in labeling. When you do want that self-report view, pair it with a matched instrument; for anxiety, our guide to the GAD-7 anxiety scale walks through one.
The practical upshot: Gottschalk-Gleser content analysis is best understood as weighted, clause-level span annotation with a fifty-year evidence base—a method built for exactly the kind of auditable, agreement-checked tagging that annotation software now makes practical. Pick it when you want to measure affect in someone’s own words, and pair it with a self-report scale when you want the fuller view.
References
- Gottschalk, L. A., & Gleser, G. C. (1969). The Measurement of Psychological States Through the Content Analysis of Verbal Behavior. University of California Press. doi:10.1525/9780520376762
- Gottschalk, L. A. (1974). A Hope scale applicable to verbal samples. Archives of General Psychiatry, 30(6), 779–785. doi:10.1001/archpsyc.1974.01760120041007
- Gottschalk, L. A., Hausmann, C., & Brown, J. S. (1975). A computerized scoring system for use with content analysis scales. Comprehensive Psychiatry, 16(1), 77–90. doi:10.1016/0010-440X(75)90024-3
- Gottschalk, L. A., & Hoigaard-Martin, J. (1986). A depression scale applicable to verbal samples. Psychiatry Research, 17(3), 213–227. (Building on Gottschalk, Winget & Gleser, 1969.)
- Gottschalk, L. A., & Lolas, F. (1989). The Gottschalk-Gleser content analysis method of measuring the magnitude of psychological dimensions: its application in transcultural research. Transcultural Psychiatric Research Review, 26(2), 83–111. doi:10.1177/136346158902600201
- Stone, P. J., Dunphy, D. C., Smith, M. S., & Ogilvie, D. M. (1966). The General Inquirer: A Computer Approach to Content Analysis. MIT Press. mitpress.mit.edu
- Tausczik, Y. R., & Pennebaker, J. W. (2010). The psychological meaning of words: LIWC and computerized text analysis methods. Journal of Language and Social Psychology, 29(1), 24–54. doi:10.1177/0261927X09351676
If you code affect from interview transcripts, Tagaroo turns weighted, clause-level schemes like the Gottschalk-Gleser depression scale into a guided, evidence-anchored annotation workflow—every weighted decision tied to the span that justifies it, with inter-rater reliability computed as your coders work.
Source & attribution: this Gottschalk-Gleser content-analysis overview is described from published sources. The original Gottschalk-Gleser scoring manual is © 1969 The Regents of the University of California (University of California Press), and this article reproduces no proprietary scoring content.
Frequently asked questions
- What is the Gottschalk-Gleser method?
- It is a content-analysis system for measuring psychological states from short samples of natural speech, developed by Louis Gottschalk and Goldine Gleser (Gottschalk & Gleser, 1969). A trained scorer reads a transcript and tags each grammatical clause against defined thematic categories—anxiety, hostility, hopelessness, and others—then applies weights and a word-count correction to produce a magnitude score. The unit of analysis is the clause, not the isolated word, which is what lets it capture who felt what about whom.
- What does the Gottschalk-Gleser depression scale measure?
- The depression scale scores depressive thematic content across seven subscales: hopelessness, self-accusation, psychomotor retardation, somatic concerns, death and mutilation, separation, and hostility directed outward (Gottschalk & Hoigaard-Martin, 1986). Each scorable clause is tagged to a subscale and weighted, and the weighted counts are combined into per-subscale and total depression scores.
- Why are some clauses weighted more heavily than others?
- The weight encodes psychological proximity. Content a speaker attributes to themselves counts more than content attributed to others or to impersonal situations, and on the self-accusation subscale an active suicidal wish carries the highest weight of all (Gottschalk & Hoigaard-Martin, 1986). The idea is that 'I am worthless' is a stronger signal of depressive affect than 'people can be hard on themselves,' so the scoring reflects that difference numerically.
- What is the Hope Scale in content analysis?
- The Hope Scale is a companion Gottschalk-Gleser scale that measures optimism about favorable outcomes from a verbal sample (Gottschalk, 1974). In validation work, Hope scores correlated negatively with Hamilton and BPRS depression ratings and predicted favorable clinical outcome and longer survival time in terminal cancer patients—a positively-keyed counterpart to the depression scale.
- Is Gottschalk-Gleser content analysis still used?
- Yes, primarily in research. The manual method remains valid but is labor-intensive, so most modern use runs through computerized scoring, an effort that began in 1975 and continues in software such as PCAD (Gottschalk, Hausmann & Brown, 1975). Its clause-level, relationship-aware design also makes it a useful historical reference point for today's NLP and LLM work on affect.
Put this into practice
Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.