phenomena
The Thought Language and Communication Scale, Explained
The Thought Language and Communication scale (TLC) is Andreasen's glossary of disordered speech. See the items, the reliability data, and how to code them.

The TLC is short for the Thought Language and Communication scale, Nancy Andreasen’s 1986 glossary that gave operational definitions to the disorders of disorganized speech—derailment, tangentiality, poverty of content, and more (Andreasen, 1986). Its purpose was reliability. Clinicians had been using “formal thought disorder” to mean wildly different things, so Andreasen replaced the fog with concrete, observable categories that two raters could actually agree on. The result is the closest thing psychiatry has to a shared dictionary for disordered speech—and, as it turns out, a ready-made annotation scheme.
What is the Thought Language and Communication scale?
The Thought Language and Communication scale is a set of operationally defined categories for rating disorders of thought, language, and communication from a person’s speech (Andreasen, 1986). Nancy Andreasen built it to fix a measurement problem: the term “formal thought disorder” had been used so loosely that two clinicians rating the same interview could not be relied on to agree.
Her solution was linguistic as much as clinical. She recommended retiring the umbrella term entirely: “Because the term ‘formal thought disorder’ has been so misunderstood and misused, it is recommended that it no longer be used,” and its parts “can be better conceptualized as ‘disorders of thought, language, and communication’” (Andreasen, 1986). Each disorder got a one-paragraph definition and verbatim examples, so a rater is matching speech against a concrete description rather than a hunch.
The original scale defines 18 scored items (plus two phonemic and semantic paraphasia items that are defined but excluded from the global rating). Tagaroo’s library models the 12 that most often carry pathological weight and are cleanest to ground in a transcript. You can see the full set of definitions on the TLC scale page.
What are the TLC items?
The core TLC items name distinct ways speech can break down—in how much is said, how ideas connect, and how words are chosen. The table below defines the twelve items Tagaroo models, each phrased so you can recognize it in a transcript.
| Item | Operational definition | Type |
|---|---|---|
| Derailment | Ideas slip off track onto obliquely related or unrelated topics, without the speaker noticing | Positive |
| Tangentiality | Replying to a question in an oblique or irrelevant way; the question is never answered | Positive |
| Circumstantiality | Very indirect, over-detailed speech that eventually does reach its point | Positive |
| Illogicality | Conclusions that don't follow—non-sequiturs and faulty inferences | Positive |
| Pressure of Speech | Increased, rapid speech that is hard to interrupt | Positive |
| Distractible Speech | Stopping mid-sentence and changing topic in response to a nearby stimulus | Positive |
| Clanging | Word choice governed by sound rather than meaning—rhyming, punning | Positive |
| Neologisms | Completely new word formations whose derivation can't be understood | Positive |
| Perseveration | Persistent, inappropriate repetition of words, ideas, or subjects | Mixed |
| Loss of Goal | Failure to follow a chain of thought to its natural conclusion; the point is never reached | Mixed |
| Poverty of Speech | Restriction in the amount of spontaneous speech; brief, unelaborated replies | Negative |
| Poverty of Content | Adequate amount of speech that conveys little information—vague, empty | Negative |
Two things make this a genuine glossary rather than a checklist. Each item is defined by what the speech does, not by what disorder it implies, which is why the same vocabulary works across schizophrenia, mania, and depression. And the definitions are deliberately narrow—Andreasen split the old catch-all “loosening of associations” into derailment and tangentiality precisely so raters would stop using one word for two different things.
Why does the positive vs negative split matter?
The TLC items divide into positive and negative thought disorder, and the division is not cosmetic—it carries diagnostic and prognostic weight. Positive FTD is speech with too much disorganized structure: derailment, tangentiality, illogicality, incoherence. Negative FTD is speech with too little: poverty of speech (alogia) and poverty of content.
The two dissociate by diagnosis. Andreasen and Grove, comparing 94 controls with 100 patients, found positive FTD more prominent in affective psychosis (especially mania), while negative FTD was more characteristic of schizophrenia and carried the worse prognosis (Andreasen & Grove, 1986). Negative thought disorder tends to persist; the positive kind, as in mania, often remits.
That enormous prevalence range, from a systematic review of 120 studies, is the strongest argument for a scale like the TLC (Roche et al., 2015). When estimates for the “same” phenomenon span from 5% to 91%, the definition and the instrument are doing more work than the biology. A shared, operational vocabulary is how you stop measuring your own method.
How do you tell derailment from tangentiality?
Derailment is drift within the speaker’s own spontaneous speech; tangentiality is an off-target reply to a question—the discriminator is where the drift starts (Andreasen, 1986). Those two, plus a couple of other confusable pairs, account for most coding disagreements, and Andreasen’s definitions settle each one cleanly.
- Derailment vs tangentiality. Both descend from the retired “loosening of associations.” Derailment is slippage within the speaker’s own spontaneous speech; tangentiality is an oblique reply to a question. The discriminator is where the drift starts: the speaker’s own train of thought (derailment) or an unanswered question (tangentiality) (Andreasen, 1986).
- Poverty of speech vs poverty of content. Poverty of speech is a deficit of quantity—too few words, brief and unelaborated. Poverty of content is enough words carrying little information: vague, over-abstract, “empty philosophizing.” One counts words; the other counts ideas per word (Andreasen, 1986).
- Circumstantiality vs loss of goal. Both wander, but circumstantiality eventually reaches the point, buried under tedious detail, while loss of goal never arrives. The discriminator is arrival: yes for circumstantial, no for loss of goal (Andreasen, 1986).
These distinctions are exactly where inter-rater reliability is won or lost. If two coders share the discriminator—“did the reply answer the question or not?”—they agree; if they each carry a private definition, they don’t. It is the same problem Andreasen set out to solve, reproduced at the level of a single ambiguous item.
How reliable is the TLC scale?
TLC reliability is high for common items and poor for rare ones, and that pattern is the most useful thing to know before you use it. In Andreasen’s own study of 113 patients, the frequent, prototypical signs were rated with strong agreement: pressure of speech κ = 0.89, incoherence κ = 0.88, derailment κ = 0.83, poverty of speech κ = 0.81 (Andreasen, 1979, 1986).
| TLC item | Kappa | Reliability |
|---|---|---|
| Pressure of speech | 0.89 | High |
| Incoherence | 0.88 | High |
| Derailment | 0.83 | High |
| Poverty of speech | 0.81 | High |
| Illogicality | 0.80 | High |
| Tangentiality | 0.58 | Fair |
| Clanging | 0.58 | Fair |
| Neologisms | 0.39 | Poor |
| Word approximations | −0.02 | None |
The rare items tell the cautionary half of the story. Neologisms managed only κ = 0.39, and word approximations came in at κ = −0.02—no better than chance (Andreasen, 1986). Andreasen’s own read was that some signs “occur so infrequently as to be of little diagnostic value.” The lesson for any coding project: an item’s reliability depends on how often it actually appears, so budget your calibration effort on the frequent items and treat the rare ones with caution. If you are computing agreement yourself, the Cohen’s kappa and inter-rater reliability guide covers why a low kappa on a rare item is often a base-rate artifact, not a coding failure.
How does the TLC feed automated speech analysis?
The TLC’s operational definitions have become the phenotype for a striking line of computational work: predicting psychosis from the structure of speech. Because Andreasen defined derailment as ideas slipping off track and poverty of content as speech that conveys little, those definitions translate almost directly into measurable language features.
In a cohort of 34 clinical-high-risk youths, a classifier built from latent-semantic-analysis coherence plus two syntactic markers predicted transition to psychosis with 100% accuracy in that sample, outperforming clinical ratings (Bedi et al., 2015). A larger follow-up across protocols generalized the finding at 83% accuracy, explicitly mapping reduced semantic coherence to derailment and tangentiality (Corcoran et al., 2018). The small samples mean these are proofs of concept, not deployable tools—but they show how directly Andreasen’s 1986 vocabulary feeds modern natural language processing, a field now reframing thought disorder as a dimensional construct running from phenomenology to neurobiology (Kircher et al., 2018). The through-line is worth stating plainly: her kappas quantified where human raters agree; the NLP work is the attempt to compute the same signs automatically.
How do you code thought disorder from a transcript?
To code the TLC from a transcript, mark the specific utterance that shows the disorder and label it with the matching item—this is instance-mode annotation, where you tag the speech that is the phenomenon rather than rating the interview as a whole. Disorganized speech is, almost by definition, a span-level event.
Consider a short synthetic exchange:
Interviewer: How have you been sleeping this week?
Subject: Sleep’s alright, mostly, though the moon was so bright, and my brother always said the tide is why the bakery shut early on Sundays.
That reply drifts from sleep to the moon to a bakery, obliquely and without the speaker noticing—a clean instance of derailment, tagged on the span “the moon was so bright … the bakery shut early.” Marking the span keeps the evidence attached to the label, so a second coder can check the call against the exact words. That is what makes the TLC coding scheme reproducible enough to compute agreement across raters, and it is the same instance-mode logic behind coding pressured speech in the Young Mania Rating Scale.
On data handling: coding interview speech means working with sensitive clinical language, so the sane default is de-identified text and a privacy-first setup. Tagaroo supports a browser-side anonymous mode, so transcript content can stay local rather than being uploaded—worth checking against your ethics approval before any real interview data touches a tool. When disorganized speech appears alongside broader psychopathology, the two thought-disorder items in the Brief Psychiatric Rating Scale (BPRS)—conceptual disorganization and unusual thought content—use the same span-level approach.
Common mistakes when coding the TLC
The most common mistake is chasing the exotic items (neologisms, clanging) while missing the common, reliable ones, but a few others recur:
- Coding content as form. A bizarre belief is a delusion (content); disorganized structure is thought disorder (form). The TLC rates form. Keep the two axes separate.
- Merging derailment and tangentiality. They are separated by their trigger, not their feel. Ask whether the drift began in spontaneous speech or in a reply to a question.
- Reading poverty of content as poverty of speech. A long, fluent, empty answer is poverty of content, not poverty of speech. Count information, not just words.
- Over-trusting rare-item ratings. With word approximations near κ = 0, a single such rating is weak evidence (Andreasen, 1986). Require clear, repeated instances before you score the rare items.
None of these are reasons to distrust the TLC—it is the field’s most carefully operationalized vocabulary for disordered speech. They are reasons to code it the way it was designed: match speech to the definition, tag the span that justifies the call, and weight your confidence by how reliable each item actually is.
The practical upshot: the Thought, Language and Communication scale works because it turned a vague label into a glossary of observable signs. Code the frequent, reliable items with confidence, anchor every rating to the utterance that shows it, and remember that a shared discriminator—not a shared intuition—is what makes two coders agree.
References
- Andreasen, N. C. (1986). Scale for the Assessment of Thought, Language, and Communication (TLC). Schizophrenia Bulletin, 12(3), 473–482. doi:10.1093/schbul/12.3.473
- Andreasen, N. C. (1979). Thought, language, and communication disorders. I. Clinical assessment, definition of terms, and evaluation of their reliability. Archives of General Psychiatry, 36(12), 1315–1321. doi:10.1001/archpsyc.1979.01780120045006
- Andreasen, N. C., & Grove, W. M. (1986). Thought, language, and communication in schizophrenia: diagnosis and prognosis. Schizophrenia Bulletin, 12(3), 348–359. doi:10.1093/schbul/12.3.348
- Roche, E., Creed, L., MacMahon, D., Brennan, D., & Clarke, M. (2015). The epidemiology and associated phenomenology of formal thought disorder: a systematic review. Schizophrenia Bulletin, 41(4), 951–962. doi:10.1093/schbul/sbu129
- Kircher, T., Bröhl, H., Meier, F., & Engelen, J. (2018). Formal thought disorders: from phenomenology to neurobiology. The Lancet Psychiatry, 5(6), 515–526. doi:10.1016/S2215-0366(18)30059-2
- Bedi, G., Carrillo, F., Cecchi, G. A., Slezak, D. F., Sigman, M., Mota, N. B., et al. (2015). Automated analysis of free speech predicts psychosis onset in high-risk youths. npj Schizophrenia, 1, 15030. doi:10.1038/npjschz.2015.30
- Corcoran, C. M., Carrillo, F., Fernández-Slezak, D., Bedi, G., Klim, C., Javitt, D. C., et al. (2018). Prediction of psychosis across protocols and risk cohorts using automated language analysis. World Psychiatry, 17(1), 67–75. doi:10.1002/wps.20491
If you code disorganized speech from clinical interviews, Tagaroo turns the TLC scale into a guided, evidence-anchored annotation workflow—with inter-rater reliability computed as your coders work.
Frequently asked questions
- What is the Thought, Language and Communication (TLC) scale?
- The TLC is a scale by Nancy Andreasen that gives operational definitions to 18 disorders of thought, language, and communication—such as derailment, tangentiality, and poverty of speech—so clinicians can rate disorganized speech reliably (Andreasen, 1986). Andreasen proposed replacing the vague term 'formal thought disorder' with these concrete, observable categories.
- What is the difference between derailment and tangentiality?
- Derailment is slippage within a person's own spontaneous speech—ideas slide onto obliquely related or unrelated topics without the speaker noticing (Andreasen, 1986). Tangentiality is an oblique or irrelevant reply to a question, where the question is never answered. The discriminator is the trigger: a question (tangentiality) versus the speaker's own train of thought (derailment).
- What is the difference between poverty of speech and poverty of content?
- Poverty of speech is a restriction in the amount of speech—brief, unelaborated replies (this is alogia, a negative symptom). Poverty of content of speech is an adequate amount of speech that conveys little information—vague, repetitive, or empty (Andreasen, 1986). One is a deficit of quantity; the other is a deficit of information at normal quantity.
- How reliable is the TLC scale?
- In Andreasen's own reliability study, common items were rated with high agreement—pressure of speech kappa 0.89, incoherence 0.88, derailment 0.83—while rare items fared poorly, such as neologisms at 0.39 and word approximations at −0.02 (Andreasen, 1986). The pattern matters: reliability is high for the frequent, prototypical signs and low for the rare ones.
- Is formal thought disorder the same as thought disorder?
- Formal thought disorder refers to the form or structure of speech—how ideas are organized and connected—rather than the content of beliefs (which is delusion). Andreasen argued the umbrella term was so misused it should be retired in favor of specific 'thought, language, and communication disorders' (Andreasen, 1986). The TLC scale is her operational vocabulary for those signs.
Put this into practice
Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.