tagaroo

methods

DSM Criteria Annotation: From Diagnosis to Taggable Spans

Turn DSM-5 and ICD-11 criteria into observable, taggable phenomena: a worked MDD-to-PHQ-9 mapping and why you annotate symptoms, not a diagnosis.

Enrique Gutiérrez14 min readUpdated July 2026
One large diagnostic-label shape unfolding into many small bracketed, color-tagged spans over an abstract transcript band, with a single coral-highlighted tag — showing a diagnosis resolved into observable annotatable phenomena.

A diagnosis is a decision a clinician reaches. An observable phenomenon is a span of text you can point to and label. DSM criteria annotation is the work in between: taking the symptom language inside a diagnostic manual and turning it into categories a coder, or an assisted agent, can attach to real words in a transcript. Done well, it yields span-level labels that map toward a criterion without ever claiming the diagnosis itself.

The gap between those two things is wider than it looks. In the STAR*D sample, 3,703 people who all met criteria for one disorder, major depressive disorder, produced 1,030 distinct symptom profiles, and about one in seven had a profile shared with no one else in the study (Fried & Nesse, 2015). A single diagnostic label collapses that variation into one word. Annotation is how you get the variation back.

What DSM criteria annotation actually means

DSM criteria annotation is the process of converting the symptom concepts inside a diagnostic manual into observable, span-taggable categories, then labeling the places in a transcript where those symptoms show up. It operates one level below the diagnosis. The manual tells a clinician when a set of symptoms, at a certain count and duration, adds up to a disorder; annotation stops at the symptom and asks a narrower question: does this span of speech express this phenomenon, yes or no, and if so, how strongly?

That narrower question is the whole point. A diagnostic criterion is written to support a categorical judgment made by a trained clinician about a whole person. A phenomenon code is written to support a repeatable labeling decision made by any coder about a single stretch of text. The two are related but not interchangeable, and treating them as the same thing is where criteria-based coding projects quietly go wrong.

PropertyA diagnosisAn annotatable phenomenon
What it isA categorical clinical judgment about a personAn observable span of text expressing one symptom
Who assigns itA qualified clinicianAny trained coder or an assisted agent
UnitThe whole person, over a time windowA span, utterance, or speaker turn
Depends onCounts, thresholds, duration, distress, rule-outsThe words present in one stretch of transcript
What you can point toAn inference, not a location in the textA start and end offset in the transcript
Reproducibility checkDiagnostic reliability studiesInter-rater agreement on the span label
Diagnosis versus annotatable phenomenon. Annotation lives entirely in the right-hand column: a labeled location in the text, not a judgment about a person (framing after APA, 2022; Fried & Nesse, 2015).

Why you annotate phenomena, not a diagnosis

You annotate phenomena because the diagnosis throws away the information an annotation set exists to preserve. A diagnostic label is a compression: it takes a rich, variable symptom presentation and reduces it to a category. That compression is useful for care and billing, but it is lossy, and the losses are exactly the observable details a coder is trying to capture.

The depression numbers make this concrete. Taking only the nine criterion symptoms of a major depressive episode, 227 distinct symptom combinations all satisfy the same diagnosis; account for the sub-symptoms folded into items like sleep, appetite, and psychomotor change and the figure rises to 16,400 (Fried & Nesse, 2015). Two patients can carry an identical diagnosis and share not one symptom.

That is the choice in front of an annotation project. If you annotate the diagnosis, you record a single label for both patients. If you annotate the phenomena, you record what is actually different.

There is a second reason, and it comes from the people who have already tried to extract symptoms from clinical text at scale. In the CRIS-CODE project, a team of psychiatrists defined 46 symptoms of severe mental illness and built NLP models to pull them from records, reaching a median F1 of 0.88; their headline observation was that “most symptoms cut across diagnoses, rather than being restricted to particular groups” (Jackson et al., 2017). Symptoms are the portable unit. A poverty-of-speech span looks the same whether it sits under a psychotic-disorder record or a mood-disorder one, which is why span-level phenomena travel across projects in a way a diagnostic category never does.

This is also the logic behind the Research Domain Criteria framework, which was launched because categorical diagnoses built from clusters of clinical symptoms often fail to line up with the underlying biology, and which studies constructs and dimensions across, rather than within, diagnostic boxes (Insel et al., 2010). You do not have to adopt RDoC to take its lesson for annotation: the measurable thing is the dimension or the symptom, and the category is a downstream summary.

A worked mapping: from an MDD criterion concept to the spans you tag

The cleanest way to see criteria-to-phenomenon annotation is to run one disorder all the way through. A major depressive episode is defined by nine symptom domains, and the PHQ-9 was built so that each of its nine items corresponds directly to one of them (Kroenke et al., 2001). That gives us a rare thing: an instrument whose items are already the observable constructs, so the criterion-to-annotation step is a matter of writing down what each span has to show.

The table below carries all nine domains through the transformation. The middle column uses the item names and codes from Tagaroo’s curated PHQ-9, so the mapping agrees with the scale page. The right column is what a coder actually tags. Every quoted span is invented, not drawn from a real interview.

MDD symptom concept (paraphrased)PHQ-9 phenomenon (code)What the annotatable span shows (synthetic)
Low or depressed moodDepressed Mood (DEP)"I feel down pretty much every day now, and it doesn't lift."
Loss of interest or pleasureAnhedonia (ANH)"Nothing I used to enjoy does anything for me anymore."
Change in appetite or weightAppetite (APP)"I've basically stopped eating; food has no appeal."
Sleep disturbanceSleep (SLP)"I'm awake at three in the morning every night, then can't get up."
Psychomotor changePsychomotor (PSY)Observed, not spoken: long pauses, markedly slowed speech (resists a text-only span)
Fatigue or low energyFatigue (FAT)"Even getting dressed wipes me out for the whole morning."
Worthlessness or guiltWorthlessness (WTH)"I feel like a burden to everyone around me."
Trouble concentrating or decidingConcentration (CNC)"I read the same paragraph five times and nothing sticks."
Thoughts of death or self-harmSuicidal Ideation (SUI)"Sometimes I think everyone would be better off without me." (flag for clinician review at any severity)
A synthetic MDD-to-PHQ-9 mapping. The nine major-depression symptom domains (paraphrased, never quoted from the manual) become nine PHQ-9 phenomena (Kroenke et al., 2001), each attached to an observable span. Codes match Tagaroo's curated PHQ-9.

Two rows do the teaching. The Psychomotor row shows a symptom that lives in what a clinician observes, not in what the subject says, so a text-only span cannot fully carry it; you either tag the interviewer’s described observation or accept that this phenomenon needs another modality. The Suicidal Ideation row shows a phenomenon that gets flagged for clinician review regardless of its severity band, because the annotation is a signal, not a rating that can be safely averaged away.

Both are the kind of boundary a criterion never spells out and a codebook has to. The item-by-item severity anchors are walked through in the PHQ-9 scoring guide.

Where the count and threshold logic lives

The mapping above tags nine phenomena; it does not produce a diagnosis, and it must not pretend to. A major depressive episode requires a minimum number of the nine symptoms to be present together over a sustained period, with at least one being depressed mood or loss of interest, plus clinically significant distress or impairment and the exclusion of other causes (APA, 2022; the count structure is laid out plainly in Fried & Nesse, 2015). None of that lives in a single utterance.

Keep the span label and the scoring rule in separate places. The annotation layer answers “is this phenomenon present in this span, and how strongly.” A separate, explicit scoring layer answers “do the labeled spans, taken together, cross the count-and-duration threshold.” Conflating the two produces codes that quietly encode a diagnosis they have no business asserting, and it is the single most common way a criteria-based coding scheme oversteps. Scoring an instrument from the labeled evidence rather than a global impression is its own discipline, covered in auditable scale scoring from the transcript.

DSM-5 and ICD-11: two manuals, two annotation problems

DSM-5-TR and the ICD-11 CDDR describe overlapping disorders but define “meeting criteria” in different styles, and the difference matters for how you operationalize each. DSM-5-TR leans on operationalized symptom counts and explicit thresholds. The ICD-11 CDDR, released in 2024, deliberately moved toward essential (required) features and a described “boundary with normality” that the clinician applies with judgment, rather than a rigid tally (WHO, 2024). ICD-11 itself was adopted in 2019 and came into effect for reporting in 2022.

For annotation, the observable spans are largely the same symptoms either way. What changes is where the diagnostic decision sits. Under a counting manual it is tempting, and wrong, to imagine the count is “in” the transcript; under the ICD-11 essential-features style it is obvious that the decision is a judgment layered on top of the spans. The second framing is the healthier one to annotate under, whichever manual you cite.

DimensionDSM-5-TR (APA, 2022)ICD-11 CDDR (WHO, 2024)
How caseness is definedOperationalized symptom counts and explicit thresholdsEssential (required) features plus a boundary with normality, applied with clinical judgment
Manual stylePolythetic criteria lists with cut-offsPrototype-style clinical descriptions, fewer rigid counts
What you annotateThe individual symptom domains, as spansThe same essential features, as spans
Where the diagnostic label livesIn the count that crosses the thresholdIn the clinician's judgment that essential features are present
Reuse rulesAPA-copyrighted; paraphrase, never pasteWHO-copyrighted (CC BY-NC-ND); paraphrase, never paste
Two manuals, two annotation problems. The observable spans overlap heavily; the manuals differ in how the diagnosis is decided, and both forbid reproducing criterion text verbatim (APA, 2022; WHO, 2024).

One rule spans both manuals and is not optional: do not reproduce criterion text. DSM-5-TR is copyrighted by the American Psychiatric Association, and the ICD-11 CDDR carries a Creative Commons licence that forbids adaptations without permission (WHO, 2024). Paraphrase the symptom concept into your own operational definition, tag observable phenomena, and cite the manual as a source. The codebook you build is your specification of observable behavior, not a copy of the manual.

The limits of DSM criteria annotation

DSM criteria annotation makes coding consistent; it does not by itself make an instrument valid, and pretending otherwise sets a project up to fail in three predictable ways. Naming the limits is what keeps the method honest.

Some criteria resist text. Symptoms defined by clinical observation, such as psychomotor change or an interviewer-rated affect, do not live cleanly in the subject’s words. A transcript-only span cannot fully justify them, which is why observational and physiological items are the first to drop out when a scale is operationalized for span annotation. Speech-borne phenomena behave better; poverty of speech from the SANS is a symptom that a span can genuinely evidence.

Operationalizing does not fix a mismatched instrument. A beautifully coded phenomenon built on a poorly chosen scale still measures the wrong thing precisely. Whether your source is a clinician-rated or a self-report instrument shapes what the spans can mean, a distinction worth settling before you code and one covered in clinician-rated versus self-report scales. Choose the instrument for the question first, then operationalize it into a codebook.

The diagnosis is never in a span. This is the limit that carries the ethical weight. Annotation can surface every phenomenon that maps toward a set of criteria and still say nothing about whether the person meets them, because the count, the duration, the distress, and the rule-outs are a clinician’s call. The tool and the coder tag observable evidence; the diagnosis stays with the clinician. Even in the CRIS-CODE work, where a team achieved good inter-annotator agreement on symptom instances, the symptoms were the deliverable, not a machine-made diagnosis (Jackson et al., 2017).

From a manual to a working codebook

The practical upshot: a diagnostic manual gives you named symptom concepts and a rule for combining them, and DSM criteria annotation is the work of turning the first half into observable, span-taggable phenomena while leaving the combining rule in a separate scoring layer. Paraphrase each criterion into a code with a definition, inclusion and exclusion rules, and synthetic examples; tag the spans; measure inter-rater agreement; and never let a span assert a diagnosis. The full mechanics of writing that codebook, item by item, are in the guide to turning a rating scale into a codebook.

Do that, and you get an annotation set that preserves what the diagnostic label discards: the specific, variable, observable symptoms that two people with the same diagnosis rarely share. Pick the instrument for the question, operationalize its symptoms into phenomena, and keep the diagnosis where it belongs.

References

  • American Psychiatric Association. (2022). Diagnostic and Statistical Manual of Mental Disorders (5th ed., text revision). American Psychiatric Association Publishing. DSM-5-TR
  • World Health Organization. (2024). Clinical Descriptions and Diagnostic Requirements for ICD-11 Mental, Behavioural and Neurodevelopmental Disorders (CDDR). Geneva: World Health Organization. WHO publication
  • Insel, T., Cuthbert, B., Garvey, M., Heinssen, R., Pine, D. S., Quinn, K., Sanislow, C., & Wang, P. (2010). Research Domain Criteria (RDoC): Toward a new classification framework for research on mental disorders. American Journal of Psychiatry, 167(7), 748–751. doi:10.1176/appi.ajp.2010.09091379
  • Fried, E. I., & Nesse, R. M. (2015). Depression is not a consistent syndrome: An investigation of unique symptom patterns in the STAR*D study. Journal of Affective Disorders, 172, 96–102. doi:10.1016/j.jad.2014.10.010
  • Jackson, R. G., Patel, R., Jayatilleke, N., Kolliakou, A., Ball, M., Gorrell, G., Roberts, A., Dobson, R. J., & Stewart, R. (2017). Natural language processing to extract symptoms of severe mental illness from clinical text: the CRIS-CODE project. BMJ Open, 7(1), e012012. doi:10.1136/bmjopen-2016-012012
  • Kroenke, K., Spitzer, R. L., & Williams, J. B. W. (2001). The PHQ-9: Validity of a brief depression severity measure. Journal of General Internal Medicine, 16(9), 606–613. doi:10.1046/j.1525-1497.2001.016009606.x

If you code symptoms from interviews, Tagaroo turns instruments like the PHQ-9 and the MADRS into guided, span-level codebooks where each item ships as an observable phenomenon with its definition, examples stay editable, and inter-rater agreement is computed as your coders work. Operationalize the criteria, tag the phenomena, and leave the diagnosis to the clinician.

Frequently asked questions

Should you annotate a diagnosis or the symptoms behind it?
Annotate the observable phenomena, not the diagnosis. A diagnosis is a categorical decision a clinician reaches after weighing counts, thresholds, duration, and rule-outs; a phenomenon like low mood or poverty of speech is a span of text a coder can point to. In one large sample, 3,703 people who all met criteria for major depressive disorder produced 1,030 distinct symptom profiles, and roughly one in seven had a profile shared with no one else (Fried & Nesse, 2015), so a single diagnostic label hides most of what an annotation set should capture. Annotate the spans that map toward criteria and leave the diagnosis to a qualified clinician.
How do you turn a DSM-5 criterion into an annotatable phenomenon?
Rewrite the criterion's symptom concept as a rule about text: give it a short code, an operational definition of what a span must express, an inclusion rule, an exclusion rule that names the neighboring code, and synthetic examples. This is the same move as turning any standard into a coding specification, done deductively from the manual rather than induced from the data. The PHQ-9 is a worked precedent: each of its nine items corresponds to one of the nine major-depression symptom domains (Kroenke et al., 2001), which is exactly the criterion-to-phenomenon mapping annotation needs.
Can you reproduce DSM-5 or ICD-11 criteria in your codebook?
No. DSM-5-TR criteria are copyrighted by the American Psychiatric Association (APA, 2022) and the ICD-11 CDDR is copyrighted by the World Health Organization (WHO, 2024, released under a CC BY-NC-ND licence that forbids adaptations without permission). Do not paste criterion text into a codebook or annotation tool. Paraphrase the symptom concept in your own words, operationalize it as an observable phenomenon, and cite the manual.
How does ICD-11 differ from DSM-5 for annotation?
DSM-5-TR uses operationalized symptom counts and explicit thresholds, while the ICD-11 CDDR describes essential (required) features and a boundary with normality that the clinician applies with judgment rather than by counting (APA, 2022; WHO, 2024). For annotation the observable spans are largely the same symptoms, but the ICD-11 style makes it even clearer that the diagnostic decision is a judgment layer sitting on top of the spans, not something a span can carry on its own.
Do PHQ-9 items map to DSM-5 criteria?
Yes. The PHQ-9 was built so that each of its nine items corresponds directly to one of the nine symptom domains used to define a major depressive episode (Kroenke et al., 2001). That makes it the cleanest available bridge from criteria to annotatable phenomena: the nine items are already the nine observable constructs you would tag, which is why the worked mapping in this article uses it.

Put this into practice

Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.