tagaroo

rating scales

Auditable Clinical Ratings: Score From the Transcript

Build an auditable clinical rating by tying every scale item to a quoted line, not a global impression. See how span-grounded scoring works.

Enrique Gutiérrez14 min readUpdated July 2026
Abstract editorial illustration of several highlighted transcript spans on the left, each linked by a thin ink line to a segment of a vertical stacked bar on the right, showing a rating-scale total built from tagged evidence rather than a fading gestalt impression.

A rating-scale total is only as trustworthy as the process that produced it. Two raters can hand you the same MADRS score of 24 and mean completely different things by it: one added up ten items after quoting the line in the interview that justified each rating, the other formed a global impression and reverse-engineered a number that felt about right. Only the first is an auditable clinical rating—a score whose every point can be traced back to the words that produced it. This piece is about how to turn the second kind of scoring into the first, and why regulators, trial managers, and reliability-obsessed methodologists should care.

What is an auditable clinical rating?

An auditable clinical rating is a score in which each item can be traced to the specific evidence that produced it. Put the standard in one line: a defensible MADRS score is one where every item points to a quoted line in the interview. The total is not a judgment handed down from the rater’s overall sense of the patient; it is the sum of item ratings, each anchored to a passage a second reader can inspect.

That distinction sounds pedantic until you try to reconstruct someone else’s score six months later. With a global impression, the number is all you have—the reasoning evaporated the moment the rater wrote it down. With span-grounded scoring, the number arrives with its own audit trail: item 1 was a 3 because of this sentence, item 7 was a 2 because of that one. The score becomes checkable, and a checkable score is a defensible one.

This is the difference between a measurement and an opinion. A well-run MADRS or Hamilton depression scale already defines anchors for each item; auditable scoring simply insists that you record which words in the transcript met each anchor before you commit to the rating.

Why do global clinical impressions drift?

Global impressions drift because they compress a whole interview into a single number without leaving a record of how. The Clinical Global Impressions scale is the purest example: it asks the clinician one question—“considering your total clinical experience, how mentally ill is the patient?”—and is completed in under a minute, capturing impressions that its own authors describe as transcending “mere symptom checklists” (Busner & Targum, 2007). That is efficient, and for tracking a single patient over time it is useful. As a research measurement, its provenance is thin: nothing in the score tells you what drove it.

The evidence that this matters is not subtle. In a meta-analysis of 136 studies across health and behavior, mechanical (formal, statistical) combination of data was about 10% more accurate than holistic clinical prediction, and the superiority held regardless of the judges’ experience (Grove et al., 2000). The finding that should stop a trial designer cold: clinical prediction did relatively worse exactly when the input was clinical interview data—the setting rating scales live in.

Watching how raters actually behave fills in the mechanism. When reviewers evaluated 104 audiotaped HAM-D interviews, 39% were conducted in ten minutes or less and most were rated fair or unsatisfactory for adequacy of the information obtained (Engelhardt et al., 2006). A cursory interview leaves the rater no choice but to fill the gaps with impression.

And when Kobak and colleagues transcribed 30 pairs of interviews and had independent assessors code why two raters disagreed, the disagreements sorted into five sources—information, observation, interpretation, criterion, and subject variance—with interpretation variance the most common (Kobak et al., 2009). Most rater disagreement is not about what the patient said; it is about what the rater made of it.

How do you build an auditable clinical rating from a transcript?

You build an auditable clinical rating by inverting the usual order: find the evidence first, then rate it. The workflow is four steps, and it is the same whether a human or an assisted tool does the coding.

  1. Read for evidence, not for a verdict. For each item on the scale, locate the span in the subject’s speech (or the interviewer’s recorded observation) that bears on it. Highlight it. If nothing in the transcript speaks to an item, that item has no evidence—which is itself a finding.
  2. Anchor the span to the item’s rating scale. Match the quoted passage to the instrument’s defined anchors. The MADRS Reported Sadness item, for instance, runs 0–6 from “occasional sadness in keeping with the circumstances” up to “continuous or unvarying sadness”; the span decides which anchor applies.
  3. Rate the item, attaching the rationale. Record the item score together with the span and a one-line justification. This is the “rationale” the trial-methods literature keeps asking for—the artifact that lets a reviewer reconstruct the decision (Kobak et al., 2009).
  4. Sum the items to a total. The total is now a derived quantity, not a first impression. For the MADRS it is the sum of ten items; for a self-report screen like the PHQ-9 it is the sum of nine.

The payoff is that the total and its parts are separable. A reviewer can accept nine of your ten MADRS items and challenge only the one whose span does not support its rating—a conversation you simply cannot have about a gestalt.

PropertyGlobal impression (gestalt)Span-grounded item scoring
What is recordedA single numberA number + the span + a rationale per item
Can a reviewer reconstruct it?No—reasoning is goneYes—each item points to a quoted line
Diagnosing rater disagreementHard; you argue about the totalItem-level; you compare the spans
Regulatory provenanceWeak (not attributable)Strong (attributable, reconstructable)
Typical failure modeInterpretation drift, halo effectsTime cost; needs a coding tool
ML/training valueLabel onlyLabel + rationale (span)
Global impression vs span-grounded scoring on the axes that decide auditability. Both yield a total; only one carries its own evidence.

A worked example: scoring depressed mood from the evidence

Consider a short synthetic exchange from a depression interview. The task is to score the relevant MADRS items from the evidence, not from an overall read of the patient.

Interviewer: How would you describe your mood over the past week?

Subject: Flat, mostly. It doesn’t lift even when something good happens—my daughter visited and I felt nothing. I’m awake by 4 most mornings and just lie there. I keep thinking things won’t get better, but I’d never do anything about it.

Coded span by span, four MADRS items have evidence, and each rating carries its quoted justification.

MADRS item (0–6)Evidence span (synthetic)Rating
Reported Sadness"Flat, mostly. It doesn't lift even when something good happens"4
Inability to Feel"my daughter visited and I felt nothing"4
Reduced Sleep"I'm awake by 4 most mornings and just lie there"3
Pessimistic Thoughts"I keep thinking things won't get better"3
Suicidal Thoughts"I'd never do anything about it"0
Synthetic MADRS coding: each item anchored to a quoted span. These five items subtotal 14; the full instrument sums all ten items (0–60). The suicide item is scored 0 here but is always flagged for clinician review regardless of score.

The subtotal for these five items is 14; a real rating would code the remaining five items (apparent sadness, inner tension, reduced appetite, concentration difficulties, lassitude) the same way and sum all ten. The point is not the number—it is that the number is now inspectable. Rate the items in the tool below and watch the total, severity band, and remission flag update as the evidence changes.

MADRS score simulator

Rate each item on the 0-6 MADRS anchors: 0 = symptom absent, 6 = most severe. Defined descriptions sit at 0, 2, 4, and 6; the odd values (1, 3, 5) mark intermediate states.

  1. 1.Apparent sadnessObserved dejection in voice, expression, and posture
  2. 2.Reported sadnessThe subject's own account of low mood
  3. 3.Inner tensionEdginess, ill-defined discomfort, dread, or panic
  4. 4.Reduced sleepReduced duration or depth vs. the subject's normal
  5. 5.Reduced appetiteLoss of appetite or of the pleasure of eating
  6. 6.Concentration difficultiesTrouble collecting or sustaining thoughts
  7. 7.LassitudeDifficulty starting and carrying out activities
  8. 8.Inability to feelReduced interest; loss of the capacity to feel
  9. 9.Pessimistic thoughtsGuilt, self-reproach, inferiority, remorse
  10. 10.Suicidal thoughtsLife not worth living; wishes for death

Total score

0 / 60

Severity band

Normal / symptom-free

Remission (<=10)

Yes

Item 10 flag

None

Severity bands follow Snaith et al. (1986); the remission target of a total of 10 or below follows Hawley et al. (2002). This tool illustrates how the MADRS is scored; it is an educational aid, not a diagnostic instrument, and a score is a measure of severity, not a diagnosis. Adjust the items above (or load the example) to see the score update.

Tagaroo’s curated MADRS exposes all ten of the published items (Montgomery & Åsberg, 1979), so the card above and the instrument agree one-to-one. The Hamilton scale is different: the published HDRS-17 has 17 scored items, while Tagaroo’s HAM-D card curates the 11 content-localizable items—the ones a subject can actually report in an interview—and omits several purely observational and objective-measurement items (anxiety-somatic, gastrointestinal and genital symptoms, weight loss) while collapsing the three insomnia items into one. If you report a HAM-D total, say the HDRS-17 total; the curated set is for span-level evidence coding, not for regenerating the 17-item sum.

For the full item walk-throughs, the MADRS scoring guide and the Hamilton Depression Rating Scale guide cover every anchor.

How does span-grounded scoring cut cross-site rater drift?

Span-grounded scoring cuts drift because it makes the largest source of disagreement—interpretation—visible and checkable. Measurement error is not a rounding footnote in trials; Kobak and colleagues argued it is a leading reason antidepressant trials fail outright, by manufacturing placebo response and burying the drug signal in noise (Kobak et al., 2007). If you cannot see why two sites score the same presentation differently, you cannot fix it.

The cost of that invisibility is quantifiable. In a study pitting site raters against blinded centralized raters on the same patients, 35% of patients admitted by site raters would have been ineligible under central raters, and mean placebo-arm change was 7.52 for site raters versus 3.18 for central raters (Kobak et al., 2010). A gap that size is enough to sink an otherwise real treatment effect.

Two findings show where the fix lives. First, competence is not seniority: across 1,241 raters scoring videotaped HAM-A, HAM-D, and YMRS interviews, clinical experience alone did not confer rating competency, whereas repeated training did (Targum, 2006). Second, calibration beats experience—rater agreement on the same interviews was highest for experienced and calibrated raters (ICC 0.93), and experienced-but-uncalibrated raters (0.55) actually trailed inexperienced ones (0.77) (Kobak et al., 2009).

Span-grounded scoring is calibration made durable: when every item cites its evidence, a central reviewer can adjudicate the interpretation directly instead of re-interviewing the patient. It is also where inter-rater work pays off, because you can compute Cohen’s kappa and related agreement coefficients item by item and see exactly which item, and which span, the coders read differently.

What do regulators expect from a defensible score?

Regulators already think in terms of provenance, even if they do not use the word “span.” The FDA’s data-integrity guidance defines the ALCOA standard: data supporting a regulated decision should be attributable, legible, contemporaneously recorded, original or a true copy, and accurate, with the metadata needed to reconstruct the activity (FDA Data Integrity guidance, 2018). A rating-scale total with no trail back to the interview is, in this vocabulary, neither attributable nor reconstructable.

The clinical-outcome-assessment framework points the same way. The FDA classifies the HAM-D and MADRS as clinician-reported outcomes (ClinROs)—the COA type chosen “if clinical judgment is required to interpret an observation”—and expects a clearly defined scoring rule plus evidence that scores are “not overly influenced by measurement error” (FDA COA guidance, 2025). The guidance carefully separates a COA, a COA score, and an endpoint, which means the score must be produced by a standardized, documented process, not improvised per rater.

Span-grounded item scoring is one concrete way to meet that expectation. It gives each ClinRO item a documented, attributable derivation, and it turns “the rater judged severity” into “the rater rated item 3 as a 4 on the basis of this quoted line.” That is the provenance an auditor can follow.

Where item-level evidence scoring falls short

Auditable scoring is not free, and pretending otherwise would undercut the point. Coding an evidence span for every item is slower than forming an impression, and without a purpose-built tool the rationales scatter across notebooks and margins where no reviewer will ever find them. The method also assumes the interview captured the evidence in the first place; a ten-minute interview that never probed sleep leaves the sleep item genuinely ungroundable, and no coding discipline can conjure evidence that was never elicited (Engelhardt et al., 2006).

Some constructs resist span-grounding by nature. Observational items—psychomotor retardation, apparent sadness read from affect rather than words—live in what the rater sees, not in the transcript, so a text-only span cannot fully justify them; this is exactly why the curated HAM-D set is smaller than the HDRS-17.

And auditability improves provenance, not validity: a perfectly documented total from a poorly chosen scale is still measuring the wrong thing. For that decision, matching the instrument to the question, compare a clinician-rated scale against a self-report screen before you commit. Auditable scoring makes a score defensible; it does not make a bad scale good.

From a number to a defensible clinical rating

The practical upshot: a rating-scale total is a claim, and a claim without evidence is worth less than one that carries its own. An auditable clinical rating attaches the evidence to the claim—every item pointing back to the line that produced it—so a reviewer, a central rater, or an auditor can check the work instead of trusting it. The literature says the payoff is real: structured, item-summed scoring beats gestalt on average (Grove et al., 2000), the drift it prevents is expensive (Kobak et al., 2010), and the provenance it produces is exactly what regulators ask for (FDA, 2018; 2025).

References

If you code depression or mania symptoms from interviews, Tagaroo turns the MADRS, the Hamilton scale, and the PHQ-9 into guided, evidence-anchored workflows where each item rating is bound to its transcript span and inter-rater reliability is computed as your coders work. Score from the transcript, not from a gestalt, and the total defends itself.

Frequently asked questions

What makes a clinical rating auditable?
A clinical rating is auditable when every item score traces back to a specific quoted span in the source material, so an independent reviewer can reconstruct the total from the evidence rather than trusting the rater's memory. That is the same standard the FDA expresses for regulated data through the ALCOA principles: data should be attributable, legible, contemporaneously recorded, original or a true copy, and accurate, with enough metadata retained to reconstruct the activity (FDA Data Integrity guidance, 2018). Span-grounded scoring makes each item attributable to its source line.
Is item-summed scoring more reliable than a global clinical impression?
On average, yes. A meta-analysis of 136 studies found mechanical (formal, statistical) combination of data was about 10% more accurate than holistic clinical judgment, and the gap widened when the input included clinical interview data (Grove et al., 2000). But reliability still depends on administration: rater agreement on the same interviews ranged from an intraclass correlation of 0.93 for experienced and calibrated raters down to 0.55 for experienced raters who were not calibrated (Kobak et al., 2009).
How does span-grounded scoring reduce cross-site drift in clinical trials?
It attacks the largest source of drift, which is how raters interpret and record the same interview. When site raters and blinded centralized raters scored the same patients, 35% of patients admitted by site raters would have been ineligible under central raters, and mean placebo-arm change was 7.52 for site raters versus 3.18 for central raters (Kobak et al., 2010). Requiring each item to point to a quoted line lets a central reviewer check the evidence, not re-run the interview.
What does the FDA expect from a defensible scale score?
For a clinician-reported outcome (ClinRO) like the HAM-D or MADRS, the FDA's clinical outcome assessment framework expects a clearly defined scoring rule and evidence that scores are not overly influenced by measurement error (FDA COA guidance, 2025). The FDA also distinguishes a COA, a COA score, and an endpoint, so the numeric score must be produced by a standardized, documented process. Span-grounded item scoring supplies that documented process at the item level.
Can you audit a MADRS total item by item?
Yes. The MADRS has 10 clinician-rated items, each scored 0 to 6 for a total of 0 to 60 (Montgomery & Åsberg, 1979). Because each item measures a distinct symptom, you can anchor every item to the passage in the subject's speech that justifies its rating, then sum the item scores to reconstruct the total. A total assembled this way carries its own provenance.

Put this into practice

Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.