rating scales
Hamilton Depression Rating Scale (HAM-D): A Coding Guide
How the Hamilton Depression Rating Scale (HAM-D) is scored and coded: the 17 items, the 0–52 range, remission and response cutoffs, and its known flaws.

The Hamilton Depression Rating Scale (HAM-D) is a clinician-rated measure of depression severity, and for more than sixty years it has been the default yardstick for antidepressant trials. The arithmetic is unglamorous: rate 17 items, add them up, land somewhere between 0 and 52. What makes the scale interesting is the gap between how heavily the field leans on that number and how much criticism its design has drawn—one major review called it “psychometrically and conceptually flawed” and still conceded it was the gold standard (Bagby et al., 2004). This guide covers the items, the scoring, the cutoffs, and why the flaws haven’t dislodged it.
What is the Hamilton Depression Rating Scale (HAM-D)?
The Hamilton Depression Rating Scale is a clinician-administered questionnaire that quantifies the severity of depressive symptoms in someone who already carries a depression diagnosis (Hamilton, 1960). Max Hamilton designed it to measure severity, not to detect or classify depression—a distinction that still trips people up. A clinician interviews the patient, then rates each item; the patient never fills it in themselves.
That clinician-rated design is the line between the HAM-D and self-report screens like the PHQ-9. It buys richer clinical observation at the cost of rater subjectivity—two clinicians can watch the same interview and score it differently, which is the reliability problem the rest of this guide keeps returning to. The scale is in the public domain and freely reproducible (Hamilton, 1960), which is part of why it spread so widely.
Several versions exist—HAM-D-6, HAM-D-21, HAM-D-24, HAM-D-29—but “the HAM-D” almost always means the 17-item HDRS-17. The 21- and 24-item forms add items meant to subtype depression (diurnal variation, paranoid and obsessional features) rather than to measure severity, and trials generally score only the first 17.
How is the Hamilton Depression Rating Scale (HAM-D) scored?
HAM-D scoring sums 17 individually rated items into a total from 0 to 52 (Hamilton, 1960). There is no reverse-scoring, but—unlike the PHQ-9’s uniform 0–3 anchors—the items do not all share a range. Nine items are rated 0–4 and eight on a compressed 0–2 scale, so a maximally severe rating on one item can contribute twice as much to the total as a maximally severe rating on another.
That mixed structure is deliberate. Hamilton reasoned that some symptoms could be graded finely while others could only be judged present, mild, or marked—but the side effect is uneven weighting. As Bagby and colleagues put it, “someone who weeps all the time can contribute 3 or 4 points on depressed mood, whereas someone who feels tired all the time can contribute only 2 points on the general somatic symptoms item” (Bagby et al., 2004). The total treats those contributions as interchangeable.
| # | HAM-D-17 item | Symptom domain | Range |
|---|---|---|---|
| 1 | Depressed mood | Sadness, hopelessness, tearfulness | 0–4 |
| 2 | Feelings of guilt | Self-reproach, ideas of punishment | 0–4 |
| 3 | Suicide | Death wishes, ideation, acts | 0–4 |
| 4 | Insomnia, early | Difficulty falling asleep | 0–2 |
| 5 | Insomnia, middle | Broken sleep | 0–2 |
| 6 | Insomnia, late | Early-morning waking | 0–2 |
| 7 | Work and activities | Loss of interest, reduced capacity | 0–4 |
| 8 | Retardation | Slowed thought, speech, movement | 0–4 |
| 9 | Agitation | Restlessness, hand-wringing | 0–4 |
| 10 | Anxiety, psychic | Subjective tension, worry | 0–4 |
| 11 | Anxiety, somatic | Physical concomitants of anxiety | 0–4 |
| 12 | Somatic symptoms, GI | Appetite loss, gut complaints | 0–2 |
| 13 | Somatic symptoms, general | Heaviness, fatigability | 0–2 |
| 14 | Genital symptoms | Libido, menstrual disturbance | 0–2 |
| 15 | Hypochondriasis | Preoccupation with health | 0–4 |
| 16 | Loss of weight | Weight change | 0–2 |
| 17 | Insight | Awareness of illness | 0–2 |
Notice that three of the seventeen items—early, middle, and late insomnia—are devoted to sleep. Sleep alone can therefore add up to 6 of the 52 points, and factor analyses consistently pull those three items onto their own dimension (Bagby et al., 2004). That’s the “insomnia over-weighting” critics point to: a patient whose sleep improves can post a meaningful HAM-D drop with little change in mood.
Work a concrete (synthetic) example. Rating depressed mood 3, guilt 2, suicide 1, the three insomnia items 2/2/1, and a scatter of 1s and 2s across the remaining items lands a total of 20—the moderate band (Zimmerman et al., 2013). Notice that the three sleep items alone contribute 5 of those 20 points: because they are 0–2 items but there are three of them, sleep can move the total as much as a severe rating on a single 0–4 mood item. The interpreter below runs that arithmetic live, so you can watch each item—and its range—move the total and the severity band.
HAM-D (HDRS-17) score interpreter
Rate each item on its own native range. Nine items run 0–4 (0 = absent … 4 = very severe) and eight run 0–2 (0 = absent, 1 = mild, 2 = marked) — so the same maximal rating counts for twice as much on some items as on others.
- 1.Depressed mood(0–4)
- 2.Feelings of guilt(0–4)
- 3.Suicide(0–4)
- 4.Insomnia, early (difficulty falling asleep)(0–2)
- 5.Insomnia, middle (broken sleep)(0–2)
- 6.Insomnia, late (early-morning waking)(0–2)
- 7.Work and activities(0–4)
- 8.Retardation (slowed thought, speech, movement)(0–4)
- 9.Agitation(0–4)
- 10.Anxiety, psychic(0–4)
- 11.Anxiety, somatic(0–4)
- 12.Somatic symptoms, gastrointestinal(0–2)
- 13.Somatic symptoms, general(0–2)
- 14.Genital symptoms(0–2)
- 15.Hypochondriasis(0–4)
- 16.Loss of weight(0–2)
- 17.Insight(0–2)
Total score
0 / 52
Severity band
No depression / remission
Sleep items (4–6)
0 pts
Item 3 (suicide)
None
Item structure and the 0–52 range follow Hamilton (1960); severity bands follow Zimmerman et al. (2013). This tool illustrates how the HAM-D is scored; it is an educational aid, not a diagnostic instrument, and the total is a rough, multidimensional index rather than a precise reading of one construct. Adjust the items above (or load the example) to see the score update.
A note on Tagaroo’s version. The curated HAM-D scale in our Scale Library codes the content-localizable items—the ones you can actually ground in what a subject says during an interview—so it collapses the three insomnia items into one and sets aside purely physiological items like weight loss and genital symptoms that a transcript rarely supports. It’s a subset built for evidence-anchored coding, not a re-scoring of the 52-point total.
What do HAM-D scores mean? Severity bands and interpretation
Hamilton never published official cutoffs, so HAM-D interpretation depends on which band scheme you adopt—and several compete. The most widely cited empirically derived scheme comes from a study of 627 outpatients with major depressive disorder (Zimmerman et al., 2013).
| HAM-D-17 total | Severity band | Typical trial use |
|---|---|---|
| 0–7 | No depression / remission | Remission target |
| 8–16 | Mild | Often below trial-entry threshold |
| 17–23 | Moderate | Common minimum for enrollment |
| ≥24 | Severe | High symptom burden |
The bands describe symptom burden; they are not themselves a clinical verdict. And because Hamilton left the cutoffs open, published band systems disagree—some set moderate depression at 18–24, others at 17–23. That ambiguity has real consequences: a trial that enrolls “moderate to severe” patients is defining its own population by the threshold it picks, which is one reason effect sizes are hard to compare across studies (Zimmerman et al., 2013).
What are the HAM-D remission and response thresholds?
In antidepressant trials, two HAM-D thresholds do most of the work: remission is a total of ≤7, and response is a ≥50% reduction from the baseline score (Frank et al., 1991). These are the numbers that decide whether a drug “worked” for a given patient, and they anchor the primary endpoints of most registration trials.
Both definitions come from a 1989 MacArthur Foundation consensus panel, not from a decisive validation study—the panel itself explicitly asked for empirical validation (Frank et al., 1991). Later work has largely confirmed the ≤7 remission cutoff against clinician global impressions in an equipercentile-linking analysis of 43 trials and 7,131 patients (Leucht et al., 2013), though some researchers argue the true asymptomatic threshold sits lower, at ≤5 or ≤6. The practical point: remission and response on the HAM-D are conventions with reasonable evidence behind them, not physical constants—cite the threshold you applied.
The SIGH-D: standardizing how the HAM-D is administered
The SIGH-D (Structured Interview Guide for the Hamilton Depression Rating Scale) is a scripted set of questions that fixes how raters elicit the information the HAM-D items depend on (Williams, 1988). The original scale shipped without a standard interview, so two raters could ask different questions, get different answers, and score the same patient differently.
That variability was measurable. Before the SIGH-D, “most of the items had only fair or poor agreement” on test-retest, and Williams built the guide precisely to close that gap (Williams, 1988). Her reliability study found that “the use of the SIGH-D results in a substantially improved level of agreement for most of the HDRS items” (Williams, 1988).
How high can that agreement climb once administration is fixed? Pooled across nearly five decades of HAM-D studies, inter-rater agreement on the total score reaches an intraclass correlation of about 0.94 (Trajković et al., 2011)—comparable to the 0.93 that the MADRS’s own structured guide achieves (Williams & Kobak, 2008)—but that ceiling depends on standardizing the interview, which is exactly what the SIGH-D was built to do.
Structured administration—via the SIGH-D or its descendant, the GRID-HAMD—is now expected in serious multi-site work, because the alternative is rater drift masquerading as clinical signal.
Why does the HAM-D still anchor antidepressant trials despite its flaws?
The HAM-D endures as the default trial endpoint not because it is the best-designed depression scale—it demonstrably is not—but because six decades of regulatory submissions, meta-analyses, and power calculations are denominated in HAM-D points. Switching scales would strand that comparability. The “gold standard” is really a common currency, and currencies are sticky.
The case against the design is well documented. Reviewing 70 studies published since 1979, Bagby and colleagues concluded that “the Hamilton depression scale is psychometrically and conceptually flawed” and that “it is time to embrace a new gold standard” (Bagby et al., 2004). Their specific charges are worth knowing:
- It isn’t unidimensional. Factor analyses recover anywhere from two to eight factors, with sleep and anxiety/agitation items reliably splitting off from a core-depression factor. A single summed score therefore blends distinct symptom dimensions (Bagby et al., 2004).
- Item weighting is uneven. The mixed 0–4 / 0–2 format lets some symptoms contribute twice what others can, without a clinical rationale for the difference (Bagby et al., 2004).
- Content is dated. The scale operationalizes a 1960 conception of depression only loosely aligned with modern DSM criteria, over-representing somatic and sleep symptoms (Bagby et al., 2004).
So why keep it? Because a benchmark’s value partly is its ubiquity. Regulators recognize it, historical trial data are expressed in it, and abandoning it would break comparisons across a literature spanning generations of antidepressants (Leucht et al., 2013). The pragmatic response has been to standardize administration (the SIGH-D), extract cleaner unidimensional subscales like the six-item Bech and Maier versions, and pair the HAM-D with a scale built to be sensitive to change—the MADRS.
HAM-D vs MADRS: which depression scale for a trial?
The Montgomery-Åsberg Depression Rating Scale (MADRS) is the HAM-D’s most common trial companion—a 10-item, clinician-rated scale that was purpose-built to be sensitive to change and to under-weight the somatic symptoms the HAM-D leans on (Montgomery & Åsberg, 1979). Where the HAM-D is a historical anchor, the MADRS is the tighter instrument for detecting treatment effects.
| HAM-D-17 | MADRS | |
|---|---|---|
| Items | 17 | 10 |
| Score range | 0–52 | 0–60 |
| Rater | Clinician | Clinician |
| Emphasis | Somatic + sleep heavy | Core mood symptoms; under-weights somatic |
| Remission cutoff | ≤7 | ≤8–9, equivalent to HAM-D ≤7 (Carmody et al., 2006) |
| Best for | Regulatory comparability, legacy data | Sensitivity to change |
Many trials run both, using the HAM-D as the recognizable primary endpoint and the MADRS for its cleaner change signal. If you’re choosing between them, the MADRS scale is often the better instrument when your question is “did this treatment move the needle,” while the HAM-D wins when you need continuity with prior results. We cover the MADRS in depth in our companion guide to the MADRS.
One number worth reconciling if you read both guides: the ≤8–9 MADRS figure in the table above is Carmody et al.’s (2006) equipercentile crosswalk—the MADRS score that lines up with HAM-D ≤7—not the MADRS’s own standalone remission cutoff. The MADRS guide headlines the more widely used ≤10 operational convention (Hawley et al., 2002). Different derivations of “remission,” not a contradiction—so cite which one you mean.
Coding the HAM-D from an interview transcript
To score a HAM-D from a recorded interview, code the passage where the subject reports each symptom, then rate that evidence on the item’s own 0–4 or 0–2 anchors—rather than ticking a form from memory. This is the workflow Tagaroo is built for, and it produces an auditable score: every item points to the exact words that justify it.
Consider a short synthetic exchange:
Interviewer: How have you been feeling in yourself this past week?
Subject: Honestly, I feel like I’ve let everyone down. Like the whole thing is my fault and I should be punished for it.
That reply supports coding item 2 (guilt) at 3—feels the present illness is a punishment, with the span “I should be punished for it” attached as the evidence. Do that across the codeable items and the total stops being a black box: a reviewer can check each rating against the transcript instead of trusting the sum.
This matters most where the HAM-D is used most—multi-site trials. When a dozen raters across several countries score the same construct, small differences in how they interpret an anchor accumulate into rater drift, and drift inflates the noise a trial has to overcome to detect a real drug effect. Evidence-anchored coding makes disagreement visible and measurable: you can compute inter-rater reliability with Cohen’s kappa across coders, spot the items where they diverge, and retrain before the drift contaminates your endpoint. That standardization problem—keeping many raters aligned on one scale—is squarely where a guided annotation workflow earns its place.
On data handling: transcript coding means working with sensitive clinical language, so the sane default is de-identified text and a privacy-first setup. Tagaroo supports a browser-side anonymous mode so transcript content can stay local rather than being uploaded—worth checking against your ethics approval before any real interview data touches a tool.
What are the most common HAM-D scoring mistakes?
The most common HAM-D mistake is treating the total as a clean, one-dimensional severity score when it isn’t. A few others recur:
- Comparing totals across band schemes. A “moderate” HAM-D in one paper may be “severe” in another because the cutoffs differ; always state which scheme you used (Zimmerman et al., 2013).
- Ignoring the subscale signal. A falling total driven entirely by the three sleep items is not the same clinical improvement as a falling total driven by mood and guilt—read the items, not just the sum (Bagby et al., 2004).
- Skipping structured administration. Scoring from an unstructured interview reintroduces exactly the rater variability the SIGH-D was built to remove (Williams, 1988).
- Confusing response with remission. A patient can meet the ≥50% response criterion and still sit well above the ≤7 remission threshold; they are different endpoints (Frank et al., 1991).
None of these are reasons to abandon the HAM-D. They are reasons to report the score with its context: the version, the band scheme, whether administration was structured, and whether the number came from a form or from evidence you can point to.
The practical upshot: the Hamilton Depression Rating Scale is easy to add up and hard to interpret cleanly. Treat the total as a rough, multidimensional index anchored to decades of trial data—not a precise reading of one construct—administer it with the SIGH-D, and, if you’re coding from transcripts, keep every rating tied to the words that earned it.
References
- Hamilton, M. (1960). A rating scale for depression. Journal of Neurology, Neurosurgery & Psychiatry, 23(1), 56–62. doi:10.1136/jnnp.23.1.56
- Williams, J. B. W. (1988). A structured interview guide for the Hamilton Depression Rating Scale. Archives of General Psychiatry, 45(8), 742–747. doi:10.1001/archpsyc.1988.01800320058007
- Williams, J. B. W., & Kobak, K. A. (2008). Development and reliability of a structured interview guide for the Montgomery-Åsberg Depression Rating Scale (SIGMA). British Journal of Psychiatry, 192(1), 52–58. doi:10.1192/bjp.bp.106.032532
- Trajković, G., Starčević, V., Latas, M., Leštarević, M., Ille, T., Bukumirić, Z., & Marinković, J. (2011). Reliability of the Hamilton Rating Scale for Depression: a meta-analysis over a period of 49 years. Psychiatry Research, 189(1), 1–9. doi:10.1016/j.psychres.2010.12.007
- Frank, E., Prien, R. F., Jarrett, R. B., Keller, M. B., Kupfer, D. J., Lavori, P. W., Rush, A. J., & Weissman, M. M. (1991). Conceptualization and rationale for consensus definitions of terms in major depressive disorder: remission, recovery, relapse, and recurrence. Archives of General Psychiatry, 48(9), 851–855. doi:10.1001/archpsyc.1991.01810330075011
- Bagby, R. M., Ryder, A. G., Schuller, D. R., & Marshall, M. B. (2004). The Hamilton Depression Rating Scale: has the gold standard become a lead weight? American Journal of Psychiatry, 161(12), 2163–2177. doi:10.1176/appi.ajp.161.12.2163
- Zimmerman, M., Martinez, J. H., Young, D., Chelminski, I., & Dalrymple, K. (2013). Severity classification on the Hamilton Depression Rating Scale. Journal of Affective Disorders, 150(2), 384–388. doi:10.1016/j.jad.2013.04.028
- Leucht, S., Fennema, H., Engel, R., Kaspers-Janssen, M., Lepping, P., & Szegedi, A. (2013). What does the HAMD mean? Journal of Affective Disorders, 148(2–3), 243–248. doi:10.1016/j.jad.2012.12.001
- Montgomery, S. A., & Åsberg, M. (1979). A new depression scale designed to be sensitive to change. British Journal of Psychiatry, 134, 382–389. doi:10.1192/bjp.134.4.382
- Carmody, T. J., Rush, A. J., Bernstein, I., Warden, D., Brannan, S., Burnham, D., Woo, A., & Trivedi, M. H. (2006). The Montgomery-Äsberg and the Hamilton ratings of depression: a comparison of measures. European Neuropsychopharmacology, 16(8), 601–611. doi:10.1016/j.euroneuro.2006.04.008
If you code depression symptoms from interviews or need to hold a dozen raters to one standard across sites, Tagaroo turns the HAM-D into a guided, evidence-anchored annotation workflow—with inter-rater reliability computed as your coders work.
Frequently asked questions
- How many items are on the HAM-D, and what is the maximum score?
- The standard version is the 17-item HAM-D (HDRS-17). Its items use two response formats—nine are rated 0–4 and eight on a compressed 0–2 scale—so the total runs from 0 to 52 (Hamilton, 1960; Bagby et al., 2004). Longer variants exist (HAM-D-21, HAM-D-24), but the 17-item form is the one used in most antidepressant trials and regulatory submissions.
- What HAM-D score counts as remission?
- A total HAM-D-17 score of ≤7 is the conventional remission threshold, and a ≥50% reduction from baseline is the conventional definition of treatment response (Frank et al., 1991). Both come from an expert consensus rather than a single validation study, and some later work argues the remission bar should be lower (≤5 or ≤6).
- How do you interpret HAM-D severity bands?
- One widely used, empirically derived scheme is: no depression 0–7, mild 8–16, moderate 17–23, and severe ≥24 (Zimmerman et al., 2013). Hamilton never published official cutoffs, so several band systems circulate—always report which one you used.
- What is the SIGH-D?
- The SIGH-D (Structured Interview Guide for the Hamilton Depression Rating Scale) is a scripted question set that standardizes how the HAM-D is administered. Williams (1988) showed it substantially improved inter-rater agreement on most items, which is why structured administration is now standard in multi-site trials.
- Is the HAM-D still the gold standard despite its known flaws?
- In practice, yes. The HAM-D has documented psychometric problems—multidimensionality and uneven item weighting among them (Bagby et al., 2004)—yet it remains the default trial endpoint because six decades of results are denominated in HAM-D points. It is a common currency more than a best-in-class measure.
Put this into practice
Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.