Rating scales
Score the 17-item Hamilton Depression Rating Scale on its native mixed item ranges, with severity band, remission threshold — and a readout of how much of the total comes from the three small sleep items.
Free · No sign-up · Runs entirely in your browser
Rate each item on its own native range. Nine items run 0–4 (0 = absent … 4 = very severe) and eight run 0–2 (0 = absent, 1 = mild, 2 = marked) — so the same maximal rating counts for twice as much on some items as on others.
Total score
0 / 52
Sum of all 17 items
Severity band
No depression / remission
Zimmerman et al., 2013
Sleep items 4–6
0 / 6
three 0–2 insomnia items
Item 3 (suicide)
None
Rated 0–4
Enter the pre-treatment total to see change from baseline and the 50% response criterion.
Remission (≤7)
Yes
Where the person is now
Plain text: every item with its range and code, the total, the band, the sleep subtotal and the remission status.
Item structure and the 0–52 range follow Hamilton (1960); severity bands follow Zimmerman et al. (2013). This is a scoring aid for education and research support, not a diagnosis: the total is a rough, multidimensional index rather than a precise reading of one construct, and the HAM-D is a semi-structured clinician rating whose reliability depends on rater training. A non-zero suicide item needs direct clinical follow-up whatever the total says; if there is immediate risk, contact local emergency services or a crisis line. Rate the items above, or load the example, to see the score update.
Reference · for the curious
The Hamilton Depression Rating Scale has been the reference instrument for measuring depression severity since Max Hamilton published it in 1960, and it has been the target of sustained methodological criticism for nearly as long. Both facts are worth holding at once: it is simultaneously the most-used clinician-rated depression scale in trials and one of the most structurally awkward.
In the standard 17-item version, nine items are rated 0–4 and eight are rated 0–2, giving a maximum of 52. That mixed range is the first thing that trips people up, and it is a common source of scoring errors — a rater who assumes a uniform 0–4 will inflate totals substantially. The scorer above enforces each item's native range.
Commonly used bands are 0–7 normal or remission, 8–16 mild, 17–23 moderate, and 24 or above severe. Remission is conventionally a total of 7 or below; response is conventionally a reduction of at least 50% from baseline. The two answer different questions — how well the patient is now, versus how much they changed — and trials normally report both, which is why this tool takes a baseline score.
Hamilton's scale exists in 17-, 21-, 24- and 29-item versions. Only the first 17 items contribute to the standard total; the extra items in longer versions cover features such as diurnal variation, depersonalisation, paranoid and obsessional symptoms, and were not intended to be summed into the severity score.
In practice, published HAM-D scores are frequently reported without stating the version, which makes them incomparable. If you are reporting a HAM-D total, say HDRS-17 (or whichever you used). If you are reading one that does not, treat the number cautiously.
Three separate items cover initial, middle and late insomnia, each scored 0–2. Sleep disturbance alone can therefore contribute 6 points, while depressed mood — the symptom the instrument is named for — maxes out at 4. This is the scale's best-known structural criticism (Bagby et al., 2004), and it has a concrete consequence: a drug that mainly improves sleep can produce a HAM-D change that reads as an antidepressant effect.
The scorer above surfaces the sleep subtotal alongside the total for exactly this reason. Two patients with a HAM-D of 18 are not necessarily comparable, and knowing how much of the total is sleep tells you something the total conceals. Composition matters as much as magnitude.
The HAM-D is a semi-structured clinician rating, which means the interview itself is part of the instrument. Unstandardised interviewing is a well-documented source of unreliability: raters who ask different questions get different answers and then score them against anchors that were written for a different question.
If you are running a study with multiple raters, this is where your reliability will be won or lost — not in the arithmetic. Establish a structured interview guide, train against shared recordings, and measure agreement before the study rather than after. Our guides to writing guidelines that reduce disagreement and running a pilot round apply directly, and the inter-rater reliability calculator will compute agreement across your raters. For ordered severity items, use weighted kappa rather than the unweighted variant — a one-point disagreement is not the same as a four-point one.
The MADRS was designed after the HAM-D and specifically to be sensitive to change, with a uniform 0–6 range across ten items and much less somatic content — one sleep item rather than three. That makes it less likely to move for reasons unrelated to mood, and it is often preferred as a primary endpoint in antidepressant trials for that reason.
The HAM-D's advantage is history: decades of trials used it, so it remains necessary for comparability with the existing literature. Our comparison of MADRS versus HAM-D covers the choice, and depression rating scales compared sets both against the PHQ-9. Converting totals between the two is possible via equipercentile linking, but only percentage change is genuinely scale-invariant — total-to-total conversions are sample-dependent and should be reported as approximations.
Bagby and colleagues reviewed the studies published since 1979 and found Cronbach's alpha ranging from 0.46 to 0.97 across samples, inter-rater reliability from 0.82 to 0.98 by Pearson correlation and 0.46 to 0.99 by intraclass correlation, and test-retest reliability from 0.81 to 0.98. That spread is not measurement noise in the reviews; it is the finding.
The pattern within it is interpretable. Inter-rater agreement is usually acceptable when raters are trained and working from a structured interview, and poor when they are not. Internal consistency is often inadequate for a different reason: the scale is multidimensional, bundling mood, sleep, anxiety, somatic and weight items into one number, so a low alpha is partly telling you the total is not measuring a single thing. Bagby's conclusion was that the instrument's psychometric weaknesses were serious enough to justify replacing it.
The practical consequence for a study is that a HAM-D total is only as trustworthy as the interview and training behind it, so report those alongside the score.
The Hamilton Depression Rating Scale is in general free use and is reproduced widely in the published literature with citation, which is why it is hosted in the Tagaroo Scale Library. Our guide to coding the HAM-D covers the items and anchors in detail.
This is a scoring aid for education and research support, not a clinical decision tool, and it cannot substitute for the trained interviewing the HAM-D depends on — the instrument is only as reliable as the semi-structured interview behind it. All calculation happens in your browser; nothing is transmitted or stored.
Far less consistently than its status suggests, and the spread is the finding. Bagby and colleagues' 2004 review of studies published since 1979 found Cronbach's alpha ranging from 0.46 to 0.97 across samples, inter-rater reliability from 0.82 to 0.98 (Pearson) and 0.46 to 0.99 (intraclass), and test-retest reliability from 0.81 to 0.98. Inter-rater agreement is generally acceptable when raters are trained and use a structured interview; internal consistency is often not, because the scale is multidimensional rather than measuring one thing. Their conclusion was that the instrument's psychometric weaknesses are serious enough to warrant replacement — which is why reporting how you trained raters matters as much as reporting the score.
In the 17-item version, nine items are rated 0–4 and eight are rated 0–2, so the maximum is 52. Commonly used bands are 0–7 normal or remission, 8–16 mild, 17–23 moderate and 24 or above severe. Several longer versions exist (21, 24 and 29 items); only the first 17 items contribute to the standard total, which is why reported HAM-D scores are not comparable across versions unless the item count is stated.
Three separate items cover initial, middle and late insomnia, each scored 0–2, so sleep disturbance alone can contribute 6 points — more than the 4-point maximum for depressed mood itself. This is the scale's best-known structural criticism (Bagby et al., 2004): a treatment that mainly improves sleep can produce a HAM-D change that looks like an antidepressant effect. This calculator shows the sleep subtotal alongside the total so the composition of a score is visible, not just its magnitude.
Remission is conventionally a total of 7 or below. Response is conventionally a reduction of at least 50% from baseline. The two criteria answer different questions — how well the patient is now, versus how much they improved — and trials normally report both.
The Hamilton Depression Rating Scale, published by Max Hamilton in 1960, is in general free use and is reproduced widely in the published literature with citation. Note that rater training matters more for the HAM-D than for self-report instruments: the scale is a semi-structured clinician rating, and unstandardised interviewing is a well-documented source of unreliability.
Written by Enrique Gutiérrez, PhD (Computer Science) — founder of Tagaroo and Associate Professor of Computer Science, working on inter-rater reliability, measurement and annotation methodology (ORCID).
How this page is checked. Scoring rules, item ranges and thresholds are transcribed from the instrument's cited primary sources and covered by automated tests that reproduce each paper's own worked examples. This page has not been reviewed by a licensed clinician, and it is a scoring aid for education and research support rather than a clinical decision tool.
Last verified: 29 July 2026. Formulas, thresholds and cited figures on this page were checked against their original sources on that date. Every calculation runs in your browser; nothing you enter is transmitted or stored.
Tagaroo lets clinicians rate the HAM-D against the transcript of the interview, linking every item score to the passage that justifies it — so a rating can be reviewed, compared between raters, and defended.