rating scales
MADRS vs HAM-D: Which Depression Scale Should You Use?
MADRS vs HAM-D compared: items, somatic weighting, sensitivity to change, score conversion, and remission cutoffs. See which depression scale to use.

The MADRS vs HAM-D choice comes down to what you are trying to measure. The Montgomery-Åsberg Depression Rating Scale (MADRS) is ten clinician-rated, mostly mood-focused items built to be sensitive to change; the Hamilton Depression Rating Scale (HAM-D / HDRS-17) is a seventeen-item, somatically heavier instrument that has been the antidepressant-trial standard for six decades. Pick the MADRS when you want a clean signal of treatment-related change; pick the HAM-D when legacy comparability with older trials matters. Both are clinician-rated severity measures, and neither is a diagnosis.
MADRS vs HAM-D at a glance
The MADRS and HAM-D measure the same construct—depression severity—but were built for different purposes, and the design choices ripple through everything else. The table below sets them side by side on the axes that decide which to use.
| Axis | MADRS | HAM-D (HDRS-17) |
|---|---|---|
| Items | 10 | 17 (scored) |
| Item range | 0–6 (uniform) | Mixed 0–4 and 0–2 |
| Total range | 0–60 | 0–52 |
| Rater | Clinician | Clinician |
| Somatic loading | Low (2 of 10 items) | High (incl. 3 insomnia items) |
| Factor structure | Unifactorial | Multidimensional |
| Designed for | Detecting change with treatment | Describing symptom severity |
| Original source | Montgomery & Åsberg, 1979 | Hamilton, 1960 |
| Remission cutoff | ≤10 | ≤7 (consensus; contested) |
| Head-to-head effect size | 0.49 | 0.53 (Khan et al., 2002) |
Each scale has a full explainer of its own: the MADRS scoring guide and the Hamilton Depression Rating Scale guide walk through every item. This page is the head-to-head.
What’s the core difference between them?
The core difference is what each scale puts on the scoreboard. The MADRS keeps eight of its ten items on core psychological symptoms—sadness, tension, loss of interest, pessimism, suicidal thoughts—and touches the body with only two (reduced sleep and reduced appetite). The HAM-D-17 spreads its seventeen items across mood, anxiety, and a broad somatic domain, and famously splits insomnia into three separate items (initial, middle, and delayed), putting up to 6 of its 52 points on sleep alone (Bagby et al., 2004).
That imbalance is the mechanism behind most of the comparison. Because the MADRS is mostly mood, its total tracks how a patient feels; because the HAM-D carries heavy somatic weight, its total also moves with sleep, appetite, and physical symptoms that antidepressants themselves affect. The design choice Montgomery and Åsberg made in 1979—derive the scale from the items that changed most with treatment—is why the MADRS reads more like a thermometer for mood and the HAM-D more like a broad symptom inventory.
Which is more sensitive to change?
The precise claim is that the MADRS was designed to be more sensitive to change, and its own 1979 derivation study found it separated antidepressant responders from non-responders better than the Hamilton scale (Montgomery & Åsberg, 1979). That result, in the original data, is real and is the reason the scale exists.
Independent head-to-head studies temper it. Comparing both scales across the same patients, Khan and colleagues found treatment effect sizes of 0.49 for the MADRS and 0.53 for the HAM-D, concluding the MADRS is “as sensitive an instrument as HAM-D” (Khan et al., 2002).
A later replication by the same group nudged the other way, with the MADRS at 0.68 ahead of the HAM-D at 0.57 (Khan et al., 2004). The fair reading is that the two are broadly comparable on raw sensitivity. The MADRS’s more defensible advantage is not a bigger number—it is that its light somatic loading keeps treatment-emergent side effects from muddying the score.
Why is the HAM-D criticized?
The standard critique is that the HAM-D-17, for all its history, is a psychometrically flawed measure of severity. In an influential review of 70 studies titled “has the gold standard become a lead weight?”, Bagby and colleagues concluded the scale is multidimensional with a factor structure that does not replicate across samples, has several items with poor item-level reliability, and offers poor content validity for a total-score interpretation (Bagby et al., 2004). The total is treated as one number, but it is not measuring one thing.
The insomnia weighting makes this concrete. With three separate sleep items, a patient’s total can climb on sleep disturbance alone, and treatment-emergent insomnia or sedation can move the score independently of mood. This is not just a theoretical worry: in the 2023 zuranolone MOUNTAIN trial, the entry criterion was formally amended from the HDRS-17 to a MADRS threshold specifically “to reduce the potential for overrepresentation of insomnia items,” with the team noting that six insomnia points weigh less on the MADRS’s higher-ceiling total than on the HDRS-17 (Clayton et al., 2023). When a modern phase-3 protocol changes its entry scale to dodge the HAM-D’s somatic weighting, the critique has teeth.
How do you convert between MADRS and HAM-D?
Convert between the scales with percentage change, not raw points, because that is the only cross-walk that travels safely. Two rigorous methods—equipercentile linking on 4,388 patients (Leucht et al., 2017) and item-response-theory equating on 1,218 patients (Carmody et al., 2006)—produce total-score equivalences that agree closely, but both authors caution that these hold at the group level and shift with the sample.
| HAM-D-17 | ≈ MADRS (equipercentile) |
|---|---|
| 7 | ≈ 9 |
| 10 | ≈ 13 |
| 20 | ≈ 26 |
| 30 | ≈ 39 |
| 40 | ≈ 52 |
The durable finding underneath is that a percentage reduction from baseline is scale-invariant: a 50% drop on the HAM-D corresponds to roughly a 50% drop on the MADRS, even though the raw totals differ (Leucht et al., 2017). So when you need to compare a MADRS trial against the vast HAM-D literature, translate the response and remission thresholds through percentage change and treat any point-for-point total conversion as an approximation with error bars.
What are the remission and response cutoffs?
Remission is conventionally a MADRS total of 10 or below and a HAM-D-17 total of 7 or below; response, on either scale, is a reduction of at least 50% from baseline. The MADRS ≤10 cutoff has an empirical basis in a study of 684 patients (Hawley, Gale & Sivakumaran, 2002). The HAM-D ≤7 cutoff, by contrast, came from a 1991 consensus task force and was not empirically derived (Frank et al., 1991).
That distinction matters, because the ≤7 threshold is now widely argued to be too lenient. A systematic review found patients scoring at or below 7 are heterogeneous, and several analyses favor a stricter cutoff nearer 4 for a genuinely near-symptom-free state (de Zwart et al., 2018). The cutoffs are roughly equivalent across scales—HAM-D-17 ≤7 maps to about MADRS ≤9–10 on the equipercentile crosswalk—so you can report them interchangeably, but state which you used and consider a stricter sensitivity analysis. For a fully self-administered screen with different thresholds again, compare the PHQ-9 scoring guide.
Which is used in clinical trials and by the FDA?
Both scales are accepted primary endpoints, and the honest picture is a split rather than a clean winner. The HAM-D-17 was the antidepressant-trial standard for decades and remains more used in the US; the MADRS is now common, especially in European registration trials and in treatment-resistant depression. The FDA’s major-depressive-disorder guidance lists the HAM-D (typically the 17-item version) and the MADRS among accepted primary endpoints.
Recent programs make the split concrete. Esketamine and AXS-05 used the MADRS as their primary outcome, with remission defined as MADRS ≤10; zuranolone and brexanolone used change on the HAM-D-17. So the accurate statement is that the MADRS is increasingly favored for modern and treatment-resistant trials, while the HAM-D-17 retains legacy dominance and regulatory continuity—not that one has replaced the other. One regulatory nuance worth knowing: the FDA generally treats the continuous change-from-baseline total as the primary endpoint on either scale, while categorical “response” and “remission” derivations are secondary outcomes whose definitions vary across programs.
Rater reliability: the factor that beats the scale choice
Whichever scale you choose, how it is administered matters more to your results than which instrument it is. Unreliable ratings manufacture placebo response and erase drug signal. In a study comparing site raters against blinded centralized raters on the same patients, 35% of patients admitted by site raters would have been ineligible under central raters, and mean placebo change was 7.52 for site raters versus 3.18 for central raters (Kobak et al., 2010). That gap is large enough to sink a trial.
The fix is structured administration plus trained raters. The structured interview guides—SIGH-D for the HAM-D (Williams, 1988) and the SIGMA for the MADRS, which pushed between-rater agreement on the total to an intraclass correlation of 0.93 (Williams & Kobak, 2008)—standardize the probes so two coders converge. And competence is not a proxy for experience: scoring 1,241 raters on videotaped interviews, Targum found rater competence improved with structured training and was not predicted by years of clinical experience (Targum, 2006). Reliability is a property of the administration system, not the instrument alone—which is exactly where inter-rater reliability work earns its keep.
How do you score each from a transcript?
To score either scale from an interview, code the passage where the subject reports each symptom, then rate that evidence on the scale’s anchors, rather than filling in a form from memory. This produces an auditable score: every item points back to the exact words that justify it. The two calculators below run the arithmetic live—rate the items and watch each total, severity band, and remission flag move.
MADRS score simulator
Rate each item on the 0-6 MADRS anchors: 0 = symptom absent, 6 = most severe. Defined descriptions sit at 0, 2, 4, and 6; the odd values (1, 3, 5) mark intermediate states.
- 1.Apparent sadnessObserved dejection in voice, expression, and posture
- 2.Reported sadnessThe subject's own account of low mood
- 3.Inner tensionEdginess, ill-defined discomfort, dread, or panic
- 4.Reduced sleepReduced duration or depth vs. the subject's normal
- 5.Reduced appetiteLoss of appetite or of the pleasure of eating
- 6.Concentration difficultiesTrouble collecting or sustaining thoughts
- 7.LassitudeDifficulty starting and carrying out activities
- 8.Inability to feelReduced interest; loss of the capacity to feel
- 9.Pessimistic thoughtsGuilt, self-reproach, inferiority, remorse
- 10.Suicidal thoughtsLife not worth living; wishes for death
Total score
0 / 60
Severity band
Normal / symptom-free
Remission (<=10)
Yes
Item 10 flag
None
Severity bands follow Snaith et al. (1986); the remission target of a total of 10 or below follows Hawley et al. (2002). This tool illustrates how the MADRS is scored; it is an educational aid, not a diagnostic instrument, and a score is a measure of severity, not a diagnosis. Adjust the items above (or load the example) to see the score update.
HAM-D (HDRS-17) score interpreter
Rate each item on its own native range. Nine items run 0–4 (0 = absent … 4 = very severe) and eight run 0–2 (0 = absent, 1 = mild, 2 = marked) — so the same maximal rating counts for twice as much on some items as on others.
- 1.Depressed mood(0–4)
- 2.Feelings of guilt(0–4)
- 3.Suicide(0–4)
- 4.Insomnia, early (difficulty falling asleep)(0–2)
- 5.Insomnia, middle (broken sleep)(0–2)
- 6.Insomnia, late (early-morning waking)(0–2)
- 7.Work and activities(0–4)
- 8.Retardation (slowed thought, speech, movement)(0–4)
- 9.Agitation(0–4)
- 10.Anxiety, psychic(0–4)
- 11.Anxiety, somatic(0–4)
- 12.Somatic symptoms, gastrointestinal(0–2)
- 13.Somatic symptoms, general(0–2)
- 14.Genital symptoms(0–2)
- 15.Hypochondriasis(0–4)
- 16.Loss of weight(0–2)
- 17.Insight(0–2)
Total score
0 / 52
Severity band
No depression / remission
Sleep items (4–6)
0 pts
Item 3 (suicide)
None
Item structure and the 0–52 range follow Hamilton (1960); severity bands follow Zimmerman et al. (2013). This tool illustrates how the HAM-D is scored; it is an educational aid, not a diagnostic instrument, and the total is a rough, multidimensional index rather than a precise reading of one construct. Adjust the items above (or load the example) to see the score update.
Consider a short synthetic exchange, coded for both scales:
Interviewer: How have you been sleeping, and how’s your mood been this week?
Subject: I’m up at 3am every night and can’t get back down. And honestly I feel like there’s no point to any of it.
On the HAM-D-17, that reply loads two of the three insomnia items plus depressed mood, so three separate items climb. On the MADRS, the same reply moves one sleep item and the reported-sadness and pessimistic-thoughts items—fewer somatic points, more mood. Coding the evidence span rather than a global impression is what lets a reviewer check each rating and lets you compute inter-rater reliability across coders. It is also what surfaces an endorsed suicide item on either scale as its own flag, independent of the total.
On data handling: transcript coding means working with sensitive language, so the sane default is de-identified text and a privacy-first setup. Tagaroo supports a browser-side anonymous mode so transcript content can stay local rather than being uploaded—worth checking against your ethics approval before any real interview data touches a tool.
MADRS vs HAM-D: which should you use?
Pick the instrument for the question, not its reputation. The evidence points to a clean rule.
- Choose the MADRS when the goal is detecting treatment-related change with a clean mood signal, when you want to minimize somatic and side-effect confounds, or when you are running a modern or treatment-resistant-depression program where it is now the default. It is unifactorial, mood-weighted, and about twice as precise as the HDRS-17 at typical severity (Carmody et al., 2006).
- Choose the HAM-D-17 when legacy comparability and regulatory continuity matter, when you are extending a program that already used it, or when you must meta-analyze against six decades of HAM-D data. If you do, pre-register remission as ≤7 while noting it is lenient, and bridge to other scales through percentage change rather than raw totals.
- Whichever you choose, invest in a structured guide, trained and re-trained raters, and reliability surveillance—that system moves your results more than the MADRS-versus-HAM-D decision itself.
The practical upshot of MADRS vs HAM-D: the MADRS is the cleaner measure of change and the HAM-D the legacy standard, they convert only through percentages and only at the group level, and the reliability of your raters outranks the choice between them. If you code depression symptoms from interviews, Tagaroo turns both the MADRS and the Hamilton scale into guided, evidence-anchored annotation workflows with inter-rater reliability computed as your coders work.
References
- Montgomery, S. A., & Åsberg, M. (1979). A new depression scale designed to be sensitive to change. British Journal of Psychiatry, 134(4), 382–389. doi:10.1192/bjp.134.4.382
- Hamilton, M. (1960). A rating scale for depression. Journal of Neurology, Neurosurgery & Psychiatry, 23(1), 56–62. doi:10.1136/jnnp.23.1.56
- Bagby, R. M., Ryder, A. G., Schuller, D. R., & Marshall, M. B. (2004). The Hamilton Depression Rating Scale: has the gold standard become a lead weight? American Journal of Psychiatry, 161(12), 2163–2177. doi:10.1176/appi.ajp.161.12.2163
- Khan, A., Khan, S. R., Shankles, E. B., & Polissar, N. L. (2002). Relative sensitivity of the Montgomery-Asberg Depression Rating Scale, the Hamilton Depression Rating Scale and the CGI in antidepressant clinical trials. International Clinical Psychopharmacology, 17(6), 281–285. doi:10.1097/00004850-200211000-00003
- Khan, A., Brodhead, A. E., & Kolts, R. L. (2004). Relative sensitivity of the MADRS, HAM-D and CGI: a replication analysis. International Clinical Psychopharmacology, 19(3), 157–160. doi:10.1097/00004850-200405000-00006
- Carmody, T. J., Rush, A. J., Bernstein, I. H., et al. (2006). The Montgomery Åsberg and the Hamilton ratings of depression: a comparison of measures. European Neuropsychopharmacology, 16(8), 601–611. doi:10.1016/j.euroneuro.2006.04.008
- Leucht, S., Fennema, H., Engel, R. R., Kaspers-Janssen, M., & Szegedi, A. (2017/2018). Translating the HAM-D into the MADRS and vice versa with equipercentile linking. Journal of Affective Disorders, 226, 326–331. doi:10.1016/j.jad.2017.09.042
- Hawley, C. J., Gale, T. M., & Sivakumaran, T. (2002). Defining remission by cut off score on the MADRS. Journal of Affective Disorders, 72(2), 177–184. doi:10.1016/S0165-0327(01)00451-7
- Frank, E., Prien, R. F., Jarrett, R. B., et al. (1991). Conceptualization and rationale for consensus definitions of terms in major depressive disorder. Archives of General Psychiatry, 48(9), 851–855. doi:10.1001/archpsyc.1991.01810330075011
- de Zwart, P. L., Jeronimus, B. F., & de Jonge, P. (2018). Empirical evidence for definitions of episode, remission, recovery, relapse and recurrence in depression. Epidemiology and Psychiatric Sciences, 28(5), 544–562. doi:10.1017/S2045796018000227
- Kobak, K. A., Leuchter, A., DeBrota, D., et al. (2010). Site versus centralized raters in a clinical depression trial. Journal of Clinical Psychopharmacology, 30(2), 193–197. doi:10.1097/JCP.0b013e3181d20912
- Williams, J. B. W. (1988). A structured interview guide for the Hamilton Depression Rating Scale (SIGH-D). Archives of General Psychiatry, 45(8), 742–747. doi:10.1001/archpsyc.1988.01800320058007
- Williams, J. B. W., & Kobak, K. A. (2008). Development and reliability of a structured interview guide for the MADRS (SIGMA). British Journal of Psychiatry, 192(1), 52–58. doi:10.1192/bjp.bp.106.032532
- Targum, S. D. (2006). Evaluating rater competency for CNS clinical trials. Journal of Clinical Psychopharmacology, 26(3), 308–310. doi:10.1097/01.jcp.0000219049.33008.b7
- Clayton, A. H., et al. (2023). Zuranolone in major depressive disorder (MOUNTAIN). Journal of Clinical Psychiatry. doi:10.4088/JCP.22m14445
- U.S. Food and Drug Administration. Major Depressive Disorder: Developing Drugs for Treatment (draft guidance for industry). FDA guidance document
Frequently asked questions
- Is the MADRS better than the HAM-D?
- Neither is universally better; they suit different jobs. The MADRS has ten mood-weighted items, is unifactorial, and shows about twice the measurement precision of the HDRS-17 at typical severity, which makes it a cleaner measure of treatment-related change (Carmody et al., 2006). The HAM-D-17 has 60-plus years of trial data behind it and remains an accepted regulatory endpoint, so it wins on legacy comparability. Choose the MADRS for a clean change signal, the HAM-D for continuity with older evidence.
- How do you convert a HAM-D score to a MADRS score?
- Convert with percentage change, not raw points. Equipercentile linking gives approximate total-score equivalences (HAM-D 10 ≈ MADRS 13, HAM-D 20 ≈ MADRS 26, HAM-D 30 ≈ MADRS 39; Leucht et al., 2017), and an IRT crosswalk agrees closely (Carmody et al., 2006). But these hold only at the group level and depend on the sample, whereas a percentage reduction from baseline transfers between the two scales far more safely.
- What are the remission cutoffs for the MADRS and HAM-D?
- Remission is conventionally a MADRS total of 10 or below (Hawley et al., 2002) and a HAM-D-17 total of 7 or below (Frank et al., 1991). The HAM-D ≤7 value was set by consensus, not derived empirically, and several analyses argue it is too lenient, favoring a stricter cutoff nearer 4 (de Zwart et al., 2018). Response, on either scale, is conventionally a reduction of at least 50% from baseline.
- Which depression scale is used in clinical trials?
- Both. The HAM-D-17 was the historical antidepressant-trial standard and the MADRS is now common, especially in Europe and in treatment-resistant depression; the US FDA accepts both as primary endpoints (FDA MDD guidance). Recent programs are split: esketamine and AXS-05 used the MADRS, while zuranolone and brexanolone used the HAM-D-17. It is not accurate to say the MADRS has replaced the HAM-D.
- Why is the HAM-D criticized?
- The main critique is that the HAM-D-17 is multidimensional and somatically loaded, so its total is a muddy index of depression severity (Bagby et al., 2004). Three separate insomnia items put up to 6 of 52 points on sleep alone, and treatment side effects can move the score independently of mood. In one 2023 trial the entry criterion was switched from the HDRS-17 to the MADRS specifically to reduce over-representation of insomnia items (Clayton et al., 2023).
Put this into practice
Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.