rating scales
GAD-7 vs HAM-A: Self-Report vs Clinician-Rated Anxiety
GAD-7 vs HAM-A compared: the 7-item self-report anxiety screen against the 14-item clinician-rated severity endpoint. See which anxiety scale to use.

The GAD-7 vs HAM-A decision comes down to the job you need the number to do. The Generalized Anxiety Disorder-7 (GAD-7) is a seven-item questionnaire the patient fills in themselves, built to screen: it flags who might have an anxiety disorder, fast and at scale (Spitzer, Kroenke, Williams & Löwe, 2006). The Hamilton Anxiety Rating Scale (HAM-A) is a 14-item instrument a trained clinician scores after an interview, built to grade severity as a trial endpoint (Hamilton, 1959).
One is a screening tool; the other is a severity endpoint. That single distinction decides almost everything else, and neither one is a diagnosis.
GAD-7 vs HAM-A at a glance
The GAD-7 and HAM-A measure the same construct—anxiety severity—but from opposite ends of the measurement chain. One is a patient’s self-report designed for reach; the other is a clinician’s graded judgment designed for precision. The table below sets them side by side on the axes that decide which to reach for.
| Axis | GAD-7 | HAM-A |
|---|---|---|
| Type | Patient self-report | Clinician-administered |
| Items | 7 | 14 |
| Item range | 0–3 (uniform) | 0–4 (uniform) |
| Total range | 0–21 | 0–56 |
| Primary job | Screening / case-finding | Severity endpoint in trials |
| Content | Mostly worry (psychic) | Psychic + heavy somatic |
| Key threshold | Cutoff ≥10 (positive screen) | Mild 8–14 / mod 15–23 / severe ≥24 |
| Administration | ~2 min, self-completed | Trained rater, ~10–15 min |
| Original source | Spitzer et al., 2006 | Hamilton, 1959 |
Each scale has a full explainer of its own: the GAD-7 anxiety scale guide walks through the seven items and the ≥10 cutoff, and the Hamilton Anxiety Rating Scale guide covers the psychic/somatic split and its contested cutoffs. This page is the head-to-head.
What’s the core difference between them?
The core difference is who does the rating and what it is for. The GAD-7 is self-report: the patient reads seven items and rates their own worry over the past two weeks, and the total exists to answer one question—should someone look harder at this person’s anxiety? The HAM-A is clinician-rated: a trained interviewer judges 14 symptoms from what the patient describes and how they present, and the total exists to grade severity finely enough to detect change in a trial.
That split cascades into everything else. Self-report buys speed, scale, and the patient’s own perspective, at the cost of depth and any clinician judgment. Clinician rating buys richer somatic detail and a wider 0–56 range sensitive to change, but it needs a trained rater and about ten minutes per interview, and it drifts if raters are not held to the anchors.
The GAD-7 is a triage instrument; the HAM-A is a measurement instrument. Reaching for one where the other belongs is the most common design error in anxiety measurement.
Which scale should you use to screen?
Use the GAD-7 to screen. It was built and validated for exactly that: across 15 US primary-care clinics and 2,740 patients, a cutoff of ≥10 gave a sensitivity of 89% and a specificity of 82% for generalized anxiety disorder against a mental-health-professional interview (Spitzer, Kroenke, Williams & Löwe, 2006). It is free, self-completed in about two minutes, and reported a Cronbach’s alpha of 0.92 in that first study, so the seven items hang together as one anxiety dimension.
The HAM-A makes a poor screen for the same reasons it makes a good endpoint: it needs a trained rater and a ten-minute interview, which does not scale to a waiting-room handout. When a national body recommends anxiety screening—as the US Preventive Services Task Force did for adults in 2023—the instrument in view is a brief self-report like the GAD-7, followed by a clinician’s assessment, not a clinician-rated scale applied to everyone (USPSTF, 2023). If your goal is to find cases quickly and route positives onward, the self-report screen is the right tool.
Which scale should you use as a severity endpoint?
Use the HAM-A as a severity endpoint. It is the older instrument—Max Hamilton published it in 1959—and it remains a common primary outcome in anxiolytic drug trials more than sixty years later, precisely because a trained rater grading 14 items on a 0–56 range produces a fine-grained severity measure that moves with treatment (Hamilton, 1959). Its 14 items split into a psychic-anxiety factor (anxious mood, tension, fears, cognitive and depressed-mood items) and a somatic-anxiety factor (muscular, cardiovascular, respiratory, gastrointestinal, and autonomic symptoms), so you can read a psychic and a somatic subscore, not just a total.
Severity, though, is where the HAM-A gets slippery, because two cutoff systems circulate and only one is empirically derived. A traditional set (≤17 mild, and upward) appears on countless clinical pages without a traceable validation source; an alternative from trial data sets mild at 8–14, moderate at 15–23, and severe at ≥24 (Matza et al., 2010).
| Severity | Traditional set | Matza et al. (2010), empirical |
|---|---|---|
| None / minimal | — | ≤7 |
| Mild | ≤17 | 8–14 |
| Moderate | 18–24 | 15–23 |
| Severe | 25–30 | ≥24 |
| Very severe | 31–56 | — |
For change over time, HAM-A trials lean on conventions rather than bands: response is usually a reduction of at least 50% from the baseline total, and remission is often set at a total of 7 or below (Bandelow et al., 2006). The GAD-7 tracks change too—its 0–21 total is repeatable—but it has no clinician-graded severity gradient, which is why the endpoint job stays with the HAM-A.
Do the GAD-7 and HAM-A agree?
They agree closely on where a patient sits, without being the same measurement. In a primary-care validation of 212 patients—half with DSM-IV generalized anxiety disorder, half matched controls—the GAD-7 and HAM-A totals correlated r=0.852, a strong convergent-validity result (Ruiz et al., 2011). In practice that means a patient who scores high on the self-report screen will usually score high on the clinician rating, and vice versa.
Here is where a high correlation becomes a trap. Convergence tells you the two totals rise and fall together; it does not license swapping one for the other, because each captures variance the other misses. The GAD-7 is almost entirely worry and psychic tension. The HAM-A adds a whole somatic half—palpitations, breathlessness, nausea, autonomic arousal—that the GAD-7 barely touches, so two patients with the same GAD-7 can have very different HAM-A somatic subscores.
Read r=0.852 as “these instruments agree on rank order,” not as “you can report a GAD-7 where a protocol asked for a HAM-A.”
The psychic/somatic overlap: a reliability trap
The reliability trap is that neither scale cleanly isolates anxiety, and the HAM-A’s somatic half is where the leak is worst. Its bodily items—cardiovascular, respiratory, gastrointestinal, autonomic—overlap directly with the side effects of the drugs a trial is testing, so a medication’s adverse effects can read as unresolved anxiety. Maier and colleagues, after testing the scale in anxiety and depression samples, put it bluntly: “the major problems with the HAM-A are that (1) anxiolytic and antidepressant effects cannot be clearly distinguished; (2) the subscale of somatic anxiety is strongly related to somatic side effects” (Maier et al., 1988).
A second overlap compounds the first. Several HAM-A psychic items—depressed mood, insomnia, cognitive difficulty—are also core depression symptoms, so in a patient with comorbid anxiety and depression the HAM-A total is partly carrying depressive variance. This is the same confound that makes the MADRS vs HAM-D comparison turn on somatic weighting: Hamilton’s anxiety and depression scales share design and overlapping items, so their totals contaminate each other.
The GAD-7 dodges the side-effect problem by barely measuring somatic symptoms. It does not escape the depression overlap, though: anxiety and depression co-occur so often that their self-reports correlate strongly whatever the item content.
The practical consequence is a measurement discipline, not a reason to abandon either scale. If reliability across raters matters for your study—and on a clinician-rated scale it always does—the somatic items are where coders disagree most, and that noise is worth quantifying with Cohen’s kappa and inter-rater reliability at the item level.
What does each scale cost to administer?
Administration cost is the axis that most often settles the choice in the real world. The GAD-7 is self-completed on paper or a screen in about two minutes, needs no trained staff, and carries a permissive reuse statement—the official GAD-7 form, developed by Spitzer and colleagues under an educational grant from Pfizer, states that no permission is required to reproduce, translate, display, or distribute it (Spitzer et al., 2006). That is what makes it feasible to screen an entire clinic population.
The HAM-A costs more on every count. It needs a clinician trained on the anchors and roughly 10–15 minutes of interview time per patient. Its quality depends on that training: unstructured administration drifts, which is why the Structured Interview Guide for the HAM-A (SIGH-A) was built to standardize the probes. The catch is that the structured guide runs “similar but consistently higher,” about 4.2 points above unstructured administration, so mixing formats within one study adds a systematic bias unrelated to the patients (Shear et al., 2001).
The cost buys depth and a change-sensitive endpoint; it also buys a training and standardization burden the GAD-7 simply does not have.
How do you score each from a transcript?
To score either scale from an interview, code the passage where the person reports each symptom, then rate that evidence on the scale’s anchors, rather than filling in a form from memory. This produces an auditable score: every item points back to the exact words that justify it, so a reviewer can check each rating and coders can compute inter-rater reliability against each other. The GAD-7 interpreter below runs the arithmetic live—rate the seven items and watch the total, the severity band, and the ≥10 screen move.
GAD-7 score interpreter
Over the last two weeks, how often has the subject been bothered by each problem? Rate each item on the standard anchors: 0 = Not at all, 1 = Several days, 2 = More than half the days, 3 = Nearly every day.
- 1.Feeling nervous, anxious, or on edge
- 2.Not being able to stop or control worrying
- 3.Worrying too much about different things
- 4.Trouble relaxing
- 5.Being so restless that it is hard to sit still
- 6.Becoming easily annoyed or irritable
- 7.Feeling afraid, as if something awful might happen
Total score
0 / 21
Severity band
Minimal
≥10 screen
Negative
Bands and the ≥10 cutoff follow Spitzer, Kroenke, Williams & Löwe (2006). This tool illustrates how the GAD-7 is scored; it is an educational aid, not a diagnostic instrument, and a score is a screen, not a diagnosis. Adjust the items above (or load the example) to see the score update.
Consider a short synthetic exchange, coded for both scales:
Interviewer: How have the last couple of weeks been—the worry, and how your body’s felt?
Subject: My mind won’t switch off, I’m bracing for bad news that never comes. And my chest is tight, my heart races, I feel sick most mornings.
On the GAD-7, that reply loads the worry items—uncontrollable worry and fear that something awful might happen—and little else, because the scale has almost no somatic content. On the HAM-A, the same reply moves the anxious-mood item and a cluster of somatic items (cardiovascular, respiratory, gastrointestinal), so the clinician-rated total climbs on the bodily detail the self-report never captured. Coding the evidence span rather than a global impression is what lets a reviewer audit each rating—and it surfaces the somatic detail as its own subscore, which is exactly where medication side effects hide.
One note on the HAM-A card above: the published HAM-A is the full 14-item scale (Hamilton, 1959), and every count in this article refers to it. Tagaroo’s library models an 11-phenomenon subset—the items a patient can describe in words. It leaves out the somatic-sensory and genitourinary items and the observer-rated behavior-at-interview item, which are scored from examination rather than report. So the card reads “11” while the scale you cite in a paper is “14”—a text-annotation scoping choice, not a different scale.
On data handling: transcript coding means working with sensitive language, so the sane default is de-identified text and a privacy-first setup. Tagaroo supports a browser-side anonymous mode so transcript content can stay local rather than being uploaded—worth checking against your ethics approval before any real interview data touches a tool. Because anxiety and depression measures are so often collected together, the same evidence-first approach carries straight over to the depression rating scales compared and to the broader question of clinician-rated vs self-report scales.
GAD-7 vs HAM-A: which should you use?
Pick the instrument for the question, not its reputation. The evidence points to a clean rule.
- Choose the GAD-7 when you need to screen or find cases at scale, when a trained rater is not available for every patient, or when you want the patient’s own perspective quickly and cheaply. It is a validated self-report screen at a ≥10 cutoff (Spitzer et al., 2006), and it doubles as a repeatable outcome you can hand out at every visit.
- Choose the HAM-A when you need a graded severity endpoint judged by a trained rater—typically a treatment trial—where the 0–56 range and somatic detail earn their cost. If you use it, fix one administration format (structured or not) across the whole study, report the psychic and somatic subscores separately, and state which cutoff system you used (Matza et al., 2010; Shear et al., 2001).
- If you use both, treat the GAD-7 as the screen and the HAM-A as the endpoint, and do not swap one for the other on the strength of their r=0.852 correlation—that agreement is rank-order, not equivalence (Ruiz et al., 2011).
The practical upshot of GAD-7 vs HAM-A: the GAD-7 is the fast self-report screen and the HAM-A the clinician-rated severity endpoint, they converge strongly without being interchangeable, and the HAM-A’s somatic and depressive overlap is a confound to manage rather than ignore. If you code anxiety symptoms from interviews, Tagaroo turns both the GAD-7 and the HAM-A into guided, evidence-anchored annotation workflows with inter-rater reliability computed as your coders work.
References
- Spitzer, R. L., Kroenke, K., Williams, J. B. W., & Löwe, B. (2006). A brief measure for assessing generalized anxiety disorder: the GAD-7. Archives of Internal Medicine, 166(10), 1092–1097. doi:10.1001/archinte.166.10.1092
- Hamilton, M. (1959). The assessment of anxiety states by rating. British Journal of Medical Psychology, 32(1), 50–55. doi:10.1111/j.2044-8341.1959.tb00467.x
- Maier, W., Buller, R., Philipp, M., & Heuser, I. (1988). The Hamilton Anxiety Scale: reliability, validity and sensitivity to change in anxiety and depressive disorders. Journal of Affective Disorders, 14(1), 61–68. doi:10.1016/0165-0327(88)90072-9
- Ruiz, M. A., Zamorano, E., García-Campayo, J., Pardo, A., Freire, O., & Rejas, J. (2011). Validity of the GAD-7 scale as an outcome measure of disability in patients with generalized anxiety disorders in primary care. Journal of Affective Disorders, 128(3), 277–286. doi:10.1016/j.jad.2010.07.010
- Matza, L. S., Morlock, R., Sexton, C., Malley, K., & Feltner, D. (2010). Identifying HAM-A cutoffs for mild, moderate, and severe generalized anxiety disorder. International Journal of Methods in Psychiatric Research, 19(4), 223–232. doi:10.1002/mpr.323
- Shear, M. K., Vander Bilt, J., Rucci, P., Endicott, J., Lydiard, B., Otto, M. W., et al. (2001). Reliability and validity of a structured interview guide for the Hamilton Anxiety Rating Scale (SIGH-A). Depression and Anxiety, 13(4), 166–178. doi:10.1002/da.1033
- Bandelow, B., Baldwin, D. S., Dolberg, O. T., Andersen, H. F., & Stein, D. J. (2006). What is the threshold for symptomatic response and remission for major depressive disorder, panic disorder, social anxiety disorder, and generalized anxiety disorder? Journal of Clinical Psychiatry, 67(9), 1428–1434. doi:10.4088/jcp.v67n0914
- Plummer, F., Manea, L., Trepel, D., & McMillan, D. (2016). Screening for anxiety disorders with the GAD-7 and GAD-2: a systematic review and diagnostic metaanalysis. General Hospital Psychiatry, 39, 24–31. doi:10.1016/j.genhosppsych.2015.11.005
- US Preventive Services Task Force. (2023). Screening for anxiety disorders in adults: US Preventive Services Task Force recommendation statement. JAMA, 329(24), 2163–2170. doi:10.1001/jama.2023.9301
If you code anxiety symptoms from interviews or track anxiety change over time, Tagaroo turns both the GAD-7 and the HAM-A into guided, evidence-anchored annotation workflows—with inter-rater reliability computed as your coders work.
Frequently asked questions
- Is the GAD-7 or the HAM-A better for anxiety?
- Neither is universally better; they do different jobs. The GAD-7 is a seven-item self-report screen built for fast case-finding (cutoff ≥10, sensitivity 89%, specificity 82%; Spitzer, Kroenke, Williams & Löwe, 2006), while the HAM-A is a 14-item clinician-rated scale built to grade severity as a trial endpoint (Hamilton, 1959). Use the GAD-7 to decide who needs a closer look and the HAM-A to measure how severe the anxiety is under a trained rater.
- Do the GAD-7 and HAM-A measure the same thing?
- They measure overlapping constructs and track each other closely: the two totals correlated r=0.852 in a primary-care validation of 212 patients (Ruiz et al., 2011). But they are not interchangeable. The GAD-7 is almost all worry, while the HAM-A carries a heavy somatic half that also registers medication side effects and depressive symptoms (Maier et al., 1988). A high score on one predicts a high score on the other without being the same number.
- Can the GAD-7 replace the HAM-A in a clinical trial?
- Not as a severity endpoint. The HAM-A is a clinician-rated severity measure built as a trial endpoint (Hamilton, 1959), while the GAD-7 is a self-report screen (Spitzer et al., 2006); anxiety trials conventionally use the clinician rating for the graded endpoint—with response set at a ≥50% reduction and remission at a total ≤7 (Bandelow et al., 2006)—and the self-report as a screen or secondary patient-reported outcome. The GAD-7 gives no clinician judgment and ceilings at 21, so it cannot stand in for the HAM-A's 0–56 severity gradient, though many programs collect both.
- What are the GAD-7 and HAM-A cutoffs?
- The GAD-7 uses a positive-screen cutoff of ≥10, with severity bands anchored at 5/10/15 (0–4 minimal, 5–9 mild, 10–14 moderate, 15–21 severe; Spitzer et al., 2006). The HAM-A has two competing systems: a traditional set with no clear derivation, and an empirical one putting mild at 8–14, moderate at 15–23, and severe at ≥24 (Matza et al., 2010). Always report which HAM-A system you used, because 'moderate' means different totals in each.
- Why does the HAM-A overlap with depression scales?
- Several HAM-A items—depressed mood, insomnia, and cognitive difficulty—are also depression symptoms, so in a patient with comorbid anxiety and depression the HAM-A total is partly carrying depressive variance (Maier et al., 1988). The self-report GAD-7 is narrower, weighted toward worry, but it still co-occurs with depression measures, which is why anxiety and depression screens are usually read together rather than in isolation.
Put this into practice
Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.