rating scales
The GAD-7 Anxiety Scale: Items, Cutoffs & Transcript Coding
The GAD-7 is a seven-item anxiety screen. See the 5/10/15 severity bands, the ≥10 cutoff, and how to code it from an interview transcript.

GAD-7 scoring is simple arithmetic: add seven items, each rated 0 to 3, for a total between 0 and 21. The judgment is in what that number licenses you to say, and the honest answer is less than most people assume. A total of 10 or more is the standard signal to look harder; at that threshold the screen catches roughly 89% of generalized-anxiety cases (Spitzer, Kroenke, Williams & Löwe, 2006). But the scale is a screen, not a diagnosis—and in a primary-care visit it almost never travels alone.
What is the GAD-7?
The GAD-7 (Generalized Anxiety Disorder-7) is a seven-item self-report questionnaire that measures the severity of anxiety symptoms over the previous two weeks (Spitzer, Kroenke, Williams & Löwe, 2006). It was built and validated in a criterion-standard study across 15 US primary-care clinics with 2,740 patients, and it has since become one of the most widely used anxiety measures in clinical and research settings.
Each item asks how often a specific problem has bothered the person, answered on a fixed four-point frequency scale. It is self-report: the person rates themselves, which separates it from clinician-rated anxiety instruments like the Hamilton Anxiety Rating Scale. The design makes it fast and cheap to administer, and it is free—the official form carries the statement that “no permission required to reproduce, translate, display or distribute,” and the item set exists in validated English and Spanish versions.
The instrument reported strong internal consistency in that first study—a Cronbach’s alpha of 0.92—meaning the seven items hang together as a single anxiety dimension rather than measuring seven unrelated things (Spitzer et al., 2006).
How does GAD-7 scoring work?
GAD-7 scoring sums the seven item ratings into one total from 0 to 21. Every item uses the same anchors: 0 = not at all, 1 = several days, 2 = more than half the days, 3 = nearly every day, all referring to the last two weeks. There is no reverse-scoring and no weighting—each item counts the same, and the total is just their sum.
Work a concrete (synthetic) example. A person who answers nervous 2, uncontrollable worry 3, excessive worry 2, trouble relaxing 2, restless 1, irritable 1, and dread 1 scores a total of 12—squarely in the moderate band. Because the ceiling is 21 rather than the PHQ-9’s 27, the same raw count of endorsed symptoms lands higher on this scale, which is a common source of confusion when the two totals are read side by side. The interactive interpreter that follows runs this arithmetic live, so you can see how each item moves the total, the severity band, and the ≥10 screen.
GAD-7 score interpreter
Over the last two weeks, how often has the subject been bothered by each problem? Rate each item on the standard anchors: 0 = Not at all, 1 = Several days, 2 = More than half the days, 3 = Nearly every day.
- 1.Feeling nervous, anxious, or on edge
- 2.Not being able to stop or control worrying
- 3.Worrying too much about different things
- 4.Trouble relaxing
- 5.Being so restless that it is hard to sit still
- 6.Becoming easily annoyed or irritable
- 7.Feeling afraid, as if something awful might happen
Total score
0 / 21
Severity band
Minimal
≥10 screen
Negative
Bands and the ≥10 cutoff follow Spitzer, Kroenke, Williams & Löwe (2006). This tool illustrates how the GAD-7 is scored; it is an educational aid, not a diagnostic instrument, and a score is a screen, not a diagnosis. Adjust the items above (or load the example) to see the score update.
One scoring detail is worth knowing: the standard form ends with an unscored functional-impairment question asking how difficult these problems have made daily life. That item informs clinical judgment but does not add to the 0–21 total.
What does each of the seven GAD-7 items measure?
Each GAD-7 item measures one facet of generalized anxiety, moving from the emotional core (nervousness, worry) through the somatic and behavioral signs (restlessness, irritability) to anticipatory dread. The first two items double as the GAD-2 short screen.
| # | What the item asks the person to report | Symptom domain |
|---|---|---|
| 1 | Feeling nervous, anxious, or on edge | Nervousness |
| 2 | Not being able to stop or control worrying | Uncontrollable worry |
| 3 | Worrying too much about different things | Excessive worry |
| 4 | Trouble relaxing | Trouble relaxing |
| 5 | Being so restless that it is hard to sit still | Restlessness |
| 6 | Becoming easily annoyed or irritable | Irritability |
| 7 | Feeling afraid, as if something awful might happen | Anticipatory dread |
The scale leans on worry more than any other content: items 2 and 3 are explicitly worry-focused, and item 1’s nervousness sits close to it, which is why the scale is sensitive to generalized anxiety specifically rather than to panic or phobias. Two people can both score 12 and look clinically different—one carrying mostly cognitive worry, another mostly somatic restlessness—so the total tells you about severity, not about the shape of the anxiety.
What do GAD-7 scores mean? Severity bands and cutoffs
GAD-7 interpretation uses four severity bands anchored at scores of 5, 10, and 15 (Spitzer et al., 2006). The bands describe symptom burden; they are not themselves a diagnosis.
| GAD-7 total | Anxiety severity | What the band signals |
|---|---|---|
| 0–4 | Minimal | Below the screening threshold |
| 5–9 | Mild | Sub-threshold symptoms; often watchful monitoring |
| 10–14 | Moderate | At or above the usual positive-screen cutoff |
| 15–21 | Severe | High symptom burden |
The single most important number is the cutoff of ≥10. In the original validation, a threshold of 10 or above “maximized combined sensitivity and specificity,” yielding a sensitivity of 89% and a specificity of 82% for generalized anxiety disorder against a mental-health-professional interview (Spitzer et al., 2006).
That cutoff is a convention, not a law of nature. A later systematic review and diagnostic meta-analysis by Plummer and colleagues found the GAD-7 also performed acceptably at a lower threshold, reporting a pooled sensitivity of 0.83 and specificity of 0.84 at a cutoff of ≥8 across the validation studies they pooled (Plummer, Manea, Trepel & McMillan, 2016). Lower the threshold and you catch more true cases at the cost of more false positives; raise it and you do the reverse. Choose the cutoff to fit the decision it feeds—a low bar for a screen that routes to further assessment, a higher bar when a positive result triggers something costly.
Is a GAD-7 score a diagnosis?
No—the GAD-7 is a screening and severity-tracking tool, and a score above the cutoff is a prompt to look closer, not a conclusion. The distinction is not pedantic. The US Preventive Services Task Force, reviewing the evidence in 2023, recommended screening adults for anxiety disorders but was explicit that a positive screen should be followed by a clinician’s assessment before any diagnosis (USPSTF, 2023).
Here is where the number and the clinical reality can diverge. Someone can post a 15 during an acute stressor that resolves in a fortnight, and someone with long-standing, partially-managed anxiety can answer conservatively and land at 8. The total is a snapshot of self-reported frequency over 14 days, filtered through how the person reads the items that day. Treat it as evidence to weigh, not a label to apply.
Its steadiest use is longitudinal: because it is quick and repeatable, the scale works well as a repeated outcome measure to track whether symptoms are moving over a course of care. A trend across several administrations usually says more than any single total.
GAD-7 vs PHQ-9: which screen, and when?
The GAD-7 and the PHQ-9 are the paired anxiety-and-depression screens of primary care, built by the same group and designed to be run together. The PHQ-9 covers depression; the GAD-7 covers generalized anxiety. They share the identical 0–3 frequency anchors and the same two-week window, so a clinic can hand a patient both on one sheet and score them the same way.
| GAD-7 | PHQ-9 | |
|---|---|---|
| Measures | Generalized-anxiety severity | Depression severity |
| Items | 7 | 9 |
| Score range | 0–21 | 0–27 |
| Positive-screen cutoff | ≥10 | ≥10 (typical) |
| Severity anchors | 5 / 10 / 15 | 5 / 10 / 15 / 20 |
| Ultra-brief form | GAD-2 (≥3) | PHQ-2 (≥3) |
| Origin | Spitzer et al., 2006 | Kroenke et al., 2001 |
Which to lead with depends on the presenting complaint: reach for the GAD-7 when worry, tension, or restlessness dominate, and the PHQ-9 when low mood or anhedonia does. In practice most primary-care workflows administer both, because anxiety and depression co-occur often enough that screening for one and ignoring the other misses a large share of cases. The two totals are read independently—a high GAD-7 with a low PHQ-9 is a meaningful pattern, not a contradiction.
GAD-7 vs GAD-2: when the two-item screen is enough
The GAD-2 is the first two GAD-7 items—feeling nervous and not being able to control worrying—used as an ultra-brief first-pass screen (Kroenke, Spitzer, Williams, Monahan & Löwe, 2007). The usual workflow is sequential: administer the GAD-2, and only give the full seven-item scale if the GAD-2 is positive.
A GAD-2 score of ≥3 is the standard positive threshold; in the Plummer meta-analysis it carried a pooled sensitivity of 0.76 and specificity of 0.81 for anxiety disorders (Plummer et al., 2016). The trade-off is that the GAD-2 gives you no severity band and no measure of change, which is why it screens but never monitors. If you need to track whether anxiety is improving, you need the full instrument.
| Measure | Items | Positive cutoff | Sensitivity / specificity | Gives a severity band? | Monitors change? |
|---|---|---|---|---|---|
| GAD-7 | 7 | ≥10 | 89% / 82% (Spitzer et al., 2006) | Yes—0–21 across four bands | Yes |
| GAD-2 | 2 | ≥3 | 0.76 / 0.81 (Plummer et al., 2016) | No | No |
Scoring the GAD-7 from an interview transcript
To score a GAD-7 from an interview, code the utterance where the person reports each symptom, then rate that evidence on the same 0–3 anchors—instead of ticking a checkbox form. This is the workflow Tagaroo is built for, and it produces an auditable score: every item points to the exact words that justify it.
Consider a short synthetic exchange:
Interviewer: Over the past couple of weeks, how has the worrying been?
Subject: It doesn’t really switch off. I’ll be doing the dishes and I’m already running through everything that could go wrong tomorrow, and I can’t steer my mind back.
That reply about runaway worry supports coding item 2 (uncontrollable worry) at 3—nearly every day, with the span “it doesn’t really switch off” attached as the evidence. Do that across the seven items and the total is no longer a black box: a reviewer can check each rating against the transcript rather than trusting the sum. It also makes the GAD-7 coding reproducible enough to compute inter-rater reliability across coders—its own discipline, covered in our guide to Cohen’s kappa and inter-rater reliability.
On data handling: transcript coding means working with sensitive language, so the sane default is de-identified text and a privacy-first setup. Tagaroo supports a browser-side anonymous mode so transcript content can stay local rather than being uploaded—worth checking against your ethics approval before any real interview data touches a tool. Because anxiety and depression screens are so often collected together, the same evidence-first approach carries straight over to depression coding, walked through in the PHQ-9 scoring guide.
What are the most common GAD-7 scoring mistakes?
The most common GAD-7 scoring mistake is reading the total as a diagnosis, but a few others recur. Watch for these:
- Confusing the ranges. The GAD-7 ceilings at 21 and the PHQ-9 at 27. Copying the PHQ-9’s 20+ “severe” anchor onto the anxiety scale is a frequent slip; the top band is 15–21.
- Ignoring the base rate. A ≥10 result means different things in a specialty anxiety clinic and in a general population—the cutoff’s accuracy is fixed, but the share of positives that are true cases follows prevalence.
- Summing an incomplete form. With blank items, the total under-counts severity. Require most items answered before scoring, and flag the rest.
- Treating a one-point change as meaningful. Small movements fall within measurement noise; look at the trend across administrations, not a single-point wobble, when you track change over time.
None of these are reasons to distrust the instrument—it is well-validated and genuinely useful. They are reasons to report the score with its context: the cutoff you used, the population, and whether the number came from a form or from evidence you can point to.
The practical upshot: GAD-7 scoring is easy, but interpretation is where the judgment lives. Add the seven items, respect the ≥10 cutoff for what it is—a screening threshold, not a diagnosis—and read the anxiety score next to its depression twin, because in the clinic they almost always arrive together.
References
- Spitzer, R. L., Kroenke, K., Williams, J. B. W., & Löwe, B. (2006). A brief measure for assessing generalized anxiety disorder: the GAD-7. Archives of Internal Medicine, 166(10), 1092–1097. doi:10.1001/archinte.166.10.1092
- Kroenke, K., Spitzer, R. L., Williams, J. B. W., Monahan, P. O., & Löwe, B. (2007). Anxiety disorders in primary care: prevalence, impairment, comorbidity, and detection. Annals of Internal Medicine, 146(5), 317–325. doi:10.7326/0003-4819-146-5-200703060-00004
- Plummer, F., Manea, L., Trepel, D., & McMillan, D. (2016). Screening for anxiety disorders with the GAD-7 and GAD-2: a systematic review and diagnostic metaanalysis. General Hospital Psychiatry, 39, 24–31. doi:10.1016/j.genhosppsych.2015.11.005
- US Preventive Services Task Force. (2023). Screening for anxiety disorders in adults: US Preventive Services Task Force recommendation statement. JAMA, 329(24), 2163–2170. jamanetwork.com
If you code anxiety symptoms from interviews or track anxiety change over time, Tagaroo turns the GAD-7 into a guided, evidence-anchored annotation workflow—with inter-rater reliability computed as your coders work.
Frequently asked questions
- What is a normal GAD-7 score?
- A GAD-7 total of 0–4 is the minimal band and is generally read as within normal limits (Spitzer, Kroenke, Williams & Löwe, 2006). Scores of 5–9 indicate mild symptoms. The total runs from 0 to 21, so 'normal' means the low end of that range—but the number is a screen, not a verdict, and belongs alongside clinical context.
- What does a GAD-7 score of 10 mean?
- A score of 10 sits at the bottom of the moderate band (10–14) and is the standard positive-screen threshold. In the original validation of 2,740 primary-care patients, a cutoff of ≥10 gave a sensitivity of 89% and a specificity of 82% for generalized anxiety disorder (Spitzer et al., 2006). A 10 means 'evaluate further,' not 'has an anxiety disorder.'
- What is the GAD-7 cutoff for anxiety?
- The conventional cutoff is a total score of ≥10 (Spitzer et al., 2006). A later diagnostic meta-analysis found the GAD-7 performed acceptably at ≥8, with a pooled sensitivity of 0.83 and specificity of 0.84 (Plummer, Manea, Trepel & McMillan, 2016). Pick the cutoff to match your goal: a lower value catches more cases, a higher value cuts false positives.
- Is the GAD-7 a diagnosis?
- No. The GAD-7 is a screening and severity measure, not a diagnostic test. The US Preventive Services Task Force (2023) recommends screening adults for anxiety disorders but is clear that a positive screen must be followed by a clinician's assessment. A score above a cutoff raises a flag; it does not confirm a disorder.
- When is the GAD-2 enough instead of the full GAD-7?
- The GAD-2 is the first two GAD-7 items used as an ultra-brief first pass; a score of ≥3 is the usual positive threshold, with a pooled sensitivity of 0.76 and specificity of 0.81 (Plummer et al., 2016). Use it to decide who gets the full GAD-7—it screens quickly but gives you no severity band.
Put this into practice
Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.