tagaroo

rating scales

PHQ-9 Scoring Explained: Items, Cutoffs, and Coding

PHQ-9 scoring made clear: what the nine items measure, the 0-3 anchors, the severity bands, the ≥10 cutoff, and how to code a PHQ-9 from an interview.

Enrique Gutiérrez12 min readUpdated July 2026
Nine short four-step item bars feeding into one horizontal severity gauge with graded bands, a single coral marker at the cutoff line — an abstract depiction of PHQ-9 items, 0–3 anchors, and severity bands.

PHQ-9 scoring is simple arithmetic: add up nine items, each rated 0 to 3, for a total between 0 and 27. The hard part is knowing what that number licenses you to say, and the honest answer is less than most people assume. A total of 10 or more is the standard signal to look harder; at that threshold the screen catches roughly 88% of major-depression cases (Levis, Benedetti & Thombs, 2019). But a PHQ-9 score is a screen, not a diagnosis, and one item on the form—thoughts of self-harm—matters more than the total it feeds.

What is the PHQ-9?

The PHQ-9 (Patient Health Questionnaire-9) is a nine-item self-report questionnaire that measures the severity of depressive symptoms over the previous two weeks, mapping one-to-one onto the nine DSM criteria for a major depressive episode (Kroenke, Spitzer & Williams, 2001). It was validated in a study of 6,000 primary-care and obstetrics-gynecology patients and has since become one of the most widely used depression measures in both clinical and research settings.

Each item asks how often a specific problem has bothered the person, answered on a fixed four-point frequency scale. It is self-report: the patient rates themselves, which is what separates the PHQ-9 from clinician-rated instruments like the Hamilton or the MADRS. That design makes it fast and cheap to administer, and it is free—the official form carries the statement that “no permission required to reproduce, translate, display or distribute.”

How does PHQ-9 scoring work?

PHQ-9 scoring sums the nine item ratings into a single total from 0 to 27. Each item uses the same anchors: 0 = not at all, 1 = several days, 2 = more than half the days, 3 = nearly every day, all referring to the last two weeks. There is no reverse-scoring and no weighting—every item counts the same, and the total is just their sum.

Work a concrete (synthetic) PHQ-9 example. A respondent who answers little interest 2, low mood 2, sleep 3, fatigue 2, appetite 1, feeling like a failure 1, concentration 2, psychomotor 1, and item 9 (thoughts of self-harm) 0 scores a total of 14—the top of the moderate band. The interactive interpreter that follows runs this arithmetic live, so you can see how each item moves the total and the severity band.

PHQ-9 score interpreter

Over the last two weeks, how often has the subject been bothered by each problem? Rate each item on the standard anchors: 0 = Not at all, 1 = Several days, 2 = More than half the days, 3 = Nearly every day.

  1. 1.Little interest or pleasure in doing things
  2. 2.Feeling down, depressed, or hopeless
  3. 3.Trouble falling or staying asleep, or sleeping too much
  4. 4.Feeling tired or having little energy
  5. 5.Poor appetite or overeating
  6. 6.Feeling bad about yourself — or that you are a failure or have let yourself or your family down
  7. 7.Trouble concentrating on things, such as reading the newspaper or watching television
  8. 8.Moving or speaking so slowly that other people could have noticed — or being so fidgety or restless that you move around a lot more than usual
  9. 9.Thoughts that you would be better off dead, or of hurting yourself in some way

Total score

0 / 27

Severity band

Minimal or none

Item 9 flag

None

Bands follow Kroenke, Spitzer & Williams (2001). This tool illustrates how the PHQ-9 is scored; it is an educational aid, not a diagnostic instrument, and a score is a screen, not a diagnosis. Adjust the items above (or load the example) to see the score update.

Two scoring wrinkles are worth knowing. First, there is a tenth, unscored question at the bottom of the form asking how difficult these problems have made daily functioning; it informs clinical judgment but does not add to the 0–27 total. Second, many scoring guides recommend that at least eight of the nine items be answered before the total is treated as valid, so a form with several blanks should be flagged rather than summed.

What does each of the nine PHQ-9 items measure?

Each PHQ-9 item measures one symptom of a major depressive episode, mapping the nine DSM criteria onto nine questions—which is why the scale doubles as both a severity meter and a symptom checklist.

#What the item asks the person to reportSymptom domain
1Little interest or pleasure in doing thingsAnhedonia
2Feeling down, depressed, or hopelessDepressed mood
3Trouble falling/staying asleep, or sleeping too muchSleep disturbance
4Feeling tired or having little energyFatigue
5Poor appetite or overeatingAppetite change
6Feeling bad about yourself, a failure, or letting others downWorthlessness / guilt
7Trouble concentrating on thingsConcentration
8Moving/speaking slowly, or being fidgety and restlessPsychomotor change
9Thoughts of being better off dead or of hurting yourselfSuicidal ideation
The nine PHQ-9 items and the DSM major-depression criteria they map to.

Items 3, 5, and 8 are each two-sided on purpose: the same score can mean insomnia or hypersomnia, appetite loss or overeating, slowing down or agitation. That bidirectionality is why the PHQ-9 total tells you about severity but not about the shape of a person’s depression—two people can both score 14 and look clinically very different.

What do PHQ-9 scores mean? Severity bands and cutoffs

PHQ-9 interpretation uses five severity bands anchored at scores of 5, 10, 15, and 20 (Kroenke, Spitzer & Williams, 2001). The bands describe symptom burden; they are not themselves a diagnosis.

PHQ-9 totalDepression severityWhat the band signals
0–4Minimal or noneBelow the screening threshold
5–9MildSub-threshold symptoms; often watchful monitoring
10–14ModerateAt/above the usual positive-screen cutoff
15–19Moderately severeSubstantial symptom burden
20–27SevereHigh symptom burden
PHQ-9 severity bands (Kroenke et al., 2001). Interpretation belongs with a clinician.

The single most important number is the cutoff of ≥10. In the individual-participant-data meta-analysis by Levis and colleagues—58 studies and 17,357 participants, the largest pooled evidence base for the instrument—a cutoff of 10 or above “maximized combined sensitivity and specificity,” with a pooled sensitivity of 0.88 and specificity of 0.85 against a semistructured diagnostic interview (Levis, Benedetti & Thombs, 2019).

That cutoff is a convention, not a law of nature. An earlier meta-analysis of 18 validation studies found the PHQ-9 had “acceptable diagnostic properties for detecting major depressive disorder for cut-off scores between 8 and 11” (Manea, Gilbody & McMillan, 2012). Lower the threshold and you catch more true cases at the cost of more false positives; raise it and you do the reverse. Choose the cutoff to fit the decision it feeds—a low bar for a screening step that routes to further assessment, a higher bar when a positive result triggers something costly or invasive.

Is a PHQ-9 score a diagnosis?

No—the PHQ-9 is a screening and severity-tracking tool, and a score above a cutoff is a prompt to look closer, not a conclusion. The distinction is not pedantic. The US Preventive Services Task Force, reviewing the evidence in 2023, put it plainly: people who screen positive should be “evaluated further for diagnosis and, if appropriate, are provided or referred for evidence-based care” (USPSTF, 2023).

Here’s where the number and the clinical reality can diverge. Someone can post a 16 during an acute stressor that resolves in a week, and someone with long-standing, partially-treated depression can answer conservatively and land at 8. The total is a snapshot of self-reported frequency over 14 days, filtered through how the person reads the items on that day. Treat it as evidence to weigh, not a label to apply.

How should PHQ-9 item 9 be handled?

Item 9—“thoughts that you would be better off dead, or of hurting yourself”—is treated as a standalone safety signal, independent of the total score. A person can post a low overall total and still endorse item 9, and that combination matters more than a high total with item 9 at zero.

The evidence for taking it seriously is direct. In a large health-system cohort, Simon and colleagues found that one-year risk of a suicide attempt climbed from 0.4% among people who answered item 9 “not at all” to 4% among those who answered “nearly every day”—a tenfold gradient—with a parallel rise in suicide-death risk (Simon et al., 2013). The item is not a suicide-risk assessment on its own, but a non-zero answer is a documented reason to follow up directly.

PHQ-9 vs PHQ-2: when the short screen is enough

The PHQ-2 is the first two PHQ-9 items—low interest and low mood—used as an ultra-brief first-pass screen (Kroenke, Spitzer & Williams, 2003). The usual workflow is sequential: administer the PHQ-2, and only give the full PHQ-9 if the PHQ-2 is positive.

PHQ-2PHQ-9
Items2 (interest, mood)9 (full DSM symptom set)
Score range0–60–27
Positive screen≥3≥10 (typical)
Accuracy for major depression83% sensitivity, 92% specificity88% sensitivity, 85% specificity
Best forRapid first-step screeningSeverity rating + monitoring change
PHQ-2 vs PHQ-9. PHQ-2 figures from Kroenke et al. (2003); PHQ-9 from Levis et al. (2019).

A PHQ-2 score of ≥3 has a sensitivity of 83% and specificity of 92% for major depression (Kroenke, Spitzer & Williams, 2003). The trade-off is that the PHQ-2 gives you no severity band and no item 9, which is why it screens but never monitors. For anxiety, the same authors built the parallel GAD-7, and the two are often administered together in primary care.

Scoring the PHQ-9 from an interview transcript

To score a PHQ-9 from an interview, code the utterance where the person reports each symptom, then rate that evidence on the same 0–3 anchors—instead of ticking a checkbox form. This is the workflow Tagaroo is built for, and it produces an auditable score: every item points to the exact words that justify it.

Consider a short synthetic exchange:

Interviewer: Over the past couple of weeks, how have you been sleeping?

Subject: Barely. I’m up until three or four most nights, and then I can’t drag myself out of bed in the morning.

That reply about sleeplessness supports coding item 3 (sleep) at 3—nearly every day, with the span “up until three or four most nights” attached as the evidence. Do that across the nine items and the total is no longer a black box: a reviewer can check each rating against the transcript rather than trusting the sum. It also makes the PHQ-9 coding reproducible enough to compute inter-rater reliability across coders—which is its own discipline, covered in our guide to Cohen’s kappa and inter-rater reliability.

On data handling: transcript coding means working with sensitive language, so the sane default is de-identified text and a privacy-first setup. Tagaroo supports a browser-side anonymous mode so transcript content can stay local rather than being uploaded—a note worth checking against your ethics approval before any real interview data touches a tool. For an orientation to how the Scale Library and this blog connect, the welcome post is the short version.

What are the most common PHQ-9 scoring mistakes?

The most common PHQ-9 scoring mistake is reading the total as a diagnosis, but a few others recur. Watch for these:

  • Ignoring the base rate. Sensitivity and specificity are fixed properties of the cutoff, but the chance a positive screen is a true case depends on how common depression is in your sample. In a low-prevalence population, most positives at ≥10 will be false positives.
  • Summing an incomplete form. With several blank items, the total under-counts severity. Require most items answered before scoring, and flag the rest.
  • Comparing scores across different administrations. A PHQ-9 read aloud by an interviewer, filled on paper, and completed on a phone app are not perfectly interchangeable; keep the mode consistent when you track change over time.
  • Treating a one-point change as meaningful. Small movements fall within measurement noise; a change of about five points is a widely used rule of thumb for a meaningful shift, though no single threshold is canonical.

None of these are reasons to distrust the PHQ-9—it is well-validated and genuinely useful. They are reasons to report the score with its context: the cutoff you used, the population, item 9, and whether the number came from a form or from evidence you can point to.

The practical upshot: PHQ-9 scoring is easy, but PHQ-9 interpretation is where the judgment lives. Add the nine items, respect the ≥10 cutoff for what it is—a screening threshold, not a diagnosis—and never let the total bury item 9.

References

  • Kroenke, K., Spitzer, R. L., & Williams, J. B. W. (2001). The PHQ-9: validity of a brief depression severity measure. Journal of General Internal Medicine, 16(9), 606–613. doi:10.1046/j.1525-1497.2001.016009606.x
  • Kroenke, K., Spitzer, R. L., & Williams, J. B. W. (2003). The Patient Health Questionnaire-2: validity of a two-item depression screener. Medical Care, 41(11), 1284–1292. doi:10.1097/01.MLR.0000093487.78664.3C
  • Manea, L., Gilbody, S., & McMillan, D. (2012). Optimal cut-off score for diagnosing depression with the Patient Health Questionnaire (PHQ-9): a meta-analysis. CMAJ, 184(3), E191–E196. doi:10.1503/cmaj.110829
  • Levis, B., Benedetti, A., & Thombs, B. D. (2019). Accuracy of Patient Health Questionnaire-9 (PHQ-9) for screening to detect major depression: individual participant data meta-analysis. BMJ, 365, l1476. doi:10.1136/bmj.l1476
  • Simon, G. E., Rutter, C. M., Peterson, D., Oliver, M., Whiteside, U., Operskalski, B., & Ludman, E. J. (2013). Does response on the PHQ-9 Depression Questionnaire predict subsequent suicide attempt or suicide death? Psychiatric Services, 64(12), 1195–1202. doi:10.1176/appi.ps.201200587
  • US Preventive Services Task Force. (2023). Screening for depression and suicide risk in adults: US Preventive Services Task Force recommendation statement. JAMA, 329(23), 2057–2067. doi:10.1001/jama.2023.9297
  • Spitzer, R. L., Kroenke, K., Williams, J. B. W., & Löwe, B. (2006). A brief measure for assessing generalized anxiety disorder: the GAD-7. Archives of Internal Medicine, 166(10), 1092–1097. doi:10.1001/archinte.166.10.1092

If you code depression symptoms from interviews or track PHQ-9 change over time, Tagaroo turns the PHQ-9 into a guided, evidence-anchored annotation workflow—with inter-rater reliability computed as your coders work.

Frequently asked questions

What is a normal PHQ-9 score?
A PHQ-9 total of 0–4 is the minimal-or-none band and is generally considered normal, or non-depressed (Kroenke, Spitzer & Williams, 2001). Scores of 5–9 indicate mild symptoms. The total runs from 0 to 27, so 'normal' means the low end of that range—but the score is a screen, not a verdict, and should be read alongside clinical context.
What does a PHQ-9 score of 10 mean?
A score of 10 sits at the bottom of the moderate band (10–14) and is the most widely used positive-screen threshold. In an individual-participant meta-analysis of 58 studies (Levis, Benedetti & Thombs, 2019), a cutoff of ≥10 gave a pooled sensitivity of 0.88 and specificity of 0.85 for major depression. A 10 flags 'evaluate further,' not 'has depression.'
What is the PHQ-9 cutoff for depression?
The conventional cutoff is a total score of ≥10, but a meta-analysis of 18 validation studies found acceptable diagnostic properties anywhere from 8 to 11 (Manea, Gilbody & McMillan, 2012). Pick the cutoff to match your goal: a lower value catches more cases (higher sensitivity), a higher value reduces false positives (higher specificity).
Is a PHQ-9 score a diagnosis?
No. The PHQ-9 is a screening and severity measure, not a diagnostic test. The US Preventive Services Task Force (2023) is explicit that people who screen positive should be 'evaluated further for diagnosis' by a clinician. A number above a cutoff raises a flag; it does not confirm major depression.
How should PHQ-9 item 9 be handled?
Item 9 (thoughts of being better off dead or of self-harm) is treated as a safety signal in its own right, independent of the total. Simon et al. (2013) found that one-year suicide-attempt risk rose from 0.4% among people answering 'not at all' to 4% among those answering 'nearly every day.' Any non-zero response warrants direct clinical follow-up.

Put this into practice

Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.