Rating scales
Score the Patient Health Questionnaire-9: nine items, a 0–27 total, the standard severity bands, and an explicit flag on item 9. For education and research support, not diagnosis.
Free · No sign-up · Runs entirely in your browser
Over the last two weeks, how often has the person been bothered by each problem? 0 = Not at all, 1 = Several days, 2 = More than half the days, 3 = Nearly every day.
Total score
0 / 27
Sum of all nine items
Severity band
Minimal or none
Kroenke et al., 2001
≥10 screen
Negative
Probable major depression
Item 9
None
Thoughts of self-harm
Enter the pre-treatment total to see change from baseline and the 50% response criterion.
Remission (<5)
Yes
Where the person is now
Remission on the PHQ-9 is conventionally a total below 5 — a “<” cut-point, unlike the “≤” thresholds used on the HAM-D and MADRS — and a change of about 5 points is generally taken as clinically meaningful. Both are reporting conventions rather than properties of the scale, so state the definitions you applied: rates are not comparable across studies that draw the lines differently.
Plain text: every item with its code and anchor, the total, the band, the threshold, the item-9 status, the remission status and — when a baseline is entered — the change from it.
Items, anchors and severity bands follow Kroenke, Spitzer & Williams (2001). This is a scoring aid for education and research support, not a diagnosis: the PHQ-9 is a screening and severity instrument, and a total belongs alongside a clinical assessment rather than in place of one. A non-zero item 9 needs direct clinical follow-up whatever the total says; if there is immediate risk, contact local emergency services or a crisis line. Rate the items above, or load the example, to see the score update.
Reference · for the curious
The Patient Health Questionnaire-9 is the most widely used depression measure in primary care, and its design explains why: nine items mapped directly onto the DSM criteria for major depressive disorder, a two-week window, a uniform 0–3 response format, and no licence fee. It takes a couple of minutes to complete and a few seconds to score.
Each item is rated 0 (not at all), 1 (several days), 2 (more than half the days) or 3 (nearly every day) for the past two weeks. The nine items sum to a total from 0 to 27. The severity bands in standard use are 0–4 minimal, 5–9 mild, 10–14 moderate, 15–19 moderately severe and 20–27 severe.
A tenth question asks how difficult the symptoms have made work, home life and relationships. It is not part of the total and is often overlooked, which is a shame: two people with the same score and different functional impairment are not in the same clinical situation.
Ten or above is the conventional cut-point for probable major depression. The meta-analysis by Manea and colleagues supports thresholds in the 8 to 11 range, with sensitivity and specificity trading off across that interval — which is a more honest description than a single magic number.
What the threshold does not do is diagnose. The PHQ-9 is a screening and severity instrument. A positive screen indicates that a clinical assessment is warranted; it does not establish that the criteria for a depressive episode are met, and it cannot distinguish major depression from the several other conditions that elevate these items. A score is a reason to ask more questions.
Item 9 asks about thoughts of being better off dead or of hurting oneself. Any non-zero response warrants direct clinical follow-up regardless of the total, because a person can endorse item 9 while scoring in a low band overall — the arithmetic buries it and the clinical significance does not diminish. The scorer above therefore flags item 9 separately rather than letting it disappear into the sum.
If you or someone you are assessing is at risk, contact local emergency services or a crisis line now. Nothing on this page is a substitute for clinical judgement or for care.
The PHQ-9 is sensitive enough to be used repeatedly, and a change of about 5 points is generally taken as clinically meaningful. Response is often defined as a 50% reduction from baseline and remission as a score below 5. When you report change, state the definition you used: response rates are not comparable across studies that draw the line differently.
The PHQ-9's brevity and zero cost come with a trade-off. It depends on the respondent's insight, willingness and reading of the items, and on somatic items — sleep, appetite, energy, concentration — that are elevated by plenty of things that are not depression. In physically ill or older populations the somatic items inflate scores in ways a clinician-rated instrument would partly avoid.
That is the trade-off, not a defect: it is the price of an instrument that takes two minutes and needs no trained rater. Our comparison of clinician-rated versus self-report scales works through when each is appropriate, and depression rating scales compared sets the PHQ-9 against the HAM-D and MADRS.
Rating the PHQ-9 from a recorded interview rather than a completed form splits the instrument in two. Some items leave traces in speech and behaviour a coder can point at: psychomotor change (item 8) is partly audible in rate and latency, and sleep, appetite and energy usually get described in concrete terms. Others — guilt and worthlessness, concentration — are internal states available only as report, so a transcript rating of them is a rating of what the person chose to say about them. Say which items you treated which way; that asymmetry, not the arithmetic, is where two coders diverge. Our guide to PHQ-9 scoring and transcript coding covers the item anchors in detail, and turning a rating scale into a codebook covers the operationalisation. If you are double-coding, the inter-rater reliability calculator will compute agreement across your raters.
In the original validation, internal consistency was high — Cronbach's alpha of 0.89 in the primary-care sample and 0.86 in the obstetrics-gynaecology sample — and agreement between the self-completed form and an independent clinician-administered interview was close. At the conventional cut-point of 10, both sensitivity and specificity for major depression were 88%.
Validity beyond the diagnostic threshold is what makes it useful as a severity measure: scores track functional status, disability days and self-reported symptom difficulty, which is why the instrument is used to follow treatment rather than only to screen once.
The PHQ-9 was developed by Robert Spitzer, Janet Williams and Kurt Kroenke with a grant from Pfizer, and no permission is required to reproduce, translate, display or distribute it. Two shorter forms are drawn from it: the PHQ-2 uses only the first two items as an ultra-brief screen, and the PHQ-4 pairs those with two GAD-7 items. Neither gives a severity score — a positive PHQ-2 is a prompt to administer the full nine items. That licensing freedom is a substantial part of why the instrument became ubiquitous, and why it is in the Tagaroo Scale Library alongside the GAD-7, with which it is usually administered.
This is a scoring aid for education and research support. It does not diagnose, it does not replace clinical assessment, and it stores nothing — every calculation happens in your browser and no responses are transmitted or saved.
In the original validation by Kroenke, Spitzer and Williams (2001), internal consistency was high — Cronbach's alpha of 0.89 in the primary-care sample and 0.86 in the obstetrics-gynaecology sample — with excellent test-retest agreement between the self-report form and an independent clinician-administered interview. At the conventional cut-point of 10 or above, sensitivity and specificity for major depression were both 88%. Construct validity is supported by strong associations with functional status, disability days and symptom-related difficulty.
Each of the nine items is rated 0 (not at all), 1 (several days), 2 (more than half the days) or 3 (nearly every day) for the past two weeks. Summing them gives a total from 0 to 27. The conventional severity bands are 0–4 minimal, 5–9 mild, 10–14 moderate, 15–19 moderately severe and 20–27 severe. A tenth item asking how difficult the symptoms make daily functioning is not part of the total.
A total of 10 or above is the threshold most commonly used to indicate probable major depression; a meta-analysis by Manea and colleagues supports cut-points in the 8–11 range. The PHQ-9 is a screening and severity instrument, not a diagnostic one: a positive screen calls for a clinical assessment against diagnostic criteria, not a diagnosis on its own. Score changes of about 5 points are generally taken as clinically meaningful.
Item 9 asks about thoughts of being better off dead or of hurting oneself. Any non-zero response warrants direct clinical follow-up regardless of the total score, because a person can endorse item 9 while scoring in a low band overall. That is why this calculator surfaces it separately rather than letting it disappear into the sum. If you or someone you are assessing is at risk, contact local emergency services or a crisis line now.
Yes. The PHQ-9 was developed by Robert Spitzer, Janet Williams and Kurt Kroenke with a grant from Pfizer, and no permission is required to reproduce, translate, display or distribute it. That freedom is a large part of why it became the most widely used depression measure in primary care.
Written by Enrique Gutiérrez, PhD (Computer Science) — founder of Tagaroo and Associate Professor of Computer Science, working on inter-rater reliability, measurement and annotation methodology (ORCID).
How this page is checked. Scoring rules, item ranges and thresholds are transcribed from the instrument's cited primary sources and covered by automated tests that reproduce each paper's own worked examples. This page has not been reviewed by a licensed clinician, and it is a scoring aid for education and research support rather than a clinical decision tool.
Last verified: 29 July 2026. Formulas, thresholds and cited figures on this page were checked against their original sources on that date. Every calculation runs in your browser; nothing you enter is transmitted or stored.
Tagaroo lets you rate the PHQ-9 and other instruments directly against a transcript, linking each item score to the words that justify it — so a rating can be audited, not just recorded.