Rating scales

PHQ-9 calculator

Score the Patient Health Questionnaire-9: nine items, a 0–27 total, the standard severity bands, and an explicit flag on item 9. For education and research support, not diagnosis.

Free · No sign-up · Runs entirely in your browser

Over the last two weeks, how often has the person been bothered by each problem? 0 = Not at all, 1 = Several days, 2 = More than half the days, 3 = Nearly every day.

  1. 1.Little interest or pleasure in doing things
  2. 2.Feeling down, depressed, or hopeless
  3. 3.Trouble falling or staying asleep, or sleeping too much
  4. 4.Feeling tired or having little energy
  5. 5.Poor appetite or overeating
  6. 6.Feeling bad about yourself — or that you are a failure or have let yourself or your family down
  7. 7.Trouble concentrating on things, such as reading the newspaper or watching television
  8. 8.Moving or speaking so slowly that other people could have noticed — or being so fidgety or restless that you move around a lot more than usual
  9. 9.Thoughts that you would be better off dead, or of hurting yourself in some way

Total score

0 / 27

Sum of all nine items

Severity band

Minimal or none

Kroenke et al., 2001

≥10 screen

Negative

Probable major depression

Item 9

None

Thoughts of self-harm

Enter the pre-treatment total to see change from baseline and the 50% response criterion.

Remission (<5)

Yes

Where the person is now

Remission on the PHQ-9 is conventionally a total below 5 — a “<” cut-point, unlike the “≤” thresholds used on the HAM-D and MADRS — and a change of about 5 points is generally taken as clinically meaningful. Both are reporting conventions rather than properties of the scale, so state the definitions you applied: rates are not comparable across studies that draw the lines differently.

Plain text: every item with its code and anchor, the total, the band, the threshold, the item-9 status, the remission status and — when a baseline is entered — the change from it.

Items, anchors and severity bands follow Kroenke, Spitzer & Williams (2001). This is a scoring aid for education and research support, not a diagnosis: the PHQ-9 is a screening and severity instrument, and a total belongs alongside a clinical assessment rather than in place of one. A non-zero item 9 needs direct clinical follow-up whatever the total says; if there is immediate risk, contact local emergency services or a crisis line. Rate the items above, or load the example, to see the score update.

Reference · for the curious

The PHQ-9: scoring, thresholds and what the total does not tell you

The Patient Health Questionnaire-9 is the most widely used depression measure in primary care, and its design explains why: nine items mapped directly onto the DSM criteria for major depressive disorder, a two-week window, a uniform 0–3 response format, and no licence fee. It takes a couple of minutes to complete and a few seconds to score.

How is the PHQ-9 scored?

Each item is rated 0 (not at all), 1 (several days), 2 (more than half the days) or 3 (nearly every day) for the past two weeks. The nine items sum to a total from 0 to 27. The severity bands in standard use are 0–4 minimal, 5–9 mild, 10–14 moderate, 15–19 moderately severe and 20–27 severe.

A tenth question asks how difficult the symptoms have made work, home life and relationships. It is not part of the total and is often overlooked, which is a shame: two people with the same score and different functional impairment are not in the same clinical situation.

The ≥10 threshold

Ten or above is the conventional cut-point for probable major depression. The meta-analysis by Manea and colleagues supports thresholds in the 8 to 11 range, with sensitivity and specificity trading off across that interval — which is a more honest description than a single magic number.

What the threshold does not do is diagnose. The PHQ-9 is a screening and severity instrument. A positive screen indicates that a clinical assessment is warranted; it does not establish that the criteria for a depressive episode are met, and it cannot distinguish major depression from the several other conditions that elevate these items. A score is a reason to ask more questions.

Item 9 stands apart

Item 9 asks about thoughts of being better off dead or of hurting oneself. Any non-zero response warrants direct clinical follow-up regardless of the total, because a person can endorse item 9 while scoring in a low band overall — the arithmetic buries it and the clinical significance does not diminish. The scorer above therefore flags item 9 separately rather than letting it disappear into the sum.

If you or someone you are assessing is at risk, contact local emergency services or a crisis line now. Nothing on this page is a substitute for clinical judgement or for care.

Tracking change over time

The PHQ-9 is sensitive enough to be used repeatedly, and a change of about 5 points is generally taken as clinically meaningful. Response is often defined as a 50% reduction from baseline and remission as a score below 5. When you report change, state the definition you used: response rates are not comparable across studies that draw the line differently.

Self-report has structural blind spots

The PHQ-9's brevity and zero cost come with a trade-off. It depends on the respondent's insight, willingness and reading of the items, and on somatic items — sleep, appetite, energy, concentration — that are elevated by plenty of things that are not depression. In physically ill or older populations the somatic items inflate scores in ways a clinician-rated instrument would partly avoid.

That is the trade-off, not a defect: it is the price of an instrument that takes two minutes and needs no trained rater. Our comparison of clinician-rated versus self-report scales works through when each is appropriate, and depression rating scales compared sets the PHQ-9 against the HAM-D and MADRS.

Coding the PHQ-9 from a transcript

Rating the PHQ-9 from a recorded interview rather than a completed form splits the instrument in two. Some items leave traces in speech and behaviour a coder can point at: psychomotor change (item 8) is partly audible in rate and latency, and sleep, appetite and energy usually get described in concrete terms. Others — guilt and worthlessness, concentration — are internal states available only as report, so a transcript rating of them is a rating of what the person chose to say about them. Say which items you treated which way; that asymmetry, not the arithmetic, is where two coders diverge. Our guide to PHQ-9 scoring and transcript coding covers the item anchors in detail, and turning a rating scale into a codebook covers the operationalisation. If you are double-coding, the inter-rater reliability calculator will compute agreement across your raters.

How reliable is the PHQ-9?

In the original validation, internal consistency was high — Cronbach's alpha of 0.89 in the primary-care sample and 0.86 in the obstetrics-gynaecology sample — and agreement between the self-completed form and an independent clinician-administered interview was close. At the conventional cut-point of 10, both sensitivity and specificity for major depression were 88%.

Validity beyond the diagnostic threshold is what makes it useful as a severity measure: scores track functional status, disability days and self-reported symptom difficulty, which is why the instrument is used to follow treatment rather than only to screen once.

Licence

The PHQ-9 was developed by Robert Spitzer, Janet Williams and Kurt Kroenke with a grant from Pfizer, and no permission is required to reproduce, translate, display or distribute it. Two shorter forms are drawn from it: the PHQ-2 uses only the first two items as an ultra-brief screen, and the PHQ-4 pairs those with two GAD-7 items. Neither gives a severity score — a positive PHQ-2 is a prompt to administer the full nine items. That licensing freedom is a substantial part of why the instrument became ubiquitous, and why it is in the Tagaroo Scale Library alongside the GAD-7, with which it is usually administered.

Scope of this tool

This is a scoring aid for education and research support. It does not diagnose, it does not replace clinical assessment, and it stores nothing — every calculation happens in your browser and no responses are transmitted or saved.

Frequently asked questions

How reliable and valid is the PHQ-9?

In the original validation by Kroenke, Spitzer and Williams (2001), internal consistency was high — Cronbach's alpha of 0.89 in the primary-care sample and 0.86 in the obstetrics-gynaecology sample — with excellent test-retest agreement between the self-report form and an independent clinician-administered interview. At the conventional cut-point of 10 or above, sensitivity and specificity for major depression were both 88%. Construct validity is supported by strong associations with functional status, disability days and symptom-related difficulty.

How is the PHQ-9 scored?

Each of the nine items is rated 0 (not at all), 1 (several days), 2 (more than half the days) or 3 (nearly every day) for the past two weeks. Summing them gives a total from 0 to 27. The conventional severity bands are 0–4 minimal, 5–9 mild, 10–14 moderate, 15–19 moderately severe and 20–27 severe. A tenth item asking how difficult the symptoms make daily functioning is not part of the total.

What PHQ-9 score indicates depression?

A total of 10 or above is the threshold most commonly used to indicate probable major depression; a meta-analysis by Manea and colleagues supports cut-points in the 8–11 range. The PHQ-9 is a screening and severity instrument, not a diagnostic one: a positive screen calls for a clinical assessment against diagnostic criteria, not a diagnosis on its own. Score changes of about 5 points are generally taken as clinically meaningful.

What does item 9 mean and why is it flagged separately?

Item 9 asks about thoughts of being better off dead or of hurting oneself. Any non-zero response warrants direct clinical follow-up regardless of the total score, because a person can endorse item 9 while scoring in a low band overall. That is why this calculator surfaces it separately rather than letting it disappear into the sum. If you or someone you are assessing is at risk, contact local emergency services or a crisis line now.

Is the PHQ-9 free to use?

Yes. The PHQ-9 was developed by Robert Spitzer, Janet Williams and Kurt Kroenke with a grant from Pfizer, and no permission is required to reproduce, translate, display or distribute it. That freedom is a large part of why it became the most widely used depression measure in primary care.

Written by Enrique Gutiérrez, PhD (Computer Science) — founder of Tagaroo and Associate Professor of Computer Science, working on inter-rater reliability, measurement and annotation methodology (ORCID).

How this page is checked. Scoring rules, item ranges and thresholds are transcribed from the instrument's cited primary sources and covered by automated tests that reproduce each paper's own worked examples. This page has not been reviewed by a licensed clinician, and it is a scoring aid for education and research support rather than a clinical decision tool.

Last verified: 29 July 2026. Formulas, thresholds and cited figures on this page were checked against their original sources on that date. Every calculation runs in your browser; nothing you enter is transmitted or stored.

Score the scale from the interview, not from memory

Tagaroo lets you rate the PHQ-9 and other instruments directly against a transcript, linking each item score to the words that justify it — so a rating can be audited, not just recorded.

Try Tagaroo free