Reliability & agreement

Screening accuracy calculator

Sensitivity is not the chance a positive result is right. Pick an instrument and a prevalence, and see the predictive values, the likelihood ratios, and the same answer as 1,000 people.

Free · No sign-up · Runs entirely in your browser

If the screen is positive, how likely is it to be right?

Not the sensitivity — that is a different question. The answer depends on how common the condition is in the people you are screening, and it moves further than most people expect. Everything runs in your browser.

The test, and who you are screening

Sensitivity

88%

95% CI 0.83–0.92

Specificity

85%

95% CI 0.82–0.88

10%
1%30%60%

Set to 10% as an illustrative primary-care figure. Use your own setting's rate: it is the input that moves the answer most.

PHQ-9 at 10 or above for major depression, against a semistructured diagnostic interview. Source: Levis, Benedetti & Thombs (2019), BMJ — 29 studies using a semistructured interview, 6725 participants (doi). See the instrument

What a result means

Positive predictive value

39.5%

Chance a positive screen is a true case

Negative predictive value

98.5%

Chance a negative screen is truly negative

False alarms per true case

1.5

How many people are worried for nothing

Likelihood ratios

+5.87 / −0.14

Independent of prevalence, unlike PPV

These predictive values come from the point estimates. The source's own 95% intervals run 0.83 to 0.92 on sensitivity and 0.82 to 0.88 on specificity, so treat the decimal place below as false precision.

Of 1,000 people screened where 10% have the condition, 223 screen positive. 88 of those have it and 135 do not, so a positive result is correct 39% of the time. 12 people with the condition screen negative and are missed.

1,000 people screened

The same result as counts rather than probabilities. Research on risk communication finds this format is understood correctly far more often than the percentages above, which is why it gets the space.

  • 88True positive
  • 135False positive
  • 12False negative
  • 765True negative

The same test at every prevalence

Sensitivity and specificity are flat lines on this chart — they do not depend on prevalence at all. The predictive values are the curves, and they are what a patient actually asks about.

0% prevalence50%100%

Positive predictive value rises with prevalence; negative predictive value (the grey line) falls. Screening a low-prevalence population is the regime where a good test still produces mostly false positives.

A positive screen is a reason for further assessment, never a diagnosis. At the settings above, 135 of every 1,000 people screened get a positive result without having the condition — and 12 who do have it are told they are negative.

Copies the predictive values, the natural-frequency breakdown, the source with its DOI and the reference standard the figures were measured against.

Reference · for the curious

Why a positive screen usually is not what people think it is

The most common error in screening is treating sensitivity as though it were the accuracy of a positive result. They are different quantities, and in the populations screening programmes actually run on, they are very far apart. A test with 88% sensitivity can be wrong about most of the people it flags, and nothing has gone wrong with the test when that happens.

This page is about the arithmetic of that gap, not about clinical practice. It explains how predictive values are computed from a test's operating characteristics and a population's base rate, using published figures for instruments we happen to cover elsewhere on the site as worked examples. It offers no guidance on whom to screen, which cutoff your service should adopt, or what to do with a result, and nothing here has been reviewed by a clinician.

PPV vs sensitivity: the direction of the conditional

Sensitivity is the probability of a positive result given the condition. Positive predictive value is the probability of the condition given a positive result. Swapping the two is a specific reasoning error, and it is common among clinicians as well as patients — the literature on it is large enough to have its own remedies.

The reason the two diverge is arithmetic rather than subtle. Screen 1,000 people where 10% have the condition, using a test at 88% sensitivity and 85% specificity, and you get 88 true positives from the 100 who have it, plus 135 false positives from the 900 who do not, because 15% of a large group is bigger than 88. The predictive value is 88 ÷ 223, or 39%. Halve the prevalence and it falls again. The instrument has not changed; the population has.

Why this tool leads with people rather than percentages

Gigerenzer and Edwards (2003) showed that the failure is largely a failure of representation. Ask the question in conditional probabilities and most people get it wrong; ask exactly the same question in natural frequencies — "of 1,000 people, how many …" — and performance improves sharply. That is why the icon array on this page is the primary output and the percentages sit above it as a summary, rather than the other way round.

The cutoff is a convention you are allowed to move

The PHQ-9's threshold of 10 is the most widely used, but Manea, Gilbody and McMillan (2012) found acceptable diagnostic properties anywhere between 8 and 11, and the GAD-7 behaves similarly — Plummer and colleagues report acceptable properties from 7 to 10. Lowering a cutoff buys sensitivity and spends specificity, which in a low-prevalence population means a large increase in false positives for a small gain in cases found. The right threshold depends on what a positive result triggers: a low bar is defensible when it routes someone to a longer conversation, much less so when it triggers something costly or stigmatising. Our full write-up of PHQ-9 scoring and interpretation works through what the total does and does not license, and the GAD-7 guide does the same for anxiety.

Screening accuracy is an agreement problem

Sensitivity and specificity describe how well a screen agrees with a reference standard — which makes them close relatives of the reliability coefficients in the rest of this section, computed on a 2×2 table with one margin treated as truth. That framing matters for two reasons. It reminds you that the reference standard is itself a measurement with its own error, so a screen can be penalised for disagreeing with an imperfect criterion. And it means the same prevalence effect that deflates PPV also destabilises kappa: when one category dominates, chance-corrected agreement collapses even though raw agreement is high. That is the kappa paradox, which the inter-rater reliability calculator flags directly, and our note on Gwet's AC1 and the kappa paradox works through a case where κ falls to 0.11 on data with 95% agreement.

If you are comparing two instruments rather than one instrument against a diagnosis, the question is usually convergence rather than accuracy — see our comparison of the GAD-7 and the HAM-A, where a correlation of about .85 still leaves the two scales non-interchangeable. The instruments themselves, with their items and anchors, are in the library: PHQ-9 and GAD-7.

What this tool does not do

It does not put a confidence interval on the predictive values. It could produce one, but it would be misleading: the preset sensitivities and specificities are pooled meta-analytic estimates with their own intervals, the prevalence you supply has no stated uncertainty at all, and combining them properly needs the covariance between sensitivity and specificity that the source papers do not report. Instead, the published intervals on the inputs are shown next to them, and they are wide — the GAD-2's sensitivity interval runs from 0.55 to 0.89, which should temper any confidence in a predictive value computed from its midpoint.

It also does not tell you where your cutoff should be, only what a given one costs, and it is not a diagnostic aid. A positive screen indicates further assessment. This is an educational and research-support tool, not a validated clinical instrument, and no output from it should be used to make a decision about a person's care.

Frequently asked questions

Why is the positive predictive value so much lower than the sensitivity?

Because they answer different questions. Sensitivity is the chance a person who has the condition screens positive; predictive value is the chance a person who screened positive has the condition. Those are only the same when the condition is universal. In a population where 10% have depression, the PHQ-9 at its standard cutoff produces 88 true positives and 135 false positives per 1,000 people screened, because the 90% without the condition are a much larger group for the 15% false-positive rate to act on. The predictive value is 39%, not 88%.

What prevalence should I use?

The rate in the population you are actually screening, which is usually not the rate in the study that validated the instrument and rarely the rate in the general population. Validation studies typically recruit from settings enriched for the condition, so their prevalence is higher than a community screening programme's and their predictive values look correspondingly better. If you do not know your rate, run the slider across a plausible range and report the span: showing that PPV moves from 20% to 55% across the range you cannot rule out is more honest than picking one number.

Does a positive PHQ-9 mean I have depression?

No. A score of 10 or more on the PHQ-9 is a signal to assess further, not a diagnosis, and the instrument's own authors describe it that way. Even in a setting where one person in ten has major depression, most people who screen positive do not have it. Diagnosis requires a clinical interview against diagnostic criteria, which is the reference standard the sensitivity and specificity figures on this page were measured against in the first place.

What are likelihood ratios, and why use them instead?

The positive likelihood ratio is sensitivity divided by one minus specificity: how much more likely a positive result is in someone with the condition than in someone without. Its advantage is that it does not depend on prevalence, so it is a stable property of the test that transfers between settings, and you can apply it to any starting probability to get a post-test one. A ratio above 10 is generally considered strong evidence and below 5 modest. The PHQ-9 at its standard cutoff has a positive likelihood ratio of about 5.9.

Where do the preset sensitivity and specificity figures come from?

Each was read from the primary publication rather than a secondary source. PHQ-9 at 10 or above: sensitivity 0.88, specificity 0.85, from Levis, Benedetti and Thombs (2019) in the BMJ, specifically the 29 studies using a semistructured interview as the reference standard. GAD-7 at 10 or above: 89% and 82%, from the original 2,740-patient validation by Spitzer and colleagues (2006). GAD-7 at 8 and GAD-2 at 3: from the pooled meta-analysis of Plummer and colleagues (2016). PHQ-2 at 3 or above: 83% and 92%, from Kroenke, Spitzer and Williams (2003). Every preset carries its DOI and the reference standard it was measured against.

Written by Enrique Gutiérrez, PhD (Computer Science) — founder of Tagaroo and Associate Professor of Computer Science, working on inter-rater reliability, measurement and annotation methodology (ORCID).

Last verified: 1 August 2026. Formulas, thresholds and cited figures on this page were checked against their original sources on that date. Every calculation runs in your browser; nothing you enter is transmitted or stored.

Keep the screen attached to the evidence for it

A total score is a summary of the answers underneath it, and the answers are usually more informative than the total. Tagaroo scores instruments against the transcript or form they came from, so an item rating stays linked to the utterance that justified it and a reviewer can see what a score was built from rather than taking the number on trust.

Try Tagaroo free