Instrument choice usually goes wrong for a mundane reason. Someone picks the scale they have seen before, and discovers late that it needs a clinician they do not have, produces one number where they needed a code on every utterance, or measures severity when the question was whether the interviewer asked leading questions. The construct is rarely the problem. The practical constraints are.
The four questions that decide it
What are you measuring? The obvious one, and the least likely to be got wrong. Who produces the rating? A self-report form, a clinician conducting an interview, or a trained coder working from a transcript are three different measurement systems with different costs and different published reliability. This is a hard constraint: an instrument validated on trained clinicians does not keep its properties when the items are handed to someone else.
What is the number for? Screening, quantifying severity, tracking change, auditing whether a professional did something correctly, assessing an account's credibility, and describing the structure of what was said are distinct jobs, and instruments are rarely good at more than two. The HAM-D is a severity and change instrument that assumes a diagnosis; using it to screen is a category error. And what does one rating attach to? One score per session, or a label on each utterance.
Rating scale vs coding scheme
That last question is the one that most often gets discovered too late, and it deserves its own name. A rating scale asks a rater to integrate everything they observed and produce a number per item — the MADRS's ten items over an interview. A coding scheme asks a coder to walk a transcript and mark every instance of a category — the TLC's signs of disordered speech, utterance by utterance. Both are legitimate measurement, both have reliability statistics, and they answer different questions.
The tell is what your research question does with the output. "How depressed was this person?" wants a scale. "Which utterances showed derailment, and where exactly?" wants a scheme. Our full treatment of which one your study needs works through the mixed cases, and evidence mode versus instance mode covers the hybrid where a scale's items are scored but each rating carries the span that justified it.
Self-report vs clinician-rated is a trade, not a hierarchy
Clinician-rated instruments are often treated as the more rigorous option and self-reports as the convenient one, which misreads the trade. A clinician can probe an ambiguous answer, weigh observed behaviour against reported experience, and notice what a form cannot ask about. A self-report is cheap enough to administer repeatedly, does not vary with rater training, and captures the person's own account without an interpreter. Neither dominates, and the choice interacts with the others: repeated measurement over weeks favours the instrument that survives being given twenty times. Our comparison of clinician-rated and self-report scales works the trade through, and depression rating scales compared applies it to one construct where you have a real choice on every axis.
Why the near misses are shown
A recommender that returns three names and hides everything else invites the reasonable suspicion that the ranking is arbitrary. More usefully, an instrument excluded for exactly one reason is often the answer to a slightly different question, and knowing which requirement blocked it tells you what to reconsider. If the MADRS fits your construct and your unit and fails only on rater availability, the actionable finding is not "use the PHQ-9" — it is that recruiting one trained rater would open up the instrument your field expects to see in a trial.
What this tool does not do
It does not rank instruments by quality. It has no opinion on which is more valid, better validated, or more sensitive to change, because those properties are population-specific and a four-question filter cannot settle them. Every result links to its originating paper precisely so you can check the validation evidence in a population resembling yours.
It also covers 25 curated instruments rather than the field. The library was assembled for transcript-based coding and clinical rating, every entry is freely reproducible with citation, and it is deliberately not exhaustive — many widely used scales are licensed and cannot be reproduced here. If nothing matches, that may mean your combination does not exist or simply that this library does not cover your area, and the tool distinguishes those two cases rather than guessing. Browse the full set in the scale library, or see the TLC for what a coding scheme's entry looks like next to the PHQ-9's.
Frequently asked questions
Which depression rating scale should I use?
It depends on who is available to rate and what the score is for, not on which scale is best. For a brief self-report screen that also bands severity and tracks change, the PHQ-9 is the default. For clinician-rated severity in a trial, the HAM-D and the MADRS are the conventional endpoints, with the MADRS often preferred where physical illness would inflate the HAM-D's somatic items. If you need to characterise how someone talks rather than rate the person, none of those apply and you want a transcript coding scheme instead. This selector asks those questions in order.
What is the difference between a rating scale and a coding scheme?
A rating scale produces one score per person or session: a clinician or respondent judges the whole picture and assigns a number per item. A coding scheme produces a label on each utterance or span: a trained coder works through a transcript marking every instance of a category. They answer different questions and are not interchangeable. A study that needs to know how severe someone's depression is wants a scale; a study that needs to know which utterances contained self-critical content wants a scheme. That is why the unit of analysis is treated as a hard constraint here rather than a preference.
Can I use a clinician-rated scale without a clinician?
Not without changing what the score means. Instruments like the HAM-D, MADRS and YMRS assume a trained clinician conducting an interview and integrating what they observe, and their published reliability figures come from raters who had that training. Handing the same items to an untrained rater or a self-report format produces a different measurement whose properties are unknown. Where a self-report alternative exists for a construct, this tool routes to it rather than suggesting a clinician-rated scale be repurposed.
Does this rank instruments by quality?
No, and that is deliberate. It ranks by fit to your design: construct, rater, purpose, unit of analysis. It makes no claim that a recommended instrument is more valid, more reliable or better validated than one it ruled out, because those properties are population-specific and cannot be settled by four questions. Every result links to its originating publication so you can check the validation evidence in a population resembling yours, which is the part no selector can do for you.
What if nothing matches my requirements?
Then the tool says so rather than relaxing a constraint quietly. Some combinations genuinely do not exist — nobody self-reports a per-utterance code on their own transcript — and some fall outside a 25-instrument library that is curated rather than exhaustive. The near-miss list is the useful output in that case: each entry fails exactly one requirement and names which, so you can see whether the blocker is your design or the library's coverage.
Written by Enrique Gutiérrez, PhD (Computer Science) — founder of Tagaroo and Associate Professor of Computer Science, working on inter-rater reliability, measurement and annotation methodology (ORCID).
Last verified: 1 August 2026. Formulas, thresholds and cited figures on this page were checked against their original sources on that date. Every calculation runs in your browser; nothing you enter is transmitted or stored.