For Clinical researchers, trial rating teams, and computational psychiatry labs

Rate the interview from what was actually said

Apply a clinician-rated scale to an interview transcript, get an item-by-item first pass with the patient's own words attached to every rating, and measure how far it sits from your trained raters. For research and rater QA—not for clinical decisions.

Why the interview itself deserves a second look

Clinician-rated scales like the HAM-D, MADRS, BPRS, and YMRS turn a semi-structured interview into item scores. Everything downstream—eligibility, change from baseline, the trial's primary outcome—depends on those scores, and no rating is better than the interview it was made from.

When they are checked, the findings are uncomfortable: interviews too brief to cover the scale, item ratings with poor inter-rater reliability, and, in one audit of trial recordings, most interviews rated fair or unsatisfactory on most dimensions of interview quality. Central review of recorded interviews is one answer, and it costs expert time.

Tagaroo gives a research team a first read before an expert listens. The agent rates each item from the transcript and quotes the passage behind the rating, so a reviewer can see at a glance where the evidence is thin, where the rating and the words disagree, and where the interview never asked. The human rating stays the rating; the agent is a second opinion you can measure.

What this costs today

39%

of 104 audiotaped baseline HAM-D interviews in two industry-sponsored trials lasted 10 minutes or less

Engelhardt et al., J Clin Psychopharmacol, 2006

70

psychometric studies reviewed in finding that many HAM-D items have poor interrater and retest reliability

Bagby et al., American Journal of Psychiatry, 2004

0.99

ICC between the live interviewer and a blinded rater scoring the same video-recorded HAM-D interviews (18 patients)

Prasad et al., Indian Journal of Psychiatry, 2009

Figures checked August 2026. Verify current rates with each source before quoting them.

How it works, end to end

  1. 01

    Bring the interview

    Import a transcript or upload the recording for speaker-attributed transcription. Check that interviewer and patient turns are labelled correctly—items are rated from the patient's speech, never the interviewer's prompts. De-identify first with the free in-browser tool.

  2. 02

    Choose the scale

    The Scale Library ships the HAM-D, MADRS, BPRS, and YMRS as evidence-mode scales: one rating per item for the whole interview, anchored to the item's own severity levels, with the supporting passage quoted. Thought-disorder items such as BPRS Conceptual Disorganization are coded utterance by utterance instead.

  3. 03

    Get the item-level first pass

    For each item the agent proposes a rating, the span it relied on, and a rationale referencing the anchor. In evidence mode an item with no supporting passage simply has no rating, so topics the interview never covered show up as gaps rather than as zeros.

  4. 04

    Review against the rater's scores

    A trained rater—or the original interviewer—rates the same transcript. Disagreements are listed at the item and span that caused them, which is where interview-quality problems tend to surface.

  5. 05

    Measure agreement

    Tagaroo reports percent agreement and Cohen's kappa per item for each pair of raters, labelled by whether the rater saw the agent's pass. For an ICC on totals, export the ratings and use the free ICC calculator.

What Tagaroo does not do here

  • This is a research and rater-quality tool. It is not a diagnostic instrument, it must not inform clinical decisions about a patient, and an agent rating is never a substitute for a trained clinician's.
  • Items that depend on observation—psychomotor agitation and retardation, blunted affect, tension, mannerisms, apparent sadness—can only be rated from what the transcript describes. A person watching the recording must rate them.
  • The library HAM-D and BPRS cover the items that can be grounded in speech, not every item on the form, so an agent pass is not a drop-in total score. Totals and severity bands are not computed in the app; the free HAM-D, MADRS, YMRS, and BPRS calculators do that from item scores.
  • The PANSS is not in the Scale Library. It is a proprietary instrument licensed by Multi-Health Systems; if you hold a licence, upload your own copy, and check that your licence permits using it this way.
  • The suicide items flag risk. Any transcript suggesting current risk needs a clinician's attention through your study's safety procedures, whatever the agent rated.
  • Identifiable clinical data is not permitted. De-identify transcripts and recordings before upload.

Frequently asked questions

Can AI score the HAM-D or MADRS from an interview transcript?

It can draft item ratings with the evidence quoted, and whether those drafts are good enough is an empirical question for your data. Tagaroo has not published agreement figures for the HAM-D or MADRS against trained raters on real interviews yet. Use it for research and rater QA, measure it against your own raters first, and never for decisions about a patient.

Which items can't be rated from a transcript?

Anything observed rather than reported: psychomotor retardation and agitation, blunted affect, tension, mannerisms and posturing, and the observed component of apparent sadness. A transcript only carries what someone said about them, so a human rater watching the recording must rate those items.

Does Tagaroo include the PANSS?

No. The PANSS is a proprietary instrument licensed by Multi-Health Systems, so we do not ship it. The BPRS, which is in the public domain and covers much of the same ground, is in the library, and the BPRS calculator shows an approximate PANSS equivalent from published linking tables. Licence holders can upload their own PANSS copy.

How would a trial team use this for rater quality?

Run the agent over transcripts of recorded rating interviews and compare its item ratings with the site rater's. Large disagreements, and items the agent could not rate because the interview never covered them, point a central reviewer at the interviews worth listening to first. It narrows the review; it does not replace it.

Is this validated?

Not for these scales on real clinical interviews. Tagaroo's companion manuscript evaluated the architecture on a different depression instrument, and the Benchmarks page lists every per-scale result and its limits. For a study, treat agreement with your own trained raters as the validation that counts.

What does it cost?

Pro is $29/month and Teams is $99/month, with no per-interview fee and invited raters annotating free on your plan. Transcription and agreement scoring require Pro or above, and students and academic staff get 50% off Pro for their first year.

Code one session and see

Start free with a transcript you already have: apply a scale, run the agent, review its calls span by span, and print the report. Your signup credits cover the first passes. Transcribing audio or video is a Pro feature at $29/month—as is agreement scoring across raters.