For Medical, nursing, and PA educators; simulation centre leads

Score the encounter, not just the checklist

Apply a validated communication rubric to a recorded standardized-patient encounter, get an AI first pass with the exact utterance behind every judgment, and give the student feedback that points at what they actually said.

Why simulation software doesn't already do this

Simulation platforms are very good at capturing encounters and handing an examiner a digital checklist. What none of them do is read the conversation: the rubric is still filled in by a human watching in real time or re-watching the recording.

That makes formative feedback expensive. Summative OSCEs get examiners because they must; practice encounters often get a grade and a sentence, because nobody has the hours to code a cohort's worth of recordings against a communication rubric.

Published work on LLMs scoring communication rubrics unaided is not reassuring — exact agreement with trained raters sits well below what an exam could use. That is the argument for a review-first design rather than against automation: the agent proposes and cites, a human decides, and you can measure how far apart the two were before you trust it with anything that counts.

What this costs today

~1,400

US programmes assessing clinical communication (MD, DO, nursing, and PA) since Step 2 CS ended

AAMC, AACN, and PA programme counts; NBME discontinuation notice, 2021

27–44%

exact item agreement reported when frontier models score a communication rubric unaided — the case for keeping a human in the loop

Published MIRS-scoring evaluation, arXiv:2501.13957

0

sim-centre platforms we found that score the encounter transcript rather than digitize a human's checklist

Review of Laerdal, CAE, EMS, Qpercom, and Fry-IT product pages

Figures checked August 2026. Verify current rates with each source before quoting them.

How it works, end to end

  1. 01

    Start from the recording or the transcript

    Upload the encounter audio or video for speaker-attributed transcription, or import a transcript your sim platform already produced. Nothing needs to be re-recorded inside Tagaroo.

  2. 02

    Apply a communication rubric

    Use ECCS for empathic-opportunity coding or OPTION for shared decision-making from the Scale Library, or upload your station's own rubric and have Tagaroo draft an agent skill per item.

  3. 03

    Get a cited first pass

    The agent proposes each item's rating with the utterance it relied on and a rationale referencing the rubric's own anchors — so a disagreeing examiner can see precisely where the machine went wrong.

  4. 04

    Examine and adjust

    Faculty accept, edit, or reject each proposal. Two examiners on the same encounter get an agreement summary and a span-level view of every disagreement, which doubles as examiner calibration evidence.

  5. 05

    Return a feedback report

    Each encounter produces a printable report with the student's own words attached to each rating — considerably more useful for debrief than a total score.

What Tagaroo does not do here

  • This is built for formative assessment and examiner calibration. We would not put an unvalidated agent pass behind a high-stakes pass/fail decision, and nothing here has been validated for that.
  • There is no MIRS, SEGUE, or Calgary-Cambridge rubric in the library. Those carry rights we have not cleared; upload your own version if your station uses one.
  • Coding runs on the transcript. Non-verbal behaviour — eye contact, posture, gesture — is outside what the agent can see, and those items still need a human watching the video.
  • There is no integration with SimCapture, LearningSpace, or SimulationIQ. Recordings and transcripts come in by upload.
  • Station-level psychometrics (borderline regression, standard setting, cohort score reports) are not part of the product.

Frequently asked questions

Can AI grade an OSCE station?

Not on its own, and we would not sell it that way. Published evaluations put unaided model agreement with trained raters far below what a summative exam requires. What works today is the reviewed pass: the agent proposes ratings with the evidence attached, faculty correct them, and you measure the gap. Use it for formative feedback and examiner calibration until you have your own agreement data.

Can software score a standardized-patient encounter automatically?

Not on its own, and we would not sell it that way. Tagaroo applies a communication rubric — ECCS, OPTION, or your own — to the encounter transcript, and the agent proposes a first-pass rating for each item with the utterance and the reasoning behind it. A faculty member reviews, corrects, and signs off; non-verbal behaviour still needs a human watching the recording. Published work puts unaided models at 27–44% exact agreement with trained raters on a 28-item communication rubric, which is why the design is review-first and why we scope it to formative scoring and examiner calibration rather than summative pass/fail.

Do we have to move our recordings into Tagaroo?

Only the ones you want coded, and you can work from transcripts alone. Your simulation platform stays the system of record for capture; Tagaroo is where the conversation gets coded and the feedback report comes from.

Which rubrics are included?

The Scale Library ships ECCS (empathic communication) and OPTION (shared decision-making), both explicitly reproducible with citation, plus the NICHD prompt-type protocol for question quality. Station-specific rubrics are usually local anyway, so the more common path is uploading yours and editing the skills Tagaroo drafts from it.

How does this help with examiner calibration?

Have two examiners code the same encounters and Tagaroo reports their agreement, with every disagreement listed at the span that caused it. That is more actionable than a summary kappa, and the same view works for onboarding a new examiner against an experienced one.

What does it cost for a course or a cohort?

Pro is $29/month per user and Teams is $99/month including three seats, with additional seats at $15. The free plan is for evaluating the workflow on a transcript you already have; transcribing recordings and agreement scoring across examiners both require Pro or above. There is no per-station or per-encounter fee, and no institutional licence to negotiate before you can try it.

Is student data safe here?

Workspace data is hosted in the EU under GDPR and is never used to train models. Treat encounter recordings as you would any student assessment record, and de-identify standardized-patient details before upload — a free in-browser de-identifier is provided.

Code one session and see

Start free with a transcript you already have: apply a scale, run the agent, review its calls span by span, and print the report. Your signup credits cover the first passes. Transcribing audio or video is a Pro feature at $29/month — as is agreement scoring across raters.