For MI researchers, trial fidelity monitors, and MI coding labs

MITI coding at the scale of a trial

Apply the MITI 4.2.1 behaviour counts to every session a trial records, not the few your coders have time for—with the agent drafting each code, your coders checking it, and agreement you can put in the fidelity section.

Why fidelity gets sampled instead of measured

A motivational interviewing trial is only interpretable if the intervention was delivered as MI. The standard way to show that is the MITI: a trained coder takes a random 20-minute segment of a session, tallies the clinician's behaviours, and rates four global dimensions. The manual stresses that the segment be truly random, especially in clinical trials, because that sample stands in for the whole intervention.

The limit is coder time. A proficient MITI coder is the product of months of training, and even then a single segment takes a fraction of a working day. So trials code a subset of sessions, report fidelity for that subset, and MI process studies that want utterance-level data across a whole corpus stay rare.

Tagaroo moves the tallying to an agent and keeps the judgement with your coders. The agent proposes a behaviour code for each clinician utterance with a clarity rating and the reason for it; coders accept or correct it; and the study reports how closely the agent agrees with each coder, so you can decide where it is good enough to extend coverage and where it is not.

What this costs today

Up to 2 h

for a proficient coder to code one 20-minute MI session, after about three months of training and supervision

Flemotomos et al., Behavior Research Methods, 2022

63%

mean adherence to NIH fidelity recommendations across 58 real-world studies of MI-based behaviour change counselling

Beck et al., British Journal of Health Psychology, 2023

21

different MI fidelity and skill tools used across 199 empirical studies

Hurlocker, Madson & Schumacher, Clinical Psychology Review, 2020

Figures checked August 2026. Verify current rates with each source before quoting them.

How it works, end to end

  1. 01

    Load the trial's sessions

    Import transcripts you already have, or upload session audio for speaker-attributed transcription. The coding targets the clinician's turns, so speaker labels matter—check them before coding. De-identify with the free in-browser tool first.

  2. 02

    Apply MITI 4.2.1 and MISC

    The Scale Library ships the nine MITI behaviour counts, each with a curated agent skill that operationalises the manual's hard boundaries: simple versus complex reflection, giving information versus persuading, persuade versus confront. Add MISC 2.x to code client change talk and sustain talk in the same session.

  3. 03

    Let the agent tally

    For each clinician utterance the agent proposes a behaviour code, the span it applies to, a clarity rating (marginal, clear, or prototypical), and a rationale quoting the criterion it used.

  4. 04

    Double-code a sample by hand

    Assign the same sessions to two trained coders working independently. Tagaroo stamps whether each coder saw the agent's draft, so the reliability figure you report is the one from raters who did not.

  5. 05

    Compare coders and agent

    Pairwise Cohen's kappa per behaviour code, for coder–coder and coder–agent pairs, with every disagreement listed at its span. That tells you which codes the agent handles and which still need a person.

  6. 06

    Export the counts

    Annotations export as CSV, so behaviour totals per session and summary scores like the reflection-to-question ratio and percent complex reflections can be computed in your analysis script.

What Tagaroo does not do here

  • The four MITI global scores—Cultivating Change Talk, Softening Sustain Talk, Partnership, Empathy—are not modelled. They are gestalt ratings of the whole segment; code them by hand or upload your own global-rating rubric.
  • MITI summary scores (reflection-to-question ratio, percent complex reflections, relational and technical composites) and the competence thresholds are not computed in the app. Export the counts and compute them.
  • The agent's MITI figures on the Benchmarks page come from a 20-instance model-written suite with a stated severity-scoring error. They show the decision rules hold on hard synthetic cases, not that the agent matches trained coders on real trial sessions—measure that on yours.
  • Tagaroo is not a certified coding service and does not replace the trained coders a trial's fidelity plan names. Use it to extend coverage where your own agreement data supports it.
  • Identifiable clinical data is not permitted. De-identify recordings and transcripts before upload.

Frequently asked questions

Is this the same as the supervision page?

The instrument is the same; the job is not. The supervision page is about feedback for a trainee. This page is about fidelity monitoring and research: coding many sessions consistently, keeping coder–coder reliability separate from agreement with the agent, and exporting counts for analysis.

Can AI code the MITI well enough for a trial?

Not on anyone's say-so, including ours. Tagaroo's published MITI suite is small and model-written, so it shows the agent applies the manual's boundaries on hard synthetic cases, not that it matches your coders on real sessions. The defensible path is to double-code a random sample by hand, compare the agent with each coder per behaviour code, and use it only where it reaches the reliability your protocol specifies.

Does Tagaroo produce MITI global scores?

No. The Scale Library models the nine behaviour counts. The four global ratings are holistic judgments of a whole segment, and we do not ship an agent for them; code them by hand, or upload your own global-rating rubric and review what the agent drafts from it.

Can we code whole sessions instead of 20-minute segments?

Yes—the agent codes whatever transcript you give it, and there is no per-minute coding fee. Keep in mind that the MITI manual recommends a random 20-minute segment and warns that global scores are harder to interpret for longer or shorter samples; behaviour counts per session are usually normalised by time or by segment in the analysis.

How do we report reliability if an AI was involved?

Report the independent coder–coder agreement as your reliability, name the coefficient and the sample, and describe the agent's role and its agreement with each coder separately. Tagaroo labels each figure by whether the raters saw the agent's draft, and the free reporting generator turns study details into a paste-ready paragraph.

What does it cost?

Pro is $29/month with 4 annotators per study; Teams is $99/month with 15. Coders you invite annotate on your plan and pay nothing, and there is no per-tape fee. Transcription and agreement scoring need Pro or above, and students and academic staff get 50% off Pro for their first year.

Code one session and see

Start free with a transcript you already have: apply a scale, run the agent, review its calls span by span, and print the report. Your signup credits cover the first passes. Transcribing audio or video is a Pro feature at $29/month—as is agreement scoring across raters.