For Psychotherapy process researchers and lab coding teams
Code the session, not a sample of it
Apply an observer-rated process measure to therapy transcripts, let the agent draft every code with the utterance behind it, and measure how closely it matches your trained coders before it touches a single analysis.
Why process studies stay small
Process research asks what happens inside a session—which interventions a therapist used, how the client responded, how the relationship moved—and answers it by having trained observers code recordings or transcripts. The coding is the study, and it is also the part that decides how many sessions the study can afford.
That cost shapes the science. Labs code a subset of sessions, a segment of each session, or a handful of cases, and the literature ends up full of small samples measured with slightly different coding manuals. And every new coder has to be trained to reliability before their ratings count.
Tagaroo is built for the deductive half of that work: your measure is fixed, and the question is whether each utterance meets it. The agent drafts the codes with a quoted span and a rationale for each, your coders review and correct them, and agreement is computed between the people and the agent—labelled by whether the raters worked independently or from the agent's draft, because those two numbers mean different things.
What this costs today
~10 h
of observational coding for each hour of recorded therapy
3.5%
of psychosocial interventions in RCTs from six leading journals that adequately addressed treatment integrity
Perepletchikova, Treat & Kazdin, J Consult Clin Psychol, 2007
0.67
inter-rater ICC between trained observer pairs rating the working alliance (WAI-SR-O) on 19 recorded sessions
Figures checked August 2026. Verify current rates with each source before quoting them.
How it works, end to end
- 01
Bring the transcripts in
Paste or import session transcripts, bring a REFI-QDA (.qdpx) project exported from NVivo, ATLAS.ti, or MAXQDA, or upload audio and video for speaker-attributed transcription. De-identify first—the free in-browser de-identifier replaces names, dates, and places with consistent pseudonyms.
- 02
Set up the coding frame
Start from a Scale Library instrument such as MISC client language or CBT cognitive distortions, or upload your own coding manual as a PDF. Tagaroo drafts an agent skill for each code—definition, examples, boundary cases, anchors—and you edit it until it says what your manual says.
- 03
Let the agent draft the codes
The agent works through each session turn by turn, proposing a code with the span it applies to, a severity or clarity rating, and a written rationale that cites the criterion it relied on.
- 04
Have your coders rate independently
Invite coders to the study. Each rates the same sessions in their own copy, and the record keeps who made each call and whether they had seen the agent's draft first.
- 05
Measure agreement before you trust it
Tagaroo reports pairwise Cohen's kappa per code, labels each figure as independent or assisted, and shows kappa again with accepted agent codes removed, so reviewing the agent's draft cannot quietly inflate reliability. Every disagreement is listed at the span that caused it.
- 06
Export for analysis
Annotations and agreement statistics export as CSV for your statistics package, and the study serializes back to .qdpx. For Krippendorff's alpha, Gwet's AC1, or an ICC, paste the exported ratings into the free calculators.
Instruments included
Each ships with items, anchors, a citation, and its reproduction terms—or bring your own rubric and Tagaroo drafts the agent skills from it.
MISC 2.x—client language
Change talk, sustain talk, and follow/neutral, the client-side codes most MI process studies are built on.
View scaleCBT cognitive distortions
Codes the distorted thinking patterns in what the client says—useful for studying cognitive change across sessions.
View scaleVR-CoDES
Emotional cues and concerns, and how the provider responds to them—an utterance-level sequence coding system.
View scaleWhat Tagaroo does not do here
- Tagaroo is a research tool. The agent's codes are a draft for trained coders to review; they are not a validated measurement on your measure until you have shown agreement on your own sessions.
- Observer-rated alliance, adherence, and competence measures are not in the Scale Library—no WAI-O, CTS-R, or CTRS. Upload your licensed copy or your own manual, and Tagaroo drafts skills from it.
- Coding runs on the transcript. Tone of voice, pauses, facial expression, and posture are outside what the agent can see, so codes that depend on them still need a human with the recording.
- The in-app statistics are pairwise Cohen's kappa and percent agreement. Sequential analyses, multilevel models, and Krippendorff's alpha happen outside Tagaroo, on the exported data.
- Identifiable clinical data is not permitted. De-identify transcripts and recordings before upload, and check that your ethics approval covers processing by an AI service hosted in the EU.
Frequently asked questions
How is this different from the supervision and training page?
Same engine, different job. Supervision is about feedback for one trainee on one tape. Process research is about producing a dataset: many sessions coded consistently on a fixed measure, with reliability you can report in a methods section. The research workflow leans on independent coders, agreement labelled by whether raters saw the agent's draft, and export for analysis rather than a printed feedback report.
Can an AI code psychotherapy sessions reliably?
Sometimes, on some codes, and you should not take anyone's word for which. Tagaroo publishes its own per-scale evaluation, including the suites where the result is weak or the test was too easy to be informative. The practical answer is to have two trained coders rate a sample of your sessions, compare each of them with the agent, and use the agent only for the codes where it reaches the agreement your field accepts.
Does reviewing the agent's draft inflate inter-rater reliability?
It can, which is why Tagaroo records it. Every annotator's exposure to the agent's output is stamped, pairs are labelled independent or assisted, and kappa is shown again with accepted agent codes removed. Report the independent figure as reliability; the assisted one measures agreement with a shared anchor.
Can we use our own coding manual?
Yes, and most process studies will. Upload the manual as a PDF or text and Tagaroo extracts the codes, definitions, and anchors into editable agent skills. The skill files stay readable, so the operational definitions can go into a supplement.
Can we import a project we already coded in NVivo or ATLAS.ti?
Yes, through the REFI-QDA exchange format. Text sources become transcripts, codes become the coding frame, and existing coding arrives as annotations. PDF, audio, and image sources inside the archive are skipped, so upload those directly.
What does it cost for a lab?
Pro is $29/month and allows 4 annotators per study; Teams is $99/month and allows 15. Invited coders annotate on the study owner's plan and pay nothing. Agreement scoring and transcription need Pro or above, and students and academic staff get 50% off Pro for their first year.
Code one session and see
Start free with a transcript you already have: apply a scale, run the agent, review its calls span by span, and print the report. Your signup credits cover the first passes. Transcribing audio or video is a Pro feature at $29/month—as is agreement scoring across raters.