For Qualitative researchers, PhD students, and coding teams using a codebook
An AI second coder you can check
Apply a deductive codebook to interview transcripts, let the agent draft every code with the excerpt and the reason, and measure its agreement with your human coders—so the AI is a coder you report on, not a black box you cite.
Why a codebook still takes weeks to apply
Deductive coding is the part of qualitative analysis that does not need new ideas: the codebook exists, and every excerpt either meets a code's definition or does not. It is also the part that takes the longest, because it has to be done line by line across every transcript—and then done again by a second coder on a sample, to show the codebook can be applied consistently.
General-purpose chatbots can apply a codebook, and recent studies find they often agree with human coders reasonably well. What they do not give you is the part a methods section needs: which excerpt each code rests on, which codes the model gets wrong, and an agreement figure computed the same way you would compute one between two people.
Tagaroo treats the agent as one more coder. It applies your codebook with the excerpt and a rationale for each code, you and your co-coder review independently, and agreement is computed per code—flagged by whether the coders saw the agent's draft first.
What this costs today
10–25%
of data units typically double-coded to estimate intercoder reliability
O'Connor & Joffe, International Journal of Qualitative Methods, 2020
0.71
mean Cohen's kappa for GPT-4o applying a human codebook (range 0 to 1.00 across codes), despite 96% mean agreement
85
errors in 2,352 coding decisions in the same study—57 codes wrongly applied and 28 missed—against the human coding
Figures checked August 2026. Verify current rates with each source before quoting them.
How it works, end to end
- 01
Bring the transcripts and the codebook
Import transcripts as text or a REFI-QDA (.qdpx) project from NVivo, ATLAS.ti, or MAXQDA, or upload interview recordings for transcription. Upload the codebook as a PDF or text and Tagaroo turns each code into an editable skill: definition, when to apply it, when not to, and examples.
- 02
Tighten the definitions first
Codebooks written for people often lean on shared context a model does not have. Run the agent on one transcript, read its rationales, and sharpen the exclusion rules where it misreads a code—the same piloting a new human coder would need.
- 03
Let the agent take the first pass
The agent codes each transcript and attaches the excerpt and a written rationale to every code, so each call can be audited rather than trusted.
- 04
Code a sample yourselves, independently
You and a co-coder code the same random sample without looking at the agent's draft. Tagaroo records whether each coder had seen it, because reviewing the AI's codes and then comparing with them measures something else.
- 05
Measure agreement code by code
The study reports percent agreement and Cohen's kappa for each code and each pair of coders, human or agent, with every disagreement listed at the excerpt that caused it. Codes with low agreement are the ones whose definition needs work.
- 06
Report it properly
Export the coding as CSV and paste it into the free IRR calculator for Gwet's AC1, Krippendorff's alpha, or Fleiss' kappa with confidence intervals, then use the reporting generator for the methods paragraph.
Instruments included
Each ships with items, anchors, a citation, and its reproduction terms—or bring your own rubric and Tagaroo drafts the agent skills from it.
CBT cognitive distortions
A worked deductive codebook: each category with its definition, boundaries, and contrastive examples.
View scaleLabov–Waletzky narrative structure
Orientation, complication, evaluation, resolution—a structural codebook for narrative interviews.
View scaleToulmin argument model
Claims, grounds, warrants, and rebuttals, for coding how interviewees justify a position.
View scaleWhat Tagaroo does not do here
- Tagaroo supports deductive coding—applying codes you have defined. It does not generate themes from the data, and it is not a substitute for the interpretive work of reflexive thematic analysis.
- In-app agreement is percent agreement and pairwise Cohen's kappa. Gwet's AC1, Krippendorff's alpha, and Fleiss' kappa are computed in the free calculators from the exported coding, not inside the study.
- There is no memoing, query builder, or matrix coding of the kind desktop QDA packages offer. If your analysis lives there, import and export through REFI-QDA and use Tagaroo for the coding pass.
- Agreement with the agent on your sample does not transfer to a different codebook or a different kind of interview. Re-check it when either changes.
- Identifiable data is not permitted. De-identify transcripts before upload, and confirm that your ethics approval and consent forms cover processing by an AI service hosted in the EU.
Frequently asked questions
Can I use AI as a second coder for inter-rater reliability?
You can compute agreement between a human and the agent, and it is useful for deciding where the agent can be trusted. Most reviewers will still expect reliability between two human coders who worked independently, so report both, clearly separated. Tagaroo labels each figure by whether the coders saw the agent's draft, which is the distinction that matters.
Which reliability statistic should I report?
It depends on the data. Cohen's kappa suits two coders on nominal codes; Gwet's AC1 behaves better when a code is rare and kappa collapses despite high agreement; Krippendorff's alpha handles more coders and missing data. The free IRR calculator computes all of them side by side with confidence intervals.
How much of my data should be double-coded?
A common guideline is 10–25% of data units, chosen randomly or by a stated criterion, but the defensible number depends on how precisely you need to estimate agreement. The reliability sample-size planner works that out from a target confidence-interval width.
Is this better than pasting transcripts into ChatGPT?
It is more auditable. Every code carries its excerpt and rationale, codes are applied from an editable skill rather than a one-off prompt, agreement is computed against your coders per code, and workspace data is hosted in the EU and never used for model training.
Can I bring my NVivo, ATLAS.ti, or MAXQDA project?
Yes, through the REFI-QDA (.qdpx) format all three export. Text sources, codes, and existing coding come across; PDF, audio, and image sources are skipped. A study exports back to .qdpx when you are done.
What does it cost for a PhD project?
The free plan lets you try the workflow on a transcript with starter credits. Agreement scoring between coders needs Pro at $29/month, and students get 50% off Pro for their first year. Coders you invite never pay.
Code one session and see
Start free with a transcript you already have: apply a scale, run the agent, review its calls span by span, and print the report. Your signup credits cover the first passes. Transcribing audio or video is a Pro feature at $29/month—as is agreement scoring across raters.