Welcome to the Tagaroo blog: annotation, done right
Field notes on annotation—inter-rater reliability, clinical rating scales, qualitative coding, and running labeling projects that actually hold up.

Annotation is the quiet foundation under most clinical and behavioral research. Whether you are coding interview transcripts, rating symptom severity, or labeling images for a model, the quality of every downstream result is capped by the quality of the labels underneath it. This blog is about closing that gap.
What this blog covers
We write about the parts of annotation that decide whether a project holds up: inter-rater reliability and how to compute it honestly, choosing the right rating scale for a question, turning a scale into a codebook, and running a labeling campaign without burning out the people doing the work.
Every post links back to the Tagaroo Scale Library, where each instrument is a live, guided annotation workflow rather than a static PDF.
Who it is for
The writing assumes you care about getting the measurement right—clinical researchers, methodologists, qualitative coders, and the data teams building labeled datasets. It is technical where it needs to be, and honest about where methods break.
Start anywhere, and pick the instrument for the question, not the other way round.
Keep reading
Reliability & AgreementConfidence Interval for Kappa, Alpha, and ICC, Done Right
Why a bare agreement number misleads, and how to build a confidence interval for kappa, alpha, and the ICC. See analytic vs bootstrap methods.
Reliability & AgreementPercent Agreement, Reconsidered: When It's Actually Fine
Raw percent agreement isn't useless. When to chance-correct, when kappa misleads, and when agreement plus a confidence interval is the honest report.
Reliability & AgreementSpan-Level Agreement: Why Kappa Fails, Use F1 and IoU
Cohen's kappa needs a fixed item set that span and NER annotation never has. Measure span-level agreement with pairwise F1 and IoU instead—see how.
Put this into practice
Tagaroo turns any rating scale or coding scheme into a guided annotation workflow—with inter-rater reliability computed as you go.