tagaroo

Reliability & Agreement

10 articles

Inter-rater reliability is where annotation projects are won or lost. If two trained raters cannot agree on what a label means, no downstream statistic can rescue the dataset—and yet IRR is routinely treated as a reporting formality rather than a design problem. These articles cover the metrics themselves—Cohen's kappa, weighted kappa, Krippendorff's alpha, the intraclass correlation coefficient—and why raw percent agreement flatters almost every project that reports it.

The metric is only half the story. The other half is the study design behind the number: how many raters per item, how to run a pilot round, what an acceptable threshold actually is in your field, how to detect coder drift in a long study, and what to do when agreement comes back low. We work through the calculations step by step, with the paradoxes (high agreement, low kappa) treated as teaching cases rather than footnotes.

If you are writing the reliability section of a paper, auditing someone else's, or deciding whether your codebook is ready for a full campaign, start here.