Reliability & AgreementConfidence Interval for Kappa, Alpha, and ICC, Done Right
Why a bare agreement number misleads, and how to build a confidence interval for kappa, alpha, and the ICC. See analytic vs bootstrap methods.
10 articles
Inter-rater reliability is where annotation projects are won or lost. If two trained raters cannot agree on what a label means, no downstream statistic can rescue the dataset—and yet IRR is routinely treated as a reporting formality rather than a design problem. These articles cover the metrics themselves—Cohen's kappa, weighted kappa, Krippendorff's alpha, the intraclass correlation coefficient—and why raw percent agreement flatters almost every project that reports it.
The metric is only half the story. The other half is the study design behind the number: how many raters per item, how to run a pilot round, what an acceptable threshold actually is in your field, how to detect coder drift in a long study, and what to do when agreement comes back low. We work through the calculations step by step, with the paradoxes (high agreement, low kappa) treated as teaching cases rather than footnotes.
If you are writing the reliability section of a paper, auditing someone else's, or deciding whether your codebook is ready for a full campaign, start here.
Reliability & AgreementWhy a bare agreement number misleads, and how to build a confidence interval for kappa, alpha, and the ICC. See analytic vs bootstrap methods.
Reliability & AgreementRaw percent agreement isn't useless. When to chance-correct, when kappa misleads, and when agreement plus a confidence interval is the honest report.
Reliability & AgreementCohen's kappa needs a fixed item set that span and NER annotation never has. Measure span-level agreement with pairwise F1 and IoU instead—see how.
Reliability & AgreementWhich ICC to use for continuous ratings: choose the model (1, 2, or 3), single vs average measures, and absolute agreement vs consistency, from one table.
Reliability & AgreementWhy Cohen's kappa collapses under skewed prevalence, and how Gwet's AC1 and AC2 chance-correction fixes it—with the formula and a worked example.
Reliability & AgreementHow to compute Krippendorff's alpha: the observed-over-expected disagreement formula for nominal, ordinal, and missing data, with a worked example.
Reliability & AgreementSet the sample size for inter-rater reliability: how many subjects and raters pin kappa or an ICC to a target CI width. See the planning tables.
Reliability & AgreementA fill-in-the-blanks methods template for how to report inter-rater reliability: coefficient, raters, unit of analysis, CI, and software, per GRRAS.
Reliability & AgreementWhich inter-rater reliability coefficient to use, by data type, rater count, and missing data—a decision guide to kappa, alpha, AC1, and ICC.
Reliability & AgreementHow to compute Cohen's kappa and weighted kappa by hand, read the result honestly, and know when to use Krippendorff's alpha, ICC, or Gwet's AC1 instead.