Reliability & agreement

Inter-rater reliability reporting generator

A coefficient is not a finding until it is reported with the things that make it interpretable. Describe your study and this writes the methods and results paragraph, then tells you which reporting elements a reviewer will still ask for.

Free · No sign-up · Runs entirely in your browser

Describe the study, get the paragraph

Fill in what you did and this writes the methods and results statement, then checks it against the GRRAS reporting checklist and tells you what a reviewer will still ask for.

The design

The statistic

The result

Your methods and results paragraph

Inter-rater reliability was assessed on 30 interview transcripts independently rated by 2 clinical psychologists (each with three or more years using the instrument), who were blinded to each other's ratings. Because the ratings were ordinal, agreement was quantified with quadratic-weighted Cohen's kappa, chosen to penalize larger severity disagreements more than adjacent ones. The unit of analysis was the per-item rating. Weighted kappa was 0.82 (95% CI 0.74–0.89), computed in R 4.4.1 (package irr 0.84.1). By the Landis and Koch (1977) benchmarks, this indicates almost perfect agreement.

Copies the paragraph plus the checklist, with unmet items still marked.

GRRAS checklist — 9 of 9

Every reporting element GRRAS asks of a methods paragraph is present.

  • Which coefficient, and why — quadratic-weighted Cohen's kappa, chosen to penalize larger severity disagreements more than adjacent ones [GRRAS 10 (statistical analysis)] (present)
  • Number of raters — 2 raters [GRRAS 6 and 11 (raters and subjects)] (present)
  • Number of subjects or units — 30 interview transcripts [GRRAS 6 and 11 (raters and subjects)] (present)
  • Rater qualifications and training — clinical psychologists (each with three or more years using the instrument) [GRRAS 12 (sample characteristics)] (present)
  • Independence and blinding — rated independently, blinded to each other's ratings [GRRAS 8 and 9 (rating process)] (present)
  • Unit of analysis — per-item rating [GRRAS 6 and 11 (subjects and objects)] (present)
  • Point estimate with a confidence interval — 0.82 (95% CI 0.74–0.89) [GRRAS 13 (statistical uncertainty)] (present)
  • Software and version — R 4.4.1 (package irr 0.84.1) [GRRAS 10 (statistical analysis)] (present)
  • Interpretation benchmark — Landis and Koch (1977) bands: almost perfect [GRRAS 13 and 14 (results and relevance)] (present)

What a reviewer will ask

  • “Which kappa or ICC?” Answered above.
  • “Over what unit?” Answered above.
  • “How many items and raters?” Answered above.
  • “Where is the confidence interval?” Answered above.
  • “Is that agreement high because the label is rare?” Report base rates and add Gwet's AC1 when prevalence is skewed (Gwet, 2008).
  • “What software?” Answered above.

Reference · for the curious

What a complete reliability report contains, and why reviewers ask for it

A reliability coefficient on its own is not a finding. The same κ = 0.82 can be a strong result or a meaningless one depending on how many subjects it came from, what the raters were, whether they conferred, and what unit it was computed over. Those facts are not garnish; they are what makes the number interpretable, and leaving them out is why reliability sections come back from review.

The six things a complete report states

A reader should be able to find each of these without inference: the coefficient and the reason you chose it; how many raters and how many subjects or units; the raters' training and qualifications; whether they rated independently and what they were blinded to; the unit each coefficient is computed over; and the point estimate with a confidence interval and the software that produced it. Then the benchmark you interpret against.

These are not a house style. They are the reporting elements the GRRAS guideline was written to standardise, distilled from its 15-item checklist for reliability and agreement studies (Kottner et al., 2011). GRRAS is the reliability equivalent of CONSORT for trials, and like CONSORT most of its force comes from being boring and specific. Our full methods template walks through the same elements in prose.

Why the interval is not optional

A kappa of 0.60 from twenty items and a kappa of 0.60 from two hundred are different claims. The first carries a 95% interval running roughly from 0.25 to 0.95 — consistent with poor agreement and with almost perfect agreement simultaneously. The second runs from about 0.49 to 0.71, which actually locates the reliability. GRRAS item 13 asks for estimates including measures of statistical uncertainty precisely because the point estimate alone cannot distinguish those two studies.

If you have not computed the interval yet, our inter-rater reliability calculator produces it alongside every coefficient in the kappa and alpha family, and the ICC calculator does the same by the exact F method for all six ICC forms. That tool also emits a one-sentence version of this paragraph as a by-product; this page owns the rest of it — the study description, the justification, the unit, the benchmark, and the audit. Our companion piece on confidence intervals for agreement covers which method suits which coefficient, including the ones that need a bootstrap.

Match the benchmark to the coefficient

Landis and Koch (1977) gave the kappa family its familiar bands — slight, fair, moderate, substantial, almost perfect. Koo and Li (2016) gave the ICC a different set — poor, moderate, good, excellent. They do not agree, and the disagreement is large enough to matter: 0.82 is "almost perfect" on the first scale and merely "good" on the second. Citing the wrong family for your coefficient is a small error that reads as carelessness.

Both sets are conventions rather than tests. Landis and Koch called their own divisions arbitrary, and treating a band as a pass mark is how a study ends up reporting that reliability was "substantial" when the interval spans three bands. Name the benchmark, report the interval, and let the reader judge.

The objection that catches people out

"Is that agreement high because the label is rare?" is the question most likely to arrive unanticipated. When one category dominates — 90% of items in a single label — two raters agree overwhelmingly by accident, kappa's chance correction subtracts nearly all of that away, and you can report a low kappa on data with 95% raw agreement. That is the prevalence paradox, and the fix is to report base rates and add Gwet's AC1 beside kappa (Gwet, 2008). Our write-up of the kappa paradox and AC1 works through a concrete case where κ collapses to 0.11 while AC1 stays at 0.80 on the same table.

This tool only raises it past about 85% skew, because below that the correction is stable and adding the caveat everywhere would train people to report an irrelevance.

Choosing the coefficient is upstream of reporting it

The generator offers only the coefficients that suit the measurement level you select, because reporting an ICC on nominal labels or a kappa on a continuous measure is a routine reviewer objection rather than a stylistic choice. If you are not yet sure which coefficient your data wants, our guide to choosing an agreement coefficient works through the decision by data type, rater count and missing data. For clinician-rated instruments in the MADRS mould, the unit of analysis question is usually the one that needs the most thought: per-item ratings and total scores give genuinely different coefficients on the same study, and only one of them is what you claimed to measure.

What this tool does not do

It does not compute a coefficient — it describes one you have already computed, and it cannot check whether your number is correct. It audits nine reporting elements rather than all fifteen GRRAS items, because the remaining six concern the title, introduction and discussion of the whole paper rather than the reliability paragraph. It is not a substitute for reading the guideline, and passing its checklist is not a claim that your study was well designed — only that it is completely described.

Frequently asked questions

How do you report inter-rater reliability in a paper?

State six things, in the order a reader expects them: the coefficient and why you chose it, how many raters and how many subjects or units, the raters' qualifications and training, whether they rated independently and what they were blinded to, the unit each coefficient is computed over, and the point estimate with its 95% confidence interval and the software that produced it. Then name the benchmark you are interpreting against. Those are the reporting elements the GRRAS guideline was written to standardise (Kottner et al., 2011), and a paragraph that covers all of them rarely comes back from review with a reliability query.

What is the GRRAS checklist?

GRRAS — Guidelines for Reporting Reliability and Agreement Studies — is a 15-item checklist published by Kottner and colleagues in 2011, covering the title, introduction, methods, results and discussion of a reliability study. It is the reliability equivalent of CONSORT for trials or PRISMA for reviews. Most of the items that decide whether a methods paragraph is complete sit in its methods and results sections, which is the nine-element subset this tool audits. It was published simultaneously in the International Journal of Nursing Studies and the Journal of Clinical Epidemiology, with identical item numbering.

Do I have to report a confidence interval for kappa?

Yes, and its absence is the single most-flagged omission in reliability reporting. A kappa of 0.60 from twenty items and a kappa of 0.60 from two hundred are not the same claim: the first carries an interval roughly from 0.25 to 0.95, the second from about 0.49 to 0.71. Only the second establishes anything. GRRAS item 13 asks for estimates of reliability and agreement including measures of statistical uncertainty, which means the interval is part of the result rather than an optional extra.

Which benchmark should I cite, Landis and Koch or Koo and Li?

Match the benchmark to the coefficient family. Landis and Koch (1977) bands — slight, fair, moderate, substantial, almost perfect — are conventional for the kappa family and for Krippendorff's alpha. Koo and Li (2016) bands — poor, moderate, good, excellent — are the standard for the ICC. The distinction matters because they disagree: a value of 0.82 reads as almost perfect on the Landis and Koch scale but only good on the Koo and Li one. Both sets are conventions rather than tests, and Landis and Koch described their own divisions as arbitrary, so cite whichever you use and do not treat the band as a pass mark.

Why does the tool ask about base rates?

Because chance-corrected agreement becomes unstable when one category dominates. If 90% of your items belong to a single label, two raters agree overwhelmingly by accident, kappa subtracts almost all of that agreement away, and you can end up reporting a low kappa on data with 95% raw agreement — the prevalence paradox. A reviewer who knows this will ask whether your high agreement is just a rare label. Reporting the base rates and adding Gwet's AC1 alongside kappa answers the question in advance (Gwet, 2008). Below about 85% skew it does not matter, which is why this tool only raises it past that point.

Written by Enrique Gutiérrez, PhD (Computer Science) — founder of Tagaroo and Associate Professor of Computer Science, working on inter-rater reliability, measurement and annotation methodology (ORCID).

Last verified: 30 July 2026. Formulas, thresholds and cited figures on this page were checked against their original sources on that date. Every calculation runs in your browser; nothing you enter is transmitted or stored.

Produce these numbers as a by-product, not a project

Every field this generator asks for — rater count, unit of analysis, the coefficient matched to the measurement level, the interval — is something Tagaroo already knows while your coders work. Double-coding runs as a normal annotation workflow and the reliability figures come out attached to the study that produced them.

Try Tagaroo free