tagaroo

inter rater reliability

How to Report Inter-Rater Reliability: A Methods Template

A fill-in-the-blanks methods template for how to report inter-rater reliability: coefficient, raters, unit of analysis, CI, and software, per GRRAS.

Enrique Gutiérrez13 min readUpdated July 2026
A methods paragraph shown as a row of bracketed blank slots being filled with green and teal blocks beside a checklist of tick-marks, with one slot marked by a coral confidence-interval bracket.

A reviewer can reject an otherwise clean reliability study over one missing sentence in the methods. The number that took weeks of double-coding to earn gets reported as “inter-rater reliability was good (κ = 0.78)” and the paper stalls, because that sentence answers none of the questions a careful reader asks: which kappa, over how many raters and items, computed on what unit, with what confidence interval, in what software. This is a copy-paste template for how to report inter-rater reliability the way the GRRAS guideline asks you to, so the writeup takes ten minutes and passes peer review the first time (Kottner et al., 2011). Everything below is synthetic and reproducible, and the fill-in-the-blanks paragraph is meant to be lifted straight into your manuscript.

What must an inter-rater reliability methods section state?

A complete inter-rater reliability methods section states six things, and a reader should be able to find each one without inference: the coefficient and the reason you chose it, how many raters and how many subjects or units, the raters’ training and qualifications, whether they rated independently, the unit each coefficient is computed over, and the point estimate with a confidence interval and the software that produced it. Miss any one and a reviewer cannot judge the number, so they ask, and the paper goes back around.

These six are not a house style. They are the reporting elements the GRRAS guideline was written to standardize, distilled from its 15-item checklist for reliability and agreement studies (Kottner et al., 2011). Gisev and colleagues (2013) frame the same discipline as a prerequisite: you cannot report agreement well until you have separated inter-rater agreement from inter-rater reliability and picked the statistic that matches your data. The template in the next section is those six elements arranged into one paragraph you can fill in.

The copy-paste inter-rater reliability methods template

Here is the fill-in-the-blanks paragraph. Replace every bracketed slot, delete the options that do not apply, and you have a methods statement that covers all six required elements in the order a reader expects them.

Inter-rater reliability was assessed on [N] [units: interview transcripts /
images / utterances / items] independently rated by [K] [rater description:
e.g. trained clinical psychologists], who were blinded to [each other's ratings
/ clinical information]. Because the ratings were [nominal / ordinal /
continuous], agreement was quantified with [coefficient: e.g. quadratic-weighted
Cohen's kappa / ordinal Krippendorff's alpha / a two-way random-effects,
absolute-agreement ICC], chosen to [rationale: e.g. penalize larger severity
disagreements more than adjacent ones]. The unit of analysis was the [per-item
rating / total score / text span]. [Coefficient] was [value] (95% CI
[lower]–[upper]), computed in [software and version: e.g. R 4.4.1, package irr
0.84.1]. By the [Landis & Koch (1977) / Koo & Li (2016)] benchmarks, this
indicates [interpretation band].

Filled in with synthetic numbers, the same paragraph reads as a finished methods statement:

That is a complete intercoder reliability writeup in five sentences. Swap the coefficient and its justification and the same frame reports Krippendorff’s alpha, an ICC, or Fleiss’ kappa without changing shape.

The GRRAS reporting checklist, distilled for a methods paragraph

GRRAS is the canonical checklist for reporting reliability and agreement studies: 15 items covering the title, introduction, methods, results, and discussion (Kottner et al., 2011). Most of the items that decide whether a methods paragraph is complete sit in the methods and results sections. The table maps each reporting element to the concrete thing you write and the GRRAS item it satisfies, so you can check your draft line by line.

Report thisConcretelyGRRAS item
Which coefficient, and why"quadratic-weighted κ, for ordinal ratings"10 (statistical analysis)
Number of raters"two raters"6 and 11 (raters/subjects)
Number of subjects / units"30 transcripts"6 and 11 (raters/subjects)
Rater qualifications and training"clinical psychologists, 3+ years"12 (sample characteristics)
Independence and blinding"rated independently, blinded"8 and 9 (rating process)
Unit of analysis"per-item rating, not total score"6 and 11 (subjects/objects)
Point estimate + confidence interval"κ = 0.82 (95% CI 0.74–0.89)"13 (statistical uncertainty)
Software and version"R 4.4.1, package irr 0.84.1"10 (statistical analysis)
Interpretation benchmark"Landis & Koch (1977) bands"13 and 14 (results, relevance)
The reporting elements a methods paragraph must state, mapped to the GRRAS 15-item checklist (Kottner et al., 2011). Item numbers follow the checklist in that paper.

Print this, hold it next to your draft, and every box should map to a phrase in your methods and results. Empty boxes are exactly the gaps a reviewer will flag.

Which coefficient, and why does the justification matter?

Naming the coefficient is half the sentence; justifying it by data type is the half people skip. The coefficient has to match the measurement level of your ratings, and a reader needs to see that you knew it. Report weighted kappa or ordinal Krippendorff’s alpha for ordered severity ratings, a plain kappa only for unordered categories, and an intraclass correlation for continuous scores such as a total scale value (Gisev et al., 2013). The reason a nominal kappa is wrong for a graded scale is that it treats a one-step slip and a four-step miss as equally bad, which understates real agreement.

The justification also has to acknowledge the coefficient’s known failure mode. Cohen’s kappa is prevalence-sensitive: when one category dominates, agreement can be high while kappa collapses toward zero, a limitation McHugh (2012) lays out and one reason she argues for conservative interpretation. If a label’s base rate is far from balanced, say so and report Gwet’s AC1 alongside kappa, because AC1 does not inflate its chance term under skew (Gwet, 2008). Our decision guide to which IRR coefficient to use walks the full data-type fork, and the two-rater mechanics live in how to compute and interpret Cohen’s kappa.

For content analysis and many-rater designs, the reporting detail that trips people is the distance function. Krippendorff’s alpha takes any number of raters, any measurement level, and missing data, but its value depends on whether you used a nominal, ordinal, or interval distance, so the metric name alone is not enough (Krippendorff, 2004). A krippendorff alpha reporting template has to include that function by name, worked through in Krippendorff’s alpha, computed by hand and in code. Hallgren’s (2012) tutorial on computing inter-rater reliability for observational data is the practical companion for producing any of these numbers before you write them up.

How many raters and subjects, and who were they?

State the actual counts, not the plan. GRRAS asks for both the number you designed for and the number you ended up with: how many raters, how many subjects or objects, and how many replicate observations were conducted (Kottner et al., 2011). “Two raters double-coded 30 transcripts” is complete; “a subset was double-coded” is not, because a reader cannot tell whether the estimate rests on 30 items or 300.

The raters themselves are a reporting element, not background color. GRRAS item 12 asks for the sample characteristics of raters, including their training and experience, because a coefficient from two expert clinicians and one from two first-week research assistants carry different weight even when the number is identical (Kottner et al., 2011). One line does it: the raters’ role, their relevant experience, and whether they trained on the codebook before the reliability sample. If a pilot round calibrated the coders first, cite it, since a pilot annotation round is where most of the reliability is actually built.

Why the unit of analysis is the detail reviewers catch

The unit of analysis is the single most-omitted element and the one that most changes the number, so state it explicitly. The same raters can look near-perfect on total scores and merely moderate on individual items, because errors that cancel in a sum are exposed one item at a time. Reporting “reliability was 0.85” without saying over what leaves that ambiguity unresolved, and reviewers know to ask.

Name the unit in plain terms: a per-item rating, a total scale score, a whole transcript, an utterance, or a text span. Content-analysis reliability is defined on the units being coded, so the unit is part of the coefficient’s definition rather than an afterthought (Krippendorff, 2004). Clinician-rated ordinal instruments make the stakes concrete. The Montgomery-Åsberg Depression Rating Scale (MADRS) and Andreasen’s Scale for the Assessment of Thought, Language and Communication (TLC) both produce per-item ordinal ratings, and whether you report agreement per item or on the summed score is a decision the reader must be told, not left to guess.

Report a confidence interval, not a bare point estimate

Every coefficient needs an interval around it, because a single value from a small reliability sample is a noisy estimate. GRRAS is explicit that authors should report estimates of reliability and agreement including measures of statistical uncertainty, which in practice means a 95% confidence interval (Kottner et al., 2011). A kappa of 0.70 with a 95% CI of 0.62 to 0.78 is a claim; a kappa of 0.70 from twelve items with a CI spanning 0.35 to 0.92 is barely more than a guess, and the interval is what tells those two apart.

High agreement does not excuse you from the interval, and it used to be the awkward case, because the standard variance formula misbehaves near the ceiling. Gwet (2008) derived variance estimators that stay well-behaved under high agreement, so a confidence interval is computable even when raters almost always agree. Modern packages report it by default, which is why the software line and the CI belong in the same sentence.

Which interpretation benchmark, and how firmly?

Cite the benchmark you are using instead of leaving “good agreement” undefined, and cite it as a convention rather than a verdict. The Landis and Koch (1977) bands are the most quoted for kappa, and their cutoffs are also, by the authors’ own admission, arbitrary divisions of a continuous scale, so a kappa of 0.61 is not meaningfully better than 0.59 despite crossing a label boundary. Report the number and the band, and let the confidence interval, not the adjective, carry the weight of the claim.

For an intraclass correlation, the benchmark and the form both need stating. Koo and Li (2016) give the standard ICC bands and, more importantly, insist you specify which of the several ICC forms you computed, since the same data yields different values under different models. Their reading is worth quoting so a reader knows your scale:

“Values less than 0.5 are indicative of poor reliability, values between 0.5 and 0.75 indicate moderate reliability, values between 0.75 and 0.9 indicate good reliability, and values greater than 0.90 indicate excellent reliability.”

Common reviewer objections, and the sentence that answers each

Most reliability-methods revisions ask for the same handful of missing facts. Pre-empt them and the writeup clears review the first time.

  • “Which kappa or ICC?” Name the exact variant and its justification: unweighted versus weighted, and for an ICC the model, type, and definition (Koo & Li, 2016).
  • “Over what unit?” State per-item, total-score, or span before the reader has to ask (Krippendorff, 2004).
  • “How many items and raters?” Give the realized counts, not just the design (Kottner et al., 2011).
  • “Where is the confidence interval?” Report the 95% CI next to every point estimate (Kottner et al., 2011; Gwet, 2008).
  • “Is that agreement high because the label is rare?” Report base rates and add Gwet’s AC1 when prevalence is skewed (Gwet, 2008).
  • “What software?” Name the package and version so the number is reproducible (Kottner et al., 2011).

How to report inter-rater reliability that survives peer review

The practical upshot: treat the methods paragraph as a checklist, not a sentence. Name the coefficient and why, the raters and their training, the subjects, the unit of analysis, and the point estimate with its confidence interval and software, and you have satisfied the reporting elements GRRAS was built to standardize (Kottner et al., 2011). Knowing how to report inter-rater reliability is mostly knowing what a careful reader will ask and answering it before they do.

In Tagaroo, double-coding runs as a live annotation workflow, and the reliability figures come out matched to each scale’s measurement level, with the rater count, unit, and interval attached, so the numbers you paste into the template are the numbers the tool already computed. Fill in the blanks, cite your benchmark, and report the interval: that is the difference between a coefficient and a defensible reliability claim.

References

  • Kottner, J., Audigé, L., Brorson, S., Donner, A., Gajewski, B. J., Hróbjartsson, A., Roberts, C., Shoukri, M., & Streiner, D. L. (2011). Guidelines for Reporting Reliability and Agreement Studies (GRRAS) were proposed. Journal of Clinical Epidemiology, 64(1), 96–106. doi.org/10.1016/j.jclinepi.2010.03.002
  • Gisev, N., Bell, J. S., & Chen, T. F. (2013). Interrater agreement and interrater reliability: key concepts, approaches, and applications. Research in Social and Administrative Pharmacy, 9(3), 330–338. doi.org/10.1016/j.sapharm.2012.04.004
  • McHugh, M. L. (2012). Interrater reliability: the kappa statistic. Biochemia Medica, 22(3), 276–282. doi.org/10.11613/BM.2012.031
  • Hallgren, K. A. (2012). Computing inter-rater reliability for observational data: an overview and tutorial. Tutorials in Quantitative Methods for Psychology, 8(1), 23–34. doi.org/10.20982/tqmp.08.1.p023
  • Koo, T. K., & Li, M. Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15(2), 155–163. doi.org/10.1016/j.jcm.2016.02.012
  • Gwet, K. L. (2008). Computing inter-rater reliability and its variance in the presence of high agreement. British Journal of Mathematical and Statistical Psychology, 61(1), 29–48. doi.org/10.1348/000711006X126600
  • Krippendorff, K. (2004). Reliability in content analysis: some common misconceptions and recommendations. Human Communication Research, 30(3), 411–433. doi.org/10.1111/j.1468-2958.2004.tb00738.x
  • Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46. doi.org/10.1177/001316446002000104
  • Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174. doi.org/10.2307/2529310

Frequently asked questions

What must an inter-rater reliability methods section include?
A complete methods paragraph states six things: which coefficient you used and why, the number of raters and the number of subjects or units, the raters' qualifications and training, whether ratings were independent, the unit of analysis, and the point estimate with a confidence interval and the software used. These map directly onto the GRRAS reporting checklist, which asks for the number of raters and subjects, the rating process, the statistical analysis, and estimates of reliability with measures of statistical uncertainty (Kottner et al., 2011).
Do I have to report a confidence interval for kappa or an ICC?
Yes. GRRAS explicitly asks authors to report estimates of reliability and agreement including measures of statistical uncertainty, which in practice means a 95% confidence interval, not a bare point estimate (Kottner et al., 2011). A single kappa or ICC value can be unstable on a small reliability sample, and the interval tells a reader how much to trust it. Gwet (2008) gives the variance formulas that make these intervals computable even when agreement is high.
What is the 'unit of analysis' in a reliability study?
The unit of analysis is the thing each reliability coefficient is computed over: a per-item rating, a total scale score, a transcript, an utterance, or a text span. It is the omission reviewers catch most often because the same raters can look highly reliable on total scores and only moderately reliable per item. Content-analysis reliability is defined on units, so naming the unit is part of reporting the coefficient at all (Krippendorff, 2004).
How do I report Krippendorff's alpha versus Cohen's kappa?
The template is the same; only the coefficient and its justification change. For two raters and complete data, name Cohen's kappa or weighted kappa (Cohen, 1960). For any number of raters, any measurement level, or missing ratings, name Krippendorff's alpha and state the distance function you used (nominal, ordinal, or interval), because alpha's value depends on it (Krippendorff, 2004). In both cases report the value, its 95% confidence interval, the unit of analysis, and the software.
Which interpretation benchmark should I cite for my coefficient?
State the benchmark you used rather than leaving 'good agreement' undefined. The Landis and Koch (1977) bands are the most cited for kappa, but the authors called the cutoffs arbitrary, so treat them as a convention, not a law. For an intraclass correlation, cite the Koo and Li (2016) bands, which read values below 0.5 as poor, 0.5 to 0.75 as moderate, 0.75 to 0.9 as good, and above 0.9 as excellent. Always pair the band with the confidence interval so the reader sees the uncertainty behind the label.

Put this into practice

Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.