tagaroo

inter rater reliability

Which ICC to Use: Choosing Model, Type, and Definition

Which ICC to use for continuous ratings: choose the model (1, 2, or 3), single vs average measures, and absolute agreement vs consistency, from one table.

Enrique Gutiérrez15 min readUpdated July 2026
One column of rating nodes read by two slightly offset rulers and a fan of dial gauges whose needles point to different values, illustrating how different ICC forms give different reliability numbers from the same data.

The intraclass correlation is really a family of coefficients, and “which ICC to use” is a question with ten defensible answers on the exact same spreadsheet of ratings. On a small synthetic table where one rater sits about four points high on a depression total, the single-measures forms range from 0.86 to 0.99 depending only on which ICC you name—not on the data, the raters, or the subjects. That spread is not a rounding artifact—it is the difference between a coefficient that punishes a systematic rater bias and one that ignores it, and reporting the wrong one either flatters raters who disagree or fails raters who agree. This guide walks the three decisions that pin down a single ICC—the model, the type, and the definition—and gives you the Shrout & Fleiss and McGraw & Wong notation to name it unambiguously.

Which ICC to use: the three decisions

Which ICC to use is decided by three properties of your study design, answered in order: the model, the type, and the definition. Get all three and you have named exactly one coefficient; skip any one and “the ICC” stays ambiguous, because the forms give materially different numbers on identical ratings (Shrout & Fleiss, 1979). None of the three is about the data values themselves—each is about how the ratings were collected and how the resulting score will be used.

Shrout and Fleiss (1979) laid out six forms from these choices, and McGraw and Wong (1996) extended the scheme to ten by separating the definition from the model more cleanly. That is why you will meet two notations for the same coefficient, and why a paper that writes only “ICC = 0.82” has told you almost nothing. The rest of this guide takes the three decisions one at a time, then maps them to both notations.

The ICC selection table

The table below maps a study design to its ICC in both the Shrout & Fleiss and McGraw & Wong notations. Find the row that matches how your raters were assigned, pick the single or average column by how the score will be used, and read off the form. This is a selection starting point, not a verdict—confirm the assumptions against the sources before you report (Koo & Li, 2016).

Your designModelShrout-Fleiss formMcGraw-Wong formDefault definition
Each subject rated by a different, randomly chosen set of ratersOne-way randomICC(1,1) / ICC(1,k)ICC(1) / ICC(k)Absolute agreement (built in)
Same raters rate every subject; raters are a random sample you want to generalize fromTwo-way randomICC(2,1) / ICC(2,k)ICC(A,1) / ICC(A,k)Absolute agreement
Same raters rate every subject; these raters are the only ones of interestTwo-way mixedICC(3,1) / ICC(3,k)ICC(C,1) / ICC(C,k)Consistency
Which ICC to use, by study design. The single form (·,1) applies when one rater scores each case; the average form (·,k) when the operational score is the mean of k raters. Notations from Shrout & Fleiss (1979) and McGraw & Wong (1996).

The one nuance the table compresses: in McGraw and Wong’s (1996) scheme, absolute agreement and consistency are a free choice you can apply to either two-way model, whereas Shrout and Fleiss (1979) coupled the two-way random form to absolute agreement and the two-way mixed form to consistency. The default column reflects that coupling, which is the common case; the definition section below shows when to override it.

Decision 1: which ICC model?

The model asks how raters were assigned to subjects, and it splits into three cases (Koo & Li, 2016). In a one-way random-effects design, each subject is rated by a different set of raters drawn at random, so no single rater appears across all subjects—common in large multi-site studies where you cannot hold the rater panel fixed. Because raters are not crossed with subjects, their systematic differences cannot be separated from noise, so the one-way ICC folds rater bias into the error term.

In a two-way design, the same panel of raters scores every subject, which lets you estimate a separate between-rater term. The two-way case then forks on one question: are these raters a random sample from a larger pool you want to generalize to, or the only raters you will ever use? Two-way random effects treats them as a sample and licenses a claim about raters in general; two-way mixed effects treats them as fixed, so the reliability estimate applies only to this specific panel and cannot be generalized to other raters (Koo & Li, 2016). The point estimates for the random and mixed forms coincide; what differs is the inference you are entitled to draw and the confidence interval.

Decision 2: single or average measures?

The type asks whether the score you will trust in practice is one rater’s rating or the average of several, and you report the ICC that matches that decision (McGraw & Wong, 1996). Single measures, written ICC(_,1), estimates the reliability of a lone rater—the right choice when a single clinician or coder will score each future case. Average measures, written ICC(_,k), estimates the reliability of the mean of k raters, which is what you should report only when the operational score is genuinely that mean.

Average measures is always the higher number, because averaging cancels part of each rater’s idiosyncratic noise—the same logic as a longer test being more reliable than a short one. That makes it tempting to quote, and easy to misuse: an ICC(_,k) of 0.92 from a three-rater consensus says nothing about how reliably a single rater will score case 401 alone. If one person will do the real rating, report single measures even when the average form looks better (Koo & Li, 2016).

Decision 3: absolute agreement or consistency?

The definition asks what counts as a disagreement, and it is where most ICC choices quietly go wrong. Absolute agreement asks whether raters assign the same score; consistency asks only whether their scores move together in an additive way, so a rater who is uniformly three points high but perfectly rank-preserving is treated as fully reliable (McGraw & Wong, 1996). The two definitions differ by a single term in the denominator—the between-rater variance.

For a single-measures, two-way analysis, the absolute-agreement form keeps a between-rater term that consistency drops:

ICC(2,1) = (BMS − EMS) / (BMS + (k − 1)·EMS + (k/n)·(JMS − EMS))

ICC(3,1) = (BMS − EMS) / (BMS + (k − 1)·EMS)

Here BMS is the between-subjects mean square, JMS the between-raters mean square, EMS the residual mean square, k the number of raters, and n the number of subjects. The extra (k/n)·(JMS − EMS) term in ICC(2,1) is exactly the systematic rater bias; consistency deletes it, which is why ICC(3,1) is never smaller than ICC(2,1) on the same table. For most inter-rater reliability studies, Koo and Li (2016) recommend absolute agreement paired with the two-way random model, because you almost always care whether raters land on the same value, not merely whether they rank cases the same way.

Two notations for the same coefficient

Shrout and Fleiss (1979) name a form by two numbers, ICC(model, raters): the first is the model (1, 2, or 3) and the second is 1 for single measures or k for average measures. McGraw and Wong (1996) instead encode the definition in a letter—A for absolute agreement, C for consistency—so ICC(A,1) and ICC(C,1) make the definition explicit where the older notation left it implied. Both describe the same underlying coefficients; they just foreground different choices.

Model + definitionShrout-Fleiss (single / average)McGraw-Wong (single / average)
One-way randomICC(1,1) / ICC(1,k)ICC(1) / ICC(k)
Two-way random, absolute agreementICC(2,1) / ICC(2,k)ICC(A,1) / ICC(A,k)
Two-way mixed, consistencyICC(3,1) / ICC(3,k)ICC(C,1) / ICC(C,k)
The Shrout & Fleiss (1979) and McGraw & Wong (1996) notations line up form for form. Naming the definition explicitly (A vs C) is why the McGraw-Wong labels are harder to misread.

The practical habit worth forming: whichever notation your software uses, translate the result into words in the write-up. psych::ICC() in R prints all six Shrout & Fleiss forms at once, and irr::icc() takes the choices as arguments, so the number you copy out is only as clear as the sentence you wrap it in.

library(irr)
# single-measures, two-way random, absolute agreement = ICC(2,1)
icc(ratings, model = "twoway", type = "agreement", unit = "single")

library(psych)
ICC(ratings)  # prints ICC1, ICC2, ICC3 and their average-measures forms

How to interpret an ICC value

Interpret an ICC against a stated band and its confidence interval, never as a lone point estimate. Koo and Li (2016) give the rule of thumb most reliability papers now cite, worth quoting in full:

“Values less than 0.5 are indicative of poor reliability, values between 0.5 and 0.75 indicate moderate reliability, values between 0.75 and 0.9 indicate good reliability, and values greater than 0.90 indicate excellent reliability.”

Two cautions come with the bands. They are rules of thumb, and stricter schemes exist, so the band is a summary rather than a pass-or-fail line. And because a small reliability study can produce a wide confidence interval, a point estimate of 0.80 with an interval from 0.55 to 0.91 spans “moderate” to “excellent”—which is why the interval, and an adequately sized sample, matter as much as the estimate itself. For planning that sample, see our guide to how many subjects and raters a reliability study needs.

A worked example: one table, four ICCs

Here is the ambiguity made concrete. Two raters score six subjects on a depression total—think a MADRS sum from 0 to 60—and Rater B is systematically about four points higher than Rater A while ranking the subjects almost identically.

SubjectRater ARater B
11014
21822
32530
43033
51217
62225
Synthetic MADRS-style totals for six subjects and two raters. Rater B runs about four points high but preserves the ordering of subjects.

The two-way ANOVA on this table gives between-subjects BMS = 112.6, between-raters JMS = 48.0, and residual EMS = 0.4 (with k = 2 raters and n = 6 subjects). Feed those into the four single-and-average forms and the same six pairs of numbers produce four different verdicts:

FormWhat it measuresValue
ICC(1,1)One-way random, single0.86
ICC(2,1)Two-way random, absolute agreement, single0.87
ICC(3,1)Two-way mixed, consistency, single0.99
ICC(2,k)Two-way random, absolute agreement, average0.93
Four ICCs computed from the same six-subject table (BMS = 112.6, JMS = 48.0, EMS = 0.4). Values are rounded; reproduce them from the formulas above.

Read the spread. The consistency form ICC(3,1) = 0.99 declares near-perfect reliability because it deletes the between-rater term and never sees Rater B’s four-point offset. The absolute-agreement form ICC(2,1) = 0.87 keeps that term and marks the same raters down for the systematic gap.

Averaging the two raters then lifts reliability from 0.87 to 0.93, since the mean cancels part of each rater’s noise. Nothing in the data changed—only which ICC you chose to name.

The lesson generalizes past this toy table. If you are validating whether two clinicians produce interchangeable HAM-D or MADRS totals, consistency is the wrong definition—a four-point systematic bias between raters is a real reliability problem, and only absolute agreement counts it (Koo & Li, 2016).

When the ICC is the wrong tool entirely

An ICC is for continuous or interval ratings—a total scale score, a reaction time, a count—so before choosing a form, confirm the data are actually continuous (Shrout & Fleiss, 1979). For unordered categories, an ICC is the wrong family: two coders labeling utterances present or absent need Cohen’s kappa, not an intraclass correlation. For ordinal severity ratings and for many-rater or missing-data designs, a distance-weighted coefficient such as Krippendorff’s alpha is the honest measure. The broader map of which coefficient fits which data type is in our guide to choosing an inter-rater reliability coefficient.

One more boundary worth naming: the ICC assumes the rating is a single number per subject. Span annotations, bounding boxes, and pixel masks are overlap problems, not continuous-score problems, and forcing an ICC onto them gives a misleading value—those tasks belong to F1, IoU, or Dice instead.

Common mistakes when choosing an ICC

Most ICC errors come from copying a default out of software rather than matching the form to the design. A short list catches the majority.

  • Reporting “ICC = 0.82” with no form. The value is uninterpretable without the model, type, and definition; state all three (Koo & Li, 2016).
  • Quoting average measures when one rater will score. ICC(_,k) describes the mean of k raters, not the lone rater you will actually use (McGraw & Wong, 1996).
  • Using consistency for inter-rater reliability. Consistency hides systematic rater bias; use absolute agreement unless a fixed offset is genuinely irrelevant (Koo & Li, 2016).
  • Picking two-way mixed but generalizing. A mixed model applies only to the raters in the study; you cannot claim it holds for other raters (Koo & Li, 2016).
  • Reporting the point estimate without a confidence interval. A small study’s interval can span two reliability bands, so the interval carries as much information as the estimate.

Which ICC to use, in one walk

Deciding which ICC to use is a three-step walk, not a coin flip. Name the model from how raters were assigned—one-way random when the panel changes across subjects, two-way random when a fixed random-sample panel rates everyone, two-way mixed when those raters are the only ones of interest. Pick the type from how the score is used—single measures for a lone rater, average measures for the mean of k. Set the definition from what a disagreement means—absolute agreement almost always for inter-rater work, consistency only when a systematic offset truly does not matter. Then write the form in full, in both words and notation, so the number can be read.

In Tagaroo, each rating scale carries its measurement level, and reliability is computed with the coefficient that fits it—an ICC with the model and definition made explicit for continuous totals like the MADRS, a distance-weighted statistic for ordinal items. Pick the ICC for the design, not the design for the ICC you already know: that is the difference between a number and a defensible reliability claim.

References

  • Shrout, P. E., & Fleiss, J. L. (1979). Intraclass correlations: uses in assessing rater reliability. Psychological Bulletin, 86(2), 420–428. doi.org/10.1037/0033-2909.86.2.420
  • McGraw, K. O., & Wong, S. P. (1996). Forming inferences about some intraclass correlation coefficients. Psychological Methods, 1(1), 30–46. doi.org/10.1037/1082-989X.1.1.30
  • Koo, T. K., & Li, M. Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15(2), 155–163. doi.org/10.1016/j.jcm.2016.02.012
  • Montgomery, S. A., & Åsberg, M. (1979). A new depression scale designed to be sensitive to change. British Journal of Psychiatry, 134(4), 382–389. doi.org/10.1192/bjp.134.4.382
  • Hamilton, M. (1960). A rating scale for depression. Journal of Neurology, Neurosurgery & Psychiatry, 23(1), 56–62. doi.org/10.1136/jnnp.23.1.56

Frequently asked questions

Which ICC should I use?
Answer three questions about your design, in order. First, the model: are the same raters used for every subject (two-way) or not (one-way random), and if two-way, are those raters a random sample you want to generalize from (random) or the only raters of interest (mixed)? Second, the type: will the reported score be one rater's rating (single measures) or the mean of k raters (average measures)? Third, the definition: does systematic rater bias count against you (absolute agreement) or only rank-order and proportional differences (consistency)? Shrout and Fleiss (1979) defined six forms from these choices; Koo and Li (2016) turn them into a selection guideline.
What is the difference between ICC(2,1) and ICC(3,1)?
Both are single-measures, two-way ICCs, and on the same data ICC(2,1) is usually the smaller of the two. ICC(2,1) is a two-way random-effects, absolute-agreement coefficient: it treats a systematic difference between raters as a source of disagreement and counts it against reliability. ICC(3,1) is a two-way mixed-effects, consistency coefficient: it removes the between-rater term, so a rater who is systematically several points high can still yield a near-perfect value as long as the ratings track each other (Shrout & Fleiss, 1979; McGraw & Wong, 1996).
Should I report single-measures or average-measures ICC?
Report the one that matches how the score is actually used. If a single rater will score each case in practice, report single measures, ICC(_,1). If the operational score is the mean of k raters, report average measures, ICC(_,k), which is higher because averaging cancels some rater noise (McGraw & Wong, 1996). The average-measures form describes the reliability of the mean, not of any one rater, so it is the wrong number to quote when a lone clinician will rate future cases (Koo & Li, 2016).
How do I interpret an ICC value?
Koo and Li (2016) give a widely used rule of thumb: values below 0.5 indicate poor reliability, 0.5 to 0.75 moderate, 0.75 to 0.9 good, and above 0.9 excellent. These are rules of thumb, not hard cutoffs, and other authors use stricter bands, so report the point estimate with its 95% confidence interval and the exact ICC form rather than leaning on a single threshold (Koo & Li, 2016).
Absolute agreement or consistency for inter-rater reliability?
For most inter-rater reliability studies, use absolute agreement, because you usually care whether different raters assign the same score, not merely whether their scores rise and fall together. Koo and Li (2016) recommend a two-way random-effects model with absolute agreement for typical inter-rater work where the results should generalize to other raters with similar training. Consistency is appropriate only when a fixed additive offset between raters is genuinely irrelevant to how the score will be used (McGraw & Wong, 1996).

Put this into practice

Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.