Reliability & agreement
Six different statistics share the name ICC, and they can differ by more than 0.1 on the same ratings. This computes all six with exact confidence intervals, and asks the four design questions that decide which one is yours.
Free · No sign-up · Runs entirely in your browser
Answer four questions about your design, paste your ratings, and get all six Shrout & Fleiss forms with exact confidence intervals — with the one that matches your study marked. Everything runs in your browser.
Did the same raters score every subject?
Chooses between a one-way and a two-way model.
Are these raters a sample from a wider population?
Random versus mixed effects. Changes what the result generalises to, not the arithmetic.
Does a constant offset between raters count against them?
The single biggest lever on the number.
Will you report the mean of all raters, or one rating?
Single versus average measures.
Report this form
ICC(2,1)
Also written ICC(A,1) (McGraw & Wong, 1996) — a two-way random-effects model, absolute agreement, single measures (ICC(2,1)).
Fixed raters, and a constant offset between them counts against them — the usual inter-rater case, and Koo and Li's default recommendation for reliability studies.
ICC(2,1) — your form
0.871
6 subjects × 2 raters
95% confidence interval
-0.011 to 0.985
Exact F method
Interpretation
Good
Koo & Li (2016) benchmarks — a convention, not a test
Mean squares
112.6 / 0.4
Between-subjects / residual
This interval spans Poor, Moderate, Good, Excellent — the study cannot yet tell those apart. That is a statement about the sample size, not about the instrument.
The spread between these rows is the reason the form has to be named when you report an ICC. They are all correct; they answer different questions.
| Form | Also called | Model | Definition | Measures | ICC | 95% CI | Band |
|---|---|---|---|---|---|---|---|
| ICC(1,1) | ICC(1) | one-way random | absolute agreement | single | 0.862 | 0.386 to 0.979 | Good |
| ICC(1,k) | ICC(k) | one-way random | absolute agreement | average | 0.926 | 0.557 to 0.989 | Excellent |
| ICC(2,1)yours | ICC(A,1) | two-way random | absolute agreement | single | 0.871 | -0.011 to 0.985 | Good |
| ICC(2,k) | ICC(A,k) | two-way random | absolute agreement | average | 0.931 | -0.023 to 0.992 | Excellent |
| ICC(3,1) | ICC(C,1) | two-way mixed | consistency | single | 0.993 | 0.950 to 0.999 | Excellent |
| ICC(3,k) | ICC(C,k) | two-way mixed | consistency | average | 0.996 | 0.975 to 1.000 | Excellent |
With fewer than about 30 subjects the interval is exact but wide — expect it to straddle two or three interpretation bands, which is a real finding about the study rather than a flaw in the calculation.
Copies a methods sentence naming the model, the definition, the measures and the interval — plus all six forms and any caveats.
Reference · for the curious
Most of the confusion around the intraclass correlation comes from treating it as one statistic. It is a family of six, they routinely disagree by more than a whole interpretation band on identical data, and choosing between them is a question about how your study was run rather than about the numbers. A paper that reports "ICC = 0.87" without naming the form has not quite told you anything.
Model: one-way, two-way random, or two-way mixed. If different subjects were rated by different raters, a rater effect cannot be separated from error and you are limited to the one-way model, ICC(1,1). If one fixed panel rated everybody, you have a two-way design. Whether you call it random or mixed depends on whether these raters stand in for a wider population of raters — and this is the subtle one, because the estimator is the same either way. What changes is how far the result generalises, not the arithmetic.
Definition: absolute agreement or consistency. Does a rater who scores everyone two points high count as disagreeing? If yes, you want absolute agreement, ICC(2,1). If you only care that raters rank subjects the same way, consistency, ICC(3,1), is the looser and higher figure. This is the biggest single lever on the number.
Type: single or average measures. If the score you will actually use is the mean of all your raters, report the ICC(·,k) form, which is higher because averaging cancels noise. If real-world use involves one rater working alone, report single measures — the average-measures figure would flatter a workflow you are not going to run.
The tool loads six subjects rated by two raters, from our guide to choosing an ICC. Rater B scores nearly everyone a few points higher than rater A, but the two rank the subjects almost identically. That is the situation where the forms diverge: absolute agreement returns 0.871, consistency returns 0.993. Both are correct. One says "these raters do not produce interchangeable scores", and the other says "they order patients the same way". Which one you should report is decided by whether the offset matters for what you do next.
Reliability studies are usually small, and the ICC's interval is correspondingly wide. On those same six subjects, the ICC(2,1) point estimate of 0.871 carries a 95% interval running from roughly −0.01 to 0.985 — from worse than useless to excellent. The honest reading is that six subjects cannot establish reliability, whatever the point estimate says.
This tool draws the interval against the Koo and Li bands for that reason: an interval crossing three boundaries is visibly not an answer. If you need the interval to sit inside a single band, that is a sample-size question to settle before collecting ratings, and our reliability sample-size calculator works out how many subjects and raters that takes. The intervals here use the exact F-distribution method from Shrout and Fleiss (1979), with Satterthwaite degrees of freedom for ICC(2,1) after Fleiss and Shrout (1978) — the one form whose interval is not a plain F ratio. Our companion piece on confidence intervals for agreement coefficients covers the equivalent problem for kappa and alpha.
The ICC partitions variance and needs ratings you can meaningfully subtract. Cohen's kappa corrects for chance agreement between categorical labels. For a continuous measurement the ICC is right; for nominal categories kappa is; and reporting either on the other's data is a routine reviewer objection.
Ordinal severity ratings sit in the genuinely arguable middle. Quadratic-weighted kappa and a two-way ICC often land within a few hundredths of each other on the same ordinal ratings, and either is defensible provided you say which you used and why. If your data are categorical, our inter-rater reliability calculator computes the kappa family, Krippendorff's alpha and Gwet's AC1 side by side, and our guide to choosing an agreement coefficient works through the decision. Clinician-rated instruments in the MADRS mould are the classic ICC case: continuous-ish totals, one fixed panel of raters, and an offset that genuinely matters.
Shrout and Fleiss numbered the forms; McGraw and Wong (1996) named them by what they measure, so ICC(2,1) became ICC(A,1) for absolute agreement and ICC(3,1) became ICC(C,1) for consistency. Both notations are in current use, and readers who know one often assume the other refers to something else. Every row of the results table carries both.
It performs a complete-case analysis: a subject missing any rater's score is dropped, with the count reported, because a two-way decomposition has no way to use a partial row. It does not fit a mixed model with random effects for both subjects and raters under missingness, which is what you would want for badly incomplete designs. It does not compute an ICC for nominal data, because that is not what the statistic is for. And the interpretation bands are a convention borrowed from Koo and Li, not a hypothesis test — a study either resolves reliability to a useful precision or it does not, and the interval is what tells you.
It depends on three things about your design, not on the data. First, did the same raters score every subject? If not, you are limited to a one-way model, ICC(1,1). If so, ask whether a constant offset between raters should count against them: if yes you want absolute agreement, ICC(2,1); if consistency is enough, ICC(3,1). Finally, decide whether the score you will actually use is one rating or the mean of all raters — the mean is more reliable than any single rating, which is what the ICC(·,k) forms report. For most inter-rater reliability work the answer is ICC(2,1), two-way random effects, absolute agreement, single measures, which is Koo and Li's (2016) default recommendation.
Whether a systematic difference between raters counts as disagreement. ICC(2,1) measures absolute agreement, so a rater who scores everyone two points higher than the others is penalised. ICC(3,1) measures consistency, so that rater is not penalised as long as they rank the subjects the same way. The gap between them is exactly the size of the rater effect, and it can be large: on the six-subject example this tool loads by default, absolute agreement gives 0.871 while consistency gives 0.993 — the difference between "good" and "excellent" on the same twelve numbers. Choose consistency only if you genuinely do not care about the offset, which usually means you will calibrate or standardise the scores later.
Koo and Li (2016) suggest below 0.5 is poor, 0.5 to 0.75 moderate, 0.75 to 0.90 good, and above 0.90 excellent. Those bands are a convention rather than a test, and the more useful habit is to read the confidence interval instead of the point estimate. A study of six subjects can return an ICC of 0.87 with an interval running from below zero to 0.99, which does not establish anything — it is consistent with the instrument being unusable and with it being excellent. If you need to place reliability in a single band, that is a sample-size requirement, and it is worth planning before you collect the ratings.
Because the differences between your subjects are smaller than the disagreement between your raters. The ICC is a ratio of between-subject variance to total variance, and the ANOVA estimator can go below zero when the between-subject term is swamped by noise. It is a real result, not a computational error, and it usually means one of two things: your subjects are genuinely too similar to distinguish, or your raters are not applying the same rules. This tool reports the negative value rather than clamping it to zero, because clamping would hide the finding.
No, and they are not interchangeable. Kappa is for categorical labels and corrects for the agreement two raters would reach by chance given their marginals. The ICC is for continuous or interval ratings and partitions variance instead. For ordinal severity ratings the choice is genuinely arguable — quadratic-weighted kappa and a two-way ICC often land close to each other — but for a continuous measurement the ICC is the right tool, and for nominal categories kappa is. Reporting an ICC on nominal labels, or a kappa on a continuous measure, is a common reviewer objection.
Written by Enrique Gutiérrez, PhD (Computer Science) — founder of Tagaroo and Associate Professor of Computer Science, working on inter-rater reliability, measurement and annotation methodology (ORCID).
Last verified: 30 July 2026. Formulas, thresholds and cited figures on this page were checked against their original sources on that date. Every calculation runs in your browser; nothing you enter is transmitted or stored.
An ICC is only as good as the ratings underneath it. Tagaroo collects those ratings against a curated instrument, keeps each one linked to the words that justify it, and computes agreement as your coders work — so the reliability figure comes out of the process rather than a separate spreadsheet exercise.
Intraclass correlation (ICC) calculator · tagaroo.ai/materials/icc-calculator · figures verified against primary sources 2026-07-30. Educational scoring aid, not a diagnosis, and not reviewed by a licensed clinician.