Reliability & agreement
Choose the coefficient you plan to report, the value you expect, and how precise the estimate has to be — and see how many items or subjects the reliability substudy needs, plus what you gain by adding raters instead.
Free · No sign-up · Runs entirely in your browser
Plan for the value you expect to see, not the one you hope for.
Share of items where the phenomenon is present. Skew inflates N sharply.
How much uncertainty you can defend. ±0.10 is a common target; ±0.20 rarely separates anything.
Items needed
196
each coded by two raters
Interval you could report
0.60 to 0.80
κ = 0.70, achieved half-width ±0.100
What that distinguishes
Substantial
Landis & Koch, 1977
With 196 items coded by two raters you could report κ = 0.70 with a 95% interval of 0.60 to 0.80. That interval falls entirely inside “substantial”, so the study can place reliability in a single band and defend it.
At 10 items the interval is ±0.44; at 196 it reaches your ±0.10 target; pushing on to 588 only narrows it to ±0.06. Half-width falls as one over the square root of N, so precision gets steadily more expensive.
What this assumes. Binary rating with the positive category expected in 50% of subjects, and both raters using it at that rate. Planned around an expected κ of 0.70: if the true value is lower, the interval will be wider than ±0.10. Large-sample (normal) interval on the kappa scale, from the Bloch & Kraemer (1989) variance. Precision, not power, is the right frame for a reliability study: reviewers ask how precisely you estimated agreement, not whether you rejected a null nobody proposed. Method: Bloch & Kraemer (1989) variance, inverted for a 95% CI half-width (Donner & Eliasziw 1992).
Reference · for the curious
"How many items do I need for my inter-rater reliability check?" has a short answer that is wrong — thirty — and a correct answer that takes a moment to set up: as many as it takes to estimate the coefficient precisely enough to support the claim you want to make. This planner implements the second answer.
Thirty subjects, or ten per category, are the numbers that circulate. The problem is not that they are arbitrary; it is that they routinely produce confidence intervals so wide the study cannot distinguish moderate agreement from excellent agreement. A kappa of 0.70 estimated from thirty items carries an interval roughly a quarter of the scale wide — compatible both with a codebook that needs rewriting and with one ready to publish. The study cost real coder time and settled nothing.
The honest figures are larger than most people expect. Pinning a kappa near 0.70 to ±0.10 takes on the order of 200 items at balanced prevalence; because required N scales with the inverse square of the half-width, tightening to ±0.05 quadruples it. Seeing that number before you commit is the point of planning.
Reporting the point estimate alone hides this, which is why so many reliability sections survive review. A single number with no interval looks equally authoritative at any sample size.
Power calculations test a null hypothesis. For reliability studies the available null — that agreement exceeds some floor — is rarely the question anyone cares about. What readers and reviewers want to know is how precisely you measured agreement.
So the planning quantity is the confidence-interval half-width: decide how much uncertainty you can defend, and solve for the number of items. A half-width of ±0.10 around an expected kappa of 0.70 is a defensible target for a codebook you intend to publish; ±0.05 is a strong claim and expensive; ±0.20 is honest for exploratory work but will not distinguish adjacent interpretation bands.
The precision of kappa depends on the expected value and on how the categories are distributed. Balanced categories are cheap; skewed ones are expensive. If the phenomenon you are coding appears in five per cent of items, you need substantially more items to pin kappa down, because most of your sample contributes almost no information about agreement on the rare category.
This is the same underlying issue as the kappa prevalence paradox that makes the coefficient collapse on rare phenomena, and it has a practical implication worth planning around: oversample the rare category in your reliability subset if you can. A reliability substudy does not have to be a random sample of your corpus, as long as you say what it was.
For continuous or ordinal-as-continuous ratings — severity scores rather than categories — the intraclass correlation is the right coefficient, and its precision depends on the expected ICC and on the number of raters per subject. Adding raters does buy precision, but the returns diminish sharply after the third, and raters are usually the scarcer resource.
Which ICC you are planning for matters too: there are several, they answer different questions, and they give different numbers on identical data. Our guide to choosing an ICC covers model, type and definition; decide that before sizing anything, because the answer changes the target.
For precision, subjects. For generalisability, raters. These are different goals and it is worth being explicit about which one you are buying.
A coefficient estimated from two coders describes those two coders. If you intend to claim that your codebook is reliable in general — that a new trained coder would reach similar conclusions — two raters cannot support that claim at any sample size, because rater is not a sampled factor in your design. Three or four raters on fewer items often supports a stronger claim than two raters on many. Our discussion of how many annotators per item works through the redundancy and cost curve.
Every sample-size figure here rests on assumptions you should state in your methods: the expected value of the coefficient, the expected category prevalence, the number of raters, and a normal approximation to the coefficient's sampling distribution. That approximation is reasonable at moderate sample sizes and gets shaky at small ones, so treat the output for very small N as an order-of-magnitude guide rather than a precise requirement.
The expected value is the assumption people find uncomfortable, because you are being asked to guess the answer before measuring it. Use a pilot. A short pilot annotation round on twenty items gives you a rough kappa to plan from, and it will surface codebook problems while they are still cheap to fix — which is usually worth more than the sample-size estimate itself.
It often is, and there are legitimate responses. Widen the target interval and say so. Collapse categories that coders cannot reliably distinguish — a three-category scheme with substantial agreement is more useful than a seven-category scheme with fair agreement. Oversample the rare category. Report percent agreement and a paradox-resistant coefficient such as Gwet's AC1 alongside kappa. What is not legitimate is running an underpowered check, reporting a bare point estimate, and letting the reader assume precision you did not have.
State the coefficient and variant, the number of items and raters, how the reliability subset was selected, the point estimate, the interval, and the software. Our IRR reporting template gives the full structure, and sample size for inter-rater reliability covers the underlying methodology in more depth.
Once you have coded the subset, compute the coefficients with the inter-rater reliability calculator, which reports intervals alongside every estimate and writes the methods sentence for you.
This planner runs entirely in your browser. Nothing you enter is sent anywhere and there is no account.
Estimating a kappa near 0.7 to a 95% confidence interval of ±0.10 needs around 200 subjects at balanced prevalence with two raters. That is far more than the widely repeated rules of thumb — 30 subjects, or 10 per category — which routinely produce intervals so wide the study cannot distinguish moderate from excellent agreement. The right number depends on the precision you need rather than any fixed rule: decide how wide an interval around your expected coefficient you can defend, then solve for N. Relaxing the target from ±0.10 to ±0.20 cuts the requirement by roughly a factor of four, because required sample size scales with the inverse square of the half-width.
For most reliability designs, adding subjects buys more precision per unit of effort than adding raters, and the returns from extra raters fall away quickly after the third. Raters matter for a different reason: a coefficient estimated from two coders describes those two coders, so more raters make the estimate generalise to the population of raters you will actually deploy. Decide raters on generalisability grounds and subjects on precision grounds.
Kappa applies to categorical ratings and its standard error depends on both the expected kappa and the marginal prevalence of the categories — skewed prevalence inflates the required N sharply. ICC applies to continuous or ordinal-as-continuous ratings, and its precision depends on the expected ICC and the number of raters per subject. They are not interchangeable: choose the coefficient from your measurement scale first, then plan the sample size for that coefficient.
For reliability work, precision is almost always the right frame. A power calculation tests a null hypothesis — usually that reliability exceeds some floor — which is rarely the scientific question. What reviewers want to know is how precisely you estimated agreement, so plan for a confidence interval of a width you can defend and report that interval alongside the point estimate.
Written by Enrique Gutiérrez, PhD (Computer Science) — founder of Tagaroo and Associate Professor of Computer Science, working on inter-rater reliability, measurement and annotation methodology (ORCID).
Last verified: 29 July 2026. Formulas, thresholds and cited figures on this page were checked against their original sources on that date. Every calculation runs in your browser; nothing you enter is transmitted or stored.
Tagaroo computes agreement continuously across your annotation team, so you can see when your estimate has stabilised instead of guessing at the sample size in advance.