Reliability & agreement
Turn a reliability coefficient into a threshold in the units of your scale: the standard error of measurement, the MDC, and a verdict on whether an observed change is larger than measurement error.
Free · No sign-up · Runs entirely in your browser
A reliability coefficient is unitless, so it cannot answer that. The standard error of measurement can, because it is in the points of your scale — and the minimal detectable change turns it into a threshold. Everything runs in your browser.
ICC(2,1)
0.912
95% CI 0.76–0.97
Mean score
36.43
15 subjects × 2 occasions
Mean change between occasions
0.07
A practice or drift effect if it is not near zero
SD of change scores
4.59
√MSE × √2 — the same error, seen from the other side
MDC95 — one person
9.00
1.96 × √2 × SEM, in scale points
Standard error of measurement
3.25
SEM = √EMS
MDC as % of the mean
24.7%
Comparable across scales
Threshold for a group mean
2.32
MDC ÷ √15 — not interchangeable with the individual one
Depends on the SEM definition
A increase of 8.80 points falls inside the range the four SEM definitions disagree over (8.57 to 9.00). Whether it counts as detectable depends on which definition you use, which is exactly why the definition has to be stated.
These are all in the literature, they all get called "the SEM", and they disagree. Pick the row you will report — Weir (2005) argues for the first — and name it in your methods. The row you select drives the threshold above.
| Definition | Formula | SEM | MDC95 | What it treats as error |
|---|---|---|---|---|
| reporting | SEM = √EMS | 3.25 | 9.00 | Weir's recommendation. Error is the residual after both subject differences and any systematic shift between occasions are removed, so a practice effect does not inflate it. |
| SEM = √WMS | 3.14 | 8.69 | The one-way error term. A systematic change between occasions counts as measurement error, which is the right choice when you cannot assume the shift will repeat. | |
| SEM = SD × √(1 − ICC) | 3.09 | 8.57 | The textbook formula with an absolute-agreement ICC. Depends on the spread of your sample, so a more heterogeneous sample gives a smaller SEM for the same instrument. | |
| SEM = SD × √(1 − ICC) | 3.19 | 8.84 | The same formula with a consistency ICC, which ignores a constant offset between occasions and therefore usually reports a slightly different error. |
With fewer than about 30 subjects the interval is exact but wide — expect it to straddle two or three interpretation bands, which is a real finding about the study rather than a flaw in the calculation.
The four SEM definitions give values from 3.09 to 3.25 on this data, so the detectable-change threshold you report moves by 0.43 points depending on a choice most papers leave unstated. Name the definition you used.
Copies the ICC, the SEM with its formula named, the MDC, the group threshold and all four definitions.
Reference · for the curious
An ICC of 0.91 tells you the instrument separates these subjects well. It does not tell you whether your patient's four-point improvement means anything, because it has no units. The standard error of measurement does have units, and the minimal detectable change turns it into the threshold the clinical question actually needs.
MDC = z × √2 × SEM. The multiplier that gets dropped is the √2, and it is there because a change score is a difference between two measurements, each carrying its own error. The variance of a difference of two independent quantities is the sum of their variances, so the error on a change is √2 times the error on a single score. Leaving it out understates the threshold by about 29% of its own value, which turns changes that are indistinguishable from noise into apparent improvement. If you are comparing your figure with a published one and they differ by roughly that much, this is usually why.
This is the part that surprises people, and it is the reason this calculator shows four numbers where others show one. Weir (2005) sets out the competing definitions: the square root of the residual mean square from a two-way ANOVA, the square root of the within-subjects mean square, and the textbook SD × √(1 − ICC) computed with either an absolute-agreement or a consistency ICC. On the fifteen-subject example loaded above they give MDC95 values from 8.57 to 9.00 points. That spread is wide enough to flip the verdict on a real patient, which is why the tool declines to answer when your observed change lands inside it.
Weir's recommendation is the residual mean square, for two reasons. It excludes any systematic shift between occasions from the error term, and unlike SD × √(1 − ICC) it does not depend on how heterogeneous your sample happens to be — recruit a wider range of severity and the ICC rises and the SEM falls, without the instrument having changed. Which ICC form you would use if you took the textbook route is its own decision, and our guide to choosing an ICC form works through it; the ICC calculator computes all six with exact intervals.
These get conflated constantly and they answer different questions. The MDC is a measurement question — how large must a change be to exceed the instrument's noise — and it comes entirely from reliability data. The MCID, the minimal clinically important difference, is a value question: how large must a change be before a patient or clinician would call it worthwhile? An MCID needs an external anchor, usually a global rating of change, and cannot be derived from reliability data however you manipulate it.
The relationship between them is diagnostic. When the MCID is larger than the MDC, the instrument can reliably detect changes at the size that matters, which is the situation you want. When the MDC is larger, the instrument cannot reliably detect a change the size people care about, and no amount of statistical treatment fixes that: you need a more reliable measure, more raters, or a different endpoint. Reporting an MCID without checking it against the MDC is how a trial ends up powered to detect something its instrument cannot see.
de Vet and colleagues (2006) draw the distinction this whole page sits on. Reliability parameters — the ICC, kappa — are ratios that describe how well an instrument distinguishes between subjects, and they depend on the population you measured. Agreement parameters — the SEM, the MDC, limits of agreement — are in the units of the scale and describe measurement error itself. You need the second kind to interpret an individual score, and only the first kind is usually reported.
The same paper is where the group-level threshold comes from: for a mean of n people the threshold is the individual MDC divided by √n, because averaging cancels error. On the example above that is 2.32 points against 9.00 for one person. Both numbers are correct and they are answers to different questions, so the tool reports them side by side and labels which is which. For categorical codes rather than continuous scores the equivalent question is chance-corrected agreement, which the inter-rater reliability calculator handles, and the sample-size planner works backwards from a target precision to the number of subjects a reliability substudy needs.
If everyone scores higher the second time, that is not measurement error — it is learning, practice, recovery or rater drift, and it changes which error term is honest. The residual mean square removes it, so it reports the error that remains once the shift is accounted for. The within-subjects term keeps it in, which is the right choice when you cannot assume the shift will repeat in the same direction for the next patient. This calculator flags a systematic shift when it finds one rather than quietly averaging over it. Structured interviews exist partly to suppress this effect, which is why instruments like the HAM-D have structured guides attached, and why our note on planning a reliability study treats the interval between occasions as a design decision rather than a logistical one.
It does not estimate an MCID, and no reliability calculator can. It does not draw Bland-Altman limits of agreement, though it reports the ingredient they are built from: the mean change between occasions and the SD of the change scores, which are the bias and the spread that plot displays. It computes the SEM from a complete-case analysis, so subjects missing an occasion are dropped and counted rather than imputed. And the MDC it produces describes measurement error only — a change larger than the MDC is real in the sense that the instrument can see it, which is a weaker claim than saying it matters.
MDC = z × √2 × SEM, where the SEM is the standard error of measurement in the units of your scale and z is 1.96 for the 95% level or 1.645 for the 90%. The √2 is there because a change score is the difference between two measurements, each carrying its own error, so the variance of the difference is twice the variance of a single measurement. Omitting it — which some published calculators do — understates the threshold by about 29%. The SEM itself comes either from the square root of the residual mean square in a test-retest ANOVA or from SD × √(1 − ICC).
They answer different questions and are not interchangeable. The MDC is a measurement question: how large must a change be before it exceeds the noise of the instrument? It comes entirely from reliability data. The MCID, the minimal clinically important difference, is a value question: how large must a change be before it matters to a patient or clinician? It requires an anchor such as a patient-reported global rating of change. A change can clear the MDC and still be trivial, and an MCID smaller than the MDC is a warning that the instrument cannot reliably detect changes of the size you care about.
Weir (2005) recommends the square root of the residual mean square from a two-way ANOVA, because it excludes any systematic shift between occasions from the error term and does not depend on how heterogeneous your sample happens to be. The common alternative, SD × √(1 − ICC), inherits both of those problems: a more variable sample produces a smaller SEM for the same instrument, and the answer changes depending on whether you use an absolute-agreement or a consistency ICC. This calculator shows all four so you can see the spread, which on typical data is several percent.
Yes. Minimal detectable change, smallest detectable change and smallest real difference are three names for the same quantity, and you will also see minimal detectable difference and the abbreviation MDC95 when the confidence level is attached. Beckerman and colleagues (2001) use SRD, de Vet and colleagues (2006) use SDC, and the rehabilitation literature mostly uses MDC. If a paper reports one without a formula, check whether the √2 is included before comparing it with your own.
Not the individual MDC, no. The threshold for a group mean of n people is the individual MDC divided by √n, because averaging reduces measurement error. That group threshold is much smaller — on the example loaded here, 2.32 points against an individual threshold of 9.00 for a sample of fifteen. Treating an individual threshold as a group one makes a study look underpowered, and the reverse mistake, using a group threshold to judge one patient, is how a change well inside measurement error gets reported as real improvement.
Written by Enrique Gutiérrez, PhD (Computer Science) — founder of Tagaroo and Associate Professor of Computer Science, working on inter-rater reliability, measurement and annotation methodology (ORCID).
Last verified: 1 August 2026. Formulas, thresholds and cited figures on this page were checked against their original sources on that date. Every calculation runs in your browser; nothing you enter is transmitted or stored.
An MDC needs the same thing a reliability study needs: the same subjects measured twice, with both sets of scores kept and comparable. Tagaroo holds transcripts, scale definitions and every rating in one place, so a second pass is a normal week of coding rather than a separate study, and the numbers come out attached to the data that produced them.
Minimal detectable change calculator · tagaroo.ai/materials/minimal-detectable-change-calculator · figures verified against primary sources 2026-08-01. Educational scoring aid, not a diagnosis, and not reviewed by a licensed clinician.