Reliability & agreement

Cronbach's alpha calculator

Paste a respondents × items matrix and get alpha with its exact confidence interval, the item analysis that tells you which item to fix, and a results paragraph you can paste into a paper.

Free · No sign-up · Runs entirely in your browser

Alpha, its interval, and which item to fix

Paste one row per respondent and one column per item. You get alpha with an exact confidence interval, the item analysis that says what to do about it, and a check for the reverse-keying error that quietly ruins most pasted matrices. Everything runs in your browser.

Your item scores

Cronbach's alpha

0.927

12 respondents × 5 items

95% confidence interval

0.832 to 0.976

Exact F method (Feldt et al., 1987)

Adequate for

Individual decisions

Nunnally's staged thresholds, not a single cutoff

Mean inter-item r

0.71

Range 0.36 to 0.91 — target .15 to .50

High enough for the applied settings Nunnally reserved .90 for, where a score drives a decision about one person.

Standardized alpha

0.926

From the correlation matrix

Standard error of measurement

1.54

Scale points, from SD 5.71

Item analysis

This is the part that tells you what to do. A low corrected item-total correlation means the item is not measuring what the rest measure; an alpha-if-deleted above your current alpha means the scale would improve without it.

ItemMeanSDItem-total rα if deletedReverse
12.081.380.970.877
22.001.410.830.906
31.921.240.860.901
42.001.210.730.925
51.921.240.670.935higher

Reversing uses the endpoints observed in your data (0 to 4), so a score of x becomes 4 − x. If your scale has anchors nobody in this sample used, reverse-score before pasting instead.

What a longer or shorter scale would give

The Spearman-Brown prophecy formula predicts alpha at a different scale length. It assumes the items you add are as good as the ones you have, so treat lengthening as an optimistic bound — and shortening as a pessimistic one, since the items you drop are usually your weakest.

Alpha at 5 items

0.927

Your scale as it stands

Items needed for α = .80

5

Nunnally's basic-research level

Items needed for α = .90

5

Required where scores drive individual decisions

With fewer than about 30 respondents the interval is exact but very wide. A point estimate on its own from a sample this size cannot distinguish an adequate scale from an unusable one, which is a fact about the study rather than the calculation.

The mean inter-item correlation is 0.71, above the .15 to .50 range Clark and Watson (1995) suggest for a coherent but non-redundant scale. Combined with a high alpha, that usually means several items are paraphrases of each other, so the scale covers a narrower slice of the construct than its length implies.

Copies alpha with its interval, the mean inter-item correlation, the full item table and any warnings — the elements a reviewer asks for.

Reference · for the curious

What alpha measures, and the four things it is routinely asked to prove

Cronbach's alpha is the most-reported statistic in questionnaire research and the most routinely over-interpreted. It answers one narrow question — how consistently a set of items rank the same people — and it is regularly asked to prove three things it cannot: that the scale measures one construct, that a cutoff of .70 has been cleared, and that the instrument is therefore fit for its purpose.

What the number actually is

Alpha is a function of two things and nothing else: the average correlation between items, and how many items there are. That second dependency is the one that catches people out. Adding items raises alpha even when the additions are mediocre, so a twenty-item scale with weak inter-item correlations can outscore a five-item scale whose items cohere tightly. Cortina (1993) worked through exactly this, which is why the mean inter-item correlation is reported here beside alpha rather than tucked away. Clark and Watson's suggested range of roughly .15 to .50 is the sanity check: below it the items have little in common, and above it they are paraphrases.

The .70 rule is a misquote

Nunnally never proposed a universal threshold. He set staged ones: about .70 suffices in the early stages of research, more is expected in established basic research, and where a score drives a decision about an individual person .90 is the minimum with .95 the desirable standard. Lance, Butts and Michels (2006) traced how that graded advice collapsed into a single number cited everywhere, and their paper is worth reading for the other three cutoffs it dismantles as well. This calculator bands alpha by intended use for that reason. The top of the range is not the best place to be either: past about .90, Streiner (2003) reads a high alpha as redundancy, meaning you are sampling a narrower slice of the construct than the item count implies.

Cronbach's alpha vs the ICC

These are the same statistic wearing different clothes. Alpha is algebraically ICC(3,k) — the two-way mixed-effects consistency coefficient for average measures — computed with items standing where raters usually stand. The calculator reports both routes to the number as a cross-check, and the confidence interval is identical in either framing.

What differs is the question. Alpha treats items as interchangeable measures of one construct within a single sitting; an inter-rater ICC treats raters that way across a single instrument. A scale can have excellent internal consistency and poor inter-rater reliability, which is a common and specific finding: the items hang together, but two clinicians reading the same interview score them differently. If that is your question, the ICC calculator computes all six forms with the same exact intervals, and our guide to choosing an ICC form works through which one your design calls for. For categorical codes rather than ratings, the inter-rater reliability calculator covers the kappa and alpha family — note that Krippendorff's alpha there is a different coefficient that happens to share the name.

Why the interval matters more than the estimate

Alpha has an exact confidence interval, derived by Feldt, Woodruff and Salih (1987), and almost no calculator reports it. On the twelve-respondent example loaded by default, alpha is .93 but the interval runs from .83 to .98 — the difference between "adequate for group research" and "usable for individual decisions". Validation studies are often run on samples that small, and a point estimate from one of them cannot distinguish those two conclusions. Our companion piece on confidence intervals for agreement covers the same argument for the coefficients in the kappa family, and the sample-size planner works the problem backwards from a target precision.

The defect this tool is built to catch

A negative or near-zero corrected item-total correlation is nearly always a reverse-keyed item that was not reverse-scored before the data went into the spreadsheet. It is the single most common defect in a pasted item matrix, it depresses alpha dramatically, and it is invisible if you only look at the headline number. Load the second example to see it: alpha reads .62, item 5 correlates −.67 with the rest of the scale, and alpha-if-deleted for that one item is .94. Flipping it with the reverse toggle recovers a .93 scale. Nothing about the instrument was wrong; one column was entered the wrong way round.

This matters beyond questionnaires. The same reversal problem appears in coding schemes where one category is defined as the absence of a behaviour, which is why our note on auditable scale scoring argues for keeping the scoring direction with the item definition rather than in the analyst's head. Instruments in our library carry their scoring direction explicitly — see the MADRS for a worked example of item-level anchors.

What this tool does not do

It does not compute McDonald's omega. Alpha equals reliability only when the items are essentially tau-equivalent — each contributing equally to the true score — and when they are not, alpha is a lower bound. Dunn, Baguley and Brunsden (2014) make the case for omega as the practical default, but omega needs a fitted factor model, which is not something a browser tool can do honestly. Approximating it badly would be worse than not offering it.

It also does not test dimensionality, and no value of alpha can. Sijtsma (2009) is the standard reference for what alpha does and does not establish, and it is blunter than most methodological papers. If you need to know whether your items measure one thing, that is a factor-analysis question, and alpha will not answer it however high it goes.

Frequently asked questions

What is a good Cronbach's alpha?

It depends on what the score will be used for, and the familiar .70 rule is a misreading of the source. Nunnally set staged thresholds: about .70 is enough for early-stage research, more is expected for established basic research, and .90 is the minimum where a score drives a decision about an individual person, with .95 the desirable standard there. Lance, Butts and Michels traced how that graded advice collapsed into one universal number. Above roughly .95, alpha usually signals redundancy rather than quality — items that are near-paraphrases of each other measure a narrower slice of the construct than their number suggests.

Does a high alpha mean my scale is unidimensional?

No, and this is the most consequential misunderstanding about the statistic. Alpha is a function of the average inter-item correlation and the number of items, so a long scale measuring two distinct things can post a comfortable alpha while a short coherent one does not. Cortina showed this directly in 1993. Unidimensionality is a question for factor analysis, not for alpha. The mean inter-item correlation, reported here alongside alpha, is the more informative number: Clark and Watson suggest roughly .15 to .50 for a scale that is coherent without being redundant.

Why is my Cronbach's alpha negative?

Almost always because an item is reverse-keyed and was not reverse-scored before the data went in. A negative alpha means the average covariance between items is negative, which is not a substantive finding about a scale that was designed to measure one thing. Check the corrected item-total correlations first: any item correlating negatively with the sum of the others is the suspect. This calculator flags those items and lets you reverse them and recompute. If no item is reverse-keyed, look for a column that does not belong to the scale.

What is the difference between Cronbach's alpha and the ICC?

Less than their separate names suggest: Cronbach's alpha is algebraically identical to ICC(3,k), the two-way mixed-effects consistency ICC for average measures, computed with items in the place of raters. The difference is what varies across the columns. Alpha treats a set of items as interchangeable measures of one construct within a single administration, while an inter-rater ICC treats a set of raters that way. This calculator reports both routes to the same number as a cross-check, and the confidence interval is the same exact F interval in either framing.

Should I use McDonald's omega instead?

Where you can fit a factor model, often yes. Alpha equals reliability only when items are essentially tau-equivalent, meaning each contributes equally to the true score; when they do not, alpha is a lower bound and can understate reliability. Omega relaxes that assumption by using the factor loadings, which is why Dunn, Baguley and Brunsden recommend it as a practical default. Omega needs a fitted factor model, which is beyond what a browser tool can honestly do, so this calculator computes alpha and says plainly what alpha does and does not establish rather than approximating omega badly.

Written by Enrique Gutiérrez, PhD (Computer Science) — founder of Tagaroo and Associate Professor of Computer Science, working on inter-rater reliability, measurement and annotation methodology (ORCID).

Last verified: 1 August 2026. Formulas, thresholds and cited figures on this page were checked against their original sources on that date. Every calculation runs in your browser; nothing you enter is transmitted or stored.

Reliability as a by-product of the work

Internal consistency is one of three reliability questions a coding project has to answer, and the other two — do two coders agree, and does the same coder agree with themselves — need the codings themselves. Tagaroo holds the transcripts, the scale definitions and the ratings in one place, so the coefficients come out attached to the study that produced them.

Try Tagaroo free