Reliability & agreement

Intercoder reliability from NVivo, ATLAS.ti and MAXQDA

NVivo reports kappa per code per file, and the usual fix—averaging those numbers in Excel—counts every file nobody coded as perfect agreement. Drop in your export and get agreement pooled the right way, the codes that need another round, and the passages to discuss.

Free · No sign-up · Runs entirely in your browser

Pooled agreement for every code, from the file your software already exports

Drop an NVivo coding comparison export or a REFI-QDA project from ATLAS.ti, MAXQDA or NVivo. You get kappa, AC1 and alpha pooled correctly across files, the codes that need another round, and a methods paragraph. Nothing is uploaded; the file is read in your browser.

Drop a .csv, .txt, .xlsx or .qdpx file here, or

The format is detected from the file. It never leaves this tab.

Mean pooled kappa across codes

0.67

Substantial · 5 codes, 6 documents

Range across codes

0.30 to 0.91

Report the range—one number hides the weak codes

Whole-codebook κ / α

0.79 / 0.79

All codes' tables summed · 98.7% raw agreement

Codes below κ = 0.60

1 of 5

Refine these definitions before the next round

Averaging the per-file kappas would have reported 0.80 instead of 0.67 — an inflation of 0.13.

The spreadsheet average scores every document where neither coder used a code as κ = 1 (that is how NVivo reports it) and weights a short memo the same as a long interview.

Agreement by code

Pooled κ Gwet's AC1 Mean of per-file κ (what a spreadsheet gives you)Bands: Landis & Koch (1977)
−0.200.20.40.60.81.0
Show the numbers as a table
CodeAgreementκAC1αPABAKMean file κ
Barriers to access98.2%0.850.980.850.960.85
Stigma98.2%0.620.980.620.960.76
Family support99.2%0.910.990.910.980.92
Coping › Humour99.7%0.301.000.300.990.77
Recommendations for services98.3%0.660.980.660.970.70

Codes marked paradox have high raw agreement but low kappa because they are rare: when a code covers a small share of the text, almost all the agreement is on "not coded", and kappa's chance correction consumes it. Report AC1 or PABAK beside kappa for these codes, and read the adjudication list rather than the coefficient.

Code × document

Kappa for each code in each document. Blank cells are documents where neither coder used the code — they carry no information about agreement, which is exactly why they must not be averaged in as κ = 1.

Interview 01Interview 02Interview 03Interview 04Interview 05Interview 06

Hover or tab to a cell for its agreement table.

Methods paragraph

Two coders (Coder A and Coder B) independently coded 6 documents with a codebook of 5 codes. Intercoder agreement was computed at the character level from NVivo's coding comparison query. For each code, the 2 × 2 agreement tables of all documents were summed, weighting each document by its size, and Cohen's kappa, Gwet's AC1 and Krippendorff's alpha were computed on the pooled table; per-document kappas were not averaged, because documents in which neither coder applied a code score a trivial kappa of 1. Pooled kappa ranged from 0.30 to 0.91 across codes (mean 0.67); across the whole codebook, κ = 0.79, AC1 = 0.99, α = 0.79, and raw agreement was 98.7%. 1 code fell below κ = 0.60 (Coping › Humour, κ = 0.30) and was refined and re-coded after discussion [edit to describe what you did]. For Coping › Humour, agreement was high but kappa low because the code was rare; AC1 is reported alongside kappa for this code (Gwet, 2008). Coefficients were computed with the Tagaroo intercoder reliability calculator (tagaroo.ai/materials/nvivo-intercoder-reliability).

Bracketed text is yours to finish. The sentence about re-coding only appears when a code fell below the threshold.

Reference · for the curious

Intercoder reliability in NVivo, ATLAS.ti and MAXQDA: what the software reports and what to publish

Intercoder reliability is the degree to which two people applying the same codebook to the same data make the same decisions. In qualitative research it is usually reported as Cohen's kappa per code, and in practice it is computed inside the analysis software—NVivo's coding comparison query, MAXQDA's intercoder agreement function, ATLAS.ti's intercoder agreement mode. Each of those tools answers a narrower question than a paper needs, and the gap is where most reporting errors come from.

What NVivo gives you, and the step it leaves to you

NVivo's coding comparison query returns one row per code per file: the file's size, a kappa, and the percentage of the file both coders coded, neither coded, and only one coded. Lumivero's own documentation is explicit that it does not calculate values for a single code across all the files, nor an overall value for the codebook, and it asks you to export the results and decide how to weight the files yourself.

The shortcut almost everyone takes is to average the kappa column. It fails for a reason that is easy to miss: when neither coder applied a code anywhere in a file, NVivo reports that row as κ = 1, because the two coders agree perfectly that none of the file belongs to the code. A code used in two interviews out of six therefore gets four automatic 1.0s in its average. In the example loaded above, the rarely used code Coping › Humour averages to 0.77 across files but has a pooled kappa of 0.30. The same shortcut treats a 2,000-character memo as equal to a 40,000-character transcript.

The correct calculation rebuilds each file's 2 × 2 agreement table from the percentage columns and the file size, adds the tables together, and computes kappa once on the total. That weights every file by how much codable content it contains, and it lets files without the code contribute only what they actually are—agreement on absence—without being scored as perfect agreement in their own right.

ATLAS.ti, MAXQDA and the REFI-QDA project file

ATLAS.ti and MAXQDA compute agreement differently from NVivo and from each other. MAXQDA's intercoder function reports segment-level percent agreement with a configurable overlap threshold and a Brennan–Prediger kappa; ATLAS.ti reports Krippendorff's alpha family. Both can be defended, but numbers produced by different software, on different units, with different chance corrections, cannot be compared or pooled.

All three packages—and QualCoder, QDA Miner and Transana—export the REFI-QDA project exchange format (.qdpx), a zip archive containing the documents' text and every coded passage with its exact character positions and the user who coded it. That is enough to recompute agreement from scratch on any unit you choose, which is what this tool does with a project file: characters (NVivo's unit), sentences, or whole documents. To use it, merge both coders' projects into one, export it as a REFI-QDA project, and pick the two coders from the list the tool shows you.

Choosing the unit of analysis

The unit decides what a disagreement is. At the character level, two coders who agree that a passage is about stigma but start their highlight a clause apart disagree on every character in that clause; kappa then measures boundary precision as much as interpretation. At the sentence level, a sentence counts as coded if any of it is, so boundary noise inside a sentence disappears—usually the fairest unit for interview transcripts. At the document level, the question is only whether each coder found the code anywhere in the interview, which suits a codebook of themes rather than of passages. None of these is the right answer in general. State the unit in your methods, because a kappa on characters and a kappa on sentences from the same coding can differ by more than 0.1.

One consequence of fine-grained units deserves a warning. When most of the text is coded by neither coder, raw agreement and Gwet's AC1 both sit near 1.0 for every code, because agreeing that text is uncoded is easy. At the character level, kappa and Krippendorff's alpha are the informative numbers; AC1 is useful beside them for rare codes, where kappa's chance correction is unstable, but it should not be the headline. Our explainer on the kappa paradox and Gwet's AC1 works through why.

What to report

A defensible intercoder reliability statement names: the number of coders and how they were trained; how many documents were double-coded and how they were chosen; the unit of analysis; the coefficient and how it was pooled; the value for each code or at least the range across codes; which codes fell short and what was done about them. The paragraph the tool writes covers the computational half of that list and leaves bracketed slots for the rest. For the complete checklist of reporting elements, the IRR reporting generator audits a draft against the GRRAS guideline, and if you have not double-coded yet, the reliability sample-size planner estimates how much you need to.

Report the range, not only a mean. A codebook with a mean kappa of 0.72 can contain a code at 0.35 that drives half the findings, and reviewers increasingly ask for per-code values. O'Connor and Joffe's (2020) guidelines on intercoder reliability in qualitative research are a good citation for why and when to report it at all; reflexive thematic analysis, for example, generally does not use it.

Using the disagreements

The coefficient tells you whether the codebook is working; the disagreements tell you what to change. With a project file, the tool lists every passage one coder applied a code to and the other did not, longest first, and a code selected in the chart filters the list. That is the agenda for a consensus meeting: most disagreements cluster around one or two definitions that need an inclusion or exclusion rule, and a single round of clarification usually moves a code from moderate to substantial. Our guide to adjudication and consensus methods compares the ways of resolving them, and why disagreement is signal, not noise makes the case for keeping a record of what was changed.

How the numbers are computed

For each code, the tool sums the four cells—coded by both, by coder A only, by coder B only, by neither—across documents, in the chosen unit. Cohen's kappa is (Po − Pe) / (1 − Pe) on that pooled table, with chance agreement from each coder's own coding rate, which is the formula NVivo documents. Gwet's AC1 uses the mean coding rate for chance agreement. Krippendorff's alpha is the nominal two-coder form, 1 − (2N − 1)(b + c) / (n0n1). PABAK is 2Po − 1. The codebook-level figures sum every code's table. All of them are the same statistics the inter-rater reliability calculator computes, cross-checked against it in the test suite.

What this tool does not do

It compares two coders at a time; for three or more, choose pairs or use Fleiss' kappa or Krippendorff's alpha on a unit-by-coder table in the general calculator. It reads text coding only—PDF regions, image areas and audio or video segments in a project file are skipped, and the tool says how many. It gives point estimates without confidence intervals, because characters and sentences within a document are not independent observations and a standard error computed as if they were would be falsely narrow. And it cannot tell you whether a disagreement is an error or a genuine ambiguity in the data; that judgement belongs to the meeting.

Frequently asked questions

How do I calculate intercoder reliability in NVivo?

Merge both coders' work into one project, then run a coding comparison query (Explore → Coding Comparison) with User group A set to the first coder and User group B to the second, and tick both Display Kappa Coefficient and Display percentage agreement. NVivo returns one row per code per file. It does not give a kappa for a code across all files, or for the whole codebook—its own help pages tell you to export the results and compute those yourself. Export the table (right-click → Export) and drop it into this tool, which pools the per-file results for you.

Why can't I just average NVivo's kappa values?

Because NVivo reports kappa = 1 for every file in which neither coder applied the code—they 'agree' that none of the file belongs to it. Averaging therefore rewards the files that say nothing about the code, and a rarely used code can go from κ = 0.30 to an apparent 0.77. Averaging also weights a short memo the same as a long interview. The defensible method is to rebuild each file's agreement table from the percentage columns and the file size, sum the tables, and compute kappa once on the total, which is what this calculator does and what NVivo itself recommends when it says to weight files by the amount of codable content.

What is an acceptable kappa for qualitative coding?

Conventions, not rules. Landis and Koch's (1977) bands call 0.61–0.80 substantial and above 0.80 almost perfect, and many qualitative methods texts treat roughly 0.60–0.70 as the minimum for a codebook ready to apply at scale, with 0.80 expected when codes feed a quantitative outcome. More useful than a single threshold: report the range across codes, name the codes that fell short, and say what you changed. O'Connor and Joffe (2020) give a practical account of when intercoder reliability is and is not appropriate in qualitative work.

Can I calculate intercoder agreement from ATLAS.ti or MAXQDA?

Yes. Both export their projects in the REFI-QDA exchange format (.qdpx), which stores every coded passage with its character offsets and the user who coded it. Export the project after merging both coders' work, drop the .qdpx here, and choose the two coders. ATLAS.ti and MAXQDA each have their own agreement features, but they report different statistics (MAXQDA's segment-level percent agreement and Brennan–Prediger kappa, ATLAS.ti's Krippendorff's alpha family), so computing all coefficients from the same file is the cleanest way to compare or report them.

Should agreement be measured by character, sentence or document?

It depends on what a coding decision is in your study. Characters are NVivo's unit and the most sensitive to boundary differences—two coders who agree a passage is about stigma but start the highlight one clause apart will disagree on those characters. Sentences are usually the fairest unit for interview transcripts, because coders rarely disagree about where a sentence is. Documents answer a coarser question: did both coders find the theme in this interview at all? Whichever you pick, state it in the methods; kappa computed on different units is not comparable.

Is my data uploaded anywhere?

No. The export is read with JavaScript in your own browser—zip archives, spreadsheets and XML included—and never sent to a server. You can disconnect from the internet after the page loads and the tool keeps working, which is worth doing if your project contains identifiable interview data.

Written by Enrique Gutiérrez, PhD (Computer Science)—founder of Tagaroo and Associate Professor of Computer Science, working on inter-rater reliability, measurement and annotation methodology (ORCID).

Last verified: 24 September 2026. Formulas, thresholds and cited figures on this page were checked against their original sources on that date. Every calculation runs in your browser; nothing you enter is transmitted or stored.

Code the transcripts and measure agreement in the same place

Exporting, merging and recomputing is the workaround for software that treats reliability as an afterthought. Tagaroo keeps every coder's decisions on the same transcript, computes agreement as you go, and can act as a second coder itself—so the adjudication list is ready before the consensus meeting, not after an afternoon in Excel.

Try Tagaroo free