tagaroo

annotation quality

Annotator Disagreement Is Signal, Not Noise to Discard

Forcing one gold label can throw away real information. See why annotator disagreement is signal, not noise, and when to reconcile versus preserve it.

Enrique Gutiérrez12 min readUpdated July 2026
Abstract tick-marks falling onto a horizontal measurement axis, forming one tight green consensus cluster, a separate teal cluster at a boundary line, and one lone coral outlier, showing disagreement as a distribution rather than a single point.

Most annotation projects treat disagreement as a defect: a number to drive down before the real work starts. That instinct is backwards. On subjective and ambiguous tasks, annotator disagreement is often the most useful signal in the dataset, because it tells you where the construct is genuinely hard, where the guideline is thin, and where reasonable experts legitimately see different things.

The research turned decisively against the single-truth assumption over the past decade. Aroyo and Welty argued that annotation rests on “an antiquated ideal of a single correct truth,” and proposed instead that “measuring annotations on the same objects of interpretation […] across a crowd will provide a useful representation of their subjectivity and the range of reasonable interpretations” (Aroyo & Welty, 2015). This piece is about how to act on that: how to read annotator disagreement as data, sort it by cause, and decide when to reconcile it versus when to keep it.

Why is annotator disagreement signal, not noise?

Annotator disagreement is signal, not noise, because on many tasks it is systematic and interpretable rather than random. When two qualified coders land on different labels, the difference usually reflects something real about the item—its position on a category boundary, an ambiguity in the language, or a legitimate difference in how each coder reads the construct. Averaging that away treats a measurement of difficulty as if it were static on the line.

Aroyo and Welty framed the problem as a set of “seven myths of human annotation,” the foundational one being that each item has exactly one correct answer that a diligent annotator will find (Aroyo & Welty, 2015). Once you drop that assumption, disagreement stops being an embarrassment to hide in an agreement statistic and becomes a second channel of information sitting right next to the labels.

Barbara Plank gave the phenomenon its now-standard name: human label variation, the observation that equally competent annotators systematically vary in their labels for reasons that are not all mistakes (Plank, 2022). The reframing matters because it changes the goal. You are no longer chasing the fiction of zero disagreement; you are trying to understand the disagreement you have.

What are the four types of annotator disagreement?

There are four broad causes of annotator disagreement, and they demand opposite responses, so sorting disagreement by cause is the whole game. Sandri and colleagues built a taxonomy for exactly this, labeling disagreements in a subjective annotation task by whether they stemmed from sloppiness, missing context, linguistic ambiguity, or genuine subjectivity (Sandri, Leonardelli, Tonelli & Jezek, 2023). The version below adapts their categories—renaming and regrouping them around the decision each one forces, and treating a codebook gap as its own under-specification case—rather than reproducing their labels verbatim.

TypeWhat it isWhat it signalsReconcile or preserve?
ErrorA coder is simply wrong under the scheme's own rules—a slip, a misread, an attention lapseA quality problem in the labeling, not the itemReconcile: correct the label
Under-specificationThe guideline doesn't cover this case, so coders fill the gap in different but defensible waysA gap in the codebook, not in the codersReconcile: fix the guideline, then re-label
AmbiguityThe item genuinely sits on a boundary between two categoriesReal item difficultyPreserve: record it as a boundary case
PerspectiveCoders read the same item through different but legitimate frames or lived experienceRater viewpoint and legitimate subjectivityPreserve: keep the distribution of labels
Four types of annotator disagreement, what each one signals, and whether it calls for reconciliation or preservation. The dividing line: error and under-specification have a right answer once fixed; ambiguity and perspective do not.

The split down the middle of that table is the useful part. Error and under-specification are quality problems—there is a correct label, and you have not reached it yet. Ambiguity and perspective are information—the spread of labels is telling you something true about the item or the construct, and forcing it to a single value would erase the finding. The same disagreement number can belong to either half, which is why a coefficient alone can never tell you what to do.

How do you tell ambiguity apart from error?

You tell ambiguity apart from error by testing whether the disagreement is reproducible. Random mistakes wash out when you re-collect labels; genuine ambiguity produces a stable pattern of disagreement that reappears with fresh annotators. If the same item keeps splitting coders the same way, the split is a property of the item, not a run of bad luck.

Pavlick and Kwiatkowski showed this directly for natural language inference. Re-collecting judgments on the same items produced consistent, reproducible distributions of disagreement, which led them to argue that models should be evaluated against “the full distribution of plausible human judgments” rather than a single collapsed label (Pavlick & Kwiatkowski, 2019). Reproducibility is the tell: a disagreement you can regenerate on demand is signal.

This is also where an agreement coefficient earns its keep, and where it stops. A low Krippendorff’s alpha flags that annotators diverge, but it cannot say whether the cause is sloppy work or a genuinely hard construct. The coefficient measures the amount of disagreement; adjudication measures the kind. Treat a low number as a prompt to open the items, not as a verdict on the coders.

What does forcing a single ‘gold’ label throw away?

Forcing a single gold label throws away the shape of the disagreement, which on subjective tasks is often the most informative part of the annotation. Majority vote keeps the winning label and discards how close the vote was, so a 5–0 item and a 3–2 item collapse to the same value even though they carry very different certainty. The confidence information is gone before modeling begins.

Standard inter-rater statistics inherit the same assumption. Cohen’s kappa and its relatives are built around the premise that each item has one true category and departures from it are error—a framing that quietly rules out the possibility that the spread is real. That is fine for objective coding and misleading for constructs where competent people legitimately differ.

Keeping the distribution instead of collapsing it is not just tidier bookkeeping. It can produce better models. Peterson and colleagues built CIFAR-10H by collecting full human label distributions over an image test set. Training on those soft labels improved robustness compared with training on single hard labels (Peterson, Battleday, Griffiths & Russakovsky, 2019).

The disagreement, preserved rather than voted away, did useful work.

The broader research program has a name for this move. Uma and colleagues survey a wide body of methods that reject the assumption that a single “gold” interpretation exists for each item, from soft-label training to disagreement-aware evaluation (Uma, Fornaciari, Hovy, Paun, Plank & Poesio, 2021). The common thread is refusing to throw the distribution away before you have used it.

A worked example: disagreement as data in a coded set

Here is a synthetic illustration—invented spans, round numbers, no real transcript data—to show the four disagreement types in one coded set. Suppose three coders are tagging speech for formal thought disorder using the TLC scale for thought, language and communication, which gives each item an operational definition with verbatim examples (Andreasen, 1986). The coders agree on most spans and split on a handful. The split spans are where the information lives.

Classify each disputed span by cause, and the right action falls out of the classification:

Synthetic span (coder labels) Why they split Type Action
“I fixed the car, cars, my brother never calls.” (Derailment / Tangentiality) Real boundary between drifting mid-thought and answering obliquely Ambiguity Preserve as a boundary case
“The meeting was fine.” coded 3 vs 0 for a symptom One coder mis-clicked the severity field Error Reconcile: correct it
“It’s all connected, you know how it is.” (Derailment / no code) Guideline never says how much vagueness counts Under-specification Fix the guideline, then re-label
A darkly humorous aside (present / absent) Coders read tone through different frames Perspective Preserve the distribution

Notice that a single agreement number would have flattened all four rows into one verdict—“the coders disagreed 4 times”—and hidden that two of those disagreements are fixable and two are findings. The same logic separates genuine label errors from legitimate ambiguity in any coded dataset: disagreement points you to the item, and a human decides which half of the table it belongs in.

The boundary cases are worth keeping visibly, not smoothing over. A run of Derailment-versus-Tangentiality splits is a precise map of where your annotation guidelines need a sharper rule or where the construct itself resists a clean cut. Either way, you learned it from the disagreement, not despite it.

When should you reconcile versus preserve disagreement?

You should reconcile disagreement when it comes from error or under-specification, and preserve it when it comes from ambiguity or perspective. The test is whether a correct label exists once the process is fixed. If clarifying the guideline or catching the slip would settle the item, reconcile it. If competent coders would still legitimately differ after every clarification, the distribution is the answer, and collapsing it manufactures false certainty.

Reconciliation is a real step with a real output. An adjudicator decides the case against a written definition and records the rationale, so the fix is reproducible rather than one reviewer’s memory. That is different from silently overwriting a minority label with the majority, which just hides the disagreement.

Take the MADRS depression scale, a ten-item clinician-rated measure designed to be sensitive to change (Montgomery & Åsberg, 1979). Disagreement clustered at a band boundary usually signals an under-specified anchor, not a bad rater. Reconcile the anchor, and the split narrows.

Preservation is the harder discipline because it means resisting the urge to report one clean number. For genuinely subjective constructs, the honest deliverable is the label distribution plus a note on why it spreads, which downstream models and readers can use as soft labels rather than fake ground truth. This is the same argument that runs through the data-centric AI case for fixing data over tuning models: the quality of what you keep about each item, disagreement included, is load-bearing for everything downstream.

Where Tagaroo fits

Tagaroo is a schema-first annotation workspace built so that disagreement is captured, not averaged away. You define a coding scheme once, collect independent reads, see inter-rater reliability on the same screen, and route disputed items to an adjudicator who records a decision and a rationale—or marks the item as a preserved boundary case. The point is to let a qualified human decide which type of disagreement each split is, instead of letting a majority vote decide silently. An AI first pass can triage where coders are likely to diverge, but it never issues a finished label and never a diagnosis.

On data handling, Tagaroo takes a de-identify-first path rather than positioning itself as a processor of identifiable records: its terms require you to strip direct identifiers before upload, its anonymous trial mode runs in the browser so trial text never leaves your machine, and the stack is EU-hosted and GDPR-oriented. It is not a medical device. If your workflow must process raw identifiable transcripts in the cloud, put that question to any vendor—including this one—before uploading; see the privacy policy for specifics.

The practical upshot

Stop scoring your project on how little the coders disagree, and start reading annotator disagreement for what it tells you. Sort each split by cause: reconcile the errors and the guideline gaps, because those have a right answer, and preserve the ambiguity and the perspective, because those are the finding. A single gold label is convenient, but on subjective work it is a summary that deletes its own evidence.

If you change one thing, change this: keep the distribution and log why each disputed item was reconciled or preserved, so your dataset records difficulty instead of hiding it. Then build your coding scheme in Tagaroo with reliability and adjudication in the loop, and let disagreement do the work it is actually good at.

References

  • Aroyo L, Welty C. Truth Is a Lie: Crowd Truth and the Seven Myths of Human Annotation. AI Magazine. 2015;36(1):15-24. doi:10.1609/aimag.v36i1.2564
  • Plank B. The “Problem” of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation. Proceedings of EMNLP 2022. 2022:10671-10682. doi:10.18653/v1/2022.emnlp-main.731
  • Uma A, Fornaciari T, Hovy D, Paun S, Plank B, Poesio M. Learning from Disagreement: A Survey. Journal of Artificial Intelligence Research. 2021;72:1385-1470. doi:10.1613/jair.1.12752
  • Sandri M, Leonardelli E, Tonelli S, Jezek E. Why Don’t You Do It Right? Analysing Annotators’ Disagreement in Subjective Tasks. Proceedings of EACL 2023. 2023:2420-2433. doi:10.18653/v1/2023.eacl-main.178
  • Pavlick E, Kwiatkowski T. Inherent Disagreements in Human Textual Inferences. Transactions of the Association for Computational Linguistics. 2019;7:677-694. doi:10.1162/tacl_a_00293
  • Peterson JC, Battleday RM, Griffiths TL, Russakovsky O. Human Uncertainty Makes Classification More Robust. Proceedings of ICCV 2019. 2019:9616-9625. arXiv:1908.07086
  • Andreasen NC. The Scale for the Assessment of Thought, Language, and Communication (TLC). Schizophrenia Bulletin. 1986;12(3):473-482. doi:10.1093/schbul/12.3.473
  • Montgomery SA, Åsberg M. A new depression scale designed to be sensitive to change. British Journal of Psychiatry. 1979;134:382-389. doi:10.1192/bjp.134.4.382

Frequently asked questions

Is annotator disagreement always a sign of a bad dataset?
No. Disagreement has several causes, and only some of them mean the data is bad. It can reflect an outright error, a gap in the guidelines, genuine ambiguity in the item, or a legitimate difference in perspective (Sandri, Leonardelli, Tonelli & Jezek, 2023). Only the first two are quality problems you fix; the last two are information about the item and the construct that a single gold label would hide (Aroyo & Welty, 2015).
What is human label variation?
Human label variation is the umbrella term for the fact that equally qualified annotators systematically assign different labels to the same item, for reasons that are not all noise (Plank, 2022). It reframes disagreement as an expected property of many natural-language and perception tasks rather than a defect to be minimized, and it groups the reasons into annotator error, task ambiguity, and genuine subjectivity.
Does forcing a single gold label lose information?
Yes, on subjective or ambiguous items. Collapsing several annotators' judgments into one majority label discards the shape of the disagreement, which on many tasks is stable and reproducible rather than random (Pavlick & Kwiatkowski, 2019). Keeping the full distribution of labels, called a soft label, can improve a model's robustness compared with training on single hard labels (Peterson, Battleday, Griffiths & Russakovsky, 2019).
When should you reconcile disagreement instead of preserving it?
Reconcile when the disagreement traces to an error or an under-specified guideline, because those have a correct answer once the rule is clarified. Preserve it when the disagreement traces to genuine ambiguity or legitimate perspective, where no single label is more correct than the distribution itself (Uma, Fornaciari, Hovy, Paun, Plank & Poesio, 2021). Adjudication is the step that decides which case you are in.
How is disagreement different from a label error?
A label error is a wrong label under the scheme's own rules; disagreement is two annotators landing differently, which may or may not involve an error. Disagreement is one signal that an error might exist, but it fires just as loudly on legitimate boundary cases (Sandri et al., 2023). Telling the two apart is an adjudication judgment, not something an agreement coefficient can decide on its own.

Put this into practice

Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.