annotator quality
Who Should Annotate Your Data? A Selection Framework
Who should annotate your data? Match annotator background to the task—expert, trained non-expert, or crowd—using a clear selection framework and matrix.

Who should annotate your data? The instinct is to reach for the most credentialed person available—a licensed clinician, a board-certified specialist, a subject-matter PhD—and assume that more expertise means better labels. For some tasks that instinct is right and non-negotiable. For others it is expensive overkill, and for a surprising number of tasks the credentialed expert is the wrong annotator, because the thing you are trying to measure is a lay reaction the expert no longer has.
The better question is not “who is most qualified” but “what does this specific task demand.” Annotator background is a task-matching decision, not a status ranking. This post gives you a framework: score the task on three axes—subjectivity, domain-knowledge requirement, and stakes—then map that profile to the right annotator background, from an aggregated lay crowd through a trained non-expert to a calibrated credentialed expert. The evidence is more interesting than “hire experts.”
Who should annotate your data? The short answer
Who should annotate your data is decided by the task, not the org chart: score the task on subjectivity, domain-knowledge requirement, and stakes, then pick the lightest annotator background that clears the bar. A lay crowd suffices for objective, learnable, low-stakes labels; a trained non-expert handles moderately subjective coding schemes; a credentialed expert is required only when the construct is specialized and the stakes are high. And for tasks that measure a lay perception, the expert is the wrong choice on purpose.
That framing matters because the default failure mode runs in both directions. Teams over-spend by putting clinicians on tasks a trained undergraduate would code just as reliably, and they under-spend by handing a specialized construct to a generic crowd that cannot see the phenomenon at all. Both are matching errors. The rest of this post is how to avoid them.
Score your task on three axes first
Before you name a person, describe the task on three continuous axes. Each one pushes the required annotator background up or down independently, so treat them as a profile, not a single score.
- Subjectivity. Is there a single defensible answer (objective) or does the label depend on the reader’s interpretation (subjective)? Word-sense disambiguation is near-objective; whether a message is “offensive” or “empathic” is not.
- Domain-knowledge requirement. Can an intelligent layperson recognize the phenomenon after reading a short guideline, or does it take years of training to even see it? Bounding a car is learnable in minutes; spotting derailment in speech is not.
- Stakes. What is the cost of a wrong label downstream—a slightly noisier sentiment model, or a mislabeled tumor boundary that propagates into a clinical tool? Higher stakes raise the bar on both accuracy and accountability.
These axes interact. High subjectivity with low stakes (rating whether a tweet is funny) is very different from high subjectivity with high stakes (rating suicidal ideation from an interview). The next section maps common profiles to a recommended background so you can place your task quickly.
The annotator selection matrix
The matrix below maps a task profile to a best-fit annotator background and the evidence behind it. Read it as a starting point to be confirmed with a pilot, not a verdict—real tasks sit between rows, and the honest move is to test two candidate pools on the same items before committing.
| Task profile | Best-fit annotator background | Why (evidence) |
|---|---|---|
| Objective, low domain knowledge, low stakes (everyday-object boxes, word sense) | Lay crowd, aggregated over several labels | ~4 non-expert labels matched expert quality across NLP tasks (Snow et al., 2008) |
| Learnable rules, moderate subjectivity, moderate stakes (sentiment, entailment, structured coding schemes) | Trained non-experts, screened and calibrated | Screening for aptitude plus training beat process-only tricks at N=68,000 (Mitra et al., 2015) |
| Perspective-laden and subjective (offensiveness, empathy, perceived clarity) | Annotators drawn from the target population; retain annotator-level labels | Annotators are not interchangeable; aggregation buries minority perspectives (Prabhakaran et al., 2021) |
| Language- or culture-specific meaning (dialect, idiom, in-group reference) | Native speakers / cultural in-group members | First language, age, and education shift labels (Al Kuwatly et al., 2020); dialect insensitivity drove racial bias (Sap et al., 2019) |
| Domain-technical, specialized construct, high stakes (thought disorder, depression severity, tumor boundaries) | Credentialed domain expert, calibrated—non-experts only against expert gold | Experts beat non-experts on biomedical segmentation (Gurari et al., 2015); professionals beat crowdworkers (Rädsch et al., 2023) |
When is a domain expert non-negotiable?
A credentialed expert is required when the construct is specialized enough that an untrained reader cannot reliably recognize it, and the stakes are high enough that errors carry real cost. This is the one profile where you should not economize on background, because the failure is not noise—it is systematically missing or inventing the phenomenon.
The clearest evidence is in medical imaging, where expertise and accuracy track together. On a benchmark of 6,148 biomedical segmentations, trained experts produced the most accurate boundaries (median overlap 0.85), ahead of crowdsourced non-experts (0.82) and well ahead of every algorithm tested (0.36) (Gurari et al., 2015). A larger study of 14,040 medical images found that professional annotators consistently outperformed Amazon Mechanical Turk crowdworkers, and that adding worked example images to the instructions helped far more than lengthening the text (Rädsch et al., 2023).
Clinical interview coding is the same story. Rating formal thought disorder on the Scale for the Assessment of Thought, Language and Communication (TLC) (Andreasen, 1986) is not something effort alone unlocks—its categories, like derailment and tangentiality, only produce reliable labels once raters are trained against a common standard. Depression severity on the ten-item, clinician-rated MADRS (Montgomery & Åsberg, 1979) demands the same calibrated reading of anchored item definitions. Forensic statement analysis with Criteria-Based Content Analysis (CBCA) (Steller & Köhnken, 1989) likewise assumes a trained rater applying published criteria, and its own authors caution it is an investigative aid rather than a standalone test.
When is a trained non-expert or crowd enough?
A trained non-expert—or an aggregated lay crowd—is enough when the task is learnable from a good guideline and the stakes tolerate a small amount of residual noise. This covers a large share of real annotation work, and defaulting to experts here wastes money without buying accuracy.
The founding result is still Snow and colleagues. Across five natural-language tasks, non-expert Mechanical Turk labels agreed closely with expert gold standards, and for affect recognition an average of just four non-expert labels per item was enough to emulate a single expert’s quality (Snow et al., 2008). For five of seven tasks, a model trained on one set of non-expert labels matched or beat one trained on a single expert—partly because pooling several non-experts cancels the idiosyncratic bias any one labeler brings.
The lever that makes non-experts reliable is not the crowd’s raw judgment but selection and training. In a controlled experiment spanning 68,000 annotations, person-centric strategies—screening workers for the relevant cognitive aptitude and training them in the coding method—outperformed process-centric tricks and control conditions on subjective coding tasks (Mitra et al., 2015). A screened, trained non-expert can beat an uncalibrated expert, which is why screening and training annotators often returns more than a credential filter does. For the broader trade-off, see expert versus crowd annotation.
When is the right annotator a layperson?
Sometimes the credentialed expert is the wrong annotator by design—specifically when the label you want is a lay or target-population reaction. If you are measuring whether ordinary readers find a message offensive, confusing, or empathic, then the ground truth is the lay perception, and an expert’s trained eye can systematically diverge from it.
Perceived offensiveness is the sharpest example. Annotators insensitive to dialect over-labeled African American English as toxic, so that tweets in AAE were up to twice as likely to be rated offensive; priming annotators to consider the dialect and likely race of the author significantly reduced that bias (Sap et al., 2019). The worker pool there was about 75% self-identified White—a background mismatch that shaped the labels. The fix was not a more senior annotator; it was a more representative one.
This is why annotator identity is a design decision on subjective tasks, not merely a competence question. Annotators draw on their lived experience, so they are not interchangeable, and flattening a non-representative pool into a majority vote quietly discards the perspectives you meant to capture (Prabhakaran et al., 2021). The practical implications: recruit annotators from the population whose judgment you are modeling, retain annotator-level labels rather than only the aggregate, and treat disagreement as signal, not noise when it splits along background lines.
Language and cultural fit is a separate axis
Language and cultural fit is its own selection axis, independent of both expertise and general subjectivity. A task can be objective in principle yet still require an annotator who shares the language, dialect, or cultural context of the material to read it correctly.
The demographic effect is measurable. Training classifiers on labels from demographically distinct annotator groups produced significantly different models, with first language, age, and education each linked to systematic differences in how the same content was labeled (Al Kuwatly et al., 2020). Idiom, sarcasm, code-switching, and in-group references are exactly the features a non-native or out-group reader misses—not for lack of effort, but for lack of the relevant cultural knowledge.
So a multilingual or cross-cultural project should treat language and cultural background as a hard requirement for the relevant items, layered on top of the subjectivity and domain axes. A native speaker with modest training will often out-code a domain expert working in a second language, because the binding constraint here is comprehension, not credentials.
A worked example: same transcripts, three annotator pools
Here is a synthetic illustration—round numbers, no real project or patient data—of how the same annotators fare on two tasks with very different profiles. Imagine 200 synthetic interview utterances labeled by three pools: a lay crowd (five workers, aggregated), a set of screened and trained non-experts (three coders), and, as a reference ceiling, a second trained clinician scored against the first.
Task A is learnable: “does this utterance contain a self-reported anxiety mention?”—a concrete, low-subjectivity judgment. Task B is a specialized construct: “does this utterance show derailment?”, the TLC thought-disorder item that takes training to recognize. Agreement is reported as Cohen’s kappa against a clinician reference.
| Annotator pool | Task A: anxiety mention (learnable) | Task B: derailment (expert construct) |
|---|---|---|
| Lay crowd (5, aggregated) | κ ≈ 0.78 | κ ≈ 0.31 |
| Trained non-experts (3, screened + trained) | κ ≈ 0.82 | κ ≈ 0.54 |
| Second clinician (reference ceiling) | κ ≈ 0.80 | κ ≈ 0.76 |
The pattern is the point, and it matches the literature. On Task A, all three pools land in the substantial-agreement range: paying for a clinician buys nothing a trained non-expert or an aggregated crowd does not already deliver (Snow et al., 2008; Mitra et al., 2015). On Task B, the crowd collapses to near-chance because untrained readers cannot see the phenomenon, trained non-experts improve but still trail, and only clinician-versus-clinician reaches reliable agreement. Same people, opposite conclusions—because the task profiles are opposite.
The hybrid model: scale with non-experts, adjudicate with experts
For many high-value projects the answer is not one annotator background but two, layered. A screened, trained non-expert pool does the first pass at volume, and a domain expert adjudicates disagreements and maintains a gold set—so you buy expert judgment only where it changes the label.
This design directly targets the trade-off the evidence describes. Aggregation and training make non-experts reliable on the learnable majority of items (Snow et al., 2008; Mitra et al., 2015), while expert adjudication protects the specialized or high-stakes minority where a non-expert would systematically err (Gurari et al., 2015). It also keeps a domain expert in the loop precisely where the data-cascade research says it matters most (Sambasivan et al., 2021). Deciding how many labels to gather before adjudication is its own question—see how many annotators per item.
Limitations: what this framework does not decide
This framework tells you which annotator background to start with; it does not settle everything, and treating it as a formula would be a mistake. Four caveats keep it honest.
First, the three axes are continuous and interacting, not switches—most real tasks sit between rows, so a short pilot on the same items with two candidate pools beats any a-priori verdict. Second, background is necessary but not sufficient: the right annotator on vague guidelines still produces poor labels, and clear rules and training often move accuracy more than a credential does, much as pay is a weak lever for accuracy on its own. Third, recruiting for representativeness on perspective-laden tasks carries real ethical and logistical constraints around collecting and handling demographic information—do it transparently and only when responsible. Fourth, once you have chosen well-matched annotators, disagreement among them is often informative rather than error, and flattening it too early discards signal.
Where Tagaroo fits
Tagaroo is built for the hybrid model this evidence points to. It is a schema-first workspace: you define a coding scheme once with anchored definitions and worked examples, an AI agent takes a first pass, and human reviewers—chosen to match the task—correct and adjudicate it, with inter-rater reliability, gold-standard checks, and annotator-level labels tracked in the workflow. That lets you route learnable items to trained non-experts and reserve credentialed experts for the specialized calls, rather than paying for the top of the org chart on every label.
That focus is also its boundary. If your task is a massive volume of objective microtasks with no specialized construct in sight, a bare crowdsourcing marketplace is the cheaper tool. Tagaroo’s strength is the guided middle—clear schemes, a model-assisted first pass, and reliability measured rather than assumed. On data handling, it takes a de-identify-first path: strip direct identifiers before upload, with an anonymous browser-side trial mode so trial data never leaves your machine; see the privacy policy for specifics.
The practical upshot
Stop asking who is most qualified and start asking what the task demands. Score it on subjectivity, domain knowledge, and stakes; take the lightest annotator background that clears the bar; and remember that “lightest” sometimes means a representative layperson rather than a credentialed expert, because on perspective-laden tasks the lay reaction is the ground truth (Sap et al., 2019; Prabhakaran et al., 2021). Reserve expertise for the specialized, high-stakes constructs where an untrained reader cannot see the phenomenon at all (Gurari et al., 2015; Sambasivan et al., 2021).
If you change one thing, change this: run a short pilot with two candidate annotator pools on the same items before you staff the project, and let the agreement data—not the credentials—decide. Then define your scheme and calibrate your raters on the MADRS in Tagaroo so the right background is built into the workflow, not assumed after the fact.
References
- Snow R, O’Connor B, Jurafsky D, Ng AY. Cheap and Fast—But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks. Proceedings of the 2008 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2008:254–263. aclanthology.org/D08-1027
- Mitra T, Hutto CJ, Gilbert E. Comparing Person- and Process-centric Strategies for Obtaining Quality Data on Amazon Mechanical Turk. Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems (CHI ’15). 2015:1345–1354. doi:10.1145/2702123.2702553
- Sap M, Card D, Gabriel S, Choi Y, Smith NA. The Risk of Racial Bias in Hate Speech Detection. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL). 2019:1668–1678. doi:10.18653/v1/P19-1163
- Al Kuwatly H, Wich M, Groh G. Identifying and Measuring Annotator Bias Based on Annotators’ Demographic Characteristics. Proceedings of the Fourth Workshop on Online Abuse and Harms. 2020:184–190. doi:10.18653/v1/2020.alw-1.21
- Prabhakaran V, Mostafazadeh Davani A, Díaz M. On Releasing Annotator-Level Labels and Information in Datasets. Proceedings of the Joint 15th Linguistic Annotation Workshop (LAW) and 3rd Designing Meaning Representations (DMR) Workshop. 2021:133–138. doi:10.18653/v1/2021.law-1.14
- Sambasivan N, Kapania S, Highfill H, Akrong D, Paritosh P, Aroyo L. “Everyone wants to do the model work, not the data work”: Data Cascades in High-Stakes AI. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21). 2021:1–15. doi:10.1145/3411764.3445518
- Gurari D, Theriault D, Sameki M, et al. How to Collect Segmentations for Biomedical Images? A Benchmark Evaluating the Performance of Experts, Crowdsourced Non-Experts, and Algorithms. IEEE Winter Conference on Applications of Computer Vision (WACV). 2015:1169–1176. doi:10.1109/WACV.2015.160
- Rädsch T, Reinke A, Weru V, et al. Labelling instructions matter in biomedical image analysis. Nature Machine Intelligence. 2023;5:273–283. doi:10.1038/s42256-023-00625-5
- Andreasen NC. The Scale for the Assessment of Thought, Language, and Communication (TLC). Schizophrenia Bulletin. 1986;12(3):473–482. doi:10.1093/schbul/12.3.473
- Montgomery SA, Åsberg M. A new depression scale designed to be sensitive to change. British Journal of Psychiatry. 1979;134:382–389. doi:10.1192/bjp.134.4.382
- Steller M, Köhnken G. Criteria-Based Content Analysis. In: Raskin DC, ed. Psychological Methods in Criminal Investigation and Evidence. Springer; 1989:217–245.
Frequently asked questions
- Who should annotate your data?
- It depends on three properties of the task: how subjective it is, how much domain knowledge it requires, and the stakes if a label is wrong. Objective, learnable, low-stakes tasks can go to a lay crowd whose labels are aggregated—about four non-expert labels matched a single expert across several NLP tasks (Snow et al., 2008). Specialized, high-stakes constructs like formal thought disorder need a calibrated credentialed expert. And perspective-laden tasks (offensiveness, empathy) need annotators from the target population, because the construct is a lay reaction the expert may no longer share (Prabhakaran et al., 2021).
- When do you need a domain expert to annotate rather than a crowd?
- When the task is technical, the construct is specialized, and the stakes are high—so an untrained reader cannot even recognize the phenomenon. On biomedical image segmentation, trained experts scored highest (median overlap 0.85), ahead of crowdsourced non-experts (0.82) and algorithms (0.36) (Gurari et al., 2015), and professional annotators consistently outperformed crowdworkers across 14,040 medical images (Rädsch et al., 2023). Rating formal thought disorder on the TLC or depression severity on the MADRS falls in this category: the categories only produce reliable labels once raters are trained and calibrated.
- Can non-experts match experts on annotation tasks?
- On learnable tasks, yes—especially when you aggregate several non-expert labels. Snow et al. (2008) found that an average of four non-expert labels per item emulated expert-level quality on affect recognition, and for five of seven tasks a single set of non-expert labels matched or beat a single expert. Screening non-experts for aptitude and training them in the coding scheme raised quality further, outperforming process-only tricks across 68,000 annotations (Mitra et al., 2015). Aggregation and training, not credentials, do the work here.
- When is the right annotator actually a layperson?
- When the label you want is a lay or target-population reaction, an expert can be the wrong annotator. For perceived offensiveness, annotators insensitive to dialect systematically over-labeled African American English as toxic, and priming them on the dialect reduced that bias (Sap et al., 2019). Because annotators draw on their own lived experience, aggregating a non-representative pool buries the very perspectives you meant to measure (Prabhakaran et al., 2021). For these tasks, recruit for representativeness, not seniority.
- Does annotator background or demographics change the labels?
- Yes, measurably, on subjective tasks. Classifiers trained on labels from demographically distinct annotator groups differed significantly, with first language, age, and education all linked to systematic label differences (Al Kuwatly et al., 2020). This is why annotator background is a design decision on perspective-laden tasks, not just a competence check—and why documenting who labeled the data, and retaining annotator-level labels, matters for fairness.
Put this into practice
Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.