tagaroo

annotator quality

Improve Annotator Quality: Screen and Train, Don't Filter

To improve annotator quality, screen candidates for aptitude and train them before coding—filtering bad labels later loses. See the evidence.

Enrique Gutiérrez14 min readUpdated July 2026
Two paths from a cluster of scattered marks: one drops pieces through a coarse filter, the other guides the same marks through aligned gates into an even, calibrated row, one mark highlighted in coral.

The most reliable way to improve annotator quality is not to filter bad labels after they arrive—it’s to screen the right people in and train them before they code a single item. In a controlled experiment with 68,000 annotations, screening workers for aptitude and training them in the coding technique beat every process-side trick the authors tested, including the much-cited Bayesian Truth Serum (Mitra, Hutto & Gilbert, 2015).

Most quality programs run backwards. They collect labels, compute agreement, drop the coders who look worst, and repeat—paying twice for work they then throw away. This post makes the opposite case: label quality is mostly a people problem, and the highest-return moves happen before and during onboarding, not in post-hoc cleanup. The evidence is specific, and so is the playbook.

Should you filter labels or improve annotator quality upstream?

To improve annotator quality, spend upstream—screen and train coders before they start—rather than treating quality as a filtering problem you solve after the labels exist. Filtering removes bad labels; it does not prevent them, and the effort that produced them is already spent. The controlled evidence favors the person-centric approach for exactly the tasks where quality is hardest to buy: subjective, judgment-heavy coding.

This is the split between process-centric and person-centric quality strategies. Process-centric tactics act on the labels after collection: aggregate many workers, apply incentives, run truth-serum scoring, filter by agreement. Person-centric tactics act on the coder before collection: select for aptitude, teach the technique, calibrate, give feedback. Both have their place, but they are not equal in strength, and most teams over-invest in the first because it feels like a knob you can turn without hiring anyone.

The person-versus-process crowdsourcing question is not rhetorical—it was measured directly, at scale, and the answer is unambiguous enough to plan around.

What the 68,000-annotation experiment actually found

The largest controlled test compared 34 quality-control strategies across 68,000 annotations and more than 280 pairwise comparisons, on subjective coding tasks of varying difficulty (Mitra, Hutto & Gilbert, 2015). The headline result: person-oriented strategies—prescreening workers for the requisite cognitive aptitudes and giving them basic training in qualitative coding methods—produced better agreement and interpretive convergence than the process-oriented alternatives.

Two findings inside that study are worth pinning down, because they contradict common practice. First, screening and training improved quality above and beyond Bayesian Truth Serum, a well-regarded process-side incentive technique—and in the head-to-head, BTS was the least effective of the strategies compared. Second, financial incentives did not reliably improve quality; their effect became negligible once stronger person-centric strategies were in play (Mitra, Hutto & Gilbert, 2015). That last point tracks the broader literature: pay buys speed and participation, not accuracy, a pattern we cover in does pay improve annotation quality.

The practical reading is not “never aggregate or incentivize.” It’s that the biggest lever is the coder, and the cheapest time to pull it is before the labels exist.

A four-step sequence to improve annotator quality

The most dependable way to improve annotator quality is a short, ordered onboarding sequence: screen, then train, then calibrate, then feed back. Each step is backed by its own evidence, and each fixes a failure the next step can’t reach on its own. Run them in order before production coding begins.

  1. Screen for the aptitude the task actually requires—reading comprehension, domain familiarity, or attention to structure—using a short qualification task on gold-labeled items. Prescreening for cognitive aptitude was half of the winning strategy in the 68,000-annotation study (Mitra, Hutto & Gilbert, 2015).
  2. Train on the coding technique with worked examples and the edge cases, not just definitions. Across 96 studies, training improved inter-coder agreement, and high-intensity training helped most (Bayerl & Paul, 2011).
  3. Calibrate by having coders label the same pilot set and reconcile their disagreements out loud, so the team converges on one reading of the hard items. This is what a pilot annotation round is for, and Cohen’s kappa is how you know it worked.
  4. Feedback early and specifically: show each coder where they diverged from the gold standard and why, before habits set. Timely task-specific feedback produced better work and improvement over time in a controlled study (Dow, Kulkarni, Klemmer & Hartmann, 2012).

Skip a step and the sequence degrades predictably. No screening, and you train people who can’t do the task; no calibration, and training produces confident coders who each learned it slightly differently; no feedback, and calibration decays the moment the hard cases stop being discussed.

Does annotator training raise inter-coder agreement?

Yes—training raises inter-coder agreement, and the effect is one of the more consistent findings in the annotation literature. In a meta-analytic investigation of 96 annotation studies spanning 346 agreement indices across three domains, whether annotators received training and the intensity of that training were among the seven factors that systematically influenced reported agreement, with more intensive training associated with higher agreement (Bayerl & Paul, 2011).

The same meta-analysis is useful for what else it found, because it points at design choices that multiply the training effect. Agreement was higher with fewer categories in the coding scheme; whether coder groups were homogeneous or mixed made no significant difference across the full sample, reaching significance only in one sub-domain (prosodic transcription) (Bayerl & Paul, 2011). Coder training quality, in other words, is not only about hours of instruction—it’s about training on a scheme that isn’t fighting the coders. Simplify the scheme, and the same training buys more.

Training sets the baseline; feedback is what keeps it from eroding. In the Shepherd experiment, crowd workers writing consumer reviews were split into no-feedback, self-assessment, and expert-feedback conditions. Both feedback conditions produced better work than no feedback and helped workers improve across tasks, with a significant main effect for assessment condition (F(1,374)=4.55, p < 0.05), and self-assessment worked nearly as well as expert review (Dow et al., 2012). The annotator training effect, then, is really two effects—an initial lift from teaching, and a compounding lift from feedback—and you want both.

How to screen and select good annotators

Selecting good annotators means testing candidates on a small, gold-labeled version of the real task and keeping the ones who perform, rather than screening on résumés or generic qualifications. The most concrete recipe comes from a controlled crowdsourcing protocol for a demanding semantic-annotation task: release a preliminary round to the whole pool, invite the workers who performed reasonably, have them study short guidelines, then run two qualification rounds of 15 items each—every round followed by detailed, personal feedback on their errors (Roit et al., 2020).

The economics are the part teams miss. That protocol cost about two hours of paid worker training and roughly half an hour of a trainer’s time per worker; the team trained 30 workers and kept the 11 who performed well (Roit et al., 2020). For that outlay it produced 25% more labeled roles than the earlier unscreened crowd, at comparable cost per item and without losing precision—non-expert output that reached expert-level quality. Screening is not a way to find people who are already perfect; it’s a filter that tells you whom to invest the training in.

Screening also changes what redundancy is for. Snow and colleagues showed that non-expert labels can match expert quality when you pool enough of them—about four non-expert labels per item to emulate an expert on an affect-recognition task—and that a small gold set can correct individual annotators’ biases (Snow et al., 2008). Redundancy and bias correction are real tools, but they work best on coders you’ve already screened and trained; piling more labels on an untrained pool mostly buys you more correlated errors. For the deeper question of who to recruit in the first place, see the ideal annotator background for your task.

Filter vs. train: which to spend on

Spend on training and screening first, and keep filtering as the safety net—because filtering acts too late and can quietly damage your dataset. The contrast below is the argument in one view: read each row as the same problem handled after the fact versus handled upfront.

Filter after the fact (process-centric)Screen and train upfront (person-centric)
Removes bad labels once they exist; the effort that produced them is wastedPrevents bad labels by selecting and teaching coders before they start (Mitra, Hutto & Gilbert, 2015)
Can over-prune—discarding hard-but-correct cases and shrinking the datasetRaises agreement broadly, especially with high-intensity training (Bayerl & Paul, 2011)
Silent on why coders diverge, so the same errors recur next batchFeedback teaches the boundary cases and coders improve over time (Dow et al., 2012)
A gate at the end; the signal arrives after you've paid for the workA gate at the start plus qualification rounds catch weak coders early (Roit et al., 2020)
Filtering versus screening-and-training on the same four failure modes. Filtering is a late, subtractive gate; screening and training are early, additive investments. Both belong in a pipeline—but they are not interchangeable.

The trap is treating them as substitutes. Aggressive agreement filtering feels like quality control, but it can strip out exactly the ambiguous items that carry the most information, and it never tells the coder what “right” looked like. Use filtering to catch drift and bad actors, not to manufacture a quality you failed to build in.

A worked example: onboarding coders for a symptom-coding task

Here is a synthetic illustration—round numbers, no real project or transcript—of the pattern the research predicts. A team needs to label 5,000 short interview utterances for whether each mentions a target symptom. They recruit six coders and, in the name of speed, skip onboarding: brief instructions, straight to production, plan to filter out whoever disagrees most at the end.

The first batch comes back at Cohen’s κ ≈ 0.52—moderate at best. They filter the two lowest coders and re-label those items, which nudges agreement to κ ≈ 0.56 and costs a full extra pass. The disagreements cluster on the genuinely hard utterances (hedged, sarcastic, or context-dependent mentions), which is precisely where the discarded labels might have been informative. They’ve paid twice and learned nothing transferable.

Now run the sequence instead. The same six coders take a 20-item qualification set; one clearly can’t do the task and is replaced. The remaining five get an hour of training on ten worked edge cases, then a calibration round on 40 shared items where they reconcile every disagreement, followed by two rounds of targeted feedback.

Production agreement lands at κ ≈ 0.78—substantial—on the first real batch, with no cleanup pass. Same budget, opposite result: the money went into the coders instead of into re-doing their work.

Where curated clinical scales fit as a coding task

Structured clinical scales are the sharpest test of this argument, because they are exactly the tasks where effort without shared definitions goes nowhere. Rating formal thought disorder on the Scale for the Assessment of Thought, Language and Communication (TLC) (Andreasen, 1986) only yields reliable labels once raters are trained and calibrated against a common standard—its categories like derailment and tangentiality are genuinely subtle. Rating depressive severity on the ten-item, clinician-rated MADRS (Montgomery & Åsberg, 1979) is not a matter of trying harder either; it’s a matter of applying anchored item definitions the same way every time.

For coding schemes like these, filtering after the fact is close to useless and upfront investment is close to everything. A raters’ handbook, a training set of exemplar interviews, and a calibration session do the work that no amount of post-hoc agreement pruning can. If you are budgeting for reliability on a subjective scheme, the money belongs in annotation guidelines that work, training, and calibration—not in a bigger cleanup crew.

When filtering still earns its keep

None of this means skip quality control after collection—it means put it in its right place, as a safety net rather than the strategy. Even well-screened, well-trained coders drift, get tired, and hit cases the training never covered. Ongoing checks catch what onboarding can’t: gold questions and honeypots flag a coder whose accuracy is slipping, reliability tracking catches gradual drift, and redundancy with adjudication resolves the genuinely ambiguous items.

Bias correction belongs here too. Modeling individual annotators against a small gold set to recalibrate their labels measurably improves quality (Snow et al., 2008)—but as a complement to good coders, not a substitute for them. The failure mode to avoid is inverting the priority: leaning on filtering and aggregation to rescue a pool you never screened or trained. That’s the backwards program from the top of this post, and it costs the most for the least.

Where Tagaroo fits

Tagaroo is built around the levers this evidence points to. You define a coding scheme once with anchored definitions and worked examples, an AI agent takes a first pass, and human reviewers screen, calibrate, and correct—with inter-rater reliability, gold-standard checks, and adjudication in the workflow rather than bolted on afterward. In short, it puts the budget where the research says quality lives: onboarding annotators, calibration, and feedback, not a bigger post-hoc filter.

That focus is also its boundary. If your only goal is to push a huge volume of trivial microtasks through the cheapest possible pool, a bare crowdsourcing marketplace is a different tool. Tagaroo’s strength is the guided middle: clear schemes, a model-assisted first pass, and reliability measured rather than assumed. On data handling, it takes a de-identify-first path—strip direct identifiers before upload, with an anonymous browser-side trial mode so trial data never leaves your machine; see the privacy policy for specifics.

The practical upshot

If you want to improve annotator quality, stop asking which labels to throw away and start asking who you let in and how you teach them. The strongest controlled evidence says the coder is the lever, and the cheapest time to pull it is before the labels exist (Mitra, Hutto & Gilbert, 2015); training raises agreement (Bayerl & Paul, 2011), feedback compounds it (Dow et al., 2012), and screening plus training pays for itself in coverage and precision (Roit et al., 2020). Filter to catch drift and bad actors—but build quality upstream, where it’s cheaper and it sticks.

If you change one thing, change the order: screen and train before the first production batch, not after the first disappointing agreement score. Then define your scheme and calibrate your raters in Tagaroo so quality is built in, not filtered out.

References

  • Mitra T, Hutto CJ, Gilbert E. Comparing Person- and Process-centric Strategies for Obtaining Quality Data on Amazon Mechanical Turk. Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems (CHI ’15). 2015:1345–1354. doi:10.1145/2702123.2702553
  • Bayerl PS, Paul KI. What Determines Inter-Coder Agreement in Manual Annotations? A Meta-Analytic Investigation. Computational Linguistics. 2011;37(4):699–725. doi:10.1162/COLI_a_00074
  • Dow S, Kulkarni A, Klemmer S, Hartmann B. Shepherding the Crowd Yields Better Work. Proceedings of the ACM 2012 Conference on Computer Supported Cooperative Work (CSCW ’12). 2012:1013–1022. doi:10.1145/2145204.2145355
  • Roit P, Klein A, Stepanov D, Mamou J, Michael J, Stanovsky G, Zettlemoyer L, Dagan I. Controlled Crowdsourcing for High-Quality QA-SRL Annotation. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL 2020). 2020:7008–7013. doi:10.18653/v1/2020.acl-main.626
  • Snow R, O’Connor B, Jurafsky D, Ng AY. Cheap and Fast—But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks. Proceedings of the 2008 Conference on Empirical Methods in Natural Language Processing (EMNLP 2008). 2008:254–263. aclanthology.org/D08-1027

Frequently asked questions

What is the best way to improve annotator quality?
Screen candidates for the aptitude the task needs, then train them in the coding technique before they touch production data. In a controlled experiment with 68,000 annotations, this person-centric approach outperformed process-side strategies—including Bayesian Truth Serum—and financial incentives did not reliably improve quality (Mitra, Hutto & Gilbert, 2015). Post-hoc filtering removes bad labels but never prevents them or teaches the coder, so it is a safety net rather than a strategy.
Does training annotators actually raise inter-coder agreement?
Yes. A meta-analysis of 96 annotation studies found that annotator training improved inter-coder agreement, with high-intensity training helping most (Bayerl & Paul, 2011). The same review found agreement was also higher with fewer categories in the scheme; whether a group's expertise was homogeneous or mixed made no significant overall difference, so scheme design is the more dependable lever.
Is screening and training worth the cost versus just hiring more annotators?
Usually. A controlled crowdsourcing protocol that screened and trained workers—about two hours of training per worker—produced 25% more labeled roles than an unscreened crowd at comparable cost and precision (Roit et al., 2020). You pay once for a coder who is right rather than repeatedly for redundant labels and cleanup.
When is filtering bad annotators still necessary?
Filtering earns its keep as a safety net even after good screening and training. Gold-standard checks, attention checks, and reliability tracking catch drift, fatigue, and the occasional adversarial worker that no onboarding prevents. The mistake is using filtering as the primary quality mechanism instead of the last line of defense; see gold questions and honeypots for how to build that net without over-pruning.
How does feedback fit into onboarding annotators?
Feedback is the step that makes training stick. In the Shepherd study, timely, task-specific feedback produced better work than no feedback and helped workers improve over time (Dow, Kulkarni, Klemmer & Hartmann, 2012). Telling coders exactly where and why they diverged from a gold standard, early in onboarding, teaches the boundary cases that written guidelines alone miss.

Put this into practice

Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.