annotator quality
Does Pay Improve Annotation Quality? What Research Shows
Does pay improve annotation quality? Higher pay lifts speed, participation, and fairness—but rarely accuracy. See what really drives label quality.

Does pay improve annotation quality? The most cited experiment on the question found something managers rarely want to hear: paying crowd workers more made them do more work, not better work. When Winter Mason and Duncan Watts varied piece-rate pay on Amazon Mechanical Turk, higher pay increased the quantity of output but left quality flat—workers who were paid more simply perceived their work as more valuable, an anchoring effect that cancelled out the extra motivation (Mason & Watts, 2010). Two decades of related evidence points the same way.
That does not make pay irrelevant. It makes pay the wrong lever for the wrong problem. This post separates what higher pay reliably does for an annotation project (speed, participation, fairness, retention) from what it does not do on its own (accuracy), and then looks at the levers that actually move quality. The evidence is nuanced, so the goal here is calibration, not a slogan.
Does pay improve annotation quality?
Higher pay does not reliably improve annotation quality on its own; it reliably improves the quantity and speed of work and matters for fairness, but accuracy responds to task design, training, and feedback rather than to the rate alone. This is the counterintuitive core of the crowdsourcing literature, and it is stable across very different tasks and platforms.
The cleanest demonstration is still Mason and Watts. Varying the piece rate on a set of Mechanical Turk tasks, they found that more money bought more completed work but not more accurate work.
Their explanation was an anchoring effect: workers paid a higher rate revised upward their sense of what the work was worth, so the higher pay felt fair rather than generous and produced no extra effort (Mason & Watts, 2010). Pay changed the volume dial; it left the quality dial where it was.
There is a telling detail in the same experiment. A quota scheme—pay a fixed amount per batch completed—produced better work for less total pay than an equivalent piece-rate scheme (Mason & Watts, 2010). The structure of the incentive moved quality where the size of it did not. That single study would be easy to dismiss, except that it agrees with a much larger body of organizational research.
| What higher pay reliably improves | What higher pay does not fix on its own |
|---|---|
| Volume: more tasks completed per session (Mason & Watts, 2010) | Accuracy: agreement with a gold standard is roughly unchanged (Mason & Watts, 2010) |
| Speed and participation: faster recruitment, less task abandonment (Ho et al., 2015) | Effort on non-effort-responsive tasks, where trying harder can't change the answer (Ho et al., 2015) |
| Quantity of output in general (corrected r ≈ 0.34) (Jenkins et al., 1998) | Quality of output (no meaningful correlation in the same meta-analysis) (Jenkins et al., 1998) |
| Fairness and retention: wages above the ~$2/hr median (Hara et al., 2018) | Attention and care absent clearer tasks and feedback (Gadiraju, Yang & Bozzon, 2017) |
What does the wider research say about pay and quality?
Beyond crowdsourcing, the largest evidence base is a meta-analysis of financial incentives and performance, and it draws the same line between quantity and quality. Jenkins, Mitra, Gupta and Shaw pooled 39 studies and found that financial incentives were correlated with performance quantity (a corrected correlation of about 0.34) but were not correlated with performance quality (Jenkins et al., 1998). Incentives make people do more; they do not, by themselves, make people do it better.
The mechanism matters for how you read this. Money is excellent at recruiting effort and time; it is poor at manufacturing the skill or the care under ambiguity that quality depends on. If a task is unclear, or an annotator has not been trained on the edge cases, no rate makes the rating correct.
That is why a pay raise so often produces a faster stream of the same errors. It’s also why quantity effects are large and quality effects are not: showing up and clicking more is effort-responsive; being right is not, until something upstream teaches the annotator what right looks like.
When does performance-based pay improve quality?
Performance-based pay improves quality only when the task is effort-responsive—when a worker who tries harder can actually produce a more accurate label. Ho, Slivkins, Suri and Vaughan tested bonus payments against flat pay across different task types and found that bonuses raised quality on effort-responsive tasks but had no quality benefit on tasks where output was insensitive to effort (Ho et al., 2015). The condition, not the bonus, does the work.
This is the important nuance that a headline like “pay doesn’t buy quality” flattens. Tie a bonus to a gold-standard accuracy check on a task where careful re-reading genuinely catches mistakes, and you can move quality. Attach the same bonus to a task where the answer is a coin-flip judgment no amount of effort resolves, and you have simply spent more money for the same distribution of answers.
The same paper notes that payment schemes are, to some degree, implicitly performance-based even without an explicit bonus—so pay and quality are not strictly independent (Ho et al., 2015); the effect just depends on whether effort can change the answer. Before you design an incentive, ask whether effort is even the binding constraint.
If not pay, what actually drives annotation quality?
If raw pay is a weak quality lever, the strong levers are upstream of the annotator: clear task design, good instructions, training, feedback, and quality control. Studying task clarity across roughly 7,000 microtasks, Gadiraju and colleagues found that unclear tasks are widespread and that low clarity degrades the quality of the work produced—workers who cannot tell what you want cannot give it to you, however well paid (Gadiraju, Yang & Bozzon, 2017). Clarity is cheaper than a raise and moves accuracy more.
A 2023 experiment makes the split concrete, and keeps us honest about the nuance. Across 307 data annotators in six conditions, annotators given clear rules were 14% more accurate than those given vague standards; a monetary incentive also raised accuracy, and the best group combined both—but clearer rules were the more cost-efficient way to buy accuracy than money was (Laux, Stephany & Liefgreen, 2023). Pay can help quality; it is just rarely the cheapest thing you can fix.
In practice, four levers do most of the work, and none of them is the hourly rate:
- Instruction and schema design. Concrete definitions, decision rules, and worked examples for the edge cases. This is the single highest-return move; see annotation guidelines that work for the pattern.
- Training and calibration. A short calibration round where annotators code the same items and reconcile differences raises agreement fast—and disagreement is signal, not noise when you use it to find the ambiguous cases your guidelines missed.
- Feedback loops. Telling annotators when they diverge from a gold standard, early and specifically, teaches the boundary cases a rate can’t.
- Quality control. Gold questions, redundant labeling with adjudication, and reliability tracking catch the errors that survive everything else.
Spend on these first. A team with sharp guidelines and a feedback loop at a modest rate will out-annotate a team with vague instructions at a premium rate, because the premium team is being paid more to be confidently wrong in the same places.
But pay still matters—for fairness, speed, and participation
None of this is an argument for paying annotators poorly. Pay is the wrong lever for accuracy, but it is the right lever for participation, speed, retention, and basic fairness—and fairness needs no instrumental justification. The wage data on microtask platforms is bleak: a data-driven analysis of MTurk earnings found a median wage of roughly $2/hour, with only about 4% of workers earning more than the US federal minimum of $7.25/hour (Hara et al., 2018). Underpaying is an ethics failure first, and a recruitment and retention failure second.
There is also a practical floor. Pay that is transparently unfair drives your best annotators away and fills the pool with rushed or adversarial work, which quietly lowers quality through the back door even though a raise wouldn’t have raised it through the front. The right mental model: fair pay removes a quality penalty and recruits effort; it does not by itself add quality on top. Get pay to fair-and-competitive, then spend your quality budget on design, training, and QA.
A worked example: what a pay bump actually buys
Here is a synthetic illustration—round numbers, no real project—of the pattern the research predicts. Suppose a team labels 4,000 short utterances for a symptom-mention scheme and, unhappy with accuracy, doubles the piece rate to fix it.
Before the raise, annotators complete about 80 items an hour and agree with a gold set at Cohen’s κ ≈ 0.55. After the raise, throughput jumps to roughly 120 items an hour and abandonment drops—but agreement barely moves, to κ ≈ 0.57, well within noise. The team paid twice as much for the same accuracy, faster. That is Mason and Watts and Jenkins in miniature: the quantity dial turned, the quality dial didn’t.
Now suppose they instead spend the same money rewriting the guidelines, adding ten worked edge-case examples, and running one calibration round where coders reconcile disagreements. Throughput holds around 80 an hour, but κ rises to 0.74—substantial agreement—because the annotators finally share a definition of the hard cases. Same budget, opposite result. The lesson isn’t “don’t pay”; it’s “pay for the thing that’s actually broken.”
Where curated scales fit as a labeling task
Structured clinical scales are a useful stress test for the pay-versus-quality question, because they are exactly the tasks where effort without shared definitions goes nowhere. Rating depressive severity from an interview with the ten-item, clinician-rated MADRS (Montgomery & Åsberg, 1979) is not a matter of trying harder; it is a matter of applying anchored item definitions the same way every time. Formal thought disorder rated on the TLC (Andreasen, 1986) is harder still—its categories only produce reliable labels once raters are trained and calibrated against a common standard.
For tasks like these, a higher rate does almost nothing and a well-anchored codebook does almost everything. If you are budgeting for reliability on a subjective coding scheme, the money belongs in guidelines, training, and adjudication—not in the per-label price. See how many annotators you actually need for the companion question of redundancy, and synthetic data versus human annotation for where cheap labels stop being safe.
Where Tagaroo fits
Tagaroo is built around the levers this evidence points to. It is a schema-first annotation workspace: you define a coding scheme once with anchored definitions and examples, an AI agent takes a first pass, and human reviewers correct it, with inter-rater reliability, gold-standard checks, and adjudication in the workflow. In other words, it puts the budget where the research says quality lives—task design, calibration, and feedback—rather than treating a higher per-label rate as the fix.
That focus is also its boundary. If your only goal is to push a huge volume of microtasks through the cheapest possible pool, a bare crowdsourcing marketplace is a different tool. Tagaroo’s strength is the guided middle: clear schemes, a model-assisted first pass, and reliability measured rather than assumed. On data handling, it takes a de-identify-first path—strip direct identifiers before upload, with an anonymous browser-side trial mode so trial data never leaves your machine; see the privacy policy for specifics.
The practical upshot
The question “does pay improve annotation quality” has a clean answer once you split quantity from quality: stop asking it, and start asking what your quality is bottlenecked on. If the bottleneck is effort or participation, pay and effort-responsive bonuses help (Ho et al., 2015). If the bottleneck is ambiguity—and on subjective coding tasks it almost always is—no rate fixes it, and clearer guidelines, training, and QA will (Gadiraju, Yang & Bozzon, 2017; Laux, Stephany & Liefgreen, 2023). Pay fairly because it is right and because it recruits good people; then spend your quality budget upstream.
If you change one thing, change this: before you raise the rate, run a short pilot, look at whether your best coders already beat your average ones, and put the difference into the codebook. Then define your scheme and calibrate your raters in Tagaroo so the quality is built in, not bought by the hour.
References
- Mason W, Watts DJ. Financial incentives and the “performance of crowds”. ACM SIGKDD Explorations Newsletter. 2010;11(2):100–108. doi:10.1145/1809400.1809422
- Ho C-J, Slivkins A, Suri S, Vaughan JW. Incentivizing High Quality Crowdwork. Proceedings of the 24th International Conference on World Wide Web (WWW ’15). 2015:419–429. doi:10.1145/2736277.2741102
- Jenkins GD, Mitra A, Gupta N, Shaw JD. Are financial incentives related to performance? A meta-analytic review of empirical research. Journal of Applied Psychology. 1998;83(5):777–787. doi:10.1037/0021-9010.83.5.777
- Gadiraju U, Yang J, Bozzon A. Clarity is a Worthwhile Quality: On the Role of Task Clarity in Microtask Crowdsourcing. Proceedings of the 28th ACM Conference on Hypertext and Social Media (HT ’17). 2017:5–14. doi:10.1145/3078714.3078715
- Hara K, Adams A, Milland K, Savage S, Callison-Burch C, Bigham JP. A Data-Driven Analysis of Workers’ Earnings on Amazon Mechanical Turk. Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (CHI ’18). 2018:1–14. doi:10.1145/3173574.3174023
- Laux J, Stephany F, Liefgreen A. Improving Task Instructions for Data Annotators: How Clear Rules and Higher Pay Increase Performance in Data Annotation in the AI Economy. arXiv preprint. 2023. arXiv:2312.14565
Frequently asked questions
- Does paying annotators more improve label quality?
- Not on its own. In a controlled crowdsourcing experiment, raising pay increased the quantity of work but not its quality—higher-paid workers simply valued their work more, an anchoring effect that left accuracy flat (Mason & Watts, 2010). A meta-analysis of 39 studies found financial incentives correlated with performance quantity but showed no correlation with quality (Jenkins et al., 1998). Pay buys speed and participation; task design, training, and feedback buy accuracy.
- If pay doesn't improve quality, why pay annotators well?
- Because pay drives participation, speed, retention, and fairness—and fairness is its own reason. A data-driven audit of Amazon Mechanical Turk found a median wage near $2/hour, with only about 4% of workers earning above the US federal minimum (Hara et al., 2018). Underpaying is an ethics problem and a recruitment problem even when it isn't an accuracy problem.
- When does performance-based pay actually improve quality?
- When the task is effort-responsive—when trying harder can actually produce a better label. Ho, Slivkins, Suri & Vaughan (2015) found that performance-based bonuses raised quality on effort-responsive tasks but not on tasks where quality was insensitive to effort. Bonuses reward care only where care changes the answer.
- What actually improves annotation quality, if not raw pay?
- Task and instruction design, training, feedback, and quality control. Unclear tasks are widespread and degrade the quality of crowd work (Gadiraju, Yang & Bozzon, 2017), and a 2023 annotation experiment found clear rules raised accuracy 14% over vague standards and were more cost-efficient than money (Laux, Stephany & Liefgreen, 2023). Clearer guidelines, worked examples, gold checks, and calibration move accuracy more reliably than a higher rate does.
- Do higher-paid annotators at least work faster?
- Yes—that is the most robust effect. Higher incentives increase output volume and recruitment speed (Mason & Watts, 2010; Jenkins et al., 1998). The trap is assuming that more work, produced faster, is also more accurate work. It usually isn't unless you change something other than the price.
Put this into practice
Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.