Study workflow

Annotation budget & throughput calculator

Six inputs to a labeling budget, two more to a calendar. See what each additional independent pass costs, what each additional coder actually buys you, and the per-label figure to hold a vendor quote against.

Free · No sign-up · Runs entirely in your browser

Six inputs set the budget; two more turn it into a timeline. Every figure updates as you type.

Cost inputs

Budget

Total cost

$50,000

6,000 labels · 1,500 h productive

Cost per item

$25.00

The figure to compare against a quote

Cost per label

$8.33

Scope-independent

Each extra pass

$15,000

One more independent pass over every item

Budget breakdown by line
Labeling2,000 × 3 passes × 12 min1,200 h$36,000
QA & adjudication+25% of labeling300 h$9,000
Fixed setupPilot, guidelines, training$5,000
Total1,500 h$50,000

Timeline

Calendar time

12.5 weeks

63 working days at 4 coders

Productive hours

1,500 h

Labeling plus QA; setup is extra

Plus setup, which headcount does not compress. Note also what this arithmetic cannot see: more coders mean more disagreements to adjudicate and more onboarding, and past a handful of people the real bottleneck becomes a single overloaded reviewer.

What each additional pass costs

Redundancy multiplies the base by a whole number, so this curve is a straight line — and each step is $15,000 at your current inputs.

  • $20,000
  • $35,000
  • $50,000
  • $65,000
  • $80,000

Snow et al. (2008) found roughly 4 non-expert passes averaged to the quality of one expert annotation on a simple affect task, which is why crowd pipelines commission four or five. Calibrated expert coders usually need two, with a third only to break ties. And Sheng et al. (2008) add the refinement that saves the money: blanket redundancy is wasteful — spend the extra passes on the ambiguous items rather than on the easy 80% everyone already agrees about.

What each additional coder buys

  • 50.0 wk
  • 25.0 wk−25.0
  • 16.7 wk−8.3
  • 12.5 wk−4.2
  • 10.0 wk−2.5
  • 8.3 wk−1.7
  • 7.1 wk−1.2
  • 6.3 wk−0.9

The saving from each added coder shrinks as you go: the first extra person halves the schedule, the eighth barely moves it. Read the small figures on the right as the weeks saved by that coder over the one before.

This is a planning estimate, not a quote. The wage you pay is not the wage a worker nets: Hara et al. (2018) tracked 2,676 Mechanical Turk workers across 3.8 million tasks and found a median effective wage of about $2.00/hour, with only around 4% clearing the $7.25 US federal minimum, even though requesters paid $11.58/hour on average. Most of the gap is unpaid time. Cost per label also spans orders of magnitude with the labeler: crowd work runs to thousands of labels per dollar, while a trained clinician applying a structured scale to a transcript is minutes of expert attention per label. Use this to size a project and interrogate a quote — then run a pilot, because that is what replaces time-per-item, disagreement rate and QA overhead with measurements instead of guesses.

Plain text: every input, the effort and cost breakdown, per-item and per-label figures, the timeline and the caveat — ready for a grant application or a planning doc.

Reference · for the curious

Costing an annotation project: the formula, the levers and the honest timeline

An annotation budget is six numbers multiplied and added in the right order. The arithmetic is not the hard part; getting the inputs honest is, and knowing which of them actually moves the total is what separates a budget from a hope.

The formula

Total cost ≈ (Items × Passes-per-item × Time-per-item × Loaded-hourly-wage) × (1 + QA-overhead) + Fixed-setup-cost.

Calendar time ≈ Total productive hours ÷ (Headcount × Productive-hours-per-day).

Two terms do most of the damage when they are wrong. Time per item multiplies against every item and every redundant pass, so an error there scales through the whole budget. Passes-per-item — redundancy — multiplies the base by a whole number. The wage feels like the obvious driver but it is usually the term you have least room to move, and the one that matters least to the total. Get the two multipliers right first.

What a labeler actually costs

The wage line hides a gap between what you pay and what the worker keeps. Hara and colleagues (2018) recorded 2,676 workers doing 3.8 million tasks on Amazon Mechanical Turk and found a median effective wage of roughly $2 an hour, with only about 4% clearing the $7.25 US federal minimum — even though requesters paid an average of $11.58 an hour. Most of the difference is unpaid time: searching for tasks, work that gets rejected, tasks started and never submitted.

That matters twice for a clinical or high-stakes project. The “loaded” wage in the formula is not the sticker rate — it has to carry benefits, unpaid overhead and idle time. And paying more is not automatically buying quality; the relationship is real but conditional, which is what the evidence on whether pay improves annotation quality untangles. Budget the wage honestly, then spend the marginal dollar where it actually moves reliability.

Cost per label spans orders of magnitude for the same reason. Snow and colleagues (2008) bought 7,000 non-expert labels for two dollars — about 3,500 labels per dollar on a simple affect task. A trained clinician applying a structured scale to an interview transcript is at the other end of that range entirely. A per-label quote is only useful with a labeler attached to it.

Redundancy is the lever people get wrong

Two independent passes double the labeling cost; three triple it. The question is never “should we double-code?” in the abstract but “how many passes does our reliability target need, and on which items?”

Snow and colleagues found roughly four non-expert labels averaged to the quality of one expert annotation, which is why crowd pipelines commission four or five. Expert clinical coders usually need fewer, because they are already calibrated — two independent coders is the common pattern, with a third to break ties. Our guide to how many annotators per item works the decision through properly, and the reliability sample-size planner tells you how many items the agreement substudy itself needs.

Sheng and colleagues (2008) add the refinement that saves real money: blanket redundancy is wasteful. Repeatedly labeling a carefully chosen set of items beats labeling everything the same number of times. Spend the extra passes where uncertainty is high rather than triple-coding the easy 80% everyone already agrees about.

QA is a line item, not a rounding error

Once two people label the same item, somebody has to resolve their disagreements — reading the conflicting cases, deciding the right label, and recording the rule so it does not recur. That is real labour, and how you resolve conflicts (majority, weighted vote, a probabilistic model, expert review) has its own cost and failure modes, laid out in adjudication and consensus methods.

There is a subtler cost Sheng and colleagues name directly: when labeling is cheap, preparing the unlabeled data can become more expensive than labeling it. Sampling, de-identification, formatting and loading all happen before a single label is applied. On low-cost tasks they can dominate — and de-identification in particular is work you should budget explicitly, which the transcript de-identifier handles for text.

Twenty-five percent is a reasonable starting assumption for a double-coded project with active adjudication. Measure your real disagreement rate in the pilot and adjust it; a scheme with high disagreement can push QA well past that.

The timeline term people fake

A productive annotation day is closer to six hours than eight. Breaks, ramp-up and the accuracy decay of sustained close reading are not optional overheads you can plan away — the vigilance decrement is a measured effect, and annotator burnout is what happens to projects that budget eight. The calculator defaults to six for that reason.

Headcount compresses the labeling phase and nothing else. It does not compress setup, it inflates adjudication, and beyond a handful of coders a single overloaded reviewer becomes the binding constraint. The headcount curve in the tool shows the arithmetic of diminishing returns; the coordination ceiling sits below where that curve flattens.

Where transcription fits

One cost sits upstream of every line in this budget: turning recordings into text in the first place. That is a separate estimate with its own drivers — audio quality, speaker count, verbatim conventions — and the transcription time and cost calculator sizes it. If your corpus is not yet transcribed, run that first and treat its output as an input here.

Protect the smallest line

The fixed setup cost is the smallest dollar figure in most versions of this budget and the one to defend hardest. Running a pilot round before you scale is where you discover that your real time per item is 15 minutes and not 10, that one scale item drives half your disagreements, and that your QA overhead should be 35% rather than 25%. The pilot does not just de-risk the project — it replaces the guesses in this calculator with measurements, which is the whole difference between a budget and a wish. Skipping it is how you end up with the compounding downstream failures described in data cascades.

The full worked example, with a side-by-side fully-human and AI-assisted budget, is in the annotation cost calculator explainer.

Scope of this tool

Every figure here is a planning estimate, not a vendor quote, and the published benchmarks it cites are ranges from specific studies of specific tasks — slide them to match your labelers and your work. Everything is computed in your browser; no inputs are transmitted or stored.

Frequently asked questions

How do you calculate the cost of an annotation project?

Multiply four numbers and add two. The base is Items × Passes-per-item × Time-per-item × Loaded-hourly-wage; then add QA and adjudication as a percentage of that base, plus a fixed setup cost for the pilot, guidelines and training. As a worked example, 2,000 transcripts × 3 passes × 12 minutes × $30/hour is 1,200 hours and $36,000 of labeling, plus 25% ($9,000) for QA and $5,000 of setup, for roughly $50,000 — about $25 per item and $8.33 per label. The single biggest lever is time per item, because it multiplies against every item and every redundant pass.

How much does data labeling cost per label?

It spans several orders of magnitude, so the figure means nothing until you name the labeler and the task. At the cheap end, Snow and colleagues (2008) collected 7,000 non-expert Mechanical Turk annotations for US$2.00 — roughly 3,500 crowd labels per dollar on a simple affect task. At the expensive end, a trained clinician reading an interview transcript and applying a structured scale is minutes of expert attention per label, which lands in dollars rather than fractions of a cent. Published crowd figures anchor the floor; treat commercial vendor pricing as something to obtain by quote rather than assume, and compare quotes on cost per label rather than on the total.

How long will an annotation project take?

Calendar time is total productive hours divided by headcount times productive hours per day — and the honest value for that last term is closer to six than eight for close-reading annotation, once breaks, ramp-up and the accuracy decay of sustained vigilance are subtracted. In the worked example above, 1,500 productive hours take about 12–13 weeks with four coders or about 6 with eight, plus one to two weeks of setup that no amount of headcount removes. Headcount compresses the labeling phase only: it leaves setup untouched and quietly inflates adjudication, and past a handful of coders a single overloaded reviewer usually becomes the real constraint.

How many independent passes per item should I budget for?

Redundancy multiplies the base cost by a whole number, so it is the second-biggest lever after time per item — two passes double the labeling cost, three triple it. Snow and colleagues (2008) found roughly four non-expert labels averaged to the quality of one expert annotation, which is why crowd pipelines often commission four or five. Calibrated expert coders usually need two, with a third only to resolve ties. Sheng and colleagues (2008) add the refinement that saves real money: blanket redundancy is wasteful, and repeatedly labeling a carefully chosen subset beats labeling everything the same number of times. Spend the extra passes on the ambiguous items.

What is the real hourly wage on Amazon Mechanical Turk?

About $2 an hour. Hara and colleagues (2018) logged 2,676 workers completing 3.8 million tasks and found a median effective hourly wage of roughly $2, with only around 4% of workers clearing the $7.25 US federal minimum — even though requesters paid an average of $11.58 per hour. The gap is unpaid time: searching for tasks, work that gets rejected, and tasks started but never submitted. For budgeting this cuts two ways. It means the figure you enter as a wage must be a loaded rate that carries that overhead, not the sticker rate you advertise; and it means crowd pricing carries an ethical question that a per-label cost comparison hides entirely.

Can preparing the data cost more than labeling it?

Yes, and on cheap tasks it usually does. Sheng, Provost and Ipeirotis (2008) observed directly that when labeling is inexpensive, "preparing the unlabeled part of the data can become considerably more expensive than labeling." Sampling, de-identification, formatting, transcription and loading all happen before a single label is applied, and none of them appear in the items × passes × time × wage formula. Two consequences for a real budget: cost per label is a misleading figure to negotiate on for a low-cost task, because the labeling is not where the money goes; and for clinical transcript work the upstream steps — transcription and de-identification — deserve their own line items rather than being folded into setup.

Does AI pre-labeling reduce annotation cost?

It lowers exactly one term — time per item — by turning annotation into review-and-correct, and the effect is large: dropping 12 minutes per item to 5 cuts the base cost in the worked example from $36,000 to $15,000 without touching redundancy or QA. But the saving is bounded and it carries a risk. Model-assisted labels still need human checks wherever the model is uncertain, and pre-labeling introduces automation bias, where a reviewer rubber-stamps a confident but wrong suggestion. Treat pre-labeling as a multiplier on the time term, keep the reliability plan intact, and re-measure agreement after you switch it on.

Written by Enrique Gutiérrez, PhD (Computer Science) — founder of Tagaroo and Associate Professor of Computer Science, working on inter-rater reliability, measurement and annotation methodology (ORCID).

Last verified: 29 July 2026. Formulas, thresholds and cited figures on this page were checked against their original sources on that date. Every calculation runs in your browser; nothing you enter is transmitted or stored.

Replace the guesses in this calculator with measurements

The inputs that decide this budget — time per item, disagreement rate, the redundancy your reliability target needs — are exactly what a pilot produces. Tagaroo loads curated scales as pre-written anchored codebooks and computes inter-rater reliability as coders work, so those numbers come out measured rather than assumed.

Try Tagaroo free