responsible data work
The Human Cost of Data Labeling: Ghost Work Behind AI
The human cost of data labeling: the underpaid, invisible ghost workers behind AI, the toll of the work, and what responsible sourcing looks like.

Behind almost every impressive AI system is a workforce most people never see: the human labelers who tag images, transcribe audio, rank chatbot answers, and read through the worst of the internet so a model does not have to. The human cost of data labeling is that this labor is often low-paid, precarious, and psychologically hard, and that it is invisible by design. When Kenyan workers labeled toxic text for a tool later used to make ChatGPT safer, they took home roughly $1.32 to $2 per hour for reading descriptions of abuse and violence all day (Perrigo, 2023).
This post looks at who these workers are, what they are paid, what the work does to them, and why they stay hidden—then at what responsible sourcing actually looks like. The aim is calibration, not outrage: real figures, primary sources, and a checklist a procurement or ethics team can use, handled carefully because the subject is people, not statistics.
What is the human cost of data labeling?
The human cost of data labeling is the gap between how AI is sold—as clean, autonomous software—and how it is actually made: by a large, dispersed, mostly low-paid workforce whose conditions and wellbeing rarely enter the conversation. That cost shows up in three linked ways: pay below a living wage, precarious and opaque employment, and, for some tasks, serious psychological strain from reviewing harmful material.
None of this is a claim that annotation is inherently exploitative. It is a claim that the dominant model for producing training data has externalized its costs onto the people least able to refuse them. Understanding that model is the first step to sourcing data differently, and it starts with a term coined to name the workforce itself.
What is “ghost work”?
Ghost work is the largely invisible, on-demand human labor that makes automated systems appear fully automatic. Anthropologist Mary L. Gray and computer scientist Siddharth Suri introduced the term in their 2019 book to describe the people who flag explicit content, clean data, and check algorithmic output behind services that feel effortless from the outside (Gray & Suri, 2019). The name captures the point: the labor is real, but the worker is meant to disappear.
Their five-year study of workers in the United States and India put rough numbers on it. About 8% of Americans had done ghost work at least once, and on the order of 20 million people worldwide took part in the on-demand economy—typically with no health benefits, no labor-law protection, and the possibility of being dropped at any time (Gray & Suri, 2019). Gray and Suri call this “automation’s last mile”: the stretch of every supposedly automated pipeline that still needs a human, quietly, to work.
The workforce has only grown since. The authors of Feeding the Machine, drawing on more than a decade of fieldwork, argue that most of the work behind modern AI is data-related labor rather than algorithm design, much of it outsourced to workers in the Global South, and describe AI less as a mind than as an “extraction machine” that runs on human effort (Muldoon, Graham & Cant, 2024). Whether or not you accept that framing, the direction is clear: more models means more labeling, and more labeling means more hidden workers.
How much are data labelers actually paid?
Data-labeling pay is frequently below a living wage, and often below the local legal minimum for other work. The most careful public estimate comes from a data-driven analysis of Amazon Mechanical Turk earnings, which found a median wage of roughly $2 per hour, with only about 4% of workers earning more than the US federal minimum of $7.25 per hour (Hara et al., 2018). That is the pay floor for a lot of the microtask economy.
Outsourcing adds a second problem on top of low pay: a markup that stays invisible to the client. In a 2023 TIME investigation, OpenAI contracted a San Francisco firm, Sama, to label descriptions of harmful content in Kenya. The contracts paid Sama about $12.50 per hour for the work, while the labelers doing it took home a wage of roughly $1.32 to $2 per hour depending on seniority and targets—between six and nine times less than the client rate (Perrigo, 2023). The arithmetic of the supply chain, using those documented figures, looks like this:
| Point in the supply chain | Approximate hourly figure (documented) |
|---|---|
| What the AI company paid the vendor | ~$12.50/hour (Perrigo, 2023) |
| What a junior labeler took home | ~$1.32/hour after tax (Perrigo, 2023) |
| What a senior quality analyst took home | up to ~$2/hour (Perrigo, 2023) |
| Local reference: Nairobi receptionist minimum at the time | ~$1.52/hour (Perrigo, 2023) |
To be fair to the full picture, the vendor disputed parts of the account, stating that workers could earn between $1.46 and $3.74 per hour after taxes and had fewer passages to label than workers described (Perrigo, 2023). The point does not hinge on the exact cent. It is that the person reading disturbing material for hours earns a small fraction of what the model’s owner pays for that same hour, and the buyer usually never sees the split. On why paying more is necessary but not sufficient, see does pay improve annotation quality.
What is the psychological toll of this work?
For content moderation and toxic-content labeling, the cost is not only financial—it is a mental-health burden that can outlast the job. Sarah T. Roberts’ ethnography of commercial content moderation describes a workforce of more than 100,000 people, invisible by design, who screen humanity’s worst uploads and carry an emotional toll that employers have been slow to acknowledge (Roberts, 2019). The Kenyan labelers in the TIME investigation, who spent shifts reading textual descriptions of extreme abuse and violence, all described being mentally scarred by the work, and said the “wellness” sessions they were offered were rare and unhelpful (Perrigo, 2023).
The research literature is careful about what we can and cannot claim here, and that care matters. A review of content-moderator wellbeing notes credible reports of PTSD-like symptoms and vicarious trauma, but also that no scientific study has yet quantified how common PTSD is among moderators specifically (Steiger et al., 2021). As a reference point, among people exposed to secondary trauma in general, about 7.8% experience lifelong symptoms and 3.6% meet full PTSD criteria within a given year (Steiger et al., 2021). The honest summary: the harm is real and well-documented in individual accounts, and its exact prevalence in this workforce is still unmeasured.
The mechanism is not mysterious. Prolonged exposure to disturbing material, combined with throughput quotas and thin support, is exactly the mix that occupational-health research associates with deteriorating mental health (Steiger et al., 2021). Which is why the design of the work—exposure limits, genuine counseling, and interfaces that reduce raw exposure—matters as much as the wage.
Why is this workforce invisible?
The invisibility is structural, not accidental. Workers are often represented to clients by numbers rather than names, bound by non-disclosure agreements, and separated by layers of outsourcing that make it hard to know who did what (Gray & Suri, 2019; Roberts, 2019). The result is a workforce that is essential to the product and absent from the story told about it.
That distance also shifts power. A study of outsourced data work in Venezuela and Argentina—210 task-instruction documents and 55 interviews—found that instructions tend to encode and normalize the worldview of the requester, while precarity and economic dependence push workers toward unquestioning obedience rather than judgment (Miceli & Posada, 2022). Workers are told to think in terms of “what the client wants,” and the threat of being dropped keeps them from pushing back on ambiguous or biased categories.
This has a quality consequence, not only an ethical one. When the people closest to the data cannot flag that a taxonomy is wrong for the community it describes, errors get baked in and propagate downstream—the data cascades that surface as failures far from the labeling step. Invisibility and low pay are not just unfair to workers; they quietly degrade the datasets everyone downstream relies on.
What does responsible sourcing of data labeling look like?
Responsible sourcing means treating data work as skilled labor with real conditions attached, not as an anonymous commodity priced only by the item. Two bodies of work turn that principle into specifics. The Partnership on AI’s guidance for buyers covers the decisions that actually shape working conditions: selecting providers, running pilots, designing tasks and instructions, setting payment terms, keeping a communication channel open, and offboarding workers without leaving them stranded (Partnership on AI, 2021). The Fairwork project scores platforms against five principles—fair pay, fair conditions, fair contracts, fair management, and fair representation—and its 2025 ratings found that most cloudwork platforms still fall short, though a record number adopted improvements after engagement (Fairwork, 2025).
Distilled into a checklist a buyer can actually apply, responsible sourcing looks like this:
- Pay a living wage, verified downstream. Confirm the take-home rate for the people doing the work, not just the vendor invoice; the gap between the two is where exploitation hides (Perrigo, 2023). Fair pay is Fairwork’s first principle (Fairwork, 2025).
- Write clear task instructions. Ambiguous tasks waste effort and push interpretive burden onto workers with no power to question it (Miceli & Posada, 2022). Clarity is a fairness measure, not only a quality one.
- Limit exposure and fund real support. For any harmful-content work, cap exposure, allow genuine opt-outs, and provide counseling that is actually accessible—not nominal sessions (Steiger et al., 2021; Perrigo, 2023).
- Be transparent about who does the work. Know your supply chain past the first vendor. Anonymity is what lets low pay and poor conditions persist unseen (Gray & Suri, 2019).
- Give workers a voice and a right to appeal. Due process for deactivation and pay decisions, plus freedom to organize, are Fairwork’s management and representation principles (Fairwork, 2025).
- Own the consent and provenance of the source data too. Fair treatment of labelers sits alongside lawful, consented source data; see data consent and licensing for annotation.
| Red flag in a data vendor | Green flag to look for |
|---|---|
| Only the per-item price is visible; take-home pay is unknown | Documented living wage after fees, verified at the worker level (Fairwork, 2025) |
| Harmful content with token or no mental-health support | Exposure caps, opt-outs, and accessible counseling (Steiger et al., 2021) |
| Opaque subcontracting; you cannot name who labels | Transparent supply chain and working conditions (Gray & Suri, 2019) |
| No route to appeal deactivation or rejected work | Due process and worker representation (Fairwork, 2025) |
Where expert, scale-guided annotation fits
Not all annotation is ghost work, and the distinction is useful for buyers deciding what to outsource cheaply and what not to. Ghost work is deskilled, anonymous piecework paid by the item under time pressure. Expert annotation is the opposite: it depends on trained judgment that cannot be compressed into anonymous micro-tasks without wrecking the result.
Rating depressive severity from an interview on the ten-item, clinician-rated MADRS is not a matter of clicking faster; it is a matter of applying anchored definitions consistently. Coding formal thought disorder on the TLC is harder still, and only becomes reliable once raters are trained and calibrated against a shared standard.
Treating that kind of work like a $1.32-per-hour microtask does not just underpay the person—it produces unusable data, because the constraint was never effort but shared expertise. This is the same reason expert annotators beat crowds on judgment-heavy tasks. The responsible move is to match the labor model to the task: reserve cheap, high-volume piecework for genuinely simple labels done fairly, and pay for expertise where the judgment is the product.
Where Tagaroo fits
Tagaroo’s approach is to shrink the volume of brute-force manual labeling rather than to source it more cheaply. It is a schema-first workspace: you define a coding scheme once with anchored definitions and examples, an AI agent takes a first pass, and human reviewers correct and adjudicate it, with inter-rater reliability tracked as you go. The intent is to spend human effort on judgment, where it is valuable and worth paying for, instead of on repetitive throughput that grinds people down.
That is a position, and it has limits worth naming. Reducing manual volume with a model raises its own questions—the models were themselves trained on labeled data, and cheaper labels are not always safer ones; see synthetic data versus human annotation for where that trade-off breaks. Tagaroo is not a crowd marketplace and does not try to be; if your goal is to push millions of anonymous micro-tasks through the lowest bidder, this is the wrong tool, and the sourcing questions above still apply to whoever you choose. On data handling, Tagaroo takes a de-identify-first path with an anonymous browser-side trial mode, so trial data never leaves your machine; the privacy policy has the specifics.
The practical upshot
The human cost of data labeling is not a side effect of AI—it is part of how AI is currently built, and it is largely a choice about sourcing rather than a law of nature. The evidence points one way: pay is often below a living wage (Hara et al., 2018), outsourcing hides a steep markup (Perrigo, 2023), the work can carry a real psychological toll (Roberts, 2019; Steiger et al., 2021), and the whole arrangement depends on keeping workers invisible (Gray & Suri, 2019; Miceli & Posada, 2022).
If you change one thing, change this: before you buy labeled data, ask what the person doing the work actually takes home, what they are exposed to, and whether they can say no. Then design the task so it needs less brute-force labor and more judgment—and turn your scheme into a guided, reviewable workflow where the human effort is the part worth paying for.
References
- Gray ML, Suri S. Ghost Work: How to Stop Silicon Valley from Building a New Global Underclass. Houghton Mifflin Harcourt; 2019. ghostwork.info
- Perrigo B. Exclusive: OpenAI Used Kenyan Workers on Less Than $2 Per Hour to Make ChatGPT Less Toxic. TIME. January 18, 2023. time.com/6247678/openai-chatgpt-kenya-workers
- Roberts ST. Behind the Screen: Content Moderation in the Shadows of Social Media. Yale University Press; 2019. doi:10.12987/9780300245318
- Steiger M, Bharucha TJ, Venkatagiri S, Riedl MJ, Lease M. The Psychological Well-Being of Content Moderators. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21). 2021. doi:10.1145/3411764.3445092
- Miceli M, Posada J. The Data-Production Dispositif. Proceedings of the ACM on Human-Computer Interaction. 2022;6(CSCW2):Article 460. doi:10.1145/3555561
- Hara K, Adams A, Milland K, Savage S, Callison-Burch C, Bigham JP. A Data-Driven Analysis of Workers’ Earnings on Amazon Mechanical Turk. Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (CHI ’18). 2018. doi:10.1145/3173574.3174023
- Muldoon J, Graham M, Cant C. Feeding the Machine: The Hidden Human Labour Powering AI. Canongate; 2024. canongate.co.uk
- Partnership on AI. Responsible Sourcing of Data Enrichment Services. 2021. partnershiponai.org/responsible-sourcing
- Fairwork. Cloudwork Principles and Cloudwork Ratings 2025. Oxford Internet Institute; WZB Berlin. fair.work/en/fw/principles/cloudwork-principles
Frequently asked questions
- What is ghost work?
- Ghost work is the largely invisible, on-demand human labor—data labeling, content moderation, and data cleanup—that makes automated systems look fully automatic. The term was introduced by anthropologist Mary L. Gray and computer scientist Siddharth Suri, who estimated that about 8% of Americans had done such work at least once and that roughly 20 million people worldwide take part in the on-demand economy, usually without benefits or job security (Gray & Suri, 2019).
- How much do data labelers actually get paid?
- Often very little. A data-driven audit of Amazon Mechanical Turk found a median wage near $2 per hour, with only about 4% of workers earning above the US federal minimum of $7.25 per hour (Hara et al., 2018). In a 2023 TIME investigation, Kenyan workers labeling toxic text for a tool used on ChatGPT took home a wage of roughly $1.32 to $2 per hour, while the client paid the outsourcing firm about $12.50 per hour for the same work (Perrigo, 2023).
- Why is content-moderation and toxic-text labeling psychologically harmful?
- Because the work means repeated, prolonged exposure to descriptions and images of abuse and violence, which can cause lasting distress and vicarious trauma (Roberts, 2019). A literature review of content-moderator wellbeing notes reports of PTSD-like symptoms and points out that among people exposed to secondary trauma generally, about 7.8% experience lifelong symptoms and 3.6% meet full PTSD criteria within a given year, while the prevalence among moderators specifically has not been quantified (Steiger et al., 2021).
- What does responsible sourcing of data labeling look like?
- It means paying at least a local living wage, writing clear task instructions, limiting exposure to harmful content with real mental-health support, being transparent about who does the work, and giving workers a way to appeal decisions. The Partnership on AI's guidance covers provider selection, pilots, task design, payment terms, and offboarding (Partnership on AI, 2021), and the Fairwork project scores platforms against five principles: fair pay, fair conditions, fair contracts, fair management, and fair representation (Fairwork, 2025).
- Is all data annotation ghost work?
- No. Ghost work describes deskilled, anonymous piecework paid by the item under heavy time pressure. Expert annotation—such as rating symptom severity on a validated clinical scale—depends on trained judgment and cannot be compressed into anonymous micro-tasks without destroying data quality. The difference is often power and visibility: task instructions frequently impose the requester's worldview on workers who have little room to push back (Miceli & Posada, 2022).
Put this into practice
Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.