A Lobsters discussion thread last week pulled us back to a forgotten question: behind the models from OpenAI, Anthropic, and Scale AI sit millions of data workers distributed across Kenya and the Philippines, often earning under $2 per hour. The thread linked to a subpage on ai-materiality-map.org with a blunt question — Where are the data workers behind AI?

What the researchers want to do is complete the AI "materiality map." Beyond compute, power, and data centers, there should be a fourth infrastructure: people. They label training corpora, run human feedback scoring (RLHF — the core step that aligns models to human preferences), filter harmful content, and tune model outputs from "readable" to "actually useful."

What this is

So-called "AI data workers" are the real laborers doing this work. A 2023 joint investigation by Time and Wired exposed that OpenAI, through outsourcing firms in Kenya, hired people to label text containing self-harm, violence, and sexual assault — at under $2 per hour, with limited psychological support. That sparked a round of discussion at the time, but it was quickly drowned out by the noise of new model releases. The value of this Lobsters thread is that research institutions are now systematically mapping AI's materiality map, aiming to prove that without these workers, any leading AI company's product today would be missing a piece of the puzzle.

Industry view

The supporting side's reasoning is direct: OpenAI, Anthropic, Scale AI, Surge AI, and Appen pay billions of dollars annually into the data annotation chain — a scale comparable to the traditional BPO (Business Process Outsourcing) industry. The difference is that BPO has clear client-vendor relationships and headcount disclosure; AI data annotation's hours, unit prices, and worker distribution are largely opaque.

The pushback is equally strong. The first view holds that labor costs account for under 5% of total training costs for top-tier large models, making the focus on data worker conditions a case of missing the point — what really consumes resources is compute and power. The second rebuttal is sharper: media repeatedly highlighting "low data worker wages" actually leads the public to believe "AI still needs humans," deflecting from the bigger issue of white-collar displacement.

Our editorial judgment is that these two things are not contradictory. In the short term, data annotation is a real industry with over a million workers, and its labor conditions are worth discussing. In the medium term, as models' self-supervised learning (Self-Supervised Learning — letting models find patterns on their own from massive text) capabilities improve, pure annotation demand will fall, but "real-time human feedback on model behavior" won't disappear — only the job type will shift.

Impact on regular people

  • For enterprise IT: When procuring GPT or Claude-style enterprise services, the vendor's "data sourcing compliance" page deserves a real read — not a marketing-deck skim.
  • For individual careers: White-collar roles aren't directly threatened, but when your company discusses "replacing a position with AI," it's worth asking one more question — where do the replaced people go next?
  • For the consumer market: Consumers won't feel it immediately, but when AI customer service gives irrelevant answers, content moderation misfires, or recommendations clearly go off the rails, there may be a direct cost from compressed human review layers behind it.