Back to home

Compare

Comparing: AI's Invisible Workforce: Millions of Annotators Power LLMs, Make Zero Headlines & AI 数据工人集体隐身 — 上百万标注员撑起大模型却极少上头条

AEN
OpenAIAnthropicScale AI·

AI's Invisible Workforce: Millions of Annotators Power LLMs, Make Zero Headlines

A Lobsters discussion thread last week pulled us back to a forgotten question: behind the models from OpenAI, Anthropic, and Scale AI sit millions of data workers distributed across Kenya and the Philippines, often earning under $2 per hour. The thread linked to a subpage on ai-materiality-map.org with a blunt question — Where are the data workers behind AI?

What the researchers want to do is complete the AI "materiality map." Beyond compute, power, and data centers, there should be a fourth infrastructure: people. They label training corpora, run human feedback scoring (RLHF — the core step that aligns models to human preferences), filter harmful content, and tune model outputs from "readable" to "actually useful."

What this is

So-called "AI data workers" are the real laborers doing this work. A 2023 joint investigation by Time and Wired exposed that OpenAI, through outsourcing firms in Kenya, hired people to label text containing self-harm, violence, and sexual assault — at under $2 per hour, with limited psychological support. That sparked a round of discussion at the time, but it was quickly drowned out by the noise of new model releases. The value of this Lobsters thread is that research institutions are now systematically mapping AI's materiality map, aiming to prove that without these workers, any leading AI company's product today would be missing a piece of the puzzle.

Industry view

The supporting side's reasoning is direct: OpenAI, Anthropic, Scale AI, Surge AI, and Appen pay billions of dollars annually into the data annotation chain — a scale comparable to the traditional BPO (Business Process Outsourcing) industry. The difference is that BPO has clear client-vendor relationships and headcount disclosure; AI data annotation's hours, unit prices, and worker distribution are largely opaque.

The pushback is equally strong. The first view holds that labor costs account for under 5% of total training costs for top-tier large models, making the focus on data worker conditions a case of missing the point — what really consumes resources is compute and power. The second rebuttal is sharper: media repeatedly highlighting "low data worker wages" actually leads the public to believe "AI still needs humans," deflecting from the bigger issue of white-collar displacement.

Our editorial judgment is that these two things are not contradictory. In the short term, data annotation is a real industry with over a million workers, and its labor conditions are worth discussing. In the medium term, as models' self-supervised learning (Self-Supervised Learning — letting models find patterns on their own from massive text) capabilities improve, pure annotation demand will fall, but "real-time human feedback on model behavior" won't disappear — only the job type will shift.

Impact on regular people

  • For enterprise IT: When procuring GPT or Claude-style enterprise services, the vendor's "data sourcing compliance" page deserves a real read — not a marketing-deck skim.
  • For individual careers: White-collar roles aren't directly threatened, but when your company discusses "replacing a position with AI," it's worth asking one more question — where do the replaced people go next?
  • For the consumer market: Consumers won't feel it immediately, but when AI customer service gives irrelevant answers, content moderation misfires, or recommendations clearly go off the rails, there may be a direct cost from compressed human review layers behind it.
BZH
OpenAIAnthropicScale AI·

AI 数据工人集体隐身 — 上百万标注员撑起大模型却极少上头条

上周 Lobsters 一条讨论帖把我们拉回到一个被遗忘的问题:OpenAI、Anthropic、Scale AI 这些公司的模型背后,是数百万分布在肯尼亚、菲律宾的数据工人,时薪常低于 2 美元。讨论链接的是 ai-materiality-map.org 的一个子页面,问句很直接——Where are the data workers behind AI?

研究人员想做的是把 AI 的「物料地图」补完。算力、电力、数据中心之外,还应该有第四种基础设施:人。它负责标注训练语料、做人类反馈打分(RLHF,让模型对齐人类偏好的核心步骤)、筛掉有害内容、把模型回答从「能看」调到「好用」。

这是什么

所谓「AI 数据工人」,就是完成上述工作的真实劳动者。2023 年《时代》和《连线》联合调查揭露,OpenAI 在肯尼亚通过外包公司雇人标注含自残、暴力、性侵的文本,时薪低于 2 美元,配套心理学支持有限。当时掀起过一轮讨论,很快被新模型发布的声量盖过。这次 Lobsters 讨论的价值在于,研究机构开始系统性地绘制 AI 物料地图,想证明没有这批人,今天任何一家头部 AI 公司的产品都会缺一块拼图。

行业怎么看

认同方的理由很直接:OpenAI、Anthropic、Scale AI、Surge AI、Appen 每年向数据标注链条支付数十亿美元,规模与传统 BPO(业务流程外包)行业相当。差别在于 BPO 有清楚的甲方乙方与员工数披露,AI 数据标注的工时、单价、人员分布基本不公开。

反驳的声音同样不小。第一种观点认为,人工成本在顶尖大模型总训练成本里占比不到 5%,把焦点放在数据工人的劳动条件上有避重就轻之嫌,真正吞噬资源的是算力和电力。第二种反驳更尖锐:媒体反复讲「数据工人的低薪」,反而让公众以为「AI 还是需要人」,转移了白领替代这个更大的议题。

我们编辑部的判断是这两件事不矛盾。短期看,数据标注是有上百万从业者的真实产业,劳动条件值得讨论;中期看,随着模型自监督学习(Self-Supervised Learning,让模型从海量文本里自己找规律)能力增强,纯标注需求会下降,但「人对模型行为的实时反馈」不会消失——只是工种会迁移。

对普通人的影响

  • 对企业 IT:采购 GPT、Claude 类企业服务时,供应商「数据来源合规」页面值得认真读,而不是当营销文案过一遍。
  • 对个人职场:白领岗位不被直接威胁,但当公司讨论「用 AI 替代某个岗位」时,可以多问一句——被替代的人接下来去哪里。
  • 对消费市场:消费者暂时无感,但 AI 客服答非所问、内容审核误判、推荐明显跑偏,背后都可能有人工审核环节被压缩的直接代价。