What this is

A hot post on Reddit's machine learning subreddit this week raised a key question—when will large models be able to train the next generation themselves? We noticed that in 2024-2025, this is no longer a thought experiment but what OpenAI, DeepSeek, and Sakana—the three leading labs—are quietly doing: letting AI train AI.

Specifics: In training OpenAI o1/o3, models are used to generate reasoning chains (text records of the inference process) as training data; China's DeepSeek R1 takes a similar path, using model-generated thought processes to replace some expensive human annotation; Japan's Sakana AI launched "AI Scientist," claiming to automate the entire machine learning research workflow at roughly $15 per run.

Key boundary: At present, this is still "AI-assisted human training of AI," not true bootstrapping (where a model autonomously iterates the next generation without human involvement). Sakana still requires human scientist review, and OpenAI has not publicly acknowledged fully departing from human feedback.

Industry view

Optimists view this as the necessary path to scaling AI research. Manual annotation is expensive and slow; using AI-generated training data can compress model iteration cycles from months to weeks. Sakana founder David Ha publicly called it an embryonic form of "research automation."

Dissent is equally forceful. A 2024 Nature study warned that training on purely synthetic data causes "model collapse"—models progressively lose coverage of long-tail content (low-frequency but real information), and outputs increasingly resemble "AI wrapped in AI." Anthropic founder Dario Amodei has publicly voiced similar concerns, arguing that without human value anchoring (aligning AI outputs with human values), the risks of this path are not yet fully understood.

Impact on regular people

For enterprise IT: In the next 12-18 months, when enterprises procure AI models, "training data provenance" will become a new due diligence item; the reliability of models trained on fully synthetic data will need independent verification.

For individual careers: Another tool iteration. Those who first understand the "AI trains AI" value chain will have an information-asymmetry advantage when evaluating AI products and making technical decisions.

For the consumer market: Declining underlying costs will pass through to consumer-facing AI application pricing; in the short term, "AI content flooding" is the more immediate user-experience problem, making it harder to distinguish authentic from synthetic information.