One news item worth our attention this week: Zhipu disclosed launching an "outer RSI loop" (Recursive Self-Improvement). RSI is shorthand for letting one AI continuously improve another AI's training process—in plain terms, "letting the model teach itself." This is the first time a Chinese major large-model company has publicly attempted this path.
What this is
Traditional large-model training runs on human teams: labeling data, tuning parameters, running experiments. RSI's premise is to hand intermediate steps over to another AI model, letting the latter auto-generate training data, design training recipes, and even reshape network architectures. What Zhipu is shipping is the "outer" variant—improvement happens outside the training loop (e.g., producing better training data), rather than letting the model modify its own internal weights.
Both OpenAI and Anthropic research this internally, but neither has publicly shipped it. By our read, Zhipu's move is relatively aggressive.
Industry view
Supporters frame this as a critical path around the "training data exhaustion" bottleneck—human text data grows slowly, but AI-generated high-quality data can scale exponentially. Chinese firms like Zhipu also face a hard practical constraint: compute capacity is limited, so they must lean on algorithmic efficiency to catch up with better-resourced competitors.
Opposition is just as sharp. The AI safety community's core concern with RSI is "goal drift": a self-improving system can, after multiple iterations, drift away from its originally set goals—and such drift is hard for humans to detect. Anthropic's safety team has long warned that pushing RSI forward without adequate "interpretability" (tools that let us see why a model makes a given decision) amounts to building another black box inside a black box. Some researchers note that outer-loop RSI is less risky than inner-loop RSI, but it still demands rigorous red-teaming and human-in-the-loop oversight.
Impact on regular people
For enterprise IT: Model iteration speed will accelerate further, meaning AI tool procurement decision windows must shrink—tools bought six months ago may be dwarfed by newer versions six months from now.
For individual careers: "AI improving AI" sounds scary, but execution still requires a large bench of human engineers for oversight and alignment. It won't strip away jobs overnight, but the job mix will keep tilting toward "supervisors" and "evaluators."
For the consumer market: The most direct effect is a generation leap in intelligent assistant products—the customer service bot you used this year may be a different beast next year.