Back to home

Compare

Comparing: Zhipu Lets AI Improve Itself — First Chinese Major LLM to Go Public With RSI & 智谱让 AI 开始自己改自己 — 中国大模型公司首次押注'自我迭代'路线

AEN
ZhipuRSISelf-Iteration·

Zhipu Lets AI Improve Itself — First Chinese Major LLM to Go Public With RSI

One news item worth our attention this week: Zhipu disclosed launching an "outer RSI loop" (Recursive Self-Improvement). RSI is shorthand for letting one AI continuously improve another AI's training process—in plain terms, "letting the model teach itself." This is the first time a Chinese major large-model company has publicly attempted this path.

What this is

Traditional large-model training runs on human teams: labeling data, tuning parameters, running experiments. RSI's premise is to hand intermediate steps over to another AI model, letting the latter auto-generate training data, design training recipes, and even reshape network architectures. What Zhipu is shipping is the "outer" variant—improvement happens outside the training loop (e.g., producing better training data), rather than letting the model modify its own internal weights.

Both OpenAI and Anthropic research this internally, but neither has publicly shipped it. By our read, Zhipu's move is relatively aggressive.

Industry view

Supporters frame this as a critical path around the "training data exhaustion" bottleneck—human text data grows slowly, but AI-generated high-quality data can scale exponentially. Chinese firms like Zhipu also face a hard practical constraint: compute capacity is limited, so they must lean on algorithmic efficiency to catch up with better-resourced competitors.

Opposition is just as sharp. The AI safety community's core concern with RSI is "goal drift": a self-improving system can, after multiple iterations, drift away from its originally set goals—and such drift is hard for humans to detect. Anthropic's safety team has long warned that pushing RSI forward without adequate "interpretability" (tools that let us see why a model makes a given decision) amounts to building another black box inside a black box. Some researchers note that outer-loop RSI is less risky than inner-loop RSI, but it still demands rigorous red-teaming and human-in-the-loop oversight.

Impact on regular people

For enterprise IT: Model iteration speed will accelerate further, meaning AI tool procurement decision windows must shrink—tools bought six months ago may be dwarfed by newer versions six months from now.

For individual careers: "AI improving AI" sounds scary, but execution still requires a large bench of human engineers for oversight and alignment. It won't strip away jobs overnight, but the job mix will keep tilting toward "supervisors" and "evaluators."

For the consumer market: The most direct effect is a generation leap in intelligent assistant products—the customer service bot you used this year may be a different beast next year.

BZH
智谱ZhipuRSI·

智谱让 AI 开始自己改自己 — 中国大模型公司首次押注'自我迭代'路线

本周值得关心的一条新闻:智谱(Zhipu)披露启动了「外部自我改进循环」(outer RSI loop)。RSI 是 Recursive Self-Improvement 的缩写,指让一个 AI 持续改进另一个 AI 的训练过程——简单说,就是「让模型自己教自己」。这是中国大模型公司首次公开尝试这条路径。

这是什么

传统大模型的训练靠人类团队:标注数据、调整参数、跑实验。RSI 的设想是,把中间某些环节交给另一个 AI 模型,让后者自动生成训练数据、设计训练方案、甚至改造网络结构。智谱这次做的是「outer」版本——改进发生在训练循环的外部(比如生成更好的训练数据),而不是让模型直接改自己的内部参数。

这条路 OpenAI 和 Anthropic 都在内部研究,但都未公开落地。智谱这一步相对激进。

行业怎么看

支持方认为,这是绕开「训练数据耗尽」瓶颈的关键路径——人类文本数据增长有限,但 AI 生成的高质量数据可以指数级扩展。智谱等中国公司还有现实压力:算力受限,必须靠算法效率追赶算力更强的对手。

反对意见同样尖锐。AI 安全圈对 RSI 的核心担忧是「目标漂移」:一个自我改进的系统,可能在多次迭代后偏离最初设定的目标,且这种偏离难以被人类察觉。Anthropic 的安全团队长期警告,没有充分的「可解释性」(也就是能看懂模型为什么这么决策的工具)之前贸然推进 RSI,等于在黑箱里再造一个黑箱。也有研究者指出,外部循环的 RSI 风险低于内部循环,但依然需要严格的红队测试和人类监督环节。

对普通人的影响

对企业 IT:模型迭代速度会进一步加快,意味着企业采购 AI 工具时,决策窗口要缩短——半年前采购的工具,半年后可能被新版本大幅超越。

对个人职场:「AI 自己改进 AI」听着吓人,但实际执行依然需要大量人类工程师做监督、做对齐。不会立刻砸掉饭碗,但岗位结构会持续向「监督者」和「评估者」倾斜。

对消费市场:最直接的影响是智能助手类产品会快速跳级——今年用过的智能客服,明年可能是另一个东西。