返回首页

对比阅读

对比阅读:OpenAI and Anthropic Hit the Brakes — Safety Upgrades Pause Release Cadence 与 OpenAI 和 Anthropic 同时踩刹车 — 安全升级正在推迟发新节奏

AEN
OpenAIAnthropicAstra·

OpenAI and Anthropic Hit the Brakes — Safety Upgrades Pause Release Cadence

OpenAI's probability of releasing a new model in August has dropped from baseline to 13% — and this week, two facts point to a single judgment: AI frontier firms are no longer just talking about safety, they're putting real money on the brakes. OpenAI paused roughly two weeks of frontier reinforcement learning training, while Anthropic simultaneously restricted release of its most advanced model, Mythos. OpenAI safety lead Mia Glaese publicly framed it: a return to training remains "quite far off."

What this is

OpenAI's Preparedness Framework received a major upgrade this week: "critical cybersecurity capabilities" have been moved from a pre-deployment evaluation metric to a hard constraint during the development process. Critical cybersecurity capabilities, in plain terms, mean a model that can develop zero-day exploits (previously undisclosed security vulnerabilities) on hardened systems or execute end-to-end attacks without human intervention. After OpenAI's internally code-named Astra model crossed this threshold, a significant portion of Astra's workloads had to migrate to higher security standards before continuing. Anthropic moved in lockstep, prioritizing cybersecurity patches before considering public release.

Industry view

Supporters argue this is the responsible call — capability advancement and safety controls are different workloads on the same infrastructure, and post-deployment remediation is far more expensive. Anthropic flagged the possibility of "recursive self-improvement" (AI iteratively upgrading itself) back in June and advocated proactively slowing down.

But VC David Sacks publicly pushed back sharply: in practice, review mechanisms raise the industry bar. Only well-capitalized frontier firms can meet the scrutiny, compliance, and safety demands; open-source models scattered globally can't be uniformly regulated. The likely outcome is not a safer field — it's a handful of companies gaining larger advantage. Worth noting: Anthropic is now valued near $1 trillion and has filed IPO paperwork, so the "opportunity cost" of voluntarily slowing down is non-zero — competitors haven't stopped.

Impact on regular people

For enterprise IT: access windows to frontier models may slip. Projects expecting new capabilities in Q3 need to re-plan timelines.

For individual careers: AI tooling iteration pace is slowing in the short term. Workflow improvements that depend on new model rollouts can breathe a little.

For consumer markets: if "safety as moat" logic hardens among frontier firms, alternatives from smaller vendors and the open-source ecosystem may get further marginalized.

来源: juejin.cn
BZH
OpenAIAnthropicAstra·

OpenAI 和 Anthropic 同时踩刹车 — 安全升级正在推迟发新节奏

OpenAI 8 月内发新模型的概率已从常态跌至 13%——本周两个事实同时指向一个判断:AI 头部公司不再是嘴上谈安全,而是真金白银地踩刹车。OpenAI 暂停了约两周的前沿强化学习训练,Anthropic 同步限制最先进模型 Mythos 的发布。安全负责人 Mia Glaese 公开定调:距离恢复"相当遥远"。

这是什么

OpenAI 的 Preparedness Framework(安全准备框架)本周迎来一次重要升级:把"关键网络安全能力"从部署后评估的前置指标,改为开发过程中的硬约束。所谓关键网络安全能力,指模型无需人工介入就能在加固系统上开发零日漏洞(即未公开的安全漏洞)或执行端到端攻击。OpenAI 内部代号 Astra 的模型触及这一阈值后,涉及 Astra 的相当一部分工作负载须迁移至更高安全标准后方可继续。Anthropic 同步行动,优先完善网络安全补丁后再考虑公开发布。

行业怎么看

支持方认为这是必要的负责任做法——能力推进与安全管控本就是同一套基础设施上的不同工作负载,等到部署后再补救代价更高。Anthropic 在 6 月文章中已预警"递归式自我改进"(AI 自行迭代升级)的可能性,主张主动放慢。

但风险投资人 David Sacks 公开提出尖锐反对:核查机制的实际效果是抬高行业门槛,能满足审查、合规、安全要求的只有资金雄厚的头部公司,开源模型分散在全球难以被统一监管。最终结果可能不是更安全,而是少数公司获得更大优势。值得留意的是,Anthropic 估值已接近 1 万亿美元并已提交 IPO 文件,主动放慢的"机会成本"并不为零——竞争对手并未停下。

对普通人的影响

对企业 IT:涉及前沿模型的接入窗口可能延后,原本预期 Q3 上线的新能力项目需要重排时间表。

对个人职场:AI 工具迭代节奏短期放缓,依赖新模型上线的工作流改进可以暂时缓口气。

对消费市场:头部公司"安全即护城河"的逻辑若固化,中小厂商与开源生态的可用替代品可能被进一步边缘化。

来源: juejin.cn