Back to home

Compare

Comparing: Agent Projects Don't Fail Because Models Are Dumb — They Fail After HTTP 200 & Agent 项目落地难不能怪模型笨 — 一次失败链排查显示,HTTP 200 之后崩的可能才刚开始

AEN
AgentObservabilityEngineering Practice·

Agent Projects Don't Fail Because Models Are Dumb — They Fail After HTTP 200

This week, an engineer recapped a fault investigation of an Agent (an AI program that can autonomously complete multi-step tasks) project on Juejin: originally, all failures were compressed into a single "PROVIDER_UNAVAILABLE" (service unavailable) message. After unpacking, however, the team found that from model invocation to content delivery, at least seven or eight intermediate stages could each fail independently—with most failures occurring after HTTP 200. This is worth caring about: when we debate why Agents don't run, the problem may never have been with the model.

What this is

The author's team's core change was introducing a "terminalBlocker" field (the last confirmed blocker) to replace the traditional "root cause" concept. The distinction: root cause implies the system knows where the problem originated, while "last blocker" only promises one thing—based on currently confirmed facts, the explicit stage at which the task finally stalled.

Two supporting practices came with it: first, record failures by stage (model invocation, stream assembly, parsing, validation, tool execution, publish validation, submission, delivery), so that downstream errors stop being dumped entirely on upstream; second, separate diagnostic records from content body—don't replicate user questions or model outputs, retain only stage status and timeline.

Industry view

The supporting voice reads this as a sign of Agent engineering maturing—from "can it run at all" to "can it run stably"—and predicts observability will eventually become a standard component.

The opposing voices deserve equal caution: this practice comes from a single team's internal refactor, not industry consensus; the standard hasn't taken shape yet. The more realistic risk: most companies will only commit to this work after being educated by failure—by which time a round of bills has already been paid. Other voices point out that "diagnostic-content separation" sounds simple, but when Agents involve compliance, auditing, and cross-team collaboration, not storing the body content may not actually be feasible.

Impact on regular people

For enterprise IT: if you're procuring or building Agents in-house, don't just ask "what's the success rate"—press on "which step fails most." That's the real metric for judging vendor honesty.

For individual professionals: when an AI tool fails, don't rush to switch tools or models. Try locating whether the problem is unclear input, the model not understanding, or wrong output format—it saves considerable communication overhead.

For the consumer market: what users perceive as "AI acting up again" is likely an engineering pipeline problem, not the model itself—and this should recalibrate our real expectations for product stability.

Source: juejin.cn
BZH
Agent可观测性工程实践·

Agent 项目落地难不能怪模型笨 — 一次失败链排查显示,HTTP 200 之后崩的可能才刚开始

一位工程师本周在掘金复盘了 Agent(能自主完成多步任务的 AI 程序)项目的故障排查:原本所有失败都被压成一句"PROVIDER_UNAVAILABLE"(服务不可用),但拆开后才发现,从模型调用到内容交付,中间至少七八个环节都可能单独崩,其中多数发生在 HTTP 200 之后。这件事值得关心:讨论 Agent 为什么跑不通,问题可能从来不在模型。

这是什么

作者团队的核心改动是引入"terminalBlocker"(最后一个确定的阻塞点)这个字段,替代传统的"根因"概念。区别在于:根因暗示系统知道问题从哪里来,而"最后阻塞点"只承诺一件事——根据当前已确认的事实,任务最后明确停在了哪一步。

配套的两点实践:一是按阶段记录失败(模型调用、流组装、解析、校验、工具执行、发布校验、提交、交付),避免下游错误全被甩给上游;二是把诊断记录和内容正文分开,不复制用户问题和模型输出,只留阶段状态与时间线。

行业怎么看

支持的声音认为,这是 Agent 工程走向成熟的标志——从"能不能跑起来"转向"能不能稳定跑",可观测性迟早会成为标准件。

反对的声音同样值得警惕:这套实践来自单一团队的内部改造,并非行业共识,标准尚未成形;更现实的风险是,多数公司会在被失败教育之后才愿意做这件事,届时账单已经付完一轮。还有观点指出,"诊断与内容分离"说起来简单,但当 Agent 涉及合规、审计、跨团队协作时,不存正文未必可行。

对普通人的影响

对企业 IT:如果在采购或自建 Agent,别只问"成功率多少",要追问"哪一步失败最多"——这才是判断供应商诚实的指标。

对个人职场:用 AI 工具遇到失败时,先别急着换工具或换模型,试着定位是输入不清楚、模型没懂、还是输出格式不对——能省下不少沟通成本。

对消费市场:用户感知到的"AI 又抽风了",背后很可能是工程链路的问题,不是模型本身——这会影响我们对产品稳定性的真实预期。

Source: juejin.cn