This week, an engineer recapped a fault investigation of an Agent (an AI program that can autonomously complete multi-step tasks) project on Juejin: originally, all failures were compressed into a single "PROVIDER_UNAVAILABLE" (service unavailable) message. After unpacking, however, the team found that from model invocation to content delivery, at least seven or eight intermediate stages could each fail independently—with most failures occurring after HTTP 200. This is worth caring about: when we debate why Agents don't run, the problem may never have been with the model.
What this is
The author's team's core change was introducing a "terminalBlocker" field (the last confirmed blocker) to replace the traditional "root cause" concept. The distinction: root cause implies the system knows where the problem originated, while "last blocker" only promises one thing—based on currently confirmed facts, the explicit stage at which the task finally stalled.
Two supporting practices came with it: first, record failures by stage (model invocation, stream assembly, parsing, validation, tool execution, publish validation, submission, delivery), so that downstream errors stop being dumped entirely on upstream; second, separate diagnostic records from content body—don't replicate user questions or model outputs, retain only stage status and timeline.
Industry view
The supporting voice reads this as a sign of Agent engineering maturing—from "can it run at all" to "can it run stably"—and predicts observability will eventually become a standard component.
The opposing voices deserve equal caution: this practice comes from a single team's internal refactor, not industry consensus; the standard hasn't taken shape yet. The more realistic risk: most companies will only commit to this work after being educated by failure—by which time a round of bills has already been paid. Other voices point out that "diagnostic-content separation" sounds simple, but when Agents involve compliance, auditing, and cross-team collaboration, not storing the body content may not actually be feasible.
Impact on regular people
For enterprise IT: if you're procuring or building Agents in-house, don't just ask "what's the success rate"—press on "which step fails most." That's the real metric for judging vendor honesty.
For individual professionals: when an AI tool fails, don't rush to switch tools or models. Try locating whether the problem is unclear input, the model not understanding, or wrong output format—it saves considerable communication overhead.
For the consumer market: what users perceive as "AI acting up again" is likely an engineering pipeline problem, not the model itself—and this should recalibrate our real expectations for product stability.