返回首页

对比阅读

对比阅读:AI Agent Gets It Right but Doubles the Cost — NVIDIA Exposes the Trail 与 AI Agent 答对题却悄悄多花一倍钱 — NVIDIA 想让企业看清它怎么跑的

AEN
NVIDIANeMo RelayHermes Agent·

AI Agent Gets It Right but Doubles the Cost — NVIDIA Exposes the Trail

Last week, NVIDIA's NeMo Relay tutorial, released September 30, made one point clear: an AI Agent getting the answer right doesn't mean it took the right path. A customer-service bot might search five times, read the same contract three times, or silently switch to a more expensive model after a failure. We've noticed this is the most easily missed source of cost when AI projects move from demo to production.

What This Is

NVIDIA calls this approach ATOF (Agent Trajectory Observability Format). At its core, it logs every Agent tool call—searches, file reads, API calls—with start and end events, unique IDs, and parent-child relationships. The validator checks "is the answer right"; the trajectory records "what actually happened." The two must be viewed separately—they cannot replace each other. The tutorial offers a counterintuitive case: the task completes, the answer is correct, yet costs suddenly double. The cause is usually the model issuing duplicate commands, or the orchestrator (the middleware that schedules tools) auto-retrying. Scanning logs in chronological order can't tell you whose fault it is.

How the Industry Sees It

Supporters see this as a course AI engineering must take. Open standards like OpenTelemetry and OpenInference are mature; enterprises will eventually need to choose a monitoring stack for Agents, just as they once chose APM (Application Performance Monitoring) tools. But the dissent is sharp: the tutorial is still written for developers, with no cost dashboards for finance or management. One architect commented: "The boss sees the bill, not the logs." A second risk: the more granular the trajectory, the higher the storage and compliance costs. In sensitive businesses, plaintext call parameters in logs could trigger cross-border data review.

Impact on Regular People

For enterprise IT: within the next year, Agent project sign-off sheets will likely add one more metric—observability. This isn't ceremony; it's cost-control infrastructure. For individual careers: engineers who can read Agent trajectories will become scarce, but even scarcer are people who can translate technical language into business questions—say, explaining why this month's bill jumped 40%. For consumer markets: consumers won't notice differences in the short term, but when an Agent fails, the difference between a 5-minute and a 5-hour diagnosis will directly shape the after-sales experience.

来源: juejin.cn
BZH
NVIDIANeMo RelayHermes Agent·

AI Agent 答对题却悄悄多花一倍钱 — NVIDIA 想让企业看清它怎么跑的

上周 NVIDIA 在 9 月 30 日发布的 NeMo Relay 教程里讲了一件事:AI Agent 答对题,不代表它走对路。一个智能客服可能搜了五次资料、把同一份合同读了三遍,或者在失败后悄悄换了一条调用更贵大模型的路径。我们注意到,这是 AI 项目从 demo 走向生产时最容易被忽视的成本来源。

这是什么

NVIDIA 把这套方法叫做 ATOF(Agent 轨迹可观测格式)。本质上,是把 Agent 每一次工具调用——搜索、读文件、调 API(应用程序接口)——按开始、结束事件记录在日志里,配上唯一编号和父子关系。验证器只管「答案对不对」,轨迹只管「实际发生了什么」,两件事分开看,不能互相替代。教程里给了一个反直觉的案例:任务最终完成、答案正确,但成本突然翻倍。原因往往是模型重复下令,或者编排器(负责调度工具的中间层)自动重试,按时间顺序扫日志分不出是谁的锅。

行业怎么看

支持方认为这是 AI 工程化必须补的一课。OpenTelemetry、OpenInference 这类开源链路标准已经成熟,企业迟早要像当年选 APM(应用性能监控)工具那样,给 Agent 选一套监控栈。但反对声音也很明确:这套教程目前仍是写给开发者的,缺少给财务和管理层的成本看板。一位架构师在评论区写道:「老板看到的是账单,不是日志。」另一层风险是,轨迹越细,存储与合规成本越高,敏感业务里调用参数含明文日志还可能触发数据出境审查。

对普通人的影响

对企业 IT:接下来一年,Agent 项目验收单上很可能多一项指标——可观测性。这不是仪式,是成本控制的基础设施。对个人职场:会读 Agent 轨迹的工程师会变得稀缺,但更稀缺的是能把技术语言翻译成业务问题的人——比如能说清为什么这个月账单涨了 40%。对消费市场:消费者短期内感觉不到差异,但当 Agent 出错时,企业能在 5 分钟而不是 5 小时定位问题,这会直接决定售后体验。
来源: juejin.cn