返回首页

对比阅读

对比阅读:After Agents Go Live: 100K Tokens Monthly, Nobody Knows Who's Spending 与 Agent 上线之后:数十万 Token 月账,没人说得清花在哪儿

AEN
AgentObservabilityOpenTelemetry·

After Agents Go Live: 100K Tokens Monthly, Nobody Knows Who's Spending

What this is

Once Agents are running in production, enterprises discover another real problem: a task takes 15 extra seconds and nobody can pinpoint why; monthly bills swell by hundreds of thousands of tokens and no department can explain who's burning them.

Traditional APM (Application Performance Monitoring) only watches HTTP status codes and response latency — it goes completely blind on Agent systems that think, call tools, and iterate across multiple rounds. The industry is now pulling in OpenTelemetry — the open-source distributed tracing standard originally built for microservices — to break every Agent execution into a Span tree: every reasoning round, every tool call, every prompt injection, every batch of token consumption, all stamped with timestamps and cost tags.

Put simply, Agent observability = a black box + a dashboard + a cost-allocation sheet for AI systems. Without it, an Agent inside an enterprise is a machine where "you can see the switch, but press it and have no idea what happens inside."

Industry view

Worth watching: this infrastructure layer is becoming its own business. Datadog, Grafana, Langfuse, Helicone — vendors across the stack are all racing for a position in Agent observability.

The optimists call this the unavoidable path from Agent PoC (proof of concept) to production: no observability, no scale — the same logic that once forced every microservice deployment onto APM.

But we also hear the other side: many enterprises haven't reached the Agent call volume that justifies fine-grained cost attribution, and this stack smells like over-engineering. One senior architect put it bluntly: "Get the business working first, then talk observability. Stacking tools too early is just technical debt."

There's a deeper concern lurking underneath: every prompt and every tool response gets logged and replayed, which means the most sensitive business data and decision logic of the enterprise gets fully exposed to the ops platform. Data compliance pressure spikes.

Impact on regular people

For enterprise IT: In the next 12–18 months, Agent ops platforms will likely become a standard procurement line for enterprise AI projects — expect a new "AI Observability" row on the budget sheet.

For individual careers: Engineers who understand distributed tracing and cost attribution will become more valuable; even ordinary business users may need to learn to read Agent consumption reports, the same way they read cloud resource bills today.

For the consumer market: If single Agent call costs can be precisely accounted for, AI product subscription pricing will get more transparent — pay-per-use and pay-per-scenario may replace flat-rate monthly subscriptions.

来源: juejin.cn
BZH
Agent可观测性OpenTelemetry·

Agent 上线之后:数十万 Token 月账,没人说得清花在哪儿

这是什么

Agent 跑起来之后,企业才发现另一个真问题:一次任务慢 15 秒不知道慢在哪,月度账单多出几十万 Token 不知道哪个部门刷的。

传统 APM(应用性能监控)只看 HTTP 状态码和响应延迟,对 Agent 这种会思考、会调用工具、会多轮迭代的系统完全失灵。业界正在搬出 OpenTelemetry——原本用于微服务的开源链路追踪标准,把每一次 Agent 执行拆成一棵 Span 树:每轮推理、每个工具调用、每段 Prompt 注入、每批 Token 消耗,都打上时间戳和成本标签。

简单说,Agent 可观测性 = 给 AI 系统装一套黑匣子 + 仪表盘 + 成本分摊表。没有它,Agent 在企业里就是一台「只看得到开关、按下去不知道里面发生了什么」的机器。

行业怎么看

值得关心的是,这套基础设施正在变成一门生意。Datadog、Grafana、Langfuse、Helicone 等国内外厂商都在抢 Agent Observability 的位置。

乐观派认为这是 Agent 从 PoC(概念验证)走向生产的必经之路,没有可观测性就没有规模化,和当年微服务必须上 APM 是一个道理。

但我们也听到另一种声音:很多企业 Agent 调用量根本没到「需要精细化成本归因」的程度,这套方案有过度工程化嫌疑。一位资深架构师直言:「先把业务跑通再谈可观测性,工具栈堆得太早反而是技术债。」

还有一层隐忧:所有 Prompt、工具回包都被记录和回放,等于把企业最核心的业务数据和决策逻辑全盘暴露给运维平台,数据合规压力陡增。

对普通人的影响

对企业 IT:未来 12-18 个月,Agent 运维平台很可能成为企业 AI 项目的标配采购项,预算单上会多出一行「AI Observability」。

对个人职场:懂链路追踪、成本归因的工程师会更值钱;普通业务人员也可能要学着看 Agent 消耗报表,就像今天看云资源账单一样。

对消费市场:Agent 单次调用成本如果能被精细核算,AI 产品订阅定价会越来越透明,按量付费、按场景付费可能取代一刀切包月。

来源: juejin.cn