返回首页

对比阅读

对比阅读:AI Agent 'Pass' ≠ Success — Pretty Numbers Are Self-Deception 与 AI Agent 跑通不等于真成功 — 一位工程师反思:漂亮数字正在骗所有人

AEN
Agent EvaluationAI EngineeringJuejin·

AI Agent 'Pass' ≠ Success — Pretty Numbers Are Self-Deception

This week on Juejin we read a long post from an Agent evaluation engineer that posed a question that stopped us cold: the model returned content, the auto-judge assigned a score, the system status shows "completed" — but did the user actually see the full answer? His judgment is direct: HTTP 200 doesn't prove the answer is correct, and "completed" doesn't prove the user saw the answer.

What this is

He broke down a single Agent request into a long chain: user query → planning → tool calls → evidence gathering → candidate answer generation → quality check → streaming send → frontend display. Each link can independently succeed or fail, but industry practice collapses them all into a single green light: "success=true".

Three specific corrections:

First, don't compute a total score. "Good language quality" cannot offset "missing critical evidence" — weighted averages hide the problem.

Second, distinguish inconclusive (insufficient evidence) from failed (evidence proves the bar wasn't met). The follow-up actions for each are completely different; treating both as 0 lets infrastructure noise pollute model rankings.

Third, fix the denominator. When only 10 of 30 questions have actually been run, writing "9 out of 10 completed samples returned an answer" is honest; writing "90% success rate" is self-deception — the unprocessed samples tend to be the complex long-context ones most likely to expose problems.

What he cares most about is a fourth point: what the judge sees ≠ what the user receives. Evaluation should distinguish three layers — the model's candidate content, the server's actually-sent content, and the client's received evidence. Only the second layer can be reliably reconstructed; without a client-side ACK (acknowledgment), don't write "sent" as "user has seen it".

Industry view

The author's stance is clear: the most dangerous thing in evaluation isn't low model scores — it's collapsing different layers of "success" into a single green light, making reports prettier and problems harder to find. "The incomplete samples are the most valuable samples" — in industry terms, that line is a direct challenge to the credibility of currently published Agent leaderboards.

Pushback exists too. Fixed denominators, fixed code versions, fixed model versions make reports "look worse" — a luxury for startups racing to ship. Separating inconclusive from failed demands dedicated infrastructure investment that most Agent vendors currently can't afford. Judging by engineering maturity, China's Agent ecosystem still trails OpenAI and Google on this front.

Impact on regular people

For enterprise IT: when procuring AI Agents, don't just look at "95% success rate" — ask how the denominator is calculated, how incomplete samples are handled, and which specific layer "completed" refers to.

For individual professionals: when using AI Agents internally to automate workflows and results look off, don't just trust the "task complete" prompt — drill down and verify whether the intermediate steps actually ran.

For consumer markets: user-facing AI assistants and similar products still market themselves with crude metrics. Within the next year or two, whoever does the evaluation engineering first may claim the real moat.

来源: juejin.cn
BZH
Agent评测AI工程掘金·

AI Agent 跑通不等于真成功 — 一位工程师反思:漂亮数字正在骗所有人

这周我们在掘金读到一篇 Agent 评测工程师的长文,他提了一个让我们停下来的问题:模型返回了内容、自动裁判给了分、系统状态显示 completed——但用户真的看到了完整答案吗?他的判断很直接:HTTP 200 不能证明回答正确,completed 也不能证明用户看到了答案。

这是什么

他把一次 Agent 请求拆成一条长链路:用户提问 → 规划 → 工具调用 → 证据整理 → 生成候选答案 → 质量检查 → 流式发送 → 前端展示。这些环节各自独立成败,但行业惯例是把它们压成一个绿灯 "success=true"。

三个具体修正:

第一,别算总分。"语言质量好"不能抵消"关键证据缺失",加权平均会把问题藏起来。

第二,区分 inconclusive(证据不足)和 failed(证据充分证明没达标)。两者后续动作完全不同;混算 0 分会让模型排名混入基础设施噪音。

第三,固定分母。原本 30 题只跑完 10 题时,写"已完成样本中 9/10 返回了答案"是诚实,写"成功率 90%"就是自我欺骗——没跑完的往往正是复杂长上下文、最容易暴露问题的样本。

他最在意第四点:裁判看到的答案 ≠ 用户收到的答案。评测应区分三层——模型候选内容、服务端实际发送内容、客户端接收证据。能可靠重建的只有第二层;如果没有客户端 ACK(确认回执),就不要把"已发送"写成"用户已看到"。

行业怎么看

作者立场很明确:评测里最危险的不是模型分数低,而是把不同层级的"成功"压成一个 Green,让报告越来越漂亮、问题越来越难找。"没完成的样本才是最值钱的样本"——这话放到行业里,等于在质疑当前大量 Agent 公开榜单的可信度。

反对意见也有。固定分母、固定代码版本、固定模型版本,会让报表"变难看",对赶着上线的初创团队是一种奢侈。把 inconclusive 和 failed 分开,意味着评测团队需要专门的基础设施投入,而绝大多数 Agent 厂商目前没有这个能力。从工程成熟度看,国内 Agent 生态在这件事上和 OpenAI、Google 仍有差距。

对普通人的影响

对企业 IT:采购 AI Agent 时别只看"成功率 95%",问清楚分母怎么算、没完成的样本怎么处理、"已完成"具体指哪一层。

对个人职场:内部使用 AI Agent 自动化流程时,若结果异常,不要只信"任务完成"提示,往下追问中间环节是否真跑通。

对消费市场:面向用户的 AI 助手等产品目前仍以粗放指标宣传。一两年内,谁先把评测工程做扎实,谁可能就拿到真正的护城河。

来源: juejin.cn