This week on Juejin we read a long post from an Agent evaluation engineer that posed a question that stopped us cold: the model returned content, the auto-judge assigned a score, the system status shows "completed" — but did the user actually see the full answer? His judgment is direct: HTTP 200 doesn't prove the answer is correct, and "completed" doesn't prove the user saw the answer.

What this is

He broke down a single Agent request into a long chain: user query → planning → tool calls → evidence gathering → candidate answer generation → quality check → streaming send → frontend display. Each link can independently succeed or fail, but industry practice collapses them all into a single green light: "success=true".

Three specific corrections:

First, don't compute a total score. "Good language quality" cannot offset "missing critical evidence" — weighted averages hide the problem.

Second, distinguish inconclusive (insufficient evidence) from failed (evidence proves the bar wasn't met). The follow-up actions for each are completely different; treating both as 0 lets infrastructure noise pollute model rankings.

Third, fix the denominator. When only 10 of 30 questions have actually been run, writing "9 out of 10 completed samples returned an answer" is honest; writing "90% success rate" is self-deception — the unprocessed samples tend to be the complex long-context ones most likely to expose problems.

What he cares most about is a fourth point: what the judge sees ≠ what the user receives. Evaluation should distinguish three layers — the model's candidate content, the server's actually-sent content, and the client's received evidence. Only the second layer can be reliably reconstructed; without a client-side ACK (acknowledgment), don't write "sent" as "user has seen it".

Industry view

The author's stance is clear: the most dangerous thing in evaluation isn't low model scores — it's collapsing different layers of "success" into a single green light, making reports prettier and problems harder to find. "The incomplete samples are the most valuable samples" — in industry terms, that line is a direct challenge to the credibility of currently published Agent leaderboards.

Pushback exists too. Fixed denominators, fixed code versions, fixed model versions make reports "look worse" — a luxury for startups racing to ship. Separating inconclusive from failed demands dedicated infrastructure investment that most Agent vendors currently can't afford. Judging by engineering maturity, China's Agent ecosystem still trails OpenAI and Google on this front.

Impact on regular people

For enterprise IT: when procuring AI Agents, don't just look at "95% success rate" — ask how the denominator is calculated, how incomplete samples are handled, and which specific layer "completed" refers to.

For individual professionals: when using AI Agents internally to automate workflows and results look off, don't just trust the "task complete" prompt — drill down and verify whether the intermediate steps actually ran.

For consumer markets: user-facing AI assistants and similar products still market themselves with crude metrics. Within the next year or two, whoever does the evaluation engineering first may claim the real moat.