A 40% error rate spike within hours of launch, customer service Agents randomly invoking tools, Coding Agents deleting the wrong files — this isn't a joke. It's the most typical Agent failure script plaguing enterprises today.

Agents are AI systems that autonomously call tools and complete tasks step by step. A technical deep-dive this week on Juejin nailed the root cause: it's not that models are too weak — evaluation systems haven't kept up. Traditional NLP tests score on "does the text look right," but Agent tests must examine "what path did it actually take, and is the final state correct" — a fundamental methodological shift from the chatbot era to the Agent era.

What This Is

The article's author, Wu Jiahao, proposes three new dimensions for Agent evaluation:

  • Final-state assertions: Regardless of what the AI claims, only check whether files were properly modified in the sandbox and database writes are correct.
  • Trajectory validity: Detect whether the Agent takes 10 rounds of meaningless retries — these "detour" behaviors.
  • Parameter precision: Check whether the tool parameters dispatched conform to business constraints.

He also includes an automated evaluation framework built on Python + LLM-as-a-Judge (using a large model as an automated judge), designed to let Agents run through CI/CD regression pipelines before launch.

Industry View

Supportive voices argue that evaluation standardization is a necessary condition for Agents to move past the demo phase. No quantification, no engineering — it's a sign of industry maturation.

The opposing view deserves more attention. A senior architect responded in a tech group: the very notion of "trajectory matching" is debatable — if only one correct path is acknowledged, it will strangle the Agent's emergent capabilities. And building a sandbox evaluation system requires serious capital: curating golden test cases, maintaining scripts, training judge models — most small and mid-sized teams can't afford it, and they'll eventually revert to the primitive "boss decides whether to ship" state.

Another measured view: the real bottleneck for Agent deployment may not be technical evaluation, but enterprise organization — how permissions are granted, how data is provisioned, who owns process changes. No evaluation framework can solve these problems.

Impact on Regular People

For enterprise IT: When selecting Agent vendors, don't just watch demos — ask explicitly, "Do you have a continuous evaluation system?" Otherwise, the blame for post-launch failures will likely fall on the procurement side.

For individual careers: The narrative that "AI will replace entry-level jobs next year" deserves more skepticism. If we can't even manage internal Agent quality — with files getting randomly deleted — talk of systematic replacement is still premature.

For consumer markets: In the short term, keep expectations low for smart customer service and AI assistants. They're more likely to show up as "fast to launch, bug-ridden in operation," with complaints and refunds exceeding expectations.