What this is

A source-code teardown on Juejin surfaces an awkward fact: inside OpenAI's Codex coding assistant, two functions both labeled "review" take in and return completely different data — exposing the real cost of AI Agents (programs that can autonomously complete multi-step tasks): easy to plug in, hard to swap out.

Specifically: the standard review goes through the review/start endpoint and returns plain text; the "con" review goes through turn/start and expects structured JSON (with fields like verdict and findings). The two paths also define "done" differently — one checks whether the current Turn has finished running, the other additionally requires the verdict to be approve. "Execution complete" and "review passed" are two different things, yet the UI likely renders both as a green checkmark.

This is the hidden cost of the protocol layer (the input/output formats and event semantics agreed upon between different Agents) — not whether you can plug it in, but whether you can swap it out.

Industry view

The pro-standardization camp argues that with Agent interfaces scattered across vendors, this is precisely the moment for Anthropic, Google and others to push MCP (Model Context Protocol, a standard that lets models uniformly invoke external tools) — protocol uniformity is the foundation of multi-Agent collaboration.

But the opposing view is worth more of our attention. One architect wrote in the comments: "If the existing executor fully meets the requirement, there's no need to build a general-purpose orchestration platform just to enable 'multi-Agent.'" Multi-Agent is the outcome, not the starting point; building a platform for "future swap-ability" upfront is usually over-engineering.

The more practical risk is semantic mismatch: a plugin treats the completed status value of 0 as "success" and the UI paints a green check — but that only means the current turn finished running, not that the code is mergeable. This kind of silent bug is the hardest to track down in production.

Impact on regular people

For enterprise IT: when evaluating an Agent, don't just watch the demo run end-to-end. First ask what inputs it accepts and whether the output can be consumed directly by downstream systems. An "Agent middleware platform" doesn't work the moment you install it.

For individual professionals: when an AI tool reports "success" but the result is wrong, more often than not it's not a model capability issue — it's a task-definition misalignment. Before you switch tools, write down clearly what you actually want it to do.

For consumer markets: AI products that tout "multi-Agent collaboration" are, for now, most likely just stitching different models into the same UI; their effects aren't directly comparable. When you see this kind of marketing, dial your expectations down a notch.