What This Is

Behind a single user query, four steps may run (retrieval, model, tool, regeneration)—but traditional logs only record the final second. This is the pain point the engineering community keeps debating: AI applications are "unobservable," and when things go wrong, you can only rely on user descriptions and engineer guesses.

Failure happens in three places. First, AI outputs are probabilistic; traditional logs rely on "expected vs. actual" comparison, but AI has no "expected." Second, multiple steps run behind one query, but a single log line only shows the final second. Third, answer quality depends heavily on "what the model saw at the time" (which documents were retrieved, which system prompt version was used)—looking at the final output alone won't tell you why it answered wrong.

The fix is unglamorous: abandon print-style free-text logs and switch to structured logs (key=value format, e.g., event=llm_call in_tokens=69), and assign each request a unique request_id (like an order number that threads through every step of the request) across all components. Grep that ID later and you can reconstruct the request's latency, token count, which step failed, and what context the model had at the time.

The author calls this "the foundation of AI application observability"—no matter how flashy the tracing platform, it's just icing on the cake; if the foundation isn't laid, going live means a black box.

Industry View

Those who agree say this hits the industry's pain point. Several AI engineering team leads echoed it on social platforms: many companies spend millions building Agents (programs that let AI autonomously complete multi-step tasks, like booking flights or writing weekly reports), then go blank when users report issues, because "there's no way to prove whether AI's answer this time was correct." Structured logging may look crude, but it feeds directly into alerts, dashboards, and regression tests—it's the threshold from "runs" to "manageable."

The dissenting view deserves attention. First, structured logs aren't a silver bullet—they tell you "what happened" but not "why"; root-causing model behavior drift requires evaluation sets and human annotation working together. Second, many AI projects in China are still stuck at the PoC (proof of concept) stage; business teams have zero incentive to do refined observability because there are no real users yet—talking about observability is "premature optimization." Third, request_id looks simple, but stitching it across multiple models and services (attaching the same ID at every hop) is heavy engineering. The author themselves gripes that OpenAI SDK's log noise can drown out business logs.

Impact on Regular People

For enterprise IT: when you've sunk hundreds of thousands into a deployed intelligent customer service bot, knowledge base, or exam-generation system, and the boss asks "why are user complaints rising?" the tech team has no answer—this is exactly what structured logging solves. Observability will eventually become a hard requirement in AI product procurement.

For individual careers: roles tied to AI applications will fork off from "prompt tuning" into new positions like "AI Ops / Observability." People who understand logs, tracing, and cost analysis will be scarcer than those who simply know how to use ChatGPT.

For consumer markets: users will increasingly feel the stability gap between AI products. With identical features, the product that can tell users "why it answered wrong" will win; the one that only says "please retry" when something breaks will lose—transparency beats cleverness.