An AI coding assistant crashes after 30 rounds of tasks — the real problem exposed isn't that models are too weak, but that engineering debt is going unpaid. A Chinese developer has just open-sourced a fix for this kind of "ambush" failure.

What This Is

The scenario: a long-running task reaches round 30, a tool returns a large block of results, and the next request gets rejected by the API (model call interface) with an "exceeded conversation length" error. His judgment: this kind of crash must be recovered by the Agent (an AI that autonomously completes multi-step tasks) itself, not bounced back to the user.

The fix is an "emergency escape hatch": identify → slim down → resend. Two details worth remembering: DeepSeek, OpenAI, and two other vendors each use different error wording for the same problem. Inconsistent identification wastes retry budget and triggers error degradation (output quality gets lowered when errors occur) — he consolidated the identification logic into one place. Emergency compression has two safeguards — a 60-second cooldown window plus a 5-time cap per session — keeping only the system prompt, one boundary placeholder, and the most recent 5 exchanges, with everything else restored from disk snapshots.

Industry View

The bullish case: we've noticed leading teams treating long-horizon tasks (multi-round Agents) as the real battleground for AI coding assistants. The model layer has been pushed to homogeneity by open-source competition (different vendors' models converge in performance) — the moat lives in engineering details. Whoever can stably run 100 rounds without crashing wins.

But we must flag the risk: this solution is essentially "building on sand." Error identification leans on keyword matching, conversation snapshots lean on the file system, compression timing leans on heuristic rules (rules decided by gut feel) — any failure in the chain can cause the Agent to hallucinate (output that sounds plausible but doesn't match facts) or lose critical context. The deeper problem: these patches are scattered across every developer who touches the issue, and nobody is building platform-level abstractions. Base models turn over every three months, so patches can失效 at any time — this isn't a technical problem, it's an ecosystem problem.

Impact on Regular People

For enterprise IT: when evaluating AI coding assistants, stability matters more than "can it write code" — a job that runs to completion is worth more than one that crashes halfway through.

For individual careers: for developers doing AI-related work, "engineering debt" is the new opportunity — anyone can tune the model layer, but error handling and context management — the dirty work — are where scarcity lives.

For consumer markets: don't expect AI assistants to reliably complete long tasks (booking flights, writing reports, running workflows) at this stage. They're good at short commands; faced with complex tasks, they'll suddenly "lose memory."