We noticed a counterintuitive judgment: LLM applications in production (products wired into large models) may have weaker defenses than 2015-era microservices. The reason isn't complicated—most teams only run functional acceptance and latency stress tests before launch, but LLM applications have uncontrollable external interfaces and unpredictable outputs. Traditional testing simply can't reach the paths that actually break.
What this is
Chaos Engineering was clearly articulated through Netflix's "monkey experiments" in the 2010s: deliberately cause damage in production to see if the system breaks. Ported to LLM applications, this becomes "poisoning AI systems."
Three concrete steps. First, map the five categories of failures AI systems will inevitably hit—provider API outages, sudden format changes in model output, token (billed by length) overrun, latency avalanches, and poisoned retrieval results. Second, design an experiment matrix along "4 dimensions × 3 intensities," injecting failures progressively from light to heavy. Third, define steady-state metrics specific to LLMs: not "the service is still alive" but "users can still get valuable responses." The entire experiment can be implemented as a 200-line Python man-in-the-middle proxy—no need to buy enterprise tooling.
Industry view
Supportive voices are direct: LLM applications' failure path combinations are far more complex than traditional services, and unit tests plus stress tests can't cover them. A common post-mortem conclusion: fallback plans are written in code but have never been triggered; backup provider API keys expired three months ago and nobody noticed.
The opposing view deserves attention. One practitioner with ten years of SRE (Site Reliability Engineering) experience believes transplanting the full complexity of chaos engineering onto AI projects is over-engineering. "For the vast majority of small and mid-sized teams, daily request volume doesn't even reach 10,000. Solid monitoring and alerting first is more realistic than building experiment matrices." His judgment: chaos engineering is a tool for the scaling phase, not the startup phase.
There's another layer of risk easily overlooked—chaos experiments themselves are destructive. Injecting failures deliberately in production with imperfect isolation can manufacture real incidents.
Impact on regular people
For enterprise IT: companies deploying AI customer service, knowledge base Q&A, and similar applications will encounter nearly all the failure profiles described here. We recommend starting with monitoring "silent degradation"—users see no error but responses get worse. This metric is harder to set than error rate.
For individual careers: limited direct relevance to your daily work, but when enterprise AI assistants occasionally "get dumber" or answer off-topic, an unvalidated failure path is likely the culprit behind the scenes.
For consumer markets: next time you open an AI product and find it giving absurd answers, don't rush to criticize the product—it may be running on some degraded chain. Understanding this saves unnecessary frustration.