What this is

DeepSeek's updated API pricing compresses "cache hit" pricing — where the AI re-reads content it has already processed — to an extremely low level: under one cent per million tokens. Non-cached input pricing also sits at roughly one-tenth that of mainstream overseas models.

For developers, this means that when high-frequency applications like Agents (AI systems that autonomously complete multi-step tasks) or RAG (Retrieval-Augmented Generation, where AI retrieves relevant documents before answering) repeatedly reuse the same context, the cost shifts from "unusable" to "use freely."

Industry view

Supporters see this as the starting point of a real AI application explosion. We agree the bottleneck for Agents has rarely been technical — it's been the cost per call. Low cache pricing gives enterprise AI assistants, knowledge-base Q&A, and other high-frequency scenarios a cost model that finally works.

But there are sober voices. One concern: DeepSeek's ultra-low pricing may not reflect true cost — over time, the vendor will either raise prices or sacrifice service quality. Another concern: once developers are locked into a specific pricing structure, switching costs rise. Neither Anthropic nor OpenAI has followed up with comparable moves, which suggests mainstream vendors don't see this as a track they need to immediately pursue.

Impact on regular people

For enterprise IT: If your company is evaluating AI projects — customer service bots, internal knowledge bases — solutions previously killed by budget deserve a fresh procurement review.

For professionals: People who regularly use AI to organize long documents or do industry research will find that "asking the same document dozens of times" no longer hurts. This is the most direct efficiency gain.

For consumer markets: More consumer-facing AI tools (AI tutors, legal/medical consultation apps) may quietly cut prices or expand usage limits over the next six months. We see consumers as the indirect beneficiaries of this price war.