Between K2 and K2.5, Moonshot poured roughly 20-25 trillion tokens into continual learning (training the model's parameters on new data after launch rather than freezing them post-deployment)—a figure that approaches the total pre-training scale of some smaller models from scratch. This week, a Reddit r/LocalLLaMA user posted asking what level K3.5 will reach. What's worth watching isn't the model itself, but which training paradigm Chinese large-model companies are betting on.

What this is

Kimi is Moonshot AI's flagship model line. K2 already shows strong code and long-context capabilities; K2.5 pushed performance further via large-scale continual learning. K3.5 still has no official release info—what we're seeing on community channels is essentially expectation management.

Industry view

The bullish camp argues that continual learning is the key variable separating the Kimi series from competitors, and that this path could lock Moonshot into a differentiated position within China's large-model tier.

But we'd flag three risks. First, r/LocalLLaMA users skew toward local-inference and open-source enthusiasts—their hype can't be directly mapped to actual enterprise demand. Second, once training tokens stack into the 20T range, marginal returns are already diminishing; more tokens doesn't necessarily mean stronger capability—what matters is return per unit of compute cost. Third, Moonshot still hasn't disclosed training specs or a timeline for K3.5, so expectation management itself is a marketing move. By contrast, DeepSeek is pursuing a "low-cost pre-training + open source" path—a fundamentally different strategy. The industry landscape is far from settled.

Impact on regular people

For enterprise IT: If K3.5 really does continue to improve in code or long-document scenarios, domestic API selection gains another comparison point, expanding procurement bargaining power.

For individual professionals: Potential upgrades for writing, coding, and research AI assistants—but in the short term, everyday users will most likely notice changes first on the Kimi web and app.

For the consumer market: The continual learning path means faster model iteration. Consumer users will face more frequent "suddenly smarter" or "suddenly dumber" shifts, requiring rebuilt usage habits.