A final training compute cost of $5.6 million is enough to support one conclusion: competition in large models is no longer just about “whoever has more money wins.” It has entered a phase where the real contest is “who can reuse capabilities more effectively.” The controversy around DeepSeek has pushed “distillation” to center stage. This is not a minor trick. It is an important method that could reshape the industry’s cost structure.

What this is

Distillation—having a smaller model learn the “knowledge structure” embedded in a larger model’s outputs—can be understood like this: the teacher solves the problem first, and the student does not just copy the answer, but also learns how the teacher makes judgments. Unlike traditional training that feeds only standard answers, distillation places more value on the probabilities, preferences, and reasoning traces produced by a large model. These “soft labels” are often more valuable than the final result itself.

Its significance is highly practical: if a student model can absorb 70% to 80% of the capability while being cheaper to run and lighter to deploy, companies can put AI into real business workflows at much lower cost.

Industry view

Supporters argue that distillation is an unavoidable step in the industrialization of AI. Not every company can afford to maintain the strongest foundation model, but many can “compress capabilities” from a strong model and build smaller models better suited to customer service, office productivity, search, and end devices.

The objections are equally clear. First, distillation could make it harder for leading model developers to earn back their investment, weakening incentives for original innovation. Second, if the training data comes from outputs of closed-source models, the boundaries around compliance and intellectual property become blurred. In other words, the more effective distillation becomes, the more urgently the industry must answer a prior question: which kinds of knowledge can be learned, and which cannot?

Impact on regular people

For enterprise IT: budgets will shift from “buying the biggest model” to “buying a deployable model stack.” Smaller, cheaper, privately deployable options will become markedly more attractive.

For individual professionals: the AI reaching many jobs will become faster, cheaper, and closer to specific tasks. We need to focus more on how to connect models to workflows, not just on whether we know how to write prompts.

For the consumer market: local AI experiences on phones, PCs, and in-car systems may spread faster. The capabilities may not be the strongest, but response speed, privacy, and price will all be more competitive than “always calling a cloud-based large model.”