返回首页

对比阅读

对比阅读:DeepSeek’s $5.6M Training Cost Pushes LLM Competition Into an Efficiency Race 与 560万美元训练成本引爆争议,蒸馏正把大模型竞争带入效率战

AEN
DeepSeekOpenAIDeepSeek-R1·

DeepSeek’s $5.6M Training Cost Pushes LLM Competition Into an Efficiency Race

A final training compute cost of $5.6 million is enough to support one conclusion: competition in large models is no longer just about “whoever has more money wins.” It has entered a phase where the real contest is “who can reuse capabilities more effectively.” The controversy around DeepSeek has pushed “distillation” to center stage. This is not a minor trick. It is an important method that could reshape the industry’s cost structure.

What this is

Distillation—having a smaller model learn the “knowledge structure” embedded in a larger model’s outputs—can be understood like this: the teacher solves the problem first, and the student does not just copy the answer, but also learns how the teacher makes judgments. Unlike traditional training that feeds only standard answers, distillation places more value on the probabilities, preferences, and reasoning traces produced by a large model. These “soft labels” are often more valuable than the final result itself.

Its significance is highly practical: if a student model can absorb 70% to 80% of the capability while being cheaper to run and lighter to deploy, companies can put AI into real business workflows at much lower cost.

Industry view

Supporters argue that distillation is an unavoidable step in the industrialization of AI. Not every company can afford to maintain the strongest foundation model, but many can “compress capabilities” from a strong model and build smaller models better suited to customer service, office productivity, search, and end devices.

The objections are equally clear. First, distillation could make it harder for leading model developers to earn back their investment, weakening incentives for original innovation. Second, if the training data comes from outputs of closed-source models, the boundaries around compliance and intellectual property become blurred. In other words, the more effective distillation becomes, the more urgently the industry must answer a prior question: which kinds of knowledge can be learned, and which cannot?

Impact on regular people

For enterprise IT: budgets will shift from “buying the biggest model” to “buying a deployable model stack.” Smaller, cheaper, privately deployable options will become markedly more attractive.

For individual professionals: the AI reaching many jobs will become faster, cheaper, and closer to specific tasks. We need to focus more on how to connect models to workflows, not just on whether we know how to write prompts.

For the consumer market: local AI experiences on phones, PCs, and in-car systems may spread faster. The capabilities may not be the strongest, but response speed, privacy, and price will all be more competitive than “always calling a cloud-based large model.”

来源: juejin.cn
BZH
DeepSeekOpenAIDeepSeek-R1·

560万美元训练成本引爆争议,蒸馏正把大模型竞争带入效率战

560万美元的最终训练算力成本,已经足够说明一个判断:大模型竞争不再只是“谁钱多谁赢”,而是进入“谁更会复用能力”的阶段。围绕 DeepSeek 的争议,把“蒸馏”推到台前;这不是小技巧,而是决定行业成本结构的重要方法。

这是什么

蒸馏(让小模型学习大模型输出中的“知识结构”)可以理解为:老师先做题,学生不只抄答案,还学习老师如何判断。和传统只喂标准答案不同,蒸馏更看重大模型给出的概率、偏好和推理痕迹,这些“软标签”往往比结果本身更有价值。

它的意义很现实:如果学生模型能学到七八成能力,但推理更便宜、部署更轻,企业就可能用更低成本把 AI 真正装进业务流程。

行业怎么看

支持者认为,蒸馏是 AI 工业化的必经之路。不是每家公司都养得起最强底座模型,但很多公司可以基于强模型“压缩能力”,做出更适合客服、办公、搜索和终端设备的小模型。

反对意见也很明确:第一,蒸馏可能让头部模型的投入更难回本,削弱原创动力;第二,如果训练数据来自闭源模型输出,合规和知识产权边界会变得模糊。换句话说,蒸馏越有效,行业越需要先回答“哪些知识可以学,哪些不能学”。

对普通人的影响

对企业 IT:预算会从“买最大模型”转向“买能落地的模型组合”。更小、更便宜、可私有化部署的方案,吸引力会明显上升。

对个人职场:很多岗位接触到的 AI,会变得更快、更便宜、更贴近具体任务。我们更需要关注“如何把模型接进流程”,而不只是会不会写提示词。

对消费市场:手机、电脑、车机上的本地 AI 体验可能更快普及。能力未必最强,但响应速度、隐私和价格,都会比“永远调用云端大模型”更有竞争力。

来源: juejin.cn