Back to home

Compare

Comparing: DeepSeek V4 at 22-28 Tokens/sec on a Mini PC — Local LLMs May Finally Be Usable & DeepSeek V4 在家用小主机跑到 22-28 字/秒,本地大模型这事可能真要开始能用了

AEN
DeepSeekAMDStrix Halo·

DeepSeek V4 at 22-28 Tokens/sec on a Mini PC — Local LLMs May Finally Be Usable

A Reddit tech user released benchmark data this week: DeepSeek's latest V4 Flash model ran at 22-28 tokens/sec on a 128GB home mini PC. That number means local LLMs have finally moved from "theoretically runnable" to "daily-use capable."

What This Is

The tested model is DeepSeek's V4 Flash — the lightweight, speed-focused variant of the V4 family. The platform was AMD's Strix Halo mini PC, heavily promoted last year, notable for unified memory shared between CPU and GPU that removes VRAM bottlenecks during generation.

The tester used an acceleration technique called speculative decoding (think of it as: a small model drafts the next passage, the large model only verifies it), pushing generation speed from ~20 tokens/sec to around 28 tokens/sec. The gain looks modest, but local deployment often lives exactly on this edge between "barely usable" and "unusably laggy."

Industry View

The call worth watching: Chinese open-source LLMs are transitioning from "big parameters" to "actually runs locally."

Bulls argue that top-tier open-source models like DeepSeek, paired with community optimization, make local deployment genuinely meaningful for indie developers, small companies, and even individuals. For data-resident scenarios (medical records, internal code, customer contracts), local is a visible path forward.

But we want to flag an often-overlooked cost ledger: the test mini PC costs roughly 15,000–20,000 RMB, plus Linux setup, command-line tuning, and model file maintenance — not something most enterprise IT teams will willingly pick up. In the short term, "runs locally" and "most people willing to run it locally" remain far apart. Hyperscaler compute revenue may actually be reinforced by this trend: local runners are the exception; the vast majority of enterprises still depend on the cloud.

Impact on Regular People

  • Enterprise IT: Highly regulated industries like finance and healthcare gain another "local LLM" option, but high cost and ops overhead keep it from going mainstream anytime soon.
  • Individual professionals: Power users willing to tinker and concerned about data privacy can start paying attention; for most, ChatGPT, ERNIE Bot, and Tongyi remain the more realistic choice.
  • Consumer market: Large-memory mini PCs are forming a new product category, currently paid for mostly by developers and enthusiasts — far from mass-market adoption.
BZH
DeepSeekAMDStrix Halo·

DeepSeek V4 在家用小主机跑到 22-28 字/秒,本地大模型这事可能真要开始能用了

一位 Reddit 技术用户这周放出实测数据:DeepSeek 最新的 V4 Flash 模型,在一台 128GB 内存的家用迷你主机上跑到 22-28 字/秒。这个数字意味着——本地大模型,从『理论上能跑』终于走到了『日常能干活』这一步。

这是什么

被测的是 DeepSeek 的 V4 Flash——V4 系列里主打速度的轻量版模型。平台是 AMD 去年力推的 Strix Halo 迷你主机,特点是 CPU 和 GPU 共用一块内存,做生成时不用再为显存捉襟见肘。 测的人用了一种叫『推测式解码』(speculative decoding,可以简单理解为:让一个小模型先猜下一段词、大模型只负责核对)的加速技术,把生成速度从每秒 20 字提到 28 字左右。提升看着不大,但本地部署往往就是这种『勉强能用』和『卡得没法用』的临界差。

行业怎么看

值得关注的判断:中国开源大模型,正在从『参数大』过渡到『本地跑得动』。 看多的人认为,DeepSeek 这类顶级开源模型配合社区优化,本地部署对独立开发者、小公司甚至个人都有了实际意义。对数据出不去的场景(病历、内部代码、客户合同),本地是一个看得见的方向。 但我们也想提醒一个容易被忽略的成本账:测试用的小主机售价约 1.5-2 万元人民币,还要装 Linux、调命令行、维护模型文件——这不是大多数企业 IT 会主动上手的事。短期内,『本地跑得动』和『大多数愿意本地跑』之间还差得远。云厂商的算力生意,反倒可能被这件事强化:能本地跑的是极少数,绝大多数企业仍要靠云。

对普通人的影响

- 企业 IT:金融、医疗这类强合规行业多了一个『本地大模型』的备选项,但成本和运维门槛偏高,短期难成主流方案。 - 个人职场:愿意折腾技术、又担心数据隐私的重度用户可以开始关注;但对绝大多数人来说,ChatGPT、文心一言、通义仍是更现实的选择。 - 消费市场:搭载大内存的迷你主机正在长出一个新品类,目前主要是开发者和极客在买单,离大众消费还远。