What this is
DeepSeek-V4.1-Flash was split-deployed by the community this week: the model is cut at layer 20, with the front half running on an Apple M5 Ultra and the back half on two RTX PRO 6000s, connected via standard gigabit/10-gigabit Ethernet. The key number is the prompt state (think of it as the model's "conversation notes") — only 0.9 KB/token, small enough that network bandwidth is not the bottleneck. Our take: DeepSeek keeps squeezing costs at the engineering level, dragging AI competition from "who has the biggest model" toward "who deploys cheapest."
Industry view
The open-source community's excitement is understandable — local large models just took another step forward. But cooler voices are also present:
- Demo is not ecosystem maturity. An M5 Ultra plus two RTX PRO 6000s is already a RMB 100,000+ combo — out of reach for typical users.
- This approach depends on a specific combination of Apple Silicon unified memory plus NVIDIA GPUs; until cross-vendor compatibility is solved, talking about "going mainstream" is premature.
- 0.9 KB/token is V4.1-Flash's selling point — not all DeepSeek models achieve it.
Deeper layer: DeepSeek's move this round points to "deployment economics" — the center of gravity in AI competition is shifting from "who has bigger parameters" to "who deploys cheaper, who covers more scenarios."
Impact on regular people
- For enterprise IT: Private-deployment cost structures could shift — full NVIDIA stacks may no longer be mandatory, with hybrid architectures emerging as a new option.
- For individual professionals: Still far from everyday use, but this hints that "running models directly on company hardware" may become more real than "companies buying API access by the call."
- For consumer market: Mac workstations will gain a sharper AI positioning, and enterprise hardware procurement combos will clearly multiply.