Back to home

Compare

Comparing: Alibaba's Qwen Flash Next Adds MTP — Local LLMs Closing In on Cloud Speed & 阿里通义千问新模型支持多 Token 预测 — 本地跑大模型速度要追上云端了

AEN
AlibabaQwenllama.cpp·

Alibaba's Qwen Flash Next Adds MTP — Local LLMs Closing In on Cloud Speed

This week, llama.cpp—the most widely used open-source local LLM inference engine—merged an update: Alibaba's Qwen new model, Qwen Flash Next, now officially supports MTP (Multi-Token Prediction). Traditional LLMs generate one token at a time; MTP lets the model predict several upcoming tokens in a single pass, boosting inference speed by 2-3x while actually reducing VRAM requirements. This is a change worth taking seriously for local AI players.

What this is

Qwen Flash Next is, according to community discussion, a quantized version (GGUF format—a packaging method that compresses models to run on consumer GPUs) of Qwen's Qwen3-Next architecture. MTP was originally a training technique proposed by Meta in the Llama 2 paper; Qwen3-Next is one of the first model architectures to use it as inference acceleration—meaning the training phase was optimized for multi-token prediction, not just patched on afterward. This update effectively opens up the acceleration pathway: local users can finally experience full speed on consumer-grade hardware.

Industry view

Community reaction is mostly positive. Many are already discussing whether to switch from Qwen3 27B, since the speed difference is visible to the naked eye. But we think it's worth hearing another perspective: the actual gains from MTP acceleration depend on the task type—long-form text generation and batch translation benefit most, while multi-turn dialogue and code completion, with their short outputs, see limited improvement. Meanwhile, the Qwen3-Next architecture is relatively new, and its ecosystem (fine-tuning tools, third-party plugins, enterprise adapters) isn't as mature as Qwen3's; direct migration carries switching costs. One developer commented: "It's genuinely fast, but first make sure your workflow actually fits this approach."

Impact on regular people

For enterprise IT: The "performance/cost" ledger for on-prem LLM deployment is being rewritten. Running cloud-grade models used to require professional cards like the A100; now, an RTX 4090 paired with MTP can potentially support mid-scale AI applications—the hardware barrier to entry is dropping significantly.

For individual professionals: The local AI experience on Mac (M-series chips) and high-end Windows laptops will become smoother. Writers, researchers, and programmers who previously found local models "too slow" should give them another try—there's now one more reason to keep data local.

For the consumer market: This is the open-source camp (Meta, Alibaba, Mistral, etc.) catching up once again with closed-source players like OpenAI. The faster and cheaper local AI becomes, the lower the dependency on cloud APIs—and pricing pressure on related subscription services will continue to transmit through.

BZH
阿里通义千问Qwen·

阿里通义千问新模型支持多 Token 预测 — 本地跑大模型速度要追上云端了

本周 llama.cpp(目前最主流的开源本地大模型推理引擎)合并了一个更新:阿里通义千问新模型 Qwen Flash Next 正式支持 MTP(Multi-Token Prediction,多 Token 预测)。传统大模型是一个字一个字往外"吐",MTP 让模型一次预测接下来几个字,推理速度可提升 2-3 倍,对显存的要求反而下降。这是一个值得本地 AI 玩家认真关注的变化。

这是什么

Qwen Flash Next 据社区讨论是通义千问 Qwen3-Next 架构的量化版本(GGUF 格式,一种把模型"压缩"到消费级显卡也能跑的封装方式)。MTP 原本是 Meta 在 Llama 2 论文里提出的训练技术,Qwen3-Next 是首批把它用作推理加速的模型架构之一——也就是说,训练阶段就为多 Token 预测做了优化,不只是事后打补丁。本次更新相当于把这条加速通路打通了:本地用户终于可以在消费级硬件上用上完整速度。

行业怎么看

社区反应偏正面。不少人已经在讨论要不要从 Qwen3 27B 切换过来,毕竟速度差距肉眼可见。但我们觉得也有必要听另一种声音:MTP 加速的实际收益取决于任务类型,长文本生成、批量翻译这种场景受益最大,而多轮对话、代码补全这种短输出场景提升有限;同时 Qwen3-Next 架构较新,生态(微调工具、第三方插件、企业适配)还不如 Qwen3 成熟,直接迁移有切换成本。一位开发者评论说:"快是真的快,但要先确认你的工作流吃不吃这套。"

对普通人的影响

对企业 IT:本地部署大模型的"性能/成本"账本正在被改写。以前要跑接近云端水平的模型往往要 A100 这种专业显卡,现在一张 RTX 4090 配合 MTP 有望撑住中等规模的 AI 应用,硬件投入门槛明显下降。

对个人职场:Mac(M 系列芯片)和高端 Windows 笔记本本地跑 AI 的体验会更顺。文字工作者、研究者、程序员如果之前嫌本地模型"慢",值得重新试一次,把数据留在本地的理由又多了一条。

对消费市场:这是开源阵营(Meta、阿里、Mistral 等)对 OpenAI 等闭源厂商的又一次追赶。本地 AI 越快、越便宜,对云端 API 的依赖就越低,相关订阅服务的定价压力会持续传导。