返回首页

对比阅读

对比阅读:Luo Fuli, 31, Tops Xiaomi's Org Chart — MiMo Hits Trillion Params in 11 Months 与 31岁罗福莉升到小米最高职级 — MiMo用11个月把模型做到万亿参数

AEN
Luo FuliXiaomiMiMo·

Luo Fuli, 31, Tops Xiaomi's Org Chart — MiMo Hits Trillion Params in 11 Months

We noted: Xiaomi promoted 31-year-old Luo Fuli to Level 22, the company's highest rank. She previously worked on V2 at DeepSeek and joined Xiaomi 11 months ago to lead the MiMo team — over the past year, MiMo scaled from 7 billion to trillion-parameter models.

MiMo's cadence this year: open-sourced reasoning model MiMo-7B in April; released V2-Flash (309B total parameters) in December; V2-Pro (total parameters broke the trillion mark) the following March; V2.6-Pro took first place among open-weight models on the Artificial Analysis intelligence index in September, with training livestreamed publicly.

What this is

Three technical points worth noting:

First, a hybrid sliding-window + full-attention architecture (Hybrid SWA — most layers only look at the most recent 128 local tokens, with a few layers seeing global context), which keeps compute costs manageable for 1M-token long context windows (enough to fit a thick book).

Second, Multi-Teacher On-Policy Distillation (multiple specialist models simultaneously teach a single student), letting one model absorb capabilities across domains like code and search.

Third, the reinforcement learning (model learns from environment-provided scores) training dashboard is livestreamed publicly on the official website. The next-generation V3 will use the proprietary HySparse2 architecture.

Beyond the model itself, Xiaomi open-sourced the coding agent MiMo Code (an AI assistant that can autonomously operate software) and the desktop client MiMo Desktop, extending into the product layer.

Industry view

Supporters: open weights, open training, first place on the intelligence index — walking the full path in one year marks the first time a Chinese foundation-model team has openly matched pace with international rivals.

But skepticism remains: Artificial Analysis is a third-party benchmark, weighty but not everything; V2.5-Pro scored only 57.2 on SWE-bench Pro (real GitHub repo fix evaluation), still a gap from top closed-source models; and the next-gen V3's HySparse2 architecture is unproven. These are real uncertainties.

Impact on regular people

For enterprise IT: open weights allow on-prem deployment, but trillion-parameter models don't run on ordinary servers — deployment costs must be calculated upfront.

For working professionals: MiMo Code and MiMo Desktop are pushing model capabilities to the desktop, and workflows like coding and research will be reshuffled.

For the consumer market: Xiaomi's model plugs directly into its full phone-auto-home ecosystem (phones, cars, smart home), making it the closest to ordinary users among this wave of foundation-model companies — but whether the ecosystem actually delivers still comes down to the users.

来源: juejin.cn
BZH
罗福莉小米MiMo·

31岁罗福莉升到小米最高职级 — MiMo用11个月把模型做到万亿参数

我们注意到:小米把31岁的罗福莉升到22级,公司最高一级。她此前在DeepSeek参与过V2研发,11个月前加入小米带MiMo团队——过去一年MiMo从70亿参数做到万亿规模。

这一年MiMo的节奏:4月开源推理模型MiMo-7B;12月发布V2-Flash(总参数309B);次年3月V2-Pro(总参数突破万亿);9月V2.6-Pro在Artificial Analysis智能指数拿下开源权重第一,训练过程公开直播。

这是什么

技术上有三件事值得记:

一是「滑动窗口+全注意力」混合架构(Hybrid SWA——大部分层只看局部最近的128个token,少数层看全局),把1M长上下文(能装一本厚书)的算力成本压低。

二是「多教师在线策略蒸馏」(Multi-Teacher On-Policy Distillation——多个专精模型同时教一个学生),让一个模型吸收代码、搜索等不同领域。

三是把强化学习(让模型从环境打分中学)训练面板放到官网公开直播。下一代V3会用自研HySparse2架构。

模型之外,配套开源了编程Agent(能自主操作软件的AI助手)MiMo Code和桌面端MiMo Desktop,往产品层铺。

行业怎么看

支持方:开源权重、开源训练、智能指数第一,一年把路走通,是中国大模型第一次在节奏上公开对线国际对手。

但也有质疑:Artificial Analysis是第三方榜单,权重不轻但不是全部;V2.5-Pro在SWE-bench Pro(真实GitHub仓库修复评测)只拿到57.2,跟顶级闭源还有距离;下一代V3的HySparse2架构还没验证。这些都是真实的不确定性。

对普通人的影响

对企业IT:开源权重可以本地部署,但万亿参数不是普通服务器能跑的,部署成本要先算清。

对个人职场:MiMo Code和MiMo Desktop开始把模型能力推到桌面,写代码、做研究这类工作流会被重新打散。

对消费市场:小米模型直接接入手车家全生态(手机、汽车、家居),是这波大模型公司里离普通用户最近的——但生态能不能落地,最终还是用户说了算。

来源: juejin.cn