返回首页

对比阅读

对比阅读:Gemma4 vs Qwen3 Split Top AI Leaderboards — A Benchmarking Trust Crisis 与 Gemma4 与 Qwen3 同台打榜,两份权威榜单给出相反答案 — AI 评测陷入信任危机

AEN
Gemma4Qwen3Artificial Analysis·

Gemma4 vs Qwen3 Split Top AI Leaderboards — A Benchmarking Trust Crisis

The two most-cited AI model evaluation platforms — Artificial Analysis and LMArena (formerly Chatbot Arena) — delivered opposite verdicts this week on the same pair of open-source models. AA ranks Alibaba's Qwen3 27B ahead of Google's Gemma4 31B; LMArena places Gemma4 nearly 20 spots above Qwen3. For developers planning to run open-source models on local machines or servers, this means "pick a model by leaderboard" is effectively broken.

What this is

The story began when a developer on Reddit's r/LocalLLaMA (a community for AI model enthusiasts running models locally) noticed that two similarly-sized open-source models — Google's Gemma4 31B and Alibaba's Qwen3 27B — landed at nearly reversed positions on the two leading benchmarks. AA's "Intelligence Index" shows Qwen winning across the board; LMArena's human blind-vote rankings place Gemma near the top, about 20 positions higher. Two methodologies — one running academic test sets automatically, the other relying on real human blind evaluation — produced entirely opposite answers.

Industry view

The local community leans toward Qwen, viewing it as more "persistent" on reasoning tasks, though it pays for this by overthinking simple problems. Gemma isn't dismissed either — it's still considered a competent model. The mainstream explanation: the Qwen team has likely optimized specifically for AA-style academic benchmarks, while the Gemma team prioritizes Arena-style "real user voting." The two optimization targets are inherently incompatible.

But there's pushback: some argue both leaderboards are vulnerable to "score gaming" — vendors can run their own Arena votes or train on data matching AA's test style. So "leaderboard divergence" may not mean "two truths" — more likely "two exam-cheating strategies." Our judgment: when the gap between two benchmarks is this large, developers should downgrade both leaderboards by one tier and run actual business-data tests rather than trusting any single number.

Impact on regular people

For enterprise IT: don't make open-source model selection decisions from a single leaderboard. Review both human evaluation and automated scores side by side, and prepare a small "business question set" to test for a week before committing.

For working professionals: if you're using closed-source products like ChatGPT, ERNIE, or Doubao, what leaderboards their vendors optimize for barely matters to you — each vendor charts its own path, and what you care about is "which product actually works."

For the consumer market: this incident will reinforce the perception that "AI models can't be directly compared." For the coming year, treat any ad claiming "Model X beats Model Y" with a heavy discount.

BZH
GemmaQwenArtificial Analysis·

Gemma4 与 Qwen3 同台打榜,两份权威榜单给出相反答案 — AI 评测陷入信任危机

两个全球最被广泛引用的 AI 模型评测平台——Artificial Analysis 和 LMArena(原 Chatbot Arena)——本周对同一对开源模型给出了相反结论:AA 把阿里 Qwen3 27B 排在 Google Gemma4 31B 前面;LMArena 把 Gemma4 排到比 Qwen3 高近 20 个位置。对于打算在本地电脑或服务器跑开源模型的开发者,这意味着「看榜单选模型」已经基本失效。

这是什么

事情起因是 Reddit 上 r/LocalLLaMA(一个本地部署 AI 模型的爱好者社区)一位开发者发现:两个尺寸相近的开源大模型——Google 的 Gemma4 31B 和阿里的 Qwen3 27B——在两大主流评测上的排名几乎颠倒。AA 的「智能指数」显示 Qwen 全面领先;LMArena 的人类盲测投票把 Gemma 排到前列,相差近 20 名。两种方法论——一种偏自动跑学术题库,一种偏真人盲评——撞出了完全相反的答案。

行业怎么看

本地社区里更倾向 Qwen,认为它在推理任务上更「执着」,但代价是处理简单问题会过度思考;Gemma 也没被贬低,仍被认为是合格的模型。一种主流解释是:Qwen 团队针对 AA 这类学术类 benchmark(标准测试题库)做了定向优化,而 Gemma 团队更重视 Arena 这种「用户真实投票」——两种优化目标天然不兼容。

但也有反对声音:有人指出两个榜单本身都存在被「刷分」的可能——模型厂商可以在 Arena 上自家测试投票,也能在 AA 题库风格的数据上做训练。所以「榜单分歧」未必是「两个真相」,更可能是「两套刷题策略」。我们更倾向的判断是:当两个评测差距大到这个程度,开发者应该把两份榜单都降权一档,回到自己的业务数据上做实测对比,而不是相信任何一个单一数字。

对普通人的影响

对企业 IT:选开源模型时不要用单一榜单做决策,最好同时看人工评估和自动跑分,并准备一个小型「业务题库」实测一周再下结论。

对个人职场:如果你用的是 ChatGPT、文心、豆包这类闭源产品,模型厂商用什么榜单优化跟你关系不大——他们各自有自己走的路,你关心的是「哪个产品好用」。

对消费市场:这件事会加剧「AI 模型没法直接比较」的认知,接下来一年里,「某某模型比某某强」的广告语,可信度都要打个折。