Back to home

Artificial Analysis

11 articles tagged with this topic

Gemma4Qwen3

Gemma4 vs Qwen3 Split Top AI Leaderboards — A Benchmarking Trust Crisis

Google's Gemma4 31B and Alibaba's Qwen3 27B received nearly opposite rankings on Artificial Analysis and LMArena — exposing systematic failure in AI b

3d ago2 min read
QwenAlibaba

Qwen 3 'Goated' at Low/Mid Compute — Alibaba Silences Overthinking Critics

Qwen 3 crushes benchmarks at low/mid compute, debunking the "overthinking" myth. The open-source ecosystem just got stronger.

Aug 212 min read
ZhipuGLM-5

GLM-5.3 Hits Artificial Analysis — China's First Open-Source Flagship Vetted

Zhipu's GLM-5.3 completes Artificial Analysis benchmarks — first Chinese open-source model to land in the global top tier with an independent score.

Aug 192 min read
Ling-3Qwen3.5

Ling-3 Tiny Outscores Qwen3.5 9B — Our Open-Source Attention Is Too Concentrated

Reddit found Ling-3 Tiny (1-2B params) outperforming Qwen3.5 9B on reasoning benchmarks. Not an isolated case — strong mid-tier open-source labs are g

Aug 182 min read
QwenAlibaba

Qwen 27B Beats Opus on Agent Index — Open Source Closes the Gap

Qwen's 27B tops Opus on Artificial Analysis's Agentic Index. Open-source small models now beat closed-source flagships on agent tasks — a bookmark-wor

Aug 182 min read
AlibabaQwen

Qwen 27B Benchmark Nears 70B — The Small-Model Card Has Been Played

Alibaba's Qwen 27B closes in on last-gen 70B scores on Artificial Analysis — the "bigger is better" scaling narrative is cracking.

Aug 172 min read
QwenTongyi Qianwen

Qwen 3.8 27B Benchmark Surges 37% — Open-Source Cracks Closed-Source Ceiling

Qwen 3.8 27B hits 52 on Artificial Analysis, up 37% from the prior version and surpassing Claude Opus 4.5 (42). The open-source vs. closed-source ceil

Aug 172 min read
alibabaQwen

Qwen3.8-27B benchmarks tie GPT-5.6 — small models can now take on big models

Artificial Analysis scored Alibaba's Qwen3.8-27B tying DeepSeek V4 and GPT-5.6 Luna Max. Enterprise private AI deployment economics may be quietly shi

Aug 172 min read
DeepSeekKimi K3

DeepSeek Extends Its Price-Performance Lead, but Cheap Isn’t a Moat

DeepSeek’s pricing shocked Reddit again, but the bigger story is China’s AI race shifting from bigger models to lower-cost usable ones.

Jul 182 min read
Moonshot AIKimi K3

Kimi K3 Raises China's Open-Model Price Ceiling With Performance

Moonshot AI’s Kimi K3 signals a shift in China’s open-model race from low prices to performance and efficiency.

Jul 162 min read
PeanutFLUX

Anonymous Peanut Hits #8 in Text-to-Image as Open-Source Race Crowds

Anonymous model Peanut hits #8 on Artificial Analysis, beating FLUX.2. Open weights promised, but safety risks and unfulfilled pledges warrant caution

May 52 min read