Artificial Analysis
11 articles tagged with this topic
Gemma4 vs Qwen3 Split Top AI Leaderboards — A Benchmarking Trust Crisis
Google's Gemma4 31B and Alibaba's Qwen3 27B received nearly opposite rankings on Artificial Analysis and LMArena — exposing systematic failure in AI b
Qwen 3 'Goated' at Low/Mid Compute — Alibaba Silences Overthinking Critics
Qwen 3 crushes benchmarks at low/mid compute, debunking the "overthinking" myth. The open-source ecosystem just got stronger.
GLM-5.3 Hits Artificial Analysis — China's First Open-Source Flagship Vetted
Zhipu's GLM-5.3 completes Artificial Analysis benchmarks — first Chinese open-source model to land in the global top tier with an independent score.
Ling-3 Tiny Outscores Qwen3.5 9B — Our Open-Source Attention Is Too Concentrated
Reddit found Ling-3 Tiny (1-2B params) outperforming Qwen3.5 9B on reasoning benchmarks. Not an isolated case — strong mid-tier open-source labs are g
Qwen 27B Beats Opus on Agent Index — Open Source Closes the Gap
Qwen's 27B tops Opus on Artificial Analysis's Agentic Index. Open-source small models now beat closed-source flagships on agent tasks — a bookmark-wor
Qwen 27B Benchmark Nears 70B — The Small-Model Card Has Been Played
Alibaba's Qwen 27B closes in on last-gen 70B scores on Artificial Analysis — the "bigger is better" scaling narrative is cracking.
Qwen 3.8 27B Benchmark Surges 37% — Open-Source Cracks Closed-Source Ceiling
Qwen 3.8 27B hits 52 on Artificial Analysis, up 37% from the prior version and surpassing Claude Opus 4.5 (42). The open-source vs. closed-source ceil
Qwen3.8-27B benchmarks tie GPT-5.6 — small models can now take on big models
Artificial Analysis scored Alibaba's Qwen3.8-27B tying DeepSeek V4 and GPT-5.6 Luna Max. Enterprise private AI deployment economics may be quietly shi
DeepSeek Extends Its Price-Performance Lead, but Cheap Isn’t a Moat
DeepSeek’s pricing shocked Reddit again, but the bigger story is China’s AI race shifting from bigger models to lower-cost usable ones.
Kimi K3 Raises China's Open-Model Price Ceiling With Performance
Moonshot AI’s Kimi K3 signals a shift in China’s open-model race from low prices to performance and efficiency.
Anonymous Peanut Hits #8 in Text-to-Image as Open-Source Race Crowds
Anonymous model Peanut hits #8 on Artificial Analysis, beating FLUX.2. Open weights promised, but safety risks and unfulfilled pledges warrant caution