What this is

An AMD Strix Halo user posted a counterintuitive test result this week: Zhipu's GLM 4.7 Flash, released roughly three years ago, beat Alibaba's newly released Qwen on tool calling (letting the AI autonomously operate external software to complete tasks) and common-sense Q&A.

This isn't a formal benchmark — it's one developer's subjective experience on their own hardware. But it points to a neglected problem: newer open-source LLMs aren't necessarily better, and that has real implications for enterprise model selection.

Industry view

Supporters will say this confirms Zhipu's early-mover advantage in the open-source ecosystem — the GLM series was one of China's first bets on the open-source path. Although it's been around for a while, capabilities like tool calling have been thoroughly engineered for real-world workloads.

The pushback deserves more attention, and we see three points. First, one user's test on specific hardware isn't a benchmark — Qwen has consistently outperformed comparable GLM models on standard eval suites. Second, the user themselves admitted "it might be a Qwen configuration issue," so we can't tell whether this is a generational gap or a usage problem. Third, and most importantly: Zhipu hasn't released a comparable-scale successor to GLM 4.7 Flash, which may signal a strategic shift toward closed-source or larger models. When enterprises pick a model, the real risk isn't who runs best today — it's whether anyone will still maintain it three years from now.

Impact on regular people

For enterprise IT: If your company is evaluating self-hosted AI (running models on your own servers rather than relying on the cloud), open-source model selection can't be driven by release date alone — a three-year-old model from one vendor may fit your hardware and workflow better than this year's release from another.

For individual careers: "Knows the latest AI tools" is depreciating fast. "Can choose between models and judge which one fits which scenario" is what pays now — this Reddit thread is a concrete case in point.

For the consumer market: Hardware like AMD Strix Halo, which lets ordinary PCs run large models locally, is maturing. Over the next year or two, AI PCs and smart devices that talk to you fully offline will multiply.