A Reddit engineer ran 429 rigorous tests under a unified benchmark, ranking 15 open-source LLMs (parameters from 8B to 35B) on Agent tasks (letting AI autonomously call tools to complete multi-step operations). Qwen3-27B took first place with 71.8%, but what concerns us more is last-place Bonsai: 70% of its failures weren't "picking the wrong tool"—they were "finishing the job but failing to stop."

What this is

The test used the developer-built Toolery benchmark: 143 real-world scenarios, three attempts each, split into easy/medium/hard/extreme tiers. All models were loaded through LM Studio (a desktop tool for running open-source models locally), with no cloud API calls.

The champion scored 93.3 on medium-difficulty tasks but plunged to 41.7% on the extreme tier—even in first place, complex multi-step tasks still fail nearly 60% of the time. Last-place Bonsai 27B scored 50.5% overall, with 103 of its 148 failures (70%) classified as "budget violations"—tool calls exceeding the scenario's allowed limit.

Industry view

Supporters argue this unified benchmark is the "mirror" the open-source ecosystem has lacked: previously, everyone relied on vendor-released leaderboards; now someone has put the models in the same ring.

Criticism exists too. Senior developers point out that Toolery is a single-person benchmark with limited coverage across 143 scenarios; the "budget violation" metric itself is debatable—in real business settings, calling a few extra tools isn't necessarily a flaw and might deliver more stable results. Another layer of skepticism: open-source tests often ignore latency, yet Bonsai happened to be the slowest model in this run (6,208 seconds total).

A deeper signal: top-tier open-source power is concentrating. Four of the top five models come from Alibaba, IBM, Mistral, or their derivatives; the middle and lower tiers are flooded with experimental models from smaller vendors of uneven quality.

Impact on regular people

[Enterprise IT] In industries with strict data compliance (finance, healthcare, government), 27B-class open-source models can now score above 70 on most Agent tasks—you don't need to hand data to the cloud to use an AI assistant.

[Individual professionals] This engineer completed all tests on a laptop using LM Studio. If you're a product manager, ops specialist, or consultant, spending a few hours weekly running local open-source models for comparison experiments can already replace parts of a ChatGPT Plus workflow.

[Consumer market] The "failure to stop" failure mode is the exact pitfall consumer AI most often stumbles into—smart speakers repeatedly waking up, outbound-voice bots stuck in loops are all variants of it. To judge whether an Agent product is mature, "knowing when to stop" matters more to ordinary users than "how accurate the answers are."