A new benchmark leaderboard this week threw out a number we think worth re-examining: on specific coding tasks, an open-source setup deployed locally on 4 RTX 6000s achieved a 100% completion rate matching Anthropic's flagship Claude Opus 5.5 — the trade-off being runtime stretching from 14 minutes to 31 minutes. But the other number is what deserves our attention: the same model, run across five "runtimes" (harnesses — think of them as the operating system for AI agents), scored anywhere from 22% to 96%.
What this is
airbench.ai launched an open leaderboard aimed at "local AI agents." Agents are AIs that autonomously complete multi-step tasks — reading emails, placing orders, writing code. The tests cover four categories: basic math, vision, computer use, and coding. Core findings: 4 RTX 6000s running glm-5.3-flash hit 100%, matching Claude Opus 5.5; a single RTX 5090 running qwen3.8-flash-next on OpenCode reached 96% and finished faster than Claude Code.
Industry view
On the surface it reads as "open source catches up to closed source," but the authors and frontline developers keep pointing to another set of numbers: the same model's average score across five harnesses climbed from 56% to 94%. The conclusion points to an underappreciated lever — much of the budget enterprises spend on model selection may be aimed at the wrong target.
Counterarguments exist. Developers note the leaderboard only tests basic tasks; in real enterprise scenarios, once harnesses face 65k long contexts, scores drop to 45–61%. On top of that, hardware depreciation, electricity, and operations headcount for local deployment may not actually be cheaper than a monthly Claude Code subscription.
Impact on regular people
For enterprise IT: the selection logic may need to be rewritten — spending big on model choice pays less than spending time on harness choice; budget structures should leave room for the runtime.
For the individual professional: the open-source ecosystem reaching this point means even if the company won't approve a Claude subscription, technically savvy people can now hit near-flagship results on their own hardware.
For the consumer market: in the short term, cloud APIs remain the default for most enterprises; but once the local cost curve keeps sliding, subscription pricing will inevitably feel pressure.