What this is
Philosopher Toby Ord dropped a sobering number on the AI industry this week that calls for a reality check: 4 AI agents processing the same task in parallel run roughly twice as fast as a single AI — but at the cost of doubling total token consumption (the unit of text AI processes, roughly equivalent to compute overhead). He calls this phenomenon "swarm scaling," essentially a new class of inference-time compute scaling (compute resources spent during model runtime, rather than during training).
The more critical finding is the "stepping on toes" parameter (coordination loss): scaling agent count by 10x yields only 3-5x performance gains, not linear — strikingly similar to the steep coordination costs that plague human organizations as headcount grows.
Industry view
Ord himself admits this isn't the conclusion he was hoping for: "I had hoped the λ value (coordination loss coefficient) would be lower, which would reduce the danger of an RSI (recursive self-improvement, i.e., AI iteratively upgrading itself) driven intelligence explosion — but the evidence suggests otherwise." In other words, multi-agent parallelism doesn't automatically serve as a safety valve against AI surpassing humans — it may actually pull the timeline forward.
But there's another reading: swarm scaling essentially trades time for compute, which is highly tempting for labs chasing fast iteration cycles — the inference-time compute track will keep gaining heat, with token spend becoming the new arms race metric. A word of caution: Ord's sample comes from specific benchmarks; real-world business scenarios may carry even higher coordination costs. Enterprises shouldn't be swept up by the "more agents is better" narrative.
Impact on regular people
- For enterprise IT: Multi-agent architectures will become standard in 2025-2026, but don't be sold on "more agents = better" — around 3 agents hits the cost-effectiveness sweet spot; beyond that, marginal returns drop sharply.
- For individual professionals: When you split work across multiple AI tools (one for writing, one for research, one for review), you're doing swarm scaling too — the more finely you divide, the higher the communication cost, and the final output may suffer.
- For consumer markets: If AI capabilities really do iterate faster through parallelism, today's AI assistants will likely be outclassed within six months by their own upgraded successors. When picking a subscription, watch how hard the vendor is burning compute.