An inference framework claiming to push open-source AI models to 70 tokens/sec (roughly "characters per second") got exposed by real-world testing: actual performance is only around 38 tokens/sec — 13% slower than its competitor. The problem isn't a technical issue — it's that the benchmark picked a task where speculative decoding (a speed-up technique) can directly copy the answer: asking the AI to write "red" 1000 times.

What this is

Gufo is an open-source large-model inference framework (the middleware layer that serves AI models externally). Its headline numbers: running Alibaba's Qwen 27B quantized version on AMD Strix Halo PC-class chips, single-user 70 tokens/sec, 8 users 122 tokens/sec.

A Reddit user re-tested using Gufo's own scripts — 70 is achievable, but only on the "write red 1000 times" task. Swap in real workloads like code, summarization, or translation, and single-user throughput drops to around 38; for 8 users, after stripping out queue time, actual throughput is only 52. Compared to competitor Halogen, Halogen is 13–18% faster on real tasks. Gufo does have strengths in long-prompt handling, and even larger advantages on repetitive tasks — those scenarios simply don't make the front page.

Industry view

The charitable read: the testing methodology itself was sound. The user ran Gufo's own scripts, and Gufo's documentation actually distinguishes between "repetitive" and "mixed" tasks — the mixed-task numbers match real-world results. There's no fraud here.

The skeptical read: the problem lies in the distribution path. The GitHub description and the top of the README only show the best-looking data point; ordinary users won't scroll down. This is the same move as foundation-model vendors cherry-picking benchmarks on leaderboards.

The risk lens: equating AI benchmarks with traditional software performance testing is fundamentally loose — AI inference workload variance far exceeds traditional software, and a single task can't stand in for a production environment. But that's exactly why cherry-picking is tempting. "Defining the product by its best-case scenario" is a systemic problem across the entire AI infrastructure industry, not a Gufo-only issue.

Impact on regular people

For enterprise IT: When evaluating locally-deployed open-source large models this year, don't take front-page numbers at face value — require vendors to run tests against your actual business data before drawing conclusions.

For individual professionals: From the model layer down to the tool layer, "screenshot speed demos" are broadly unreliable — when picking AI writing or coding assistants, lean on sustained, real-world evaluations.

For consumer markets: Consumer AI products (translation pens, learning machines) advertising "characters per second" or "accuracy rates" may have cherry-picked their tasks too — when you see the words "real-world test," give it slightly more weight.