Back to home

Compare

Comparing: Claimed 70 tokens/sec fails real test — AI benchmark inflation strikes again & 号称 70 tokens/秒实测打回原形 — AI 跑分注水这事又来了

AEN
GufoHalogenQwen·

Claimed 70 tokens/sec fails real test — AI benchmark inflation strikes again

An inference framework claiming to push open-source AI models to 70 tokens/sec (roughly "characters per second") got exposed by real-world testing: actual performance is only around 38 tokens/sec — 13% slower than its competitor. The problem isn't a technical issue — it's that the benchmark picked a task where speculative decoding (a speed-up technique) can directly copy the answer: asking the AI to write "red" 1000 times.

What this is

Gufo is an open-source large-model inference framework (the middleware layer that serves AI models externally). Its headline numbers: running Alibaba's Qwen 27B quantized version on AMD Strix Halo PC-class chips, single-user 70 tokens/sec, 8 users 122 tokens/sec.

A Reddit user re-tested using Gufo's own scripts — 70 is achievable, but only on the "write red 1000 times" task. Swap in real workloads like code, summarization, or translation, and single-user throughput drops to around 38; for 8 users, after stripping out queue time, actual throughput is only 52. Compared to competitor Halogen, Halogen is 13–18% faster on real tasks. Gufo does have strengths in long-prompt handling, and even larger advantages on repetitive tasks — those scenarios simply don't make the front page.

Industry view

The charitable read: the testing methodology itself was sound. The user ran Gufo's own scripts, and Gufo's documentation actually distinguishes between "repetitive" and "mixed" tasks — the mixed-task numbers match real-world results. There's no fraud here.

The skeptical read: the problem lies in the distribution path. The GitHub description and the top of the README only show the best-looking data point; ordinary users won't scroll down. This is the same move as foundation-model vendors cherry-picking benchmarks on leaderboards.

The risk lens: equating AI benchmarks with traditional software performance testing is fundamentally loose — AI inference workload variance far exceeds traditional software, and a single task can't stand in for a production environment. But that's exactly why cherry-picking is tempting. "Defining the product by its best-case scenario" is a systemic problem across the entire AI infrastructure industry, not a Gufo-only issue.

Impact on regular people

For enterprise IT: When evaluating locally-deployed open-source large models this year, don't take front-page numbers at face value — require vendors to run tests against your actual business data before drawing conclusions.

For individual professionals: From the model layer down to the tool layer, "screenshot speed demos" are broadly unreliable — when picking AI writing or coding assistants, lean on sustained, real-world evaluations.

For consumer markets: Consumer AI products (translation pens, learning machines) advertising "characters per second" or "accuracy rates" may have cherry-picked their tasks too — when you see the words "real-world test," give it slightly more weight.

BZH
GufoHalogenQwen·

号称 70 tokens/秒实测打回原形 — AI 跑分注水这事又来了

一个号称让开源 AI 模型跑到 70 tokens/秒(粗略理解为"字/秒")的推理框架,被实测打回原形:真实任务下只有 38 左右,比竞品还慢 13%。问题不在技术,而是 benchmark 选了一道投机解码(一种加速技巧)能直接抄答案的题——让 AI 写 1000 遍"red"。

这是什么

Gufo 是开源大模型推理框架(让 AI 模型对外提供服务的中间层),主推数据:在 AMD Strix Halo PC 级芯片上跑阿里 Qwen 27B 量化版,单用户 70 tokens/秒,8 用户 122 tokens/秒。

Reddit 用户用 Gufo 自家脚本复测,70 确实能跑——但测试题是"写 1000 遍 red"。换成代码、总结、翻译等真实任务,单用户跌到 38 左右,8 用户去掉排队时间后实际只有 52。和竞品 Halogen 对比,真实任务下 Halogen 反而快 13-18%。Gufo 在长提示词处理上有优势,重复任务优势更大,只是首页没展示这些场景。

行业怎么看

正面:测试方法本身规范。用户用了 Gufo 自家脚本,Gufo 文档其实分了"重复题"和"混合题",混合题数字与实测吻合,问题不在造假。

质疑:问题在传播路径。GitHub 描述和 README 顶部只展示最好看那格,普通用户不会往下翻。这和大模型厂商在榜单上"挑题刷分"是同一种动作。

风险视角:把 AI 跑分等同于传统软件性能测试本身就不严谨,AI 推理负载波动远大于传统软件,单题不代表生产环境。但正因如此,挑题刷分才有诱惑力。"用最佳场景定义产品"是整个 AI 基础设施行业的系统性问题,不是 Gufo 一家的事。

对普通人的影响

对企业 IT:今年评估本地部署开源大模型时,别直接采信首页数字,要求对方在你们真实业务数据上跑一遍再下结论。

对个人职场:从模型层到工具层,"截图秀速度"普遍不可信,挑 AI 写作、编程助手时看长期真实评测更靠谱。

对消费市场:消费级 AI 产品(翻译笔、学习机)的"每秒多少字""准确率多少"同样可能挑过题,看到"实测"二字可多信一分。