What this is

A web tool called human-tps surfaced this week on Reddit's r/LocalLLaMA, a community for enthusiasts who run large models locally. Developer HomoAgens1 borrowed the "tokens per second" (i.e., "typing speed") metric that AI labs use to benchmark model output, then turned it on humans — letting visitors test how fast they can "generate." His own best score: 2 tokens/second.

The cross-reference is the punchline: that speed beats the output rate of a 70B (70-billion-parameter) large model running on a laptop CPU, yet trails an RTX 4090 GPU running an 8B (8-billion-parameter) model by roughly 76×. In other words, your fingers outpace a laptop CPU's LLM output, but lag absurdly behind a prosumer GPU.

The developer admits this is a lark — the site closes with "so for open it while waiting for your LLM to answer" (open it while you wait for your AI to respond). Still, the post shot to the top of the community's front page.

Industry view

The local-AI crowd is sharing it as a meme, but it surfaces a real fault line: the hardware generation gap. Industry consensus says models are shrinking — 8B and 4B are already capable — yet the compute required for inference (running a trained model to answer questions) hasn't dropped. Models advance, consumer hardware lags, and that's the root reason "running AI locally" has stayed niche.

Pushback: critics will note this comparison is meaningless for cloud AI — enterprise GPU clusters deliver tokens-per-second orders of magnitude higher than a 4090, so consumer hardware bottlenecks don't represent the industry as a whole. Others question the metric itself: tokens/sec only measures "typing speed" and ignores the real latency of comprehension, reasoning, and external tool calls. Using it to grade AI performance is itself a shortcut.

Impact on regular people

For enterprise IT: On-premise AI looks cheap on paper, but GPU servers, data-center cooling, and power are the hidden line items. Don't size the decision around software licensing alone.

For working professionals: When "AI feels slow" in daily tasks, the culprit usually isn't the model — it's the device. Right now, cloud AI services (ChatGPT, Tongyi, Wenxin, etc.) are more cost-effective than local deployment.

For the consumer market: AI hardware startups — AI PCs, personal AI boxes — serve real demand, but the price-performance ratio still trails simply calling cloud APIs. In the short term, don't pay a premium for "local AI."