On an M1 MacBook Air, Liquid's LFM2.5 running an Agent (letting the model decide on its own to call tools or search the web) hit 200 tokens/sec (tokens: the smallest unit of text processed by large models) and used only 2.5GB of memory — outpacing Tsinghua-affiliated MiniCPM5 (162 tokens/sec, 3.8GB) on both speed and efficiency. The model with more parameters lost. What this exposes is the real divergence in the local small-model track.
What this is
LFM2.5 comes from US startup Liquid AI; MiniCPM5 is led by Tsinghua-affiliated team OpenBMB (backed by ModelBest). Both shipped updates intensively across 2024-2025, targeting "good enough" models that run on regular laptops and mini-PCs.
This week's benchmark by Reddit user parepeg was simple: have the model call web search to answer "what's the weather today" — type questions, then compare speed and stability. LFM2.5 won on generation speed, memory footprint, and English stability; MiniCPM5 often answered English prompts in Chinese — a bug for English Agent workflows, but arguably a feature for Chinese users.
Industry view
Supporters argue that small models capable of running Agents mean enterprises can deploy AI to edge devices (local endpoints that don't depend on the cloud), cutting compute costs and protecting data privacy. Some practitioners back Liquid's technical path on the grounds that "lightweight models have a clearer commercialization path than brute-forcing ever-larger models."
Pushback comes in three lines. First, a single Reddit user's test can't represent all scenarios — switch hardware or prompt and the result could flip. Second, MiniCPM5's "Chinese-first" design is an absolute advantage in the Chinese market; LFM2.5's English win means little to Chinese users. Third, this test only covered simple tasks like "asking about the weather"; 2B-class models still struggle with complex Agents (multi-step, long-term memory), and we shouldn't over-generalize the conclusion.
Impact on regular people
For enterprise IT: deploying a local Agent assistant inside a small team now costs no more than a Mac mini — cloud APIs are no longer mandatory. But selection should follow business language: MiniCPM5 for Chinese workloads, LFM2.5 for English.
For individual professionals: people doing cross-border business, content creation, or industry research can run models locally so sensitive data never leaves the device. But small models hit a ceiling — for any key decision, you'll still need to fall back on cloud-based large models.
For consumer markets: embedding small-model Agents inside phones, earbuds, and car infotainment systems now has engineering footing. But mainstream "on-device AI" still needs chips and models to break through on both sides — that's not a 2025 story.