What This Is
This week, a developer on Juejin stress-tested two domestic models—MiniMax 3.1 and Space Bunny—on three high-difficulty Agent tasks (autonomous multi-step complex coding): a cyber version of Along the River During the Qingming Festival, an astronomical mechanical clock, and a Super Mario replica. Conclusion: Space Bunny is slightly better, but both lag noticeably behind OpenAI and Anthropic flagships.
In detail: on the cyber Qingming scroll, MiniMax took 56 minutes, Space Bunny 44 minutes. Both attempted to render spatial depth, with Space Bunny showing finer detail but misplacing the bridge's spatial relationship. The game replica was worse—MiniMax struggled for 37 minutes to produce a blank page (console full of syntax errors), while Space Bunny ran but with low-fidelity characters and maps. The author's words: "Spending an hour or two unwrapping a beautiful box, you find garbage inside. The Claude flagship next door—you unwrap it and there's actually a refined gift inside."
More notably: Space Bunny topped OpenRouter (an AI model API aggregator marketplace) with 28.4T tokens, surpassing DeepSeek, GLM, and Xiaomi MiMo. The key variable: it's free.
Industry View
The bull case: both used the standard "multi-round iteration + screenshot verification" Agent workflow, indicating their underlying architectures have caught up with the mainstream. Recognizing task complexity, breaking it into steps, and validating output—that's progress in itself.
But we lean toward the bear case: traffic-first ≠ capability-first. OpenRouter's top rank comes from free-tier scraping, not from being good. Space Bunny beats MiniMax functionally (buttons work, lunar phase timing is correct), but left basic UI bugs like a flying clock face unfixed; MiniMax was more egregious—to take a screenshot, it duplicated Chrome four times, burning tokens when a single command would have sufficed. Compared to the Claude flagship's ability to replicate classic games at 90%+ fidelity, the gap is a visible generational divide.
A deeper layer: the free strategy buys short-term token volume, not PMF (product-market fit). MiniMax previously loudly claimed "world's #1 coding agent," got slapped down, then went silent for a long time. This quiet 3.1 update is pragmatic—but silence doesn't fix capability gaps. Under heavy fire from Anthropic and OpenAI, domestic players collectively going quiet is a rational choice; the question is how long that silence lasts before they ship something real.
Impact on Regular People
- For enterprise IT: Free models suit PoC (proof of concept) and prototype stages, but production environments, complex business logic, and long-chain tasks still need budget reserved for paid flagships. Cost control cannot come at the expense of stability.
- For individual careers: Programmers using AI to build games, visualization tools, and multi-step projects can't yet let domestic models run unsupervised. On important projects, reserve time for review and correction.
- For consumer markets: Free AI products will keep multiplying. Today's free is tomorrow's paid-conversion funnel—or a test of "no longer free." Consumers should remember free in the short term but judge by real capability long term.