What This Is

This week on r/LocalLLaMA (a developer community focused on locally deployed large models), one user shared test results: running Alibaba's open-source Qwen 3.8 on an RTX 5090 (currently the top consumer GPU, retailing around RMB 16,000) for 3D modeling tasks, comparing it against Anthropic's latest Claude, and concluding it is "miles behind."

This is not an isolated case. In code generation, complex Agent (AI assistants capable of autonomously completing multi-step tasks) workflows, and specialized domain Q&A — the "real work" scenarios — local open-source models still show a significant gap behind top cloud closed-source models. Qwen, DeepSeek, GLM and other Chinese open-source models have closed the distance on general conversation, but on tasks requiring long-chain reasoning and domain expertise, cloud models still lead decisively.

Industry View

The prevailing voice in the local-deployment community is: "the progress is real, and so is the gap." The Qwen 3 series and DeepSeek V3 iterated quickly through 2025, with standard test scores (benchmarks) approaching the top tier. But in real-world developer testing, there is still meaningful distance between "high benchmark scores" and "actually usable."

The counterarguments deserve equal weight. One data engineer who has long run open-source models told us that what enterprises truly agonize over is not "can we run it locally," but "is what comes out of local deployment worth our staff using." Tasks like 3D modeling tolerate zero errors — one wrong parameter can void an entire blueprint. This is a scenario open-source models currently cannot support.

The more grounded assessment: over the next 12–18 months, local models will become increasingly capable on general tasks like "writing emails and organizing materials," but in scenarios where mistakes are not allowed — "professional engineers, production lines" — the moat (a short-term advantage competitors cannot easily close) of cloud APIs remains deep.

Impact on Regular People

For enterprise IT: data-sensitive industries (finance, healthcare, manufacturing) still face a binary choice — swap capability for cloud APIs, or swap data security for local models. There is no perfect answer today, so budgets should be planned for both.

For working professionals: local models are already sufficient for everyday writing and information organization. For deliverables to clients or bosses that carry real stakes, we still recommend going with top-tier cloud versions — the time saved is worth more than the subscription fee.

For consumers: an RTX 5090 costs RMB 16,000; a year of cloud subscription may run less than one-tenth of that. Unless you have sustained high-frequency demand, the math may not work out.