Not available in English yet
Qwen3.8 27B 在 12GB 显存上跑不顺:MoE 才是正解
Related Reading
More on #MoE
Tencent Stacks Model from 295B to 770B in 6 Weeks — China's Open-Source Sprint
Tencent's Hy4 open-weights preview: 770B total, 49B active, 1M context, text-only. 2.6x scale in 6 weeks — but 49B active sets real compute cost.
llama.cpp Has 50 PRs Pending — Local AI No Longer Needs a High-End GPU
Open-source llama.cpp has 50+ performance PRs pending merge, some claiming 3x CPU inference speedup. Local LLM deployment is shedding its dependence o
Why AI Assistants 'Forget'? Three-Layer Memory: Why More Memory Is Riskier
Agent doc Ch.9: AI memory splits into 3 layers. Remembering more = riskier than less. Most assistants only ship layer one — the 'toy vs. tool' line.
30 Structured Questions Expose AI's Real Level — Firms Write Their Own Tests
What you grade AI with decides what you measure—and miss. Juejin's Chen Yingbo: structured evaluation sets are enterprise AI's new infrastructure.
Qwen and Wanxiang Adopt OpenAI Protocol — Self-Hosted AI Gateways Grow Up
OctaFuse 2.8 splits subscriptions and top-ups, and brings four Chinese image models onto the OpenAI protocol — making multi-model, multi-vendor AI sta
Qwen 1-bit is still 6x slower — companies eyeing local LLMs should wait
Qwen's 1-bit quantization runs 6x slower than the 4-bit version with ~70% accuracy. Local LLM deployment isn't ready to replace cloud APIs.