open-source LLMs
8 articles tagged with this topic
DeepSeek Hits 67 token/s on Two $9K Mini Boxes — Local LLM Floor Is Caving In
Reddit user hit 67-84 token/s on DeepSeek V4 Flash with a 1M-token context window on two ~$9K NVIDIA DGX Sparks. The local-LLM cost barrier is collaps
Qwen Runs Locally — The "AI-Must-Be-Cloud" Assumption Cracks
Alibaba's Qwen3.8-Flash-Next lands in llama.cpp, letting regular PCs run AI locally without cloud APIs—open-source's quiet win over closed vendors.
DeepSeek as Brain, Qwen as Assistant — China's Open-Source LLMs Split by Role
3 Chinese open-source LLMs tested on 4 DGX Spark cards. GLM cut for slowness. DeepSeek leads, Qwen supports — multi-model orchestration replaces singl
Ornith 1.5 Shipped with an Untrained MTP Head — Open-Source AI's QC Problem
Ornith 1.5 shipped with an untrained MTP head. Not a bug — a symptom of QA gaps in open-source AI that any cost-cutting enterprise should heed.
Reddit's LocalLLaMA Community Hit by Content Pollution — As LLMs Spread, Niche Communities Struggle
Reddit's r/LocalLLaMA, a hub for open-source LLM discussion, is drowning in AI-generated low-quality posts. The moderation ban is failing—and it signa
A 35B Model Claims to Match Trillion-Param Giants — Reddit Isn't Buying It
A US team fine-tuned Qwen 3.5 into a 35B model claiming DeepSeek Pro-level performance. Reddit found vague benchmarks and possible contamination.
DeepSeek Flash Quantization: Why Has No One Systematically Benchmarked It Yet?
Quantized versions of DeepSeek Flash 0731 have leaked into the community, but nobody has run systematic benchmarks to measure how much capability is a
Four Years, Eight GPUs: A Local AI Cluster That Cloud APIs Already Outprice
Reddit user documents a 4-year local AI cluster evolution from gaming rigs to 4× RTX 6000 Pro Max Q + 4× 3090. Cloud APIs already undercut it.