Back to home

open-source LLMs

8 articles tagged with this topic

DeepSeekNVIDIA

DeepSeek Hits 67 token/s on Two $9K Mini Boxes — Local LLM Floor Is Caving In

Reddit user hit 67-84 token/s on DeepSeek V4 Flash with a 1M-token context window on two ~$9K NVIDIA DGX Sparks. The local-LLM cost barrier is collaps

4h ago2 min read
QwenAlibaba

Qwen Runs Locally — The "AI-Must-Be-Cloud" Assumption Cracks

Alibaba's Qwen3.8-Flash-Next lands in llama.cpp, letting regular PCs run AI locally without cloud APIs—open-source's quiet win over closed vendors.

2d ago2 min read
DeepSeekQwen

DeepSeek as Brain, Qwen as Assistant — China's Open-Source LLMs Split by Role

3 Chinese open-source LLMs tested on 4 DGX Spark cards. GLM cut for slowness. DeepSeek leads, Qwen supports — multi-model orchestration replaces singl

2d ago2 min read
OrnithMTP

Ornith 1.5 Shipped with an Untrained MTP Head — Open-Source AI's QC Problem

Ornith 1.5 shipped with an untrained MTP head. Not a bug — a symptom of QA gaps in open-source AI that any cost-cutting enterprise should heed.

Aug 202 min read
RedditLocalLLaMA

Reddit's LocalLLaMA Community Hit by Content Pollution — As LLMs Spread, Niche Communities Struggle

Reddit's r/LocalLLaMA, a hub for open-source LLM discussion, is drowning in AI-generated low-quality posts. The moderation ban is failing—and it signa

Aug 102 min read
QwenDeepSeek

A 35B Model Claims to Match Trillion-Param Giants — Reddit Isn't Buying It

A US team fine-tuned Qwen 3.5 into a 35B model claiming DeepSeek Pro-level performance. Reddit found vague benchmarks and possible contamination.

Aug 92 min read
DeepSeekFlash 0731

DeepSeek Flash Quantization: Why Has No One Systematically Benchmarked It Yet?

Quantized versions of DeepSeek Flash 0731 have leaked into the community, but nobody has run systematic benchmarks to measure how much capability is a

Aug 82 min read
RTX 6000 ProRTX 3090

Four Years, Eight GPUs: A Local AI Cluster That Cloud APIs Already Outprice

Reddit user documents a 4-year local AI cluster evolution from gaming rigs to 4× RTX 6000 Pro Max Q + 4× 3090. Cloud APIs already undercut it.

Aug 82 min read