DGX Spark
8 articles tagged with this topic
DeepSeek Hits 67 token/s on Two $9K Mini Boxes — Local LLM Floor Is Caving In
Reddit user hit 67-84 token/s on DeepSeek V4 Flash with a 1M-token context window on two ~$9K NVIDIA DGX Sparks. The local-LLM cost barrier is collaps
Maxing AI 'Thinking Depth' Hurts Results — DGX Spark Local Test Warns Enterprises
A Reddit developer tested DeepSeek/Qwen on four DGX Sparks: 'deep thinking' mode lowers scores and doubles runtime — a direct cost warning for AI infe
Two DGX Sparks Hit 181 tok/s Concurrent: Local Multi-Agent Is Now Viable
Two NVIDIA DGX Sparks + Qwen models hit 181 tok/s concurrent across 9 parallel AI Agents. Local multi-Agent moves from demo to real work.
DeepSeek as Brain, Qwen as Assistant — China's Open-Source LLMs Split by Role
3 Chinese open-source LLMs tested on 4 DGX Spark cards. GLM cut for slowness. DeepSeek leads, Qwen supports — multi-model orchestration replaces singl
124B model ran stably 7 min on desktop — local AI crosses usability threshold
Reddit test: 124B Ling model held 35.7 tok/s for 15K+ tokens on a single NVIDIA DGX Spark — local LLMs shifting from geek toy to enterprise option.
Two Flags Nearly Double Small Model Throughput — But the Hidden Compatibility Trap Matters More
InclusionAI's Ling-3.0-flash INT4 hits 38.7 tok/s on DGX Spark with two config tweaks — but default vLLM silently breaks V3 architecture, producing fl
16 Nvidia DGX Spark Units Clustered for LLMs — Enterprise Compute Focus Shifts to VRAM
Reddit user clusters 16 Nvidia DGX Spark units, runs 434GB LLM. Unified memory validated. Inference bottlenecks shift from compute to VRAM — new path
Running Gemma 4 26B-A4B on vLLM: Community Troubleshooting Notes
Developers report mixed results deploying Gemma 4 26B-A4B on vLLM, with INT4 quants too slow on DGX Spark GB10.