Back to home

DGX Spark

8 articles tagged with this topic

DeepSeekNVIDIA

DeepSeek Hits 67 token/s on Two $9K Mini Boxes — Local LLM Floor Is Caving In

Reddit user hit 67-84 token/s on DeepSeek V4 Flash with a 1M-token context window on two ~$9K NVIDIA DGX Sparks. The local-LLM cost barrier is collaps

4h ago2 min read
NVIDIADGX Spark

Maxing AI 'Thinking Depth' Hurts Results — DGX Spark Local Test Warns Enterprises

A Reddit developer tested DeepSeek/Qwen on four DGX Sparks: 'deep thinking' mode lowers scores and doubles runtime — a direct cost warning for AI infe

14h ago2 min read
QwenNVIDIA

Two DGX Sparks Hit 181 tok/s Concurrent: Local Multi-Agent Is Now Viable

Two NVIDIA DGX Sparks + Qwen models hit 181 tok/s concurrent across 9 parallel AI Agents. Local multi-Agent moves from demo to real work.

1d ago2 min read
DeepSeekQwen

DeepSeek as Brain, Qwen as Assistant — China's Open-Source LLMs Split by Role

3 Chinese open-source LLMs tested on 4 DGX Spark cards. GLM cut for slowness. DeepSeek leads, Qwen supports — multi-model orchestration replaces singl

2d ago2 min read
inclusionAILing

124B model ran stably 7 min on desktop — local AI crosses usability threshold

Reddit test: 124B Ling model held 35.7 tok/s for 15K+ tokens on a single NVIDIA DGX Spark — local LLMs shifting from geek toy to enterprise option.

Aug 132 min read
inclusionAILing-3.0-flash

Two Flags Nearly Double Small Model Throughput — But the Hidden Compatibility Trap Matters More

InclusionAI's Ling-3.0-flash INT4 hits 38.7 tok/s on DGX Spark with two config tweaks — but default vLLM silently breaks V3 architecture, producing fl

Aug 92 min read
NvidiaDGX Spark

16 Nvidia DGX Spark Units Clustered for LLMs — Enterprise Compute Focus Shifts to VRAM

Reddit user clusters 16 Nvidia DGX Spark units, runs 434GB LLM. Unified memory validated. Inference bottlenecks shift from compute to VRAM — new path

May 12 min read
Gemma 4vLLM

Running Gemma 4 26B-A4B on vLLM: Community Troubleshooting Notes

Developers report mixed results deploying Gemma 4 26B-A4B on vLLM, with INT4 quants too slow on DGX Spark GB10.

Apr 62 min read