DeepSeek
30 articles tagged with this topic
DeepSeek Hits 67 token/s on Two $9K Mini Boxes — Local LLM Floor Is Caving In
Reddit user hit 67-84 token/s on DeepSeek V4 Flash with a 1M-token context window on two ~$9K NVIDIA DGX Sparks. The local-LLM cost barrier is collaps
Speculative decoding is becoming standard — open-source LLMs now predict ahead
Reddit users spotted speculative decoding working on local GPUs—AI instantly outputting phrases via MTP. The local inference cost curve is being quiet
Maxing AI 'Thinking Depth' Hurts Results — DGX Spark Local Test Warns Enterprises
A Reddit developer tested DeepSeek/Qwen on four DGX Sparks: 'deep thinking' mode lowers scores and doubles runtime — a direct cost warning for AI infe
DeepSeek's Agent Framework Exposes Five Keys as Chinese Devs Crack the Source
DeepSeek Agent framework has 5 event-dispatch modes decoded in an 8,000-word Chinese source tutorial - a shift from API calls to source-reading.
DGX Overheats on AI — What It Means When NVIDIA's Flagship Needs User Fans
User added active cooling to NVIDIA DGX after DeepSeek runs overheated it. NVIDIA's flagship needs DIY cooling under sustained load, puncturing plug-a
AI Spots Bugs in 10 Minutes — Open Source's 30-Year Embargo Must Be Rebuilt
AI coding assistants shrink vulnerability discovery from days to 10 minutes, breaking open source's decades-old embargo. Maintainers and enterprises m
Zhipu GLM-5.3 Scores Higher Without New Architecture — Chinese LLMs Turn Inward
Zhipu's GLM-5.3 hits stronger benchmarks with architecture identical to 5.2 — a training-only upgrade worth watching while peers chase new designs.
NVIDIA Cuts Open-Source Deployment to Two Commands — Convenience Is New Business
NVIDIA's TensorRT Model Connect deploys open-source LLMs in two commands. As GPUs become abundant, "making AI run" is itself a new business.
8GB GPUs Can Now Run 70B Models — Quantization Crushes Local AI Deployment Costs
8GB consumer GPUs couldn't fit 130GB model weights; now quantization runs 7B models on 3.5GB. The real story isn't specs — AI deployment may finally l
GLM, Qwen Catch Closed-Source Leaders on Agent Benchmarks in Just Two Months
GLM and Qwen now match top closed-source models on Agent Arena Code — a result unthinkable two months ago.
DeepSeek as Brain, Qwen as Assistant — China's Open-Source LLMs Split by Role
3 Chinese open-source LLMs tested on 4 DGX Spark cards. GLM cut for slowness. DeepSeek leads, Qwen supports — multi-model orchestration replaces singl
User Claims Qwen Beats DeepSeek — But the Post Has Zero Content
A Reddit post claims Qwen3.8-Flash-Next beats DeepSeek V4 Pro-no benchmarks, no tables. We unpack the anxiety behind this empty post.
DeepSeek Agent Tutorial Hits Chapter 10 — China's LLMs Now Chase Developers
DeepSeek's Harness Agent has a 10-chapter Chinese tutorial. The shift: top LLM firms are moving from benchmark wars to developer ecosystem grabs.
DeepSeek V4 Vision Weights Delayed — China's Open-Source AI Lead Is Shrinking
DeepSeek rewrote rules with open-sourced V3 and R1. Now V4 Vision weights are silent. Question: how long can China's open-source lead last?
DeepSeek Open-Sources Harness — A New Agent Training Paradigm
DeepSeek open-sources Harness for self-evolving agents. A training paradigm shift—Chinese labs chase Silicon Valley's agent path at lower cost.
Zhipu Open-Sources Ox Alpha vs DeepSeek — Open Source Becomes China's LLM Default
Z.AI confirms Ox Alpha is GLM's next gen, open-sourced tonight. After DeepSeek, a second Chinese AI firm bets open source for ecosystem.
DeepSeek Adds Vision at Lowest Domestic Price—But Page-Cloning Trails K3
DeepSeek launched deepseek-v4-flash-vision-exp at ¥0.05/million input tokens—lowest domestic price. Closes Agent gap, but page-cloning lags K3.
DeepSeek Open-Sources Its Agent OS — And That Matters More Than the Code
DeepSeek's Cordis Agent framework dissected in 9 chapters — first time a top Chinese AI firm has exposed production Agent infra at this depth.
$10K Mac Studio M5 Max for Local LLMs? Reddit Math: Cloud Wins
A Reddit user did the math: $10K on a Mac Studio M5 Max equals up to 100B cloud tokens. Local deployment's cost moat is being eroded by pay-as-you-go.
An Engineer Says AI Thinks Slow, Copies Fast — He's Only Half Right
UK engineer Michael Gomes Vieira argues AI inference is slow but distillation is fast. Chinese model teams already use this — but copying speed doesn'
2.8T-Parameter Model Crammed Into Gaming PC — Runs, But Too Slow to Use
Engineer runs 711GB, 2.8T-param Kimi K3 on RTX 4070 Ti + 32GB RAM — normal speed 1-2 tok/s. Dual SSD parallelism pushes DeepSeek from 1.14 to 8.08 tok
DeepSeek Agent Framework Reverse-Engineered — LLM Race Moves to App Layer
DeepSeek's dsh Agent framework just got fully reverse-engineered. China's LLM firms are pushing from models to app infrastructure—real impact for deve
RTX 3090 Runs a 489-Step Local Agent — Cloud LLMs Aren't Big Tech's Privilege
Open-source community quantizes Alibaba's Qwen 3 8B to run a 489-step Agent on a 2020 RTX 3090. Local AI agents finally affordable for SME IT budgets.
DeepSeek Harness After One Week: Open-Source Foundation, Not Claude Code Clone
DeepSeek's Harness got 95K stars in one week. Testers warn it's not a Claude Code clone but an open-source base users must assemble themselves.
OpenCode Frozen by DeepSeek Folder — The Shoddy Reality Under AI Tool Halos
OpenCode CLI froze due to a DeepSeek CLI folder forming a junction loop. A revealing look at AI tool ecosystem fragility.
DeepSeek Reveals dsh Agent Architecture: Same Code, Local or Cloud via Part Swap
DeepSeek's dsh splits Agents into Definition, Provider, Consumer. Swap a part, local code becomes cloud sandbox—real framework bet from Chinese AI fir
EvoX Swarm Mode Lifts Accuracy from 26% to 71%, Putting Architecture in Focus
EvoX, EvoMap's desktop Agent, uses task splitting and context isolation to lift accuracy to 71%, suggesting architecture matters as much as model capa
6GB VRAM, Local AI Coding: A Developer's Plea Exposes Cloud's Real Cost
A Reddit programmer asks: 6GB VRAM, 64GB RAM for local AI coding with sub-minute responses. Behind it: the real ledger of cloud subscription vs local
Newbie asks which upcoming models to wait for — we mapped H2's release calendar
A LocalLLaMA newbie asked what upcoming models to watch. Comments were thin. We mapped H2's release calendar for the top labs.
DeepSeek Model Hits 25 token/s on $7K Mac — Local AI Catches the Cloud
DeepSeek's latest model hits 25.8 tokens/s locally on an M2 Ultra Mac — smaller than official quant. Chinese open-source LLMs are now viable.