GPU
14 articles tagged with this topic
AMD Leaps ROCm 7→10 in One Month — A Market-Grabbing Cadence
AMD's ROCm jumps from 7.14 to 10.0 in just one month—the fastest iteration in a decade. Signals potential cracks in NVIDIA's AI compute monopoly.
GPU 堆到几十万张后,瓶颈变成了网线 — NVIDIA 想重做这件事
NVIDIA's Spectrum-X tackles AI training's network bottleneck. With 100,000+ GPUs per model, traditional Ethernet breaks. AI infrastructure power reshu
Browsers Can Now Call GPUs Directly — WebGPU Opens the Last Mile for Web AI
WebGPU lets web JS call local GPU compute directly. Major browsers shipped support over the past two years, opening the path for browser-native AI inf
Running LLMs at Home Hits a Chassis Wall — Local AI Is Far from Plug-and-Play
A 3090+3070 LLM build hits a PCIe spacing wall on a B450 board—a symptom of consumer hardware never being designed for multi-GPU AI.
NVIDIA Moves Asset Clustering to GPUs — The Quant Compute Arms Race Continues
NVIDIA's AdaptGrow moves financial product clustering to GPUs, cutting a core quant task by an order of magnitude. Our read: finance's compute arms ra
NVIDIA turns GPU docs into AI-callable modules — LLMs now compete on real work
NVIDIA's CUDA MCP lets AI search GPU docs and write optimized code. The takeaway: LLMs will compete on real-world work, not just intelligence.
DeepSeek Runs on 16 Consumer GPUs — Frontier Model Deployment Barrier Crumbles
Reddit shows DeepSeek running locally on 16 RTX 5060 Ti at 60% of pro GPU cost. Hardware cost curve for local frontier models is plummeting fast.
NVIDIA's Multi-GPU UMAP: 50x Faster Research, Deeper Enterprise Lock-In
NVIDIA parallelizes UMAP across multiple GPUs, slashing research data analysis from hours to minutes — but small teams get priced out as enterprise NV
Used AI hardware up 8% in a week — compute crunch hits individual buyers
eBay Australia: AI compute hardware up 8% in a week. Individual buyers now squeezed as supply-demand imbalance spreads from enterprise to consumer tie
Same GPU Cluster, New Scheduling Order: 33% Utilization Boost — Don't Celebrate Yet
A Dharma-AI Hugging Face post shows rearranging task scheduling on the same GPU cluster boosted utilization 33 points — no new hardware required.
NVIDIA Dynamo 1.0: Rewriting the AI Inference Ledger for the Agent Era
NVIDIA pushes Dynamo inference platform to 1.0—a tech upgrade on the surface, but a billing shift underneath: from 'GPU hours' to 'Agent tasks'.
Nvidia's Real Moat Isn't Faster Chips — It's 20 Years of Software Lock-in
Nvidia's trillion-dollar AI dominance came not from faster chips, but a 20-year software stack since 2006 nobody dares replace.
RTX 5090's 3x successor: not until 2029–2038, and 1200W may break first
Reddit LocalLLaMA user calculated 3x RTX 5090 performance won't arrive until 2029–2038. Power draw—potentially 1200W—may break first.
GPU Agent Utilization at 30-40%: Purpose-Built Inference Chip Window Opens
YC finds GPU Agent utilization at only 30-40%. Purpose-built inference chips offer an opportunity, but ecosystem lock-in and evolving demand remain ha