Back to home

GPU

14 articles tagged with this topic

AMDROCm

AMD Leaps ROCm 7→10 in One Month — A Market-Grabbing Cadence

AMD's ROCm jumps from 7.14 to 10.0 in just one month—the fastest iteration in a decade. Signals potential cracks in NVIDIA's AI compute monopoly.

1d ago2 min read
NVIDIASpectrum-X

GPU 堆到几十万张后,瓶颈变成了网线 — NVIDIA 想重做这件事

NVIDIA's Spectrum-X tackles AI training's network bottleneck. With 100,000+ GPUs per model, traditional Ethernet breaks. AI infrastructure power reshu

5d ago2 min read
WebGPUBrowser

Browsers Can Now Call GPUs Directly — WebGPU Opens the Last Mile for Web AI

WebGPU lets web JS call local GPU compute directly. Major browsers shipped support over the past two years, opening the path for browser-native AI inf

6d ago2 min read
LocalLLaMARTX 3090

Running LLMs at Home Hits a Chassis Wall — Local AI Is Far from Plug-and-Play

A 3090+3070 LLM build hits a PCIe spacing wall on a B450 board—a symptom of consumer hardware never being designed for multi-GPU AI.

Aug 222 min read
NVIDIAQuantitative Trading

NVIDIA Moves Asset Clustering to GPUs — The Quant Compute Arms Race Continues

NVIDIA's AdaptGrow moves financial product clustering to GPUs, cutting a core quant task by an order of magnitude. Our read: finance's compute arms ra

Aug 212 min read
NVIDIACUDA

NVIDIA turns GPU docs into AI-callable modules — LLMs now compete on real work

NVIDIA's CUDA MCP lets AI search GPU docs and write optimized code. The takeaway: LLMs will compete on real-world work, not just intelligence.

Aug 202 min read
DeepSeekNVIDIA

DeepSeek Runs on 16 Consumer GPUs — Frontier Model Deployment Barrier Crumbles

Reddit shows DeepSeek running locally on 16 RTX 5060 Ti at 60% of pro GPU cost. Hardware cost curve for local frontier models is plummeting fast.

Aug 202 min read
NVIDIAUMAP

NVIDIA's Multi-GPU UMAP: 50x Faster Research, Deeper Enterprise Lock-In

NVIDIA parallelizes UMAP across multiple GPUs, slashing research data analysis from hours to minutes — but small teams get priced out as enterprise NV

Aug 182 min read
LocalLLaMANVIDIA

Used AI hardware up 8% in a week — compute crunch hits individual buyers

eBay Australia: AI compute hardware up 8% in a week. Individual buyers now squeezed as supply-demand imbalance spreads from enterprise to consumer tie

Aug 182 min read
Dharma-AIGPU

Same GPU Cluster, New Scheduling Order: 33% Utilization Boost — Don't Celebrate Yet

A Dharma-AI Hugging Face post shows rearranging task scheduling on the same GPU cluster boosted utilization 33 points — no new hardware required.

Aug 172 min read
NVIDIADynamo 1.0

NVIDIA Dynamo 1.0: Rewriting the AI Inference Ledger for the Agent Era

NVIDIA pushes Dynamo inference platform to 1.0—a tech upgrade on the surface, but a billing shift underneath: from 'GPU hours' to 'Agent tasks'.

Aug 172 min read
NVIDIAJensen Huang

Nvidia's Real Moat Isn't Faster Chips — It's 20 Years of Software Lock-in

Nvidia's trillion-dollar AI dominance came not from faster chips, but a 20-year software stack since 2006 nobody dares replace.

Aug 142 min read
NVIDIARTX 5090

RTX 5090's 3x successor: not until 2029–2038, and 1200W may break first

Reddit LocalLLaMA user calculated 3x RTX 5090 performance won't arrive until 2029–2038. Power draw—potentially 1200W—may break first.

Aug 132 min read
YCGPU

GPU Agent Utilization at 30-40%: Purpose-Built Inference Chip Window Opens

YC finds GPU Agent utilization at only 30-40%. Purpose-built inference chips offer an opportunity, but ecosystem lock-in and evolving demand remain ha

May 52 min read