AMD
26 articles tagged with this topic
Community Qwen3.8 Quantization Saves 30GB — Local LLM Bar Drops Again
Community dev agentionai ships a custom quantized Qwen3.8-Flash-Next, 20–30GB smaller than mainstream versions at comparable quality. Local LLM bar dr
Ubuntu 26.04 + AMD ROCm 10.0: Local AI Advances, NVIDIA Still Safe
Ubuntu 26.04 LTS ships with Kernel 7.0 and AMD ROCm 10.0, sparking r/LocalLLaMA benchmarks. AMD challenges NVIDIA's pricing—but the real bottleneck li
AMD Leaps ROCm 7→10 in One Month — A Market-Grabbing Cadence
AMD's ROCm jumps from 7.14 to 10.0 in just one month—the fastest iteration in a decade. Signals potential cracks in NVIDIA's AI compute monopoly.
$6,500 Used GPUs Hit Top AI Speed — Solo Dev Rewrites Inference Economics
GitHub's curvedinf pushed Qwen3 27B from 15 to 972 token/s (65x) on $6,500 of used AMD MI100 GPUs. Enterprise AI's hardware bar just dropped.
Lemonade SDK Unifies 15 AI Engines to Make Local LLMs Plug-and-Play
AMD-led open-source Lemonade unifies 15 AI engines behind one API across NVIDIA, AMD, Apple, Qualcomm hardware—making local LLM deployment viable for
$35K AMD DIY Rig Runs Qwen 122B Locally, 2x Faster — Private AI Now PC-Tier
A Reddit user built a $5K AMD-based local AI workstation (Framework Strix Halo + R9700) running Qwen 122B at 2x speed. The real story: workstation-gra
Qwen 27B Long-Context Speed Drops 65% — Local Players Distrust Benchmarks
AMD's 51.8 tokens/sec Qwen3 27B claim is ideal-only, Reddit user finds; real long-context drops 75→26 (–65%). Local LLMs now judged on stability.
GMKtec Unveils IFA Berlin Mini PC — Local LLMs Enter Workstation Class
GMKtec launches an AMD Ryzen AI Max+ PRO 395 mini PC at IFA Berlin 2026—a sign local LLMs are going pro, though pricing and software raise questions.
llama.cpp Fork Turns Retired AMD Server GPUs into AI Inference Rigs
Reddit user milpster and Zhipu AI's GLM team release a llama.cpp fork optimized for AMD GFX906, letting Mi50, Mi60, and Radeon VII run LLMs locally.
Consumer GPU Hits 153 tok/s on a 27B Model — Local AI Cost Inflection Arrives
A Reddit developer ran 159 experiments: Qwen3.8-27B on consumer hardware beats dual-3090 cloud servers on HumanEval. Just swapping the chat template s
Used P40 and MI50 Can Still Run Local LLMs — But This Path Is Narrowing
Local AI tinkerers ask: can a $150 used data center GPU still run LLMs in 2025? The hardware market is quietly redrawing the entry threshold.
DeepSeek V4 at 22-28 Tokens/sec on a Mini PC — Local LLMs May Finally Be Usable
We see DeepSeek V4 Flash hit 22-28 tokens/sec on AMD Strix Halo. Chinese open-source LLMs shifting from 'big' to local — but hardware costs stay high.
Used AI hardware up 8% in a week — compute crunch hits individual buyers
eBay Australia: AI compute hardware up 8% in a week. Individual buyers now squeezed as supply-demand imbalance spreads from enterprise to consumer tie
Qwen 27B Hits 19 tok/s on a Laptop — Open-Source LLMs Are Finally Usable Locally
Alibaba's open-source Qwen 27B hits 19 tok/s on a 128GB ROG laptop, autonomously building an HTML flight simulator via Agent mode.
AMD GPUs Run Local LLMs 50% Faster — but Barriers Remain
Llama.cpp on ROCm 7.14 makes AMD Radeon 780M ~50% faster on dense models. AMD's first usable local AI option — limited to dense models, requires sourc
European GPU Prices Surge 19% in a Month as AI Compute Spills Into Consumer Hardware
Consumer GPU prices up 19.2% in a month across 9 EU countries, per PriceSquirrel—three straight weeks of hikes. PC gamers feel it; the real cause is A
Local LLMs Finally 'Run' on Laptops — But Three Gaps Keep Them Off Your Work PC
Qwen 3.8 adds MTP (30–60% faster). Strix Halo and M4/M5 unified memory run 70B models on $2K laptops — but 'runnable' isn't 'work-ready.'
Intel Razor Lake AX Hits 512GB/s Bandwidth — Local LLM Hardware Ceiling Cracks
Intel's Razor Lake AX mobile processor surfaces in HWiNFO logs with 512GB/s memory bandwidth — nearly 2× Strix Halo. Local LLM hardware ceiling may cr
Radeon 7600 Runs 35B MoE Model — Local AI's Hardware Barrier Just Dropped
A Reddit user ran Qwen's 35B MoE model at 21 t/s on a ~$280 Radeon 7600 + 64GB DDR4 + Ryzen 5600. Local AI's hardware barrier is dropping fast.
Build a 48GB Home AI Server: AMD Undercuts NVIDIA by $650 — But the Software Bill Is Not So Simple
Overseas builders are running the numbers: three AMD consumer GPUs deliver 48GB VRAM for local LLM inference at one-third the price of an NVIDIA build
AMD MI25 Used Card at €100: Is Local LLM Viable?
AMD MI25 used cards go for €80-100: 16GB VRAM but inference speed near a decade-old mid-range GPU. We run the numbers on whether local LLMs are worth
Qwen3.5 Runs Fast on AMD GPUs, Slow on NVIDIA — A Bug Nobody Understands Yet
The same Qwen3.5-9B model runs 32% slower than baseline on NVIDIA via CUDA, but 77%-128% faster on AMD via Vulkan. One unverified report, but local-LL
AMD Strix Halo Rumored at 192GB: Local LLM Hardware Bottleneck is Loosening
AMD's next-gen Strix Halo rumored with 192GB unified memory can run 122B LLMs locally. Breaking this memory bottleneck reshapes enterprise private AI
AMD's 128GB Halo Box Prototype Challenges Apple Mac's Local LLM Dominance
AMD's Halo Box prototype (Ryzen 395 + 128GB) gives x86 Mac Studio-rivaling local LLM capacity. We see the local AI inference hardware landscape shifti
AMD In-House AI Mini PC in June: Chipmaker Building Systems is a Major Signal
AMD's in-house Ryzen AI 395 mini PC (June, Lenovo OEM) shows local AI inference moving from concept to product as chipmakers pivot from parts to syste
一个 Reddit 帖子揭示的真相:本地跑 AI 大模型,硬件门槛比厂商说的要高得多
A user's 24GB AMD mini PC could only allocate 8GB VRAM to AI. The fix isn 't simple—and that gap exposes a wider industry problem .