A Reddit user this week got a 35B-parameter-class large model running on two secondhand AMD MI50 GPUs: 32GB of VRAM, total hardware cost under $300. We think this is worth attention because it means data-center GPUs (specialized compute accelerators) being phased out of the mainstream market are quietly becoming a low-cost option for running AI locally.

What this is

The poster built a dual-MI50 platform. The MI50 is a 2018 AMD data-center product, now discontinued — 16GB versions go for under $150 each on the secondhand market. Two cards together give 32GB of HBM2 high-bandwidth memory and roughly 2 TB/s of bandwidth. He used llama.cpp (an open-source local large-model inference tool), via its Vulkan build, to run a batch of 27B–35B open-source models.

The standout result was on MoE models (Mixture of Experts — architectures that activate only a subset of sub-networks per inference). Qwen3 35B-A3B hit 983 tokens/s at the prompt stage (the number of text fragments generated per second) and about 47 tokens/s at generation; Qwen3 27B, a dense model (one where all parameters participate every pass), ran around 167/17.

Industry view

Supporters read this as a "compute-equalization" signal: retired data-center cards no longer have to be scrapped for parts — they can run mainstream open-source models, which clearly matters for budget-constrained small teams and individual hobbyists.

The pushback is real. Veteran users point out that the MI50 lacks high-speed multi-GPU interconnect protocols like NVLink, so two-card setups rely on PCIe (the motherboard expansion bus), and scaling beyond four cards is effectively a non-starter. AMD's software ecosystem — ROCm and the rest — lags NVIDIA's noticeably, and when things break you often have to dig through source code and patch drivers yourself. "Cheap hardware" does not equal "hardware you can actually use" — a lesson the industry has relearned many times.

Impact on regular people

  • For enterprise IT: The hardware bar for self-hosted inference (running the model locally instead of through a cloud API) drops from thousands of dollars to a few hundred, so testing and small-scale production no longer depends entirely on cloud services.
  • For individual professionals: Tech enthusiasts can now experience local large models at extremely low cost, but there's still a clear gap between "it runs" and "it runs reliably" — and the value for non-technical roles is limited.
  • For the consumer market: No near-term shift in the consumer GPU landscape — gaming and video editing will stay on NVIDIA and AMD mainstream cards — but we may see a crop of small vendors offering "local AI workstation" build-outs.