This week, a technical post on Reddit caught our editorial team's eye: a tinkerer modded a $280 (about 2,000 RMB) second-hand mining FPGA chip bought on eBay and successfully ran Alibaba's Tongyi Qianwen Qwen3.5 9B model, generating at roughly 3 tokens per second (a token is the smallest unit a model processes—one Chinese character or roughly half an English word). Slow, but it runs.

What This Is

The FPGA (a reprogrammable hardware chip originally used for crypto mining) runs $280 per card, paired with 8GB of HBM2 high-speed memory. Technically, the author quantized the model to INT4 (cutting parameter precision from 32-bit to 4-bit, shrinking the model to one-eighth of its original size) and used two cards in parallel to run the 9B model.

Measured performance: the 9B model generates ~3 tokens/sec, with prefill at ~6 tokens/sec; the 27B model on a dual-chip version ($375) is estimated at 3–8 tokens/sec, and hits 25 tokens/sec with 4 chips in parallel. Local deployment, no cloud API fees, no data uploaded to third parties—that's why this matters.

Industry View

Supporters see this as a signal of "AI democratization at the edge": running a 27B model used to require tens of thousands of dollars in NVIDIA GPUs; now, in theory, a few hundred dollars gets it done. For data-sensitive industries—finance, healthcare, legal, government—on-premises LLMs are no longer out of reach.

The pushback is equally sharp. First, 3 tokens/sec is unusable for daily office work. Second, FPGA programming has an extremely high barrier to entry—you need to know hardware description languages and RTL design, and a typical IT team cannot replicate this. Third, cloud inference prices are still falling fast, with Volcano Engine and Alibaba Cloud both engaged in price wars; on total cost of ownership, self-hosting may not actually beat them.

Impact on Regular People

For enterprise IT: Industries with strict data compliance requirements (finance, healthcare, central SOEs) should reassess their on-premises strategy—dropping the hardware bar from the million-dollar tier to the hundred-thousand tier is now realistic.

For individual careers: Regular employees don't need to worry short-term—cloud APIs remain the best price-performance choice. But engineers with a combined "hardware + AI" background will become more valuable.

For consumer markets: Within the next 1–2 years, we may see sub-thousand-yuan home AI box products, but whether the experience can catch up with the cloud remains an open question.