Back to home

Compare

Comparing: Three Used Mining GPUs Run AI 3x Faster — A Cheap Inflection Point for Local LLMs & 有人用三张旧矿卡跑 AI 提速三倍 — 本地大模型的廉价拐点到了

AEN
llama.cppRTX 3060Flash Next·

Three Used Mining GPUs Run AI 3x Faster — A Cheap Inflection Point for Local LLMs

This week, a post on r/LocalLLaMA caught our attention: user MD_Reptile assembled a local AI inference rig (running models on your own computer to generate text on the spot, rather than calling a cloud API) using three RTX 3060 12GB cards (old mining GPUs, in the ~$150 range on the second-hand market). Running an inference engine called Flash Next, it hit 38-40 tokens/sec.

The baseline for comparison is llama.cpp — currently the most mainstream open-source tool for running large models locally — which delivers only 13.2 tokens/sec on the same hardware. The gap is nearly 3x. The key is Flash Next's strata low-bit quantization (compressing model parameters to extreme low precision to save VRAM; IQ3 means each parameter uses only about 3 bits), effectively packing larger models into the 3060's 12GB VRAM while running faster.

What This Is

In short: old mining GPUs + a new inference engine let a roughly $400-class machine deliver speeds approaching small cloud models. Tokens/sec is the standard unit for measuring LLM generation speed; 40 t/s is roughly 1.5x human reading speed, making for a smooth conversational experience.

Why this matters: local inference has long been stuck on two problems — expensive hardware and slow speeds. If three old cards can solve both, the hardware threshold and electricity costs drop together.

Industry View

The local inference community's reaction is excited but cautious. One camp sees this as a genuine inflection point — when three used cards worth a few hundred dollars can deliver cloud-small-model speeds, the hardware cost of enterprise in-house AI inference will plummet, and the demand for keeping sensitive data on-premises becomes easier to meet. Pricing pressure on cloud inference APIs will follow.

The counter-arguments are equally concrete: first, this is a single-point test, not a full benchmark suite (standardized performance test set); second, the 3060's 12GB VRAM is still narrow for genuinely useful models, and complex tasks can easily blow out memory; third, Flash Next remains a relatively niche tool, with uncertain long-term maintenance and stability. In other words, cheaper hardware doesn't mean ready-to-replace-cloud today.

Impact on Regular People

For enterprise IT: If local inference pricing continues trending down over the next year, private deployment solutions (AI systems where data never leaves the company network) for mid-sized companies will become cost-effective again — worth asking vendors for fresh quotes.

For individual professionals: Tinkering with local LLMs is still an engineer and hardcore enthusiast game today. But within a year, "running a usable AI assistant on your own laptop" may become something ordinary office workers can do.

For the consumer market: Used RTX 3060 mining cards have already been stuck in the GPU market. This trend may turn them from "mining scrap" into "AI starter cards." Readers planning a build or upgrade soon should watch prices.

BZH
llama.cppRTX 3060Flash Next·

有人用三张旧矿卡跑 AI 提速三倍 — 本地大模型的廉价拐点到了

本周 r/LocalLLaMA 上一篇帖子引起我们注意:用户 MD_Reptile 用三张 RTX 3060 12GB(旧矿卡,二手市场千元级)组了一台本地 AI 推理机(让模型在你自己的电脑上现场生成文字,而不是调用云端 API),跑一个叫 Flash Next 的推理引擎,速度达到 38-40 token/秒。

对比基线是 llama.cpp——目前本地跑大模型最主流的开源工具,同硬件下只有 13.2 token/秒。差距接近三倍。关键在于 Flash Next 用的 strata 低比特量化(把模型参数压缩到极低精度以省显存,IQ3 意味着每个参数只用约 3 bit),相当于在 3060 的 12GB 显存里塞进更大的模型,还跑得更快。

这是什么

简单说就是:旧矿卡 + 新推理引擎,让一台三千块级别的机器跑出接近云端小模型的速度。token/秒是衡量大模型生成速度的常用单位,40 t/s 大致是人类阅读速度的 1.5 倍,对话体验已经很流畅。

这件事之所以值得关心,是因为本地推理一直卡在两个问题上:硬件贵、速度慢。如果三张旧卡就能解决,硬件门槛和电费成本会同步下降。

行业怎么看

本地推理社区的反应是兴奋但审慎。一派认为这是真正的拐点——当三张几百块的二手卡能跑出接近云端小模型的速度,企业自建 AI 推理的硬件成本会骤降,敏感数据不出内网的诉求也更易满足。云端推理 API 的定价压力也会随之而来。

反对意见同样具体:第一,这是单点测试,没有跑完整 benchmark(标准性能测试套件);第二,3060 的 12GB 显存对真正能用的模型仍然偏窄,复杂任务容易爆显存;第三,Flash Next 还是相对小众的工具,长期维护和稳定性未知。换句话说,硬件便宜了不等于现在就能替代云端。

对普通人的影响

对企业 IT:如果未来一年本地推理价格曲线继续下行,部分中等规模企业的私有化部署方案(数据不出公司内网的 AI 系统)会重新变得划算,值得让供应商给出新报价。

对个人职场:现在折腾本地大模型仍然是工程师和硬核玩家的游戏。但一年内,「在自己笔记本上跑一个可用的 AI 助手」可能变成普通职场人也能做到的事。

对消费市场:二手 3060 矿卡本来就在显卡市场里滞销,这个趋势可能让它们从「矿渣」变成「AI 入门卡」。近期有装机或升级打算的读者可以留意价格。