What this is

A user posted on Reddit's LocalLLaMA forum: they are running llama.cpp on an Nvidia P40 (a 2016 datacenter-grade professional GPU with 24GB of VRAM, currently priced between a few hundred and just over 1,000 RMB on the second-hand market), loading a 35B-parameter Qwen-series model. IQ is one of llama.cpp's quantization methods, more aggressively compressing the model to a smaller footprint than standard Q4.

Test results: generation speed fluctuates between 37 and 83 tok/s (tokens of text emitted per second). The model's reading of the input — prefill, the "reading" step where the model processes your question — starts at 600 characters per second, dropping to 300–400 with long inputs. Subjectively, the user reports it does not feel slower than standard Q4 quantization.

However, two-year-old references they found claim IQ quantization is clearly slower on P40. The question is whether that conclusion still holds.

Industry view

Arguments supporting "IQ is slower": IQ quantization requires CPU-side real-time dequantization (restoring compressed data back into usable numbers) during inference, and P40's older architecture is not friendly to this — theoretically a speed hit.

The counterargument — and where this user's tests lean — is that the data behind that two-year-old discussion is likely outdated. Newer llama.cpp has shipped significant optimizations. And with only 24GB of VRAM, P40 is already at the edge running a 35B model; speed was never going to be pretty.

A risk worth flagging: P40 is currently the most-discussed "lowest-barrier hardware for local LLMs," but stability, power draw (around 250W per card), and driver compatibility all lack long-term data. The second-hand market is a mixed bag, and enterprise self-builds should proceed with caution.

Impact on regular people

For traditional enterprise IT: budget-constrained departments are indeed experimenting with retired server GPUs to build local AI sandboxes, and P40 is the most-discussed option — but it remains an informal procurement category, not a sanctioned one.

For working professionals: unless there is a clear data-privacy or offline requirement, the price-performance of cloud APIs still vastly outweighs the hassle of local tinkering.

For the consumer market: most stories about local AI on consumer hardware revolve around the RTX 4090 and RTX 5090. The P40 path is developer territory, still far from mainstream.