1 article tagged with this topic
A Reddit user squeezed a quantized Qwen onto two RTX 5060ti GPUs at 30 chars/sec — local AI's practical tipping point. Runnable, but not useful.