1 article tagged with this topic
A Reddit user ran 35B LLMs on Nvidia's 2017 V100 via a vLLM fork, matching new consumer chips. AI inference costs may be looser than most assume.