1 article tagged with this topic
DeepSeek's latest model hits 25.8 tokens/s locally on an M2 Ultra Mac — smaller than official quant. Chinese open-source LLMs are now viable.