1 article tagged with this topic
Reddit user runs Qwen 30B MoE on RTX 3050 6GB, hitting 30 tokens/sec at 90k context. Enterprise IT should reassess local LLM costs.