1 article tagged with this topic
SHADOW-250M: 250M params in 60 MB, 400 tokens/s on CPU. If reproducible, local AI costs drop fast — but single source, no benchmarks yet.