small models
6 articles tagged with this topic
Context Is Burning Cash — Why 0.8B Models Are Easing 70B's Load
A Reddit LocalLLaMA thread: Qwen 0.8B compresses conversation history, hands off to a 70B model for inference. Task-matched routing is quietly reshapi
Liquid AI Ships 2.6B Local Model — The Small-Model Path Is Getting Real
Liquid AI releases 2.6B-parameter LFM 2.5 in GGUF format, runnable on laptops. Paired with Phi, Llama, and Qwen, "local AI" is quietly becoming an ent
Liquid's 2.6B Model Pushes Local AI Closer to the Mainstream
Liquid AI ships a 2.6B parameter model hitting 260 tokens/sec on an RTX 3090. Limited capability, but enough for daily chores — and it lowers the bar
Gemma 4 Per-Layer Embeds: Knowledge-Reasoning Split, Hope or Hype
Gemma 4's per-layer embeddings spark debate: Can knowledge and reasoning scale separately? If so, 2B models could hold 20B knowledge, redefining local
764 Experiments Reveal: Three Counterintuitive Pitfalls in Small Model Deployment
Copying GPT-4 prompting standards for small AI models causes accuracy to crash 64%—forcing enterprise leaders planning private AI deployment to reasse
The Birth of 9B Local AI Analyst: The Hour of Labor Cost Restructuring for Data Departments
A 9B-parameter local model fine-tuned with LoRA achieves 89.7% autonomous completion of data analysis workflows, fundamentally rewriting the business