Back to home

small models

6 articles tagged with this topic

Qwen千问

Context Is Burning Cash — Why 0.8B Models Are Easing 70B's Load

A Reddit LocalLLaMA thread: Qwen 0.8B compresses conversation history, hands off to a 70B model for inference. Task-matched routing is quietly reshapi

Aug 232 min read
Liquid AILFM 2.5

Liquid AI Ships 2.6B Local Model — The Small-Model Path Is Getting Real

Liquid AI releases 2.6B-parameter LFM 2.5 in GGUF format, runnable on laptops. Paired with Phi, Llama, and Qwen, "local AI" is quietly becoming an ent

Aug 192 min read
Liquid AILFM 2.6B

Liquid's 2.6B Model Pushes Local AI Closer to the Mainstream

Liquid AI ships a 2.6B parameter model hitting 260 tokens/sec on an RTX 3090. Limited capability, but enough for daily chores — and it lowers the bar

Aug 92 min read
GemmaGoogle

Gemma 4 Per-Layer Embeds: Knowledge-Reasoning Split, Hope or Hype

Gemma 4's per-layer embeddings spark debate: Can knowledge and reasoning scale separately? If so, 2B models could hold 20B knowledge, redefining local

May 32 min read
local deploymentprompt engineering

764 Experiments Reveal: Three Counterintuitive Pitfalls in Small Model Deployment

Copying GPT-4 prompting standards for small AI models causes accuracy to crash 64%—forcing enterprise leaders planning private AI deployment to reasse

Apr 112 min read
local AIdata analysis

The Birth of 9B Local AI Analyst: The Hour of Labor Cost Restructuring for Data Departments

A 9B-parameter local model fine-tuned with LoRA achieves 89.7% autonomous completion of data analysis workflows, fundamentally rewriting the business

Apr 102 min read