We have noticed an unsettling thesis circulating among investors: AI's future may not depend on larger models (LLMs, large language models), but on smaller models (SLMs, small language models, typically with parameters in the single-digit billions, capable of running locally on phones and PCs). If this is true, the hundreds of billions of dollars Microsoft, Google, Amazon, and Meta have sunk into data centers over the past two years may face repricing.
What this is
"Hyperscalers" refers specifically to Microsoft, Google, Amazon, and Meta. Over the past two years their strategies have aligned tightly: build larger data centers, stockpile more GPUs (specialized chips for AI training), train larger models, then sell "AI compute subscriptions" to enterprise customers.
The thesis thrown out this week by the investment column Klement on Investing is that this path could be upended by the "smallification" trend. SLMs are already capable of handling most everyday tasks and can run locally on phones, PCs, even factory equipment—enterprises no longer need to pay the cloud for every inference (the process of getting a model to produce an answer). That cuts the legs out from under the core logic on which cloud vendors monetize.
Industry view
Evidence backing the "small model camp" is mounting: Microsoft itself has shipped Phi-4; Apple's Apple Intelligence runs on local models of roughly 3 billion parameters; Meta released the 1B and 3B versions of Llama 3.2; and Microsoft Research has published a paper specifically titled "Small Language Models are the Future of Agentic AI."
But the counterarguments are equally hard-hitting. OpenAI, Anthropic, and xAI are still pushing frontier models ever larger; for complex reasoning, long-chain tasks, and code generation, "small models" still can't quite beat GPT-4-class systems. Andrew Ng and other industry voices have repeatedly stressed recently that the bottleneck for AI project deployment is not that models are too small, but engineering and data governance. OpenAI CEO Sam Altman has likewise emphasized that scaling laws—the rule of thumb that "more parameters plus more data yields a stronger model"—have not yet plateaued.
Our read of the situation: models may be "getting smaller," but frontier research is still "getting bigger"—two forces coexisting, and no one can declare a winner.
Impact on regular people
For enterprise IT: beyond cloud-based large-model APIs, local small models have become a new option. Data-sensitive industries (healthcare, finance, government) stand to benefit especially—running AI while keeping data on internal networks.
For individual professionals: phones and computers will become increasingly "locally intelligent"—tasks like drafting emails and organizing notes won't all need to route through the cloud, and will work offline.
For consumer markets: AI applications will become cheaper and more widespread, but the privacy boundary will shift—your chat logs and files will increasingly stay on-device rather than being uploaded to cloud vendors.