A long-time self-hosting enthusiast basnijholt on Reddit's r/LocalLLaMA published a blog post doing the math: self-hosting AI is actually not cheaper than subscribing to an API (a pay-per-call cloud interface).
His comparison is blunt — a consumer-grade GPU (graphics processor, the core hardware behind AI compute) server capable of running open-source models like Qwen 3.8 27B requires roughly ¥15,000 in upfront hardware, plus monthly electricity, SSD (solid-state drive) depreciation, and the occasional hardware failure. Annualized, the cost easily exceeds ¥2,400. Meanwhile Claude's Opus 5.5 costs $240 per year (about ¥1,700) for a subscription — essentially buying "no compute to manage, just results" convenience. The kicker: he admits all of this and still chooses to self-host.
What this is
"Self-hosting" means downloading an AI model and running it on your own machine, rather than calling a cloud vendor's API. In the Chinese context, it is often translated as "private deployment" (私有化部署) — an unavoidable option when banks, hospitals, and large-enterprise legal departments evaluate AI. The core conclusion of basnijholt's blog: on hardware-plus-electricity alone, self-hosting is almost always more expensive than the API. But he still self-hosts — and the reason is not money.
Industry view
Supporters see this as the price of data sovereignty. Customer emails, internal chats, location histories — these assets cannot be coaxed out the door by any API terms of service. Zero-retention promises (where vendors commit to not storing user data) sound reassuring, but legal and compliance teams don't necessarily accept them. This is a hidden cost that never appears on a cloud vendor's price sheet.
The counterargument is just as direct. Consumer-grade GPUs double in performance every 18 months — today's RTX 4090 won't keep up with new models in two years. Electricity, noise, and heat aren't friendly to home users. When something breaks, there's no SLA (service-level agreement, the vendor's written commitment on availability) to fall back on. An even sharper point: APIs and local models aren't even the same kind of product. A $200 subscription gets you today's strongest model; a 27B open-source model running locally is several tiers behind. Building the case around "saving money" simply doesn't hold up.
For enterprises, the more practical question is this: the true TCO (total cost of ownership, including hardware, labor, operations, and upgrades) of private deployment is often underestimated by sales and overestimated by CIOs (Chief Information Officers). Over a three-year horizon, the gap with annual subscription fees isn't as wide as you'd think.
Impact on regular people
For enterprise IT: Manufacturing, finance, and healthcare managers evaluating private deployment need to fold hardware depreciation, operations headcount, and model upgrade costs into a three-year total — not just take the vendor's quote at face value.
For working professionals: For ordinary white-collar workers and indie developers, ChatGPT or Claude subscriptions remain the best value. Unless your work involves sensitive data, there's no reason to mess with local deployment.
For consumer markets: The marketing pitch for home AI boxes and AI PCs deserves scrutiny — hardware costs are high, model updates are slow, and the value proposition usually loses to a monthly subscription.