An independent engineer ran the numbers all the way down: an 8-GPU H200 server runs about $370K, with pure hardware payback between 14 and 36 months — the key variable isn't price, it's utilization.

What This Is

The user published the full calculation on r/LocalLLaMA. The core figures:

  • Hardware: 8-GPU HGX H200 server quoted at $320K–$420K, median $370K
  • Rental: 34 vendors' on-demand rate of $4.40/GPU-hour (Sept 18), or about $35.20/hour for the full box
  • Pure hardware payback: 14.4 months at 100% utilization, 24 months at 60%, 36 months at 40%

But hardware is only the tip of the iceberg. Four hidden costs remain: power and cooling (the author was quoted above budget by a colo provider), depreciation (we'd recommend halving any assumption — secondary-market datacenter hardware holds thin residual value), ops time, and idle hours. His final call: stay above 60% utilization for two consecutive years and self-hosting wins; below 40%, renting wins. A third path: offload idle compute to an offtake network (external compute marketplace) to sell capacity to other teams and hedge cost.

Industry View

Once we lay the math flat, three counter-views deserve our attention:

Depreciation is widely underestimated. Consumer GPUs typically retain about 30% residual value after three years — and dedicated cards like the H200 fare worse. You're not selling to a typical buyer but to another AI team running the same math, so residual value is structurally compressed.

Utilization is hard to stabilize. Training is bursty (one big fine-tune saturates the cluster for two weeks, then it sits idle); inference is steadier. Small teams rarely sustain 60% utilization.

Hardware iteration is too fast. B200 and GB300 are already on the roadmap. Equipment bought today at $370K will see book value cut in half within 18 months, while cloud providers' next-gen silicon will already be live. Payback math must include "generational depreciation."

Pro-self-hosting voices also exist: organizations with stable model roadmaps, strict data-sovereignty requirements, or latency sensitivity (finance, healthcare) won't find renting cheap either — but only if they actually have someone on staff who can run a cluster.

Impact on Regular People

For enterprise IT: A $300K+ self-built AI cluster remains a major CAPEX for mid-sized companies. Unless the business can stably sustain above 60% utilization, on-demand rental is the more controllable choice.

For individual careers: AI infrastructure cost decisions are entering the mid-market CFO's scope. People who can clearly account for "compute economics" are becoming the new scarce role — a viable transition path for IT or finance backgrounds.

For consumer markets: The AI products you use daily almost certainly run on rented compute. Watching H100/H200 rental prices is effectively forecasting whether subscription fees will rise or fall next year.