Is Self-Hosting an Open Weight Model Cheaper?
Open weights make self-hosting sound obviously cheaper. Converting the comparison into sustained throughput shows why it usually is not.
For almost every team asking, no: self-hosting an open weight model is not cheaper than calling an API, and it is not close. The reason is that a rented GPU bills every second whether you use it or not, while an API bills only tokens. That turns the comparison into a utilisation question, and once you express the break-even as sustained tokens per second rather than monthly dollars, most teams discover they are three orders of magnitude below it.
Open Weight Model Hosting Versus an API: Two Cost Shapes
An API cost is linear in usage. Ten times the tokens costs ten times as much, and zero tokens costs zero.
A self-hosted cost is a step function in capacity. You rent a machine by the hour and the bill is identical whether it is saturated or idle. Every unused second is money you have already spent.
That is the whole argument. Self-hosting wins when you can keep the hardware busy, and loses, often spectacularly, when you cannot.
Real Numbers, September 2026
For GPU rental, IntuitionLabs' comparison across providers puts the median on-demand H100 at $3.36 per GPU-hour as of 20 September 2026, across 38 providers. Budget marketplaces list from roughly $2.10, hyperscalers run considerably higher, and the cheapest listings tend to be oversubscribed capacity rather than something you would put production on.
For the API side, take an open weight model with public pricing. Xiaomi's MiMo-V2.6 Flash, a 310 billion parameter model with 15 billion active, is priced at $0.14 per million input tokens and $0.28 per million output, with the larger MiMo-V2.6 Pro at $0.435 and $0.87, per its model page. Those are the same weights you would be self-hosting, which makes this an unusually clean comparison.
The Break-Even, in Tokens Per Second
One H100 at the $3.36 median, running continuously, costs about $80.64 a day and roughly $2,420 a month.
At $0.28 per million output tokens, $2,420 of API spend buys about 8.6 billion output tokens. Spread across a month that is roughly 3,300 output tokens every second, sustained, day and night, before your self-hosted setup has broken even against a single rented GPU.
For the larger model the arithmetic gets worse rather than better, because a 1.02 trillion parameter model does not fit on one card. At one byte per parameter it needs somewhere north of a terabyte of memory for weights alone, so at least 13 H100s before you account for key-value cache, activations or any redundancy. Call it $31,000 a month at the same median rate. At $0.87 per million output tokens that is about 36 billion output tokens, or roughly 13,900 tokens per second sustained.
Setup | Monthly hardware cost | Equivalent API tokens | Sustained rate to break even |
|---|---|---|---|
1 H100, smaller model | about $2,420 | about 8.6B output tokens | about 3,300 tokens/sec |
13 H100s, trillion-parameter model | about $31,000 | about 36B output tokens | about 13,900 tokens/sec |
These are back-of-envelope figures using the median rental rate and a one-byte-per-parameter assumption, and they ignore serving efficiency in both directions. They are not a quote. They are a sanity check, and as a sanity check they are decisive: if your product is serving a few hundred users a day, you are not within a factor of a thousand of that line.
When Self-Hosting Actually Wins
Cost is usually the wrong reason to self-host. These are the right ones:
Data residency or contractual terms that forbid sending data to a third party. This is the most common genuine driver and cost barely enters into it.
A latency floor you cannot reach over the public internet, typically for on-premises or edge workloads.
Sustained high volume, meaning you genuinely are near the throughput numbers above.
Heavy fine-tuning where you need the modified weights under your control.
Predictability, where a fixed monthly bill is worth more to you than a lower variable one.
Availability is worth a line of its own. A model you host cannot be deprecated, reprised, or restricted underneath you, which is the substance of the open versus closed trade-off rather than the price tag people usually lead with.
The Costs the Comparison Usually Omits
The GPU bill is the visible number. The rest is not:
Idle time. Traffic is diurnal. A machine sized for your peak sits mostly empty overnight, so effective utilisation of 20% to 30% is normal and triples your real per-token cost.
Operations. Someone has to handle drivers, serving frameworks, batching configuration, upgrades and the pager. That is a part-time role you may not have budgeted.
Redundancy. One GPU is not a production deployment. A second machine, or a fallback API key, or accepting downtime, each has a price.
Model upgrades. The API provider's improvements arrive for free. Yours arrive when you redo the deployment work.
Batching is the one lever that genuinely moves these numbers in your favour. Serving frameworks that batch concurrent requests can push a card's throughput several times above what sequential serving achieves, which lowers the effective break-even. It only helps if you have concurrent requests to batch, though, and a product with bursty single-user traffic does not, so batching tends to reward exactly the teams who were already past the line.
A useful move before any of this is to check whether you have an efficiency problem rather than a pricing problem, because caching, shorter prompts and smaller models routinely take more off a bill than changing where the model runs. After that, compare what the same workload costs across providers, and only then price hardware. The hardware question is also getting more interesting as inference silicon diversifies beyond GPUs.
Frequently Asked Questions
What about renting by the second instead of the month?
Serverless GPU platforms that bill per second genuinely help with idle time, but they reintroduce cold starts, and loading tens or hundreds of gigabytes of weights is not fast. They suit bursty batch work far better than interactive traffic.
Does running a smaller model on cheaper hardware change the answer?
It shifts the break-even down, sometimes a lot. A small model on a consumer card you already own can be genuinely cheap. The analysis above concerns frontier-tier open weight models, where the hardware floor is high.
Is buying GPUs outright cheaper than renting?
Over a multi-year horizon at high utilisation, often yes. It also converts a variable cost into capital expenditure plus hosting, power and depreciation, which is a different business decision rather than a cheaper version of the same one.
What is the cheapest way to start?
Use the API for the same open weight model. You keep portability, because moving to your own hardware later is a deployment change and not a rewrite, and you find out what your real token volume is before committing to capacity. That sequencing question sits inside the broader monetisation picture.
How did this land?
About the author

Growth & SEO Lead
Manuele covers distribution: SEO, content strategy, and how AI-built products find their first thousand users. He tests everything he recommends.


