Self-Hosting vs a Managed API: Find Your Break-Even
Running open weights on rented GPUs has a break-even volume. Here is how to compute yours from throughput and utilisation, and why most teams never reach it.
Self-hosting an open-weights model looks cheaper because the GPU bill is a fixed number and the API bill is not. The comparison only works if you convert both to the same unit: cost per million tokens, at your actual utilisation.
That last clause is where nearly every self-hosting business case falls apart. Here is the calculation, done honestly.
Convert GPU hours into a per-token price
Three inputs:
hourly_gpu_cost — what you pay per GPU-hour
throughput — output tokens/second the deployment sustains
utilisation — fraction of paid hours doing useful work
Then:
cost_per_million_output_tokens
= hourly_gpu_cost / (throughput x 3600 / 1e6) / utilisation
For the first input, market rates as of 2026 for an NVIDIA H100 span a wide range: specialist GPU clouds advertise roughly $1.50 to $3.50 per hour, with a market median commonly quoted in the $2.29 to $3.12 band, while the large hyperscalers land nearer $7 on AWS and around $12 on Azure. That is a five-to-eight-fold spread for the same silicon, and it is the single biggest variable in the whole model.
Throughput is yours to measure. It depends on the model, quantisation, batch size, sequence length and serving stack. Do not take a vendor benchmark number — run your own traffic shape.
The worked comparison
Take a mid-size open model served on two H100s at $2.50 per GPU-hour, sustaining 900 output tokens per second aggregated across concurrent requests.
hourly cost = 2 x $2.50 = $5.00/hour
tokens per hour = 900 x 3600 = 3,240,000
raw cost per 1M output tokens = $1.54
At 100% utilisation, $1.54 per million output tokens. Compare that against DeepSeek's published V4 Pro rate of $0.87 per million output tokens, verified late July 2026, and the self-hosted deployment is already more expensive than a managed endpoint serving a comparable open model — before you have paid a single engineer.
Now apply real utilisation. A deployment sized for peak weekday traffic, idle overnight and at weekends, commonly runs at 20 to 35% of capacity:
at 30% utilisation: $1.54 / 0.30 = $5.13 per 1M output tokens
at 20% utilisation: $1.54 / 0.20 = $7.70 per 1M output tokens
You are now above Anthropic's listed $5 per million input token rate for Claude Opus 5 and in the same neighbourhood as its $25 output rate only once you factor in that frontier models and mid-size open models are not doing the same job.
Utilisation is not a detail. It is the term that decides the answer.
The costs the GPU invoice does not include
Engineering time. Someone builds the serving stack, tunes batching, handles model updates, debugs OOMs at 3am, and maintains the deployment forever. At a fully loaded US developer cost — median salary around $132,270 per BLS data, plus 30–40% overhead — half an engineer is roughly $90,000 a year, or $7,500 a month. That is 1.5 million tokens a day of managed API spend at Opus-tier rates, every day, purely in salary.
Redundancy. One node is not a production deployment. Two regions or two nodes doubles the fixed cost and halves your utilisation again unless traffic doubles too.
Headroom for spikes. A managed API absorbs a traffic spike; your cluster queues or drops. Sizing for the spike means paying for the spike all month.
Model churn. Open-weights releases arrive frequently. Every migration is evaluation work, serving-stack work, and prompt re-tuning. On a managed API that is a string change.
Where self-hosting genuinely wins
It is not never. The cases are specific:
- Sustained high volume with a flat profile. Batch pipelines running near capacity around the clock push utilisation towards 80% or more, which is where the fixed-cost model finally pays.
- Data residency or contractual isolation that no managed provider will satisfy. Here cost is not the deciding variable at all.
- Fine-tuned or heavily modified weights that no endpoint serves.
- Predictable latency requirements where you need to control queueing rather than share it.
- Owned hardware already amortised for another purpose, where marginal GPU cost is power rather than rent.
Notice what is not on that list: "we use a lot of tokens". Volume alone is handled better by negotiating a managed contract than by building an inference team.
Compute your break-even before you build anything
Set the two costs equal and solve for monthly volume:
fixed_monthly = (gpu_count x hourly_rate x 730) + amortised_engineering
break_even_tokens = fixed_monthly / managed_price_per_token
With two H100s at $2.50/hour and half an engineer:
gpus: 2 x $2.50 x 730 = $3,650/month
engineering: = $7,500/month
fixed total = $11,150/month
against a managed open-model endpoint at $0.87 per 1M output:
break-even = 11,150 / 0.87 = 12.8 billion output tokens/month
Roughly 430 million output tokens a day. If that number is far above your traffic — and for most teams it is by two or three orders of magnitude — the business case is not close, and no amount of tuning closes it.
Re-run the same calculation against a frontier-model API and the break-even drops sharply, because the price per token is much higher. But then you are comparing a mid-size open model against a frontier model, which is a capability comparison, not a cost one. Keep the two questions separate.
The middle option people skip
Between "call a frontier API" and "run your own cluster" sits a managed endpoint serving open weights — someone else's GPUs, someone else's serving stack, per-token billing, open model. It captures most of the price advantage of open weights with none of the operational cost, and for teams whose actual motivation was "frontier models cost too much for this workload", it is usually the correct answer.
The same is true of flat-rate access to a managed gateway: it fixes your monthly number the way self-hosting does, without the cluster. It is a good fit when usage is heavy and continuous, and a poor one when usage is light — in which case metered per-token billing remains the cheapest thing available to you.
A decision rule
Self-host when you can honestly answer yes to all four: utilisation above 70% sustained, an existing platform team with GPU serving experience, a volume above your computed break-even with room to spare, and a reason the workload cannot leave your infrastructure. Anything less than all four, and the managed path is cheaper once engineering time is priced.
Common questions
How do I convert a GPU hourly rate into a cost per million tokens?
Divide the hourly cost by measured tokens per hour, then divide again by your utilisation fraction. Tokens per hour is sustained tokens per second times 3,600. Utilisation is the term that usually decides the answer.
What utilisation do self-hosted deployments actually achieve?
Deployments sized for weekday peaks and idle overnight commonly run at 20-35%. That triples to quintuples the effective per-token cost against the naive calculation, which is why paper business cases fail in production.
Is self-hosting ever cheaper than a managed API?
Yes, for sustained high-volume batch workloads at high utilisation, for fine-tuned weights no endpoint serves, and where data residency forces it. Volume alone is usually better addressed by a managed contract than by building an inference team.