Input vs Output Token Pricing: Which One Actually Bills You
Output tokens cost several times more per token, yet input usually dominates the bill. Here is how to compute your own blended rate and act on it.
Every provider charges more per output token than per input token, usually by a factor of four to six. The natural conclusion is that output is where the money goes. For chat workloads that is often true. For anything agentic it is emphatically false, and acting on the wrong assumption sends optimisation effort in the wrong direction.
What matters is not the unit prices. It is the ratio between them multiplied by the ratio of volumes in your actual traffic.
The two ratios
Define them plainly:
price_ratio = output_price_per_token / input_price_per_token
volume_ratio = input_tokens_used / output_tokens_used
Output dominates your bill when price_ratio > volume_ratio. Input dominates when the reverse holds. That is the whole analysis.
As of August 2026, published list prices give price ratios in a narrow band. Anthropic lists Claude Opus 5 at $5 input and $25 output per million tokens — a ratio of 5. Claude Haiku 4.5 sits at $1 and $5, also 5. OpenAI lists GPT-5.6 Sol at $5 and $30, a ratio of 6. DeepSeek published V4 Pro at $0.435 cache-miss input and $0.87 output, verified late July 2026 — a ratio of 2.
So price ratios cluster between 2 and 6. Volume ratios, by contrast, span three orders of magnitude.
Where different workloads land
Long-form generation. A short brief producing a lengthy document: perhaps 500 input tokens and 3,000 output. Volume ratio 0.17, well below any price ratio, so output dominates overwhelmingly. Cutting output length is the only lever that matters.
Classification and extraction. A 4,000-token document producing a 60-token label. Volume ratio 67, far above any price ratio, so input dominates by roughly an order of magnitude. Shortening the response is pointless; shrinking or caching the document is everything.
Chat. Roughly balanced early in a conversation, tilting towards input as history accumulates. A ten-turn conversation is usually already input-dominated.
Agentic coding. The extreme case. Because the model is stateless, every step resends the full conversation, and tool output enters context and stays. Volume ratios of 20 to 100 are normal. Input dominates by a wide margin, and no amount of output trimming changes the picture.
Computing your own blended rate
The single most useful number is the effective cost per request:
blended_cost = (in_tokens x in_price) + (out_tokens x out_price)
input_share = (in_tokens x in_price) / blended_cost
Worked for an agentic coding request averaging 45,000 input and 900 output tokens, at $5 and $25 per million:
input: 45,000 / 1e6 x $5 = $0.2250
output: 900 / 1e6 x $25 = $0.0225
blended cost per request = $0.2475
input share = 91%
Ninety-one percent. If you halved every response length in this workload you would save four and a half percent. If you cut input by a third — through caching, tighter tool output, or better file selection — you would save thirty percent.
Now the same calculation for a summarisation endpoint averaging 1,200 input and 800 output:
input: 1,200 / 1e6 x $5 = $0.0060
output: 800 / 1e6 x $25 = $0.0200
blended cost per request = $0.0260
input share = 23%
Here the advice inverts completely. Capping output length is the dominant lever and input work barely registers.
Run this per endpoint, not for your account as a whole. Aggregate figures average away exactly the distinction you need.
Two things that distort the comparison
Cache reads are priced separately, and cheaply. Anthropic prices cache reads at roughly 10% of the base input rate, with cache writes at 1.25x for the five-minute TTL. OpenAI publishes a similar structure for GPT-5.6, with cached input at $0.50 against $5 uncached for Sol. This means your effective input price is not the list price — it is a weighted average of cached and uncached reads. A workload with an 80% cache hit rate has an effective input price near a quarter of list, which can flip a marginal workload from input-dominated to output-dominated.
Long-context tiers. Some providers charge a premium above a context threshold. OpenAI's published GPT-5.6 Sol pricing rises from $5/$30 to $10/$45 for long-context requests. If your workload straddles that boundary, the input side of your calculation has a step function in it and averages will mislead.
Reasoning tokens count as output
Models that produce internal reasoning bill those tokens at the output rate even when the text is not returned to you. On a workload with heavy reasoning and short visible answers, your measured output token count may be several times the length of what you actually see.
This is worth checking before concluding that output is negligible. Read the usage fields your provider returns rather than counting characters in the response.
What to do with the answer
If input dominates — the usual case for agents and retrieval:
- Cache the stable prefix. This is the largest single win available.
- Truncate tool output before it enters context, and say that you did.
- Select files rather than dumping directories.
- Compact conversation history before it reaches the window limit.
If output dominates — long-form generation, code synthesis from short briefs:
- Ask for the diff rather than the whole file.
- Set
max_tokensas a genuine backstop, not a formality. - Specify length in the prompt; models comply well with explicit limits.
- Check whether reasoning depth settings are producing more than the task needs.
And if you are on flat-rate access, the whole calculation stops mattering for cost and continues to matter for latency and quality. Fewer input tokens still means faster responses and better attention, which is a reason to keep doing the work even when it no longer shows up on an invoice.
Common questions
Should I optimise input tokens or output tokens first?
Compute the input share of your blended cost per request for each endpoint. Whichever side exceeds roughly 70% is where effort pays. For agentic workloads that is almost always input; for long-form generation it is output.
Do cached tokens change which side dominates?
Yes. Cache reads are typically around a tenth of the base input rate, so a high hit rate can cut your effective input price by three-quarters or more, which is enough to flip a borderline workload towards output-dominated.
Are reasoning tokens billed as input or output?
As output, at the output rate, even when the text is not returned to you. Read the usage fields in the API response rather than measuring the visible answer, or you will undercount output substantially.