Per-Token vs Flat-Rate AI Pricing: Run the Numbers Yourself
Metered inference makes sense until you start using agents. Here is the arithmetic that decides which billing model is cheaper for how you actually work.
Per-token billing is the default across the industry, and for good reason: it is honest about the underlying cost structure. Inference genuinely costs more when you use more.
It also has a failure mode that shows up the moment you move from chat to agents, and most people discover it on an invoice rather than a spreadsheet.
Where the cost actually comes from
Two facts about how models work drive everything:
Models are stateless. They do not remember your previous message. Every request resends the full conversation.
Agents loop. A single instruction becomes many requests, each carrying everything before it.
Put those together and input tokens grow roughly with the square of the number of turns. Ten turns is not ten times one turn — it is closer to fifty, because turn 10 re-reads turns 1 through 9.
A worked example
Say a task starts with 8k tokens of context — a system prompt, tool definitions, a couple of files — and each turn adds 1.5k as the agent reads output and writes code.
Turn 1: 8,000 in
Turn 5: 14,000 in
Turn 10: 21,500 in
Turn 15: 29,000 in
Cumulative input over 15 turns: ~275,000 tokens
That is one bug fix. Not one day — one bug. At $3 per million input tokens that is about $0.83, which sounds fine until you notice a productive day is twenty or thirty of those, and that is before output tokens.
The number that surprises people is not the total. It is the ratio: agentic coding uses somewhere between ten and fifty times the tokens of the equivalent chat conversation, for work that feels about the same size.
The behavioural cost nobody prices in
The direct cost is manageable. The indirect cost is worse.
When the meter is visibly running, developers optimise for it. They cut context to save tokens — and get worse answers. They avoid letting the agent explore. They stop mid-task when it looks expensive. They do not build the eval set, because a hundred throwaway runs feels wasteful.
Every one of those is a rational response to metering, and every one makes the tool less useful. You end up paying for a capability and then rationing yourself out of it.
When each model wins
Per-token is cheaper when:
- Usage is spiky — heavy some weeks, nothing others.
- Workloads are short and self-contained: classification, summarisation, extraction.
- You are running production inference where volume is predictable and you can optimise prompts against a known unit cost.
- You need a specific model no flat-rate provider carries.
Flat-rate is cheaper when:
- You use agentic coding tools daily.
- You want predictable spend — a number for the finance conversation, not a range.
- You are experimenting, benchmarking, or learning, where most runs are discarded.
- Multiple tools share one key across a working day.
Work out your own break-even
Do not guess. If you are on a metered provider, pull last month's actual token usage and divide:
break_even_days = flat_rate_price / (your_daily_token_cost)
If a pass costs $50 for 30 days and you currently spend $4 a day, you break even in under two weeks and everything after is upside. If you spend $0.40 a day, stay metered — you are not the customer flat rate is for, and any honest provider will tell you so.
The one adjustment worth making: if you have been rationing, your historical usage understates what you would use without a meter. That suppressed demand is real, and it is usually where the productivity gain hides.
What to check before committing
Flat-rate offers vary enormously in what they actually give you. Before you buy anything, including ours, confirm:
- Which models, named plainly. If a provider is vague about what a given alias serves, treat that as the answer.
- Rate limits. "Unlimited" with an undisclosed throttle is not unlimited.
- Context window actually available to you, which is sometimes lower than the model's headline number.
- API compatibility. An OpenAI-compatible endpoint means your existing tools work with a base-URL change.
- What happens at the end of the term — does access simply stop, or does it silently renew?
The honest summary: if you are a heavy agentic user, flat rate almost certainly saves money and definitely removes a tax on curiosity. If you are a light or bursty user, metered is genuinely the better deal and you should stay there.
Common questions
Why is agentic coding so much more expensive than chat?
Models are stateless, so each step in an agent loop resends the whole conversation. Input tokens grow roughly quadratically with turn count, and one instruction can trigger a dozen or more requests.
How do I estimate my token usage before switching?
Most metered dashboards export per-day token counts. Take a representative fortnight, compute a daily average cost, and divide the flat-rate price by it. Remember that rationing suppresses your historical numbers.
Is prompt caching enough to fix metered agent costs?
It helps a lot where supported, since the stable prefix of a conversation gets discounted. It does not remove the quadratic growth, and cache hits depend on prefixes staying byte-identical — which agent loops frequently break.