Token Accounting: Explaining AI Costs to Finance
Finance teams need cost drivers, allocation and controls, not a lecture on transformers. Here is how to translate token usage into terms a budget owner can act on.
The conversation usually goes badly in the same way. An engineer explains tokens, context windows and agent loops. The finance partner nods, then asks the only question they actually needed answered: is this number going to do that again next month, and who owns it?
Token accounting is the translation layer. It maps a metering unit nobody outside engineering understands onto the three things finance always needs — a driver, an allocation, and a control.
The one paragraph explanation of a token
Use this, not the tokenizer lecture. A token is a billing unit roughly equivalent to three-quarters of an English word, or about three characters of source code. You are charged separately for tokens sent to the model and tokens generated by it, usually at different rates, with output priced higher. Because the model has no memory, every follow-up message resends the whole conversation and is charged again.
That last sentence is the one that matters, and it is the one that explains almost every surprising invoice. Finance is used to units that are consumed once. Tokens in a conversation are consumed repeatedly.
Cost per task is the only unit worth reporting
Cost per request is meaningless when one piece of work is forty requests. Cost per user is too coarse to act on. Cost per task is the unit that maps to business value, because a task is something someone asked for.
To produce it you need a task or session identifier attached to every API call, propagated through your agent loop. Then:
cost_per_task = Σ over requests in task of
(input_tokens_uncached × input_rate)
+ (input_tokens_cached × cached_rate)
+ (output_tokens × output_rate)
Report the median and the 95th percentile, never just the mean. Agent workloads have long tails — a small number of runaway sessions frequently account for a large share of spend, and the mean hides them while the p95 makes them obvious.
Separate the four cost buckets
Aggregate spend tells you nothing. Split it at source:
- Uncached input. The expensive default. Grows with conversation length and tool output.
- Cached input. Stable prefixes — system prompts, tool definitions — billed at a discount where the provider supports it. Track the hit rate as a KPI.
- Output. Usually the highest unit rate but the smallest volume. Reasoning-heavy models change this; thinking tokens are typically billed as output.
- Failed and retried calls. Requests that errored, timed out, or produced output you discarded. This bucket is invisible unless you build it, and it is rarely zero.
The reason to split them is that each has a different remedy. A bad cached-input ratio is an engineering fix worth doing this week. A high output share is a model-selection question. A large retry bucket is a reliability problem masquerading as a cost problem.
Allocation: who gets charged
Finance will ask how to split the bill across teams or products. There are three workable schemes, in increasing order of effort.
- Per-key attribution. Issue one API key per team or per environment and let the provider dashboard do the split. Crude, near-zero effort, and good enough for most organisations under fifty engineers.
- Metadata tagging. Attach a team or cost-centre tag to each request and aggregate in your own logs. Requires a logging pipeline but survives shared services.
- Task-level chargeback. Cost per task rolled up to the product feature that triggered it. This is the only scheme that supports a real margin conversation for customer-facing AI features.
Pick the cheapest one that answers the question being asked. Building task-level chargeback for a twelve-person startup is a way to spend a fortnight producing a spreadsheet nobody reads.
Three metrics to put on the monthly report
Resist the temptation to report token counts. Nobody can interpret them. Report these instead:
- Cost per completed task, median and p95, trended month over month. Falling is good even if total spend rises, because it means adoption is outpacing unit cost.
- Cache hit rate on input tokens. A single number between 0 and 1 that directly moves the bill. It is the closest thing to a gross-margin lever engineering controls.
- Share of spend on the most expensive model. If 80% of your spend goes to a frontier model doing work a cheap model could do, that is a routing problem with a known fix.
Each of these is a number a non-engineer can hold an opinion about, which is the entire point.
Controls finance will ask for, and what to say
"Can we cap it?" Yes — set hard spend limits per project key at the provider, plus a token ceiling per task enforced in your own code. Both are necessary. The provider cap stops the month; the task ceiling stops the incident.
"Can we make it predictable?" Partly. Per-token billing is genuinely variable and no amount of accounting changes that. Flat-rate access converts the line to fixed, which is the honest reason to consider it — a flat-rate pass such as ours trades a possible saving for a definite number. If your usage is light or spiky, metered remains cheaper and you should say so plainly rather than buying predictability you do not need.
"Is it capitalisable?" Almost never. Inference spend for running a service is operating cost, in the same bucket as cloud compute. Treat it that way and the conversation gets much shorter.
A one-page format that works
Give the budget owner this, monthly, and most of the friction disappears:
- Total spend, versus forecast, versus cap.
- Split into the four buckets — uncached input, cached input, output, retries.
- Cost per completed task, median and p95, with last month for comparison.
- Top three consumers by team or feature.
- One sentence naming the biggest driver of any variance, and what is being done about it.
The discipline that makes this possible is unglamorous: log input, output and cached tokens separately, tag every request with a task ID and an owner, and record the model that actually served the request rather than the alias you asked for. Everything above falls out of those three fields. Without them, every cost conversation is a reconstruction exercise.
Common questions
How do I explain tokens to a finance team in one sentence?
A token is a billing unit worth about three-quarters of a word, charged separately for what you send and what the model generates, and because the model has no memory every follow-up message resends and re-pays for the whole conversation.
Should AI inference spend be capitalised or expensed?
In almost all cases it is operating expense, in the same category as cloud compute for running a service. Treating it as opex from the start avoids a long and usually unsuccessful argument with your auditors.
What is the single most useful AI cost metric to report?
Cost per completed task, reported as median and 95th percentile. It maps to something the business asked for, and the gap between median and p95 exposes runaway sessions that a mean would hide.