Forecasting AI Spend Without Guessing
Cost & Pricing

Forecasting AI Spend Without Guessing

Most AI budget forecasts are a headcount multiplied by a hopeful number. Here is a model that decomposes spend into drivers you can actually measure and control.

Ask an engineering manager to forecast next quarter's AI spend and you usually get one of two answers: last month multiplied by a growth guess, or seat count multiplied by a list price. Both are wrong in the same way — they treat inference as a subscription when it behaves like a utility with a variable load.

A forecast is useful when it tells you which lever to pull when the number comes in high. That requires decomposing spend into drivers, not extrapolating a total.

The unit-economics model

Every dollar of inference spend factors into the same four terms. Write them down explicitly:

monthly_cost =
    active_users
  × tasks_per_user_per_day
  × tokens_per_task
  × price_per_token
  × working_days

That is it. Every cost surprise you will ever have is one of those five terms moving. The value of writing it this way is that each term has a different owner and a different remedy:

  • active_users — hiring and adoption. Grows in steps, not smoothly.
  • tasks_per_user_per_day — behaviour. Rises sharply as people get good at the tools.
  • tokens_per_task — engineering. The one you actually control day to day.
  • price_per_token — procurement. Falls over time, but only if you re-evaluate.
  • working_days — roughly 21 a month. Do not forecast on 30.

Why the naive forecast is always low

Three effects compound, and none of them appear in a linear extrapolation.

Adoption is S-shaped, not linear. The first month of a rollout is dominated by people trying it once. The third month is dominated by people who have restructured their workflow around it. A 3x jump between month one and month three is normal and is not a spike.

Token intensity per task rises as capability rises. When agents get better at multi-step work, people give them multi-step work. The task that was a 5k-token completion becomes a 400k-token agent session, because it now succeeds.

Context accumulates within a task. Models are stateless, so every turn resends the conversation. Input tokens grow roughly with the square of turn count. A task that goes from 6 turns to 12 does not double in cost — it roughly quadruples.

Put those together and a forecast built on "last month plus 20%" will be wrong by a multiple, not a percentage.

A worked forecast

Take a team of 20 engineers rolling out an agentic coding tool. Rather than one number, build three scenarios by varying only the behavioural terms.

Shared assumptions
  working_days       = 21
  blended price      = P per 1M tokens (use your own contract)

Conservative   8 engineers active, 3 tasks/day,  120k tokens/task
Central       15 engineers active, 6 tasks/day,  250k tokens/task
Heavy         20 engineers active, 10 tasks/day, 450k tokens/task

Compute monthly tokens for each:

Conservative:  8 × 3  × 120k × 21 =    60.5M tokens
Central:      15 × 6  × 250k × 21 =   472.5M tokens
Heavy:        20 × 10 × 450k × 21 = 1,890.0M tokens

The spread between conservative and heavy is roughly 31x. That is the honest answer, and it is far more useful than a single point estimate, because it tells finance that this is a variable-cost line requiring a cap, not a fixed cost requiring an approval.

Multiply each by your own blended per-token price to get currency. Deliberately leaving the price symbolic is the point — published prices move, and your contract, cache-hit rate and model mix all shift the blend. Recompute quarterly.

Instrument before you forecast

A forecast is only as good as the telemetry underneath it. The minimum viable instrumentation is three fields logged per request:

  • Input tokens, output tokens, and cached input tokens separately — cache hits are often billed at a large discount, and blending them hides your best optimisation.
  • A task or session identifier, so you can aggregate cost per task rather than per request. Cost per request is a meaningless unit when one task is 40 requests.
  • The model actually served, not the alias requested. Routing and fallbacks silently change your blended price.

With those three you can answer the only question that matters when a bill spikes: did more people use it, did each person do more, or did each task get more expensive? Those have completely different responses.

Controls that actually hold

Forecasts drift. Controls stop the drift becoming an incident.

  1. Hard spend caps per project key, set below the budget, with alerting at 50% and 80%. A cap you have never hit is a cap you do not know works.
  2. A token ceiling per task enforced in your own code. Fail loudly when a single session exceeds it rather than discovering a runaway loop on the invoice.
  3. Truncate tool output. The dominant cost in agent sessions is almost always tool results carried forward. Capping search results at 50 rows changes the whole trajectory.
  4. Route by difficulty. Cheap models for classification, extraction and summarisation; expensive models only where the reasoning is the product.
  5. Re-tender annually. Per-token prices fall. A contract you have not revisited is a contract you are overpaying on.

When a flat rate simplifies the forecast

If your spread between conservative and heavy scenarios is more than about 5x, the forecast is not really a forecast — it is a range with a promise to explain the variance later. Flat-rate access converts that variable line into a fixed one, which is worth something on its own even when the expected value is similar.

That is the honest case for a flat-rate pass like ours: not that it is always cheaper, but that it is predictable, and predictability has real value in a budget conversation. If your usage sits at the conservative end and stays there, metered billing will cost you less and you should keep it.

The takeaway: forecast drivers, not totals. Publish a range with named assumptions, instrument the three fields that let you attribute variance, and put a cap on the number before you need one.

Common questions

How far ahead can I realistically forecast AI spend?

One quarter with a stated range, and only if adoption is already past the early ramp. Beyond that, model prices, model mix and usage behaviour all move enough that a point estimate is theatre. Forecast the drivers and refresh the numbers monthly.

Why did our AI bill triple without headcount changing?

Almost always token intensity per task, not user count. Agent loops resend the full conversation each turn, so input tokens grow roughly quadratically with turn count, and better tools invite longer tasks. Log tokens per task to confirm.

Should AI spend be a fixed cost or a variable cost line?

It is genuinely variable under per-token billing, and finance should treat it that way with caps and alerting. Flat-rate arrangements convert it to fixed, which simplifies planning but only pays off if your usage is consistently heavy.

Similar articles

Forecasting Next Quarter's AI Spend as a Range
Cost & Pricing
Cost & Pricing·12 min read

Forecasting Next Quarter's AI Spend as a Range

Why extrapolating token usage gives the wrong number, which drivers actually move an AI bill, and how to build a defensible range instead of a point estimate.

Read
Token Accounting: Explaining AI Costs to Finance
Cost & Pricing
Cost & Pricing·8 min read

Token Accounting: Explaining AI Costs to Finance

Finance teams need cost drivers, allocation and controls, not a lecture on transformers. Here is how to translate token usage into terms a budget owner can act on.

Read
Setting an AI Budget for a Small Engineering Team
Cost & Pricing
Cost & Pricing·8 min read

Setting an AI Budget for a Small Engineering Team

A bottom-up method for budgeting model spend on a team of five to twenty, including the buffer to hold, the caps to set, and the alerts that matter.

Read