What AI Actually Costs Per Developer Per Month
Cost & Pricing

What AI Actually Costs Per Developer Per Month

There is no single number, but there is a reliable way to derive yours. Work through the token arithmetic behind light, typical and heavy agentic usage.

The question comes up in every budget review, and the honest answer is unsatisfying: it depends by a factor of about fifty. A developer who asks a model three questions a day and a developer who runs agents continuously are not on the same order of magnitude, and averaging them produces a number that describes nobody.

What you can do is derive your own figure from three inputs you can measure in an afternoon. That is more useful than any benchmark average, because the whole point of the number is to decide what to buy.

The three inputs

Everything reduces to these:

monthly_cost = tasks_per_day × tokens_per_task × price_per_token × 21

Twenty-one is working days. Do not forecast on thirty; people do not use developer tools at the weekend in any volume that matters.

tasks_per_day is the easy one — count what you actually asked for, not what you sent. A task is a unit of work someone wanted done, whether it took one request or forty.

tokens_per_task is where the variance lives, and it depends almost entirely on how you work rather than on what you are building:

One-shot question, no context          ~1k
Question with two files attached       ~8k
Chat session over a module            ~40k
Short agent run (5 turns)            ~120k
Long agent run (15+ turns)         ~400k-1M

price_per_token should be your blended rate — input, cached input and output weighted by your actual mix — not a headline number from a pricing page. Published prices also move often enough that any figure written into an article is a snapshot; take yours from last month's invoice divided by last month's tokens.

Three worked profiles

Use an illustrative blended rate of $3 per million tokens to keep the arithmetic legible. Substitute your own; the ratios are the durable part, not the currency.

Light user. Uses a model as a better search engine. Five questions a day, mostly single-shot with a file attached.

5 × 8k × 21 = 840k tokens/month  ≈ $2.50

Typical user. Chat sessions for design work, occasional agent runs for well-scoped changes. Eight tasks a day, mixed.

8 × 60k × 21 = 10.1M tokens/month  ≈ $30

Heavy agentic user. Agent-first workflow, several concurrent sessions, tool output flowing back on every turn. Twelve tasks a day at 350k each.

12 × 350k × 21 = 88.2M tokens/month  ≈ $265

The spread between the light and heavy profile is over a hundredfold, on the same illustrative price and the same team. This is why "cost per developer" as an industry benchmark is close to meaningless, and why the only useful version of the number is the one you derive.

Why agentic usage dominates everything

Two properties of how models work explain the entire gap.

Models are stateless. Every request resends the full conversation, so a fifteen-turn agent loop pays for turn one fifteen times. Input tokens grow roughly with the square of turn count.

Agents generate their own context. Tool output — file contents, test logs, search results — arrives in large chunks and stays in the conversation forever after. In a typical agent session, tool output is the single largest cost component, usually larger than the system prompt, the user instructions and the model output combined.

The practical consequence is that the highest-leverage cost control is not switching model or negotiating a rate. It is truncating tool results and compacting conversation history, both of which change the trajectory of every subsequent turn in the session.

The costs nobody puts in the spreadsheet

Two effects reliably distort the number in opposite directions.

Suppressed demand. If your team has been watching a meter, your historical usage understates what they would use without one. People stop tasks early, skip second attempts, and do not build eval sets because a hundred discarded runs feels wasteful. That gap is real usage you are not measuring, and it is usually where the productivity is.

Waste. Failed runs, retried calls, sessions abandoned halfway, and agents that looped fruitlessly all cost full price. If you are not logging errored and discarded requests as their own bucket, your effective cost per completed task is higher than you think — often by 10 to 20 percent.

Turning the number into a decision

Once you have a monthly figure per profile, the buying decision is arithmetic rather than argument.

  1. Compute the number for each profile on your team, not a team average.
  2. Compare each against flat-rate options at the same capability level.
  3. Add a premium for predictability if the variance between your low and high months exceeds about 3x — a fixed number has real value in a budget even at equal expected cost.
  4. Adjust the metered figure upward for suppressed demand if your team has been rationing.

Light users should stay metered; at a few dollars a month, no flat rate will beat it and any honest provider will tell you so. Heavy agentic users are usually better off on a flat rate, which is the segment our own pass is built for. Typical users sit in the genuinely ambiguous middle, where the deciding factor is usually whether the spend is volatile enough to be annoying.

Measure yours this week

Three fields logged per request gets you everything above: input tokens, output tokens and cached input tokens separately, plus a task identifier and the model that actually served the request. Aggregate by task, report the median and the 95th percentile, and you will have a number that is defensible, actionable, and specific to how your team actually works.

Common questions

What is a typical monthly AI cost per developer?

There is no useful single figure. Derived from token arithmetic, light users land in the low single-digit dollars, mixed users in the tens, and heavy agentic users in the hundreds. The spread is over a hundredfold, so derive your own rather than using a benchmark.

Why does agentic coding cost so much more than chat?

Models are stateless, so each agent turn resends the entire conversation, making input tokens grow roughly quadratically with turn count. Tool output is also carried forward permanently and is usually the largest single cost component in a session.

What is the fastest way to cut AI cost per developer?

Truncate tool output and compact long conversation histories. Both reduce the payload carried into every subsequent turn, so the saving compounds across the session. Model switching and rate negotiation are smaller levers by comparison.

Similar articles

Budgeting for Experimentation Without a Surprise Bill
Cost & Pricing
Cost & Pricing·12 min read

Budgeting for Experimentation Without a Surprise Bill

Experiments produce the largest unexpected AI invoices because nobody set a ceiling. How to fund trying things: separate ledgers, time-boxes and kill switches.

Read
FinOps for AI Teams: What Transfers From Cloud and What Does Not
Cost & Pricing
Cost & Pricing·9 min read

FinOps for AI Teams: What Transfers From Cloud and What Does Not

Cloud FinOps practice mostly transfers to LLM spend, but rightsizing and reserved capacity do not. What to keep, what to drop, and what to instrument first.

Read
Free Tier Strategies: Getting Real Answers Without Paying
Cost & Pricing
Cost & Pricing·12 min read

Free Tier Strategies: Getting Real Answers Without Paying

Free tiers are evaluation budgets, not production capacity. What they are for, where the limits bite, and how to fit a real model evaluation inside one.

Read