The Hidden Costs of AI Coding Tools
The model invoice is the smallest line item. Here is the full list of what agentic coding actually costs, and a method for pricing each part yourself.
The invoice from your model provider is the visible part of what AI coding costs. It is rarely the largest part, and it is almost never the part that causes the awkward conversation six months in.
What follows is the list of line items that do not appear on that invoice, with a way to put a number on each. The argument is not that these tools are expensive — for most teams they are obviously worth it. The argument is that a comparison built on the API bill alone is comparing the wrong two things.
Why the invoice understates the total
Metered billing captures tokens that were successfully generated. It does not capture the tokens you paid for and threw away, the human time spent supervising the loop, or the second tool you bought because the first one lacked a model you needed.
Those costs are real, they are measurable, and they are all denominated in the same currency once you convert engineer-time into money. Do that conversion once and the whole picture becomes comparable.
The conversion rate: US Bureau of Labor Statistics data puts the median software developer salary around $132,270 a year. Benefits, payroll tax and overhead typically add 30–40%, so a fully loaded hour lands somewhere near $85–$95 for a median developer on a 2,000-hour year. Substitute your own figure — the method matters more than the constant.
Cost 1: generations you paid for and discarded
Every retry pays full input cost again. So does every run you cancel because the agent went the wrong way, and every generation that keeps streaming after the developer has already navigated away.
This is invisible in an aggregate token count because a discarded generation looks identical to a useful one. It becomes visible the moment you log a per-request outcome flag — accepted, retried, abandoned — alongside the token count.
Teams that instrument this usually find the discard rate sits somewhere between a fifth and a third of total spend, though it varies enormously by workload. You will not know yours until you measure it, and the measurement is a few lines of logging.
Cost 2: the tool you are paying for twice
A common pattern: a per-seat subscription for an editor integration, plus metered API access for scripts and CI, plus a second subscription because one team wanted a different model. Three bills, overlapping capability, and nobody owns the overlap.
Audit this annually. The question to ask for each line is not "is this useful" — everything is useful — but "what breaks if we cancel it this month". Anything where the honest answer is "one person would be mildly inconvenienced" is a candidate.
Cost 3: automated runs nobody is watching
Agents wired into CI are the classic source of a surprise bill, because the usual cost signal — a developer noticing the meter — is absent. A review agent on every pull request, a triage agent on every issue, a nightly documentation pass: each is cheap per run and none of them stop.
Price these deliberately:
monthly_cost = runs_per_day x 30 x avg_tokens_per_run x price_per_token
A PR review agent on a repo with 40 merges a day, averaging 60k input and 4k output tokens per review, at $5 per million input and $25 per million output, comes to roughly:
input: 40 x 30 x 60,000 = 72M tokens -> $360
output: 40 x 30 x 4,000 = 4.8M tokens -> $120
total ~$480/month
Those per-token figures are Anthropic Claude Opus 5 list rates as of August 2026 and will change; the structure of the calculation will not. Note how the input side dominates despite output costing five times more per token — that ratio inversion is the single most reliable feature of agentic workloads.
Cost 4: review and rework
Generated code still needs reading. If an AI-assisted change takes 12 minutes to review where a hand-written one took 8, and your team merges 200 PRs a month, that is roughly 13 extra engineer-hours — over $1,000 a month at a loaded rate, comfortably more than the model spend in the example above.
This cuts both ways. If the tooling reduces the number of review cycles per PR, the same arithmetic runs in your favour. The honest version of the calculation tracks review cycles per merged change, before and after, rather than assuming a direction.
Cost 5: the evaluation you did not build
This one is a cost of not spending. Without an eval set, you cannot tell whether a cheaper model would do a given job, so you route everything to the most capable option and pay for it forever.
Building a small evaluation harness — fifty real tasks from your own repository, with a pass criterion — is a few days of work and it pays for itself the first time it lets you move a high-volume, low-difficulty workload down a tier. Treat the absence of one as a recurring line item.
Assembling the real number
Per developer, per month, add up:
- Metered or flat-rate model access — the visible number.
- Discarded generations — your measured discard rate applied to that number.
- Subscription overlap — the seats you would not re-buy today.
- Automation — total CI and background agent spend, divided by headcount.
- Review delta — change in review hours per month, times your loaded hourly rate.
The result is usually two to four times the invoice. That is not a scandal; it is what total cost of ownership looks like for any tool. It just needs to be the number you compare against, rather than the one on the receipt.
The one structural decision worth making
Once you have that number, the choice between metered and flat-rate access becomes arithmetic rather than argument. Flat-rate access — ours included — makes sense when a developer uses agents daily and the meter is causing them to ration. It is a bad deal when usage is light or bursty, and you should stay metered in that case.
What flat rate genuinely removes is one specific hidden cost: the time developers spend deciding whether a run is worth the money. That deliberation is unmeasured, unbudgeted, and for heavy users it is not small.
Common questions
How do I measure how much I spend on discarded generations?
Log an outcome flag on every request alongside its token counts: accepted, retried, or abandoned. Sum tokens by flag at the end of a month. Aborting upstream requests when a user navigates away also stops you paying for output nobody reads.
Is a per-seat subscription cheaper than metered API access?
It depends entirely on daily usage. Divide the seat price by your measured daily token cost to get a break-even in days. Under roughly a third of the month, metered wins; well over it, the subscription does.
Why does input cost dominate when output tokens are priced higher?
Because agent loops resend the whole conversation on every step. Volume beats unit price: you might use fifty times more input tokens than output tokens, which swamps a five-times-per-token difference.