Setting an AI Budget for a Small Engineering Team
A bottom-up method for budgeting model spend on a team of five to twenty, including the buffer to hold, the caps to set, and the alerts that matter.
Most small teams budget for model spend by picking a round number and revisiting it after the first overage. That works until the month somebody starts running agents overnight.
Here is a bottom-up method instead. It takes about an hour, produces a number you can defend, and — more usefully — produces the two or three controls that stop the number being wrong in an expensive direction.
Build the estimate from usage archetypes, not headcount
Dividing a guess by headcount produces a guess. Instead, sort your team into usage archetypes and estimate each separately, because the spread between them is enormous.
Three archetypes cover most teams:
- Occasional — chat-style questions, the odd code explanation. A handful of requests a day, short contexts.
- Assisted — an editor integration running most of the working day, mostly single-file edits and completions.
- Agentic — long-running tasks, multi-file changes, tool loops. Orders of magnitude more tokens than the other two.
The gap between occasional and agentic is not a factor of two or three. It is routinely a factor of fifty or more, because agent loops resend accumulated context on every step while a chat turn does not.
Get a per-archetype daily figure from real data
You need one number per archetype: average tokens per developer per day, split into input and output. Every metered provider exports this. Take a representative fortnight — avoiding a release week and avoiding a holiday — and average it.
If you have not started yet and have no data, run a two-week pilot with two volunteers rather than estimating. Published averages will not match your codebase, your prompt sizes, or your tooling.
Convert to money with your provider list price. As of August 2026, Anthropic lists Claude Opus 5 at $5 per million input tokens and $25 per million output; Claude Haiku 4.5 at $1 and $5. OpenAI lists GPT-5.6 Sol at $5 input and $30 output, with cached input at $0.50. These change; check before you commit a figure to a spreadsheet.
A worked budget for a team of eight
Say the team is two occasional, four assisted, two agentic, and the fortnight of data gives these daily averages per person:
Occasional: 150k input, 15k output per dev/day
Assisted: 1,200k input, 90k output per dev/day
Agentic: 9,000k input, 400k output per dev/day
At $5 per million input and $25 per million output, and 21 working days:
Occasional: (0.15 x 5) + (0.015 x 25) = $1.13/day
x 2 devs x 21 days = $47
Assisted: (1.2 x 5) + (0.09 x 25) = $8.25/day
x 4 devs x 21 days = $693
Agentic: (9.0 x 5) + (0.40 x 25) = $55.00/day
x 2 devs x 21 days = $2,310
Human subtotal = $3,050/month
Two agentic developers cost more than three times the other six combined. That is the normal shape of this distribution, and it is why headcount-based budgeting fails.
Add automation separately, then a buffer
Background agents — PR review, issue triage, nightly jobs — are a separate line because they scale with repository activity rather than headcount. Estimate them as runs per day times tokens per run, and note that they will keep running whether or not anyone is looking.
Then add a buffer. Not a round 20% — derive it. Take your fortnight of daily figures and find the ratio between the busiest day and the median day. For agentic usage that ratio is commonly 2.5–4x. Your buffer needs to cover a month where several people have a busy fortnight simultaneously:
buffer = (p90_daily_total / median_daily_total - 1) x 0.5
The 0.5 is because a whole month rarely runs at the peak rate. With a p90/median ratio of 3, that gives a 100% buffer on the agentic line — which sounds absurd until the first time an agent gets stuck in a retry loop over a weekend.
Controls, in order of usefulness
1. A hard cap per API key. Not per team — per key. If your provider supports spend limits, set one on every key at roughly 2x its expected monthly figure. This is the only control that actually stops a runaway.
2. A daily alert, not a monthly one. A monthly threshold tells you about a problem three weeks after it started. Alert on daily spend crossing 1.5x the median daily figure, delivered somewhere a human reads within hours.
3. Per-feature attribution. Tag every request with which tool or workflow issued it. When the number moves, you want to know which line moved, not just that the total did.
4. A named owner. One person whose job includes looking at the number weekly. Budgets without an owner do not survive contact with a busy quarter.
When flat rate makes the budget easier — and when it does not
The arithmetic above produces a range, and finance departments dislike ranges. Flat-rate access converts the agentic line into a fixed number, which removes both the buffer calculation and the runaway risk for that segment.
Whether it saves money is a separate question with a simple test: divide the flat-rate price by that developer's measured daily cost. In the example above, an agentic developer at $55 a day breaks even against a $50 monthly pass in one day, which makes the decision obvious. An occasional developer at $1.13 a day would need 44 days to break even on the same pass, so they should stay metered — and any provider that tells you otherwise is selling rather than advising.
The subtlety worth flagging: if your agentic developers have been rationing because the meter is visible, their historical figures understate real demand. Budget from behaviour you want, not behaviour the meter produced.
Review cadence
Re-run the whole exercise quarterly, and immediately after any of these: a new agent framework, a model change, a new background automation, or a team member moving between archetypes. Each of those can move the number by more than the buffer covers.
Common questions
How much should a small team budget per developer per month?
There is no single figure — the spread between light and agentic users is routinely fifty-fold. Measure a representative fortnight per usage archetype and build up from there rather than using a per-head average.
What buffer should I hold on top of the estimate?
Derive it from your own variance. Take the ratio of your p90 daily spend to your median daily spend, subtract one, and halve it. For agentic workloads that commonly lands near 100% on that line item.
What is the single most effective cost control?
A hard spend cap on each individual API key. Alerts tell you after the fact; a cap is the only mechanism that stops a retry loop or a misconfigured background job before it finishes the month.