Cost Per Pull Request: A Unit Metric Worth Tracking
Total AI spend tells you nothing actionable. Cost per merged pull request ties inference to delivered work and exposes exactly where the money goes.
Monthly AI spend is a number that goes up. It does not tell you whether that is good, bad, or a symptom of something breaking. A unit metric does — and for software teams the natural unit is the pull request.
Cost per merged PR ties inference spend to a thing the business already counts. It moves when something real changes, it is comparable across teams and quarters, and it decomposes cleanly into terms you can act on.
Define it before you measure it
Three definitional choices, and they matter more than the instrumentation.
Merged or attempted? Track both. Cost per attempted PR includes work that was abandoned; cost per merged PR is the delivery figure. The ratio between them is your waste rate, and it is often the most interesting number of the three.
Which spend counts? All inference attributable to the change: the developer's interactive session, any agent runs, the automated review pass, and CI-triggered analysis. Not the developer's unrelated exploration that day.
What window? From branch creation to merge. Post-merge fixes attach to their own PRs, which is correct — it means a change that needed three follow-ups shows up as four PRs rather than being quietly amortised.
Instrumenting it
You need a correlation key that travels from the API request to the PR. Three approaches, in increasing order of fidelity.
Branch-scoped keys or request tags. The cleanest. Most providers accept a metadata field on each request; populate it with the branch name. Agent harnesses and CI jobs know their branch, so this covers automation completely.
Session correlation. If the tooling issues a session identifier, log the mapping from session to branch at session start. Slightly lossy when a developer works across branches in one session.
Time-window attribution. Assign a developer's spend to whichever branch they had checked out at the time. Crude, but it works with no tooling changes and gets you a usable first number in an afternoon.
Start with the third. The insight from a rough number available this week beats a precise one available next quarter.
A worked decomposition
One month, one repository. 140 PRs opened, 118 merged.
Interactive developer sessions $1,840
Autonomous agent runs $960
Automated PR review (every PR) $410
CI analysis on failures $120
total $3,330
cost per attempted PR = 3,330 / 140 = $23.79
cost per merged PR = 3,330 / 118 = $28.22
waste rate = 22 / 140 = 15.7%
Now the decomposition, which is where the value is:
Interactive per merged PR $15.59 (55%)
Agent runs per merged PR $8.14 (29%)
Review per merged PR $3.47 (12%)
CI per merged PR $1.02 (4%)
Immediately visible: the automated review pass is 12% of spend and runs on every PR including the 22 that never merged. That is $76 a month spent reviewing work that was thrown away — small, but it is the kind of thing that is completely invisible in a total.
Also visible: interactive sessions dominate. Any optimisation aimed at the agent runs is aimed at 29% of the number.
What a healthy figure looks like
There is no benchmark, and anyone offering one is guessing. What there is, is a comparison you can make yourself.
A US developer at a fully loaded rate near $90 an hour — median salary around $132,270 per BLS data, plus 30–40% overhead — costs roughly $90 for one hour of work on a PR. Against that, $28 of inference per merged PR needs to save about 19 minutes of engineer time to break even.
break_even_minutes = (cost_per_merged_pr / loaded_hourly_rate) x 60
= (28.22 / 90) x 60
= 18.8 minutes
State it that way when the number is challenged. Nineteen minutes per PR is a claim people can assess from experience.
Reading the metric when it moves
A rising cost per merged PR has four common causes, and they need different responses.
PRs got bigger. Check median lines changed alongside the cost. If both rose, unit cost is stable and nothing is wrong. This is the most common false alarm, and it is why you should always plot change size next to cost.
The waste rate rose. More abandoned work, same spend per attempt. Look at why PRs are being abandoned; the cost metric is reporting a process problem, not a spending problem.
Turns per task rose. Agents taking longer to reach the same result. Usually a degraded tool description, a changed prompt, or a model swap. This is the one that most often goes undiagnosed because it looks like a cost problem.
A new automation was added. Check whether anything started running on every PR recently. Background automation is the standard source of a step change in the number.
A falling cost per merged PR is not automatically good either. If it fell because developers stopped using the tooling, you have saved money and lost the benefit. Track usage rate alongside it.
Segment before you conclude
An aggregate figure hides most of what is useful. Segment by:
- Repository or service. A legacy service with sprawling context routinely costs several times a well-factored one, and that comparison is a genuine argument for refactoring.
- Change type. Dependency bumps, bug fixes and feature work have very different profiles. Mixing them produces an average that describes nothing.
- Interactive versus autonomous. These respond to completely different interventions.
The most common finding when teams first segment: one repository or one workflow accounts for a majority of spend. Optimise that and ignore the rest.
Where the metric misleads
It undercounts work that never becomes a PR — investigation, debugging that ends in a config change, architectural exploration. On teams where a lot of AI assistance goes into understanding rather than shipping, cost per PR will make the tooling look worse than it is.
It also penalises good practice. A team that splits work into small reviewable PRs will show a lower cost per PR than a team shipping large ones, without being more efficient. Normalise by lines changed if you need to compare teams; use raw cost per PR only for tracking one team over time.
Making the number stable
If your unit cost swings month to month, the underlying distribution is heavy-tailed and a handful of expensive runs are moving the average. Report the median cost per PR alongside the mean, and consider a turn cap on agent runs — that bounds the tail without touching typical work.
Teams on flat-rate access get a different version of this metric: the numerator is fixed, so cost per merged PR falls automatically as throughput rises. That is genuinely useful for capacity planning and genuinely useless for spotting a runaway agent, so keep tracking turns per task separately. Flat rate removes the cost signal, not the underlying problem.
Common questions
How do I attribute AI spend to a specific pull request?
Tag every API request with the branch name via the provider metadata field, or map session identifiers to branches at session start. A rough time-window attribution using the checked-out branch is enough to get a first number quickly.
What is a reasonable cost per merged pull request?
There is no benchmark worth trusting. Convert your figure into break-even minutes by dividing it by your fully loaded hourly engineer cost and multiplying by 60, then judge whether that much saved time is plausible.
Why did my cost per PR rise without spend changing?
Almost always because fewer PRs merged. Check the waste rate — attempted versus merged — before investigating spend. Plot median change size alongside cost so that bigger PRs do not read as a regression.