Evaluating Agent Reliability: Past the Single Pass Rate
An agent that works four times in five is not 80 percent useful. Here is how to measure agent reliability in a way that predicts production behaviour.
Demos succeed. Agents fail intermittently. The gap between those two facts is the whole problem with evaluating agents, and a single headline pass rate hides it almost perfectly.
An agent that completes a task 90 percent of the time sounds strong. Run it as an eight-step workflow where every step must succeed and you are down near 43 percent end to end. Reliability compounds badly, so it has to be measured directly rather than inferred.
Measure consistency, not just success
The standard code-generation metric is pass@k: at least one of k attempts succeeded. It rewards a model for occasionally getting it right, which is the right framing when a human picks the best candidate.
It is the wrong framing for an autonomous agent, because nobody is picking. The τ-bench work made this concrete by introducing pass^k — the probability that all k attempts succeed. That figure decays exponentially, so a 90 percent per-attempt success rate lands around 57 percent at k equals 8.
The original τ-bench results were sobering: leading function-calling agents solved under half the tasks, and pass^8 in the retail domain sat below 25 percent. Whatever the current numbers, the shape of the finding holds. Run every eval case at least five times and report the strict metric alongside the friendly one.
Build the task set from your own failures
Public benchmarks tell you which model is generally stronger. They do not tell you whether your agent works, because your tools, your data and your users are not in them.
The highest-value eval set is a collection of real traces that went wrong. Pull them from production, strip anything sensitive, write down what the correct outcome would have been, and freeze them as cases. Thirty of these are worth more than a thousand synthetic tasks.
Keep the distribution honest. Include easy cases so you notice regressions on things that used to work, ambiguous cases where the right move is to ask rather than guess, and cases where the correct behaviour is to refuse or escalate. An eval set of only hard tasks optimises for a model that never admits uncertainty.
Outcome evaluation versus trajectory evaluation
Two questions, both worth asking.
Did it end in the right state? This is the stronger signal and it should be checked programmatically wherever possible — compare final database state, run the test suite, diff the produced file. τ-bench does exactly this, comparing the end state against an annotated goal state rather than grading the conversation.
Did it get there sensibly? Trajectory evaluation looks at the path: which tools were called, in what order, whether policy was respected, how much was wasted. An agent that reaches the right answer after fourteen redundant tool calls is a cost and latency problem waiting to become an outage.
Outcome checks catch correctness. Trajectory checks catch the drift that precedes correctness failures, usually a release or two earlier.
Use a judge, but calibrate it
Some criteria have no programmatic check — was the explanation accurate, was the tone appropriate, did it follow the policy in spirit. An LLM judge is a reasonable tool here, with two conditions.
First, calibrate against human labels. Score fifty cases by hand, run the judge on the same fifty, and measure agreement. A judge you have not calibrated is a random number generator with good prose. Second, prefer narrow binary questions over holistic scores. "Did the response cite a source for every factual claim" is answerable; "rate quality from 1 to 10" produces mush that clusters around 7.
Use code checks where code checks work. Reserve the judge for what genuinely needs semantic judgement, and route low-confidence or disagreeing cases to a human queue.
Metrics that belong on the scorecard
- Task success rate at k equals 1, and consistency across k attempts.
- Steps to completion. Median and p95. A rising p95 means the agent is flailing on the tail.
- Tokens and cost per completed task. Per task, not per request — an agent that halves its success rate while halving its per-call cost got more expensive.
- Tool error rate and, separately, recovery rate after an error.
- Escalation rate. How often it asks for help, and how often that was the right call.
- Human intervention rate in production. The most honest reliability number you have.
Recovery rate deserves emphasis. Injecting a failure into a tool and observing whether the agent adapts, retries sensibly or gives up predicts real-world behaviour better than almost any clean-path metric, because production is mostly non-clean paths.
Close the loop with production
Offline evals are for regression testing. When you change a prompt or swap a model, you run the frozen set and confirm nothing broke — that belongs in CI with a threshold that blocks a merge.
Online monitoring is for discovery. Sample production traces continuously, run cheap deterministic checks on all of them and a judge on the subset that needs one, and route anomalies to review. Every confirmed failure becomes a new offline case.
That flywheel — trace, label, deduplicate, add to the versioned set, gate in CI, monitor again — is the entire practice. Teams that run it improve steadily. Teams that rely on spot checks discover their regressions from users.
A starting point
- Collect 30 real tasks, including 5 that should be refused or escalated.
- Define a programmatic success check for each. If you cannot, the task is underspecified.
- Run each 5 times. Report pass@1 and the all-5-succeeded rate.
- Log steps, tokens and cost per completed task alongside success.
- Inject a tool failure into 5 cases and measure recovery.
- Wire it into CI with a threshold, and add every production failure to the set.
That fits in a day of work and will tell you more about your agent than any leaderboard.
Common questions
What is the difference between pass@k and pass^k?
pass@k asks whether at least one of k attempts succeeded, which suits a human picking the best candidate. pass^k asks whether all k succeeded, which is what an unattended agent actually needs.
How many eval cases do I need?
Thirty real cases drawn from your own failures beat a thousand synthetic ones. Grow the set from production incidents rather than trying to enumerate scenarios up front.
Can I trust an LLM judge?
Only after calibrating it. Score around fifty cases by hand, measure agreement with the judge, and keep judge questions narrow and binary. Use deterministic checks wherever one exists.