Stop Reading Benchmarks. Build a 20-Task Eval Instead.
Public leaderboards tell you almost nothing about your codebase. Here is how to build a private eval set in an afternoon and actually know which model to use.
Every model launch comes with a chart. The new model wins. It always wins — nobody publishes the benchmark where they lost.
The problem is not that the numbers are fake. It is that they are measuring something that is not your job. A model that tops a competitive-programming benchmark may be mediocre at reading a sprawling legacy codebase and making a careful two-line change without breaking anything.
The fix is unglamorous and takes about an afternoon: build a small eval set from your own work.
Why public benchmarks mislead
- Contamination. Public test sets leak into training data. A model may have effectively memorised problems it is being scored on. Your private tasks cannot leak.
- Distribution mismatch. Benchmarks favour self-contained puzzles. Real work is "change this without breaking the 40 things that depend on it".
- Aggregate scores hide variance. A model with a good average can fail catastrophically on the one category you care about.
- Vendor framing. Whoever publishes the chart chose the axes.
None of this means benchmarks are worthless. They are a decent filter for which three models to try. They are a terrible basis for a final decision.
Building the eval set
Target 20 tasks. Fewer and noise dominates; more and you will not keep it current.
Pull them from git history. This is the trick that makes it fast. Your merged PRs are real tasks with known-good solutions and, often, tests that verify them.
git log --oneline --merges -n 100
git show <sha> --stat
For each candidate, write down: the prompt you would have given, the files involved, and how you will check the result. Aim for a spread:
- ~6 bug fixes with a failing test that should go green
- ~4 small features touching two or three files
- ~3 refactors where behaviour must not change
- ~3 "explain this" tasks over unfamiliar code
- ~2 tasks you expect to be genuinely too hard
- ~2 tasks with a deliberately ambiguous or under-specified brief
Those last two categories matter more than people expect. The hard ones show you whether a model fails loudly or fabricates confidently. The ambiguous ones show you whether it asks a clarifying question or barrels ahead on a guess — which, in an agent loop, is the difference between a useful teammate and a liability.
Scoring without kidding yourself
Resist a 1–10 quality score. You will not apply it consistently across sessions. Use three buckets:
- Pass — tests green, diff is something you would merge.
- Salvageable — right idea, needed edits.
- Fail — wrong, or would have shipped a bug.
Record two more things per run, because they drive real cost:
- Turns to completion in an agent loop. A model that gets there in 4 turns beats one that needs 18, even at equal quality.
- Recovery behaviour. When a test fails, does it read the error and adjust, or repeat itself louder?
Running it fairly
A few controls stop you fooling yourself:
- Same harness. Same tool definitions, same system prompt, same repo state for every model. Reset with
git stashbetween runs. - Blind the results. Save outputs to files named by hash, score them later, reveal which model produced what afterwards. Expectation bias is real and large.
- Three runs on your ten most important tasks. Sampling is stochastic. A single run tells you very little about a close call.
- Fixed temperature. Whatever you use in production.
The part everyone skips
Re-run it. An eval set is only valuable if it is current — new model versions ship constantly, and your codebase drifts. Half an hour each quarter keeps it honest, and it turns "which model should we use" from an argument into a lookup.
This is also the only way to answer the question that actually matters: not "which model is best" but "is the cheaper one good enough for this". Usually, for most of your tasks, it is — and knowing precisely where it stops being good enough is worth more than any leaderboard.
A note on flat-rate access
Building an eval is exactly the workload that metered pricing punishes: dozens of full agent runs, most of them thrown away. It is worth doing regardless, but it is noticeably less painful when the meter is not running — which is one of the more legitimate arguments for flat-rate inference during development.
Common questions
How many tasks do I really need?
Twenty is a good target. Below about ten, run-to-run variance swamps the signal. Above thirty, most teams stop maintaining it, and a stale eval is worse than none.
Should I automate the scoring with another model?
For pass/fail on tasks with tests, use the tests. LLM-as-judge is reasonable for explanation quality but inherits the judge’s biases, so spot-check it against your own grading before trusting it.
Can I just use SWE-bench?
It is a useful filter for narrowing a shortlist, but it is public and therefore contaminable, and its task distribution is not your codebase. Use it to pick candidates, then use your own eval to pick the winner.