Stop Reading Benchmarks. Build a 20-Task Eval Instead.
Models

Stop Reading Benchmarks. Build a 20-Task Eval Instead.

Public leaderboards tell you almost nothing about your codebase. Here is how to build a private eval set in an afternoon and actually know which model to use.

Every model launch comes with a chart. The new model wins. It always wins — nobody publishes the benchmark where they lost.

The problem is not that the numbers are fake. It is that they are measuring something that is not your job. A model that tops a competitive-programming benchmark may be mediocre at reading a sprawling legacy codebase and making a careful two-line change without breaking anything.

The fix is unglamorous and takes about an afternoon: build a small eval set from your own work.

Why public benchmarks mislead

  • Contamination. Public test sets leak into training data. A model may have effectively memorised problems it is being scored on. Your private tasks cannot leak.
  • Distribution mismatch. Benchmarks favour self-contained puzzles. Real work is "change this without breaking the 40 things that depend on it".
  • Aggregate scores hide variance. A model with a good average can fail catastrophically on the one category you care about.
  • Vendor framing. Whoever publishes the chart chose the axes.

None of this means benchmarks are worthless. They are a decent filter for which three models to try. They are a terrible basis for a final decision.

Building the eval set

Target 20 tasks. Fewer and noise dominates; more and you will not keep it current.

Pull them from git history. This is the trick that makes it fast. Your merged PRs are real tasks with known-good solutions and, often, tests that verify them.

git log --oneline --merges -n 100
git show <sha> --stat

For each candidate, write down: the prompt you would have given, the files involved, and how you will check the result. Aim for a spread:

  • ~6 bug fixes with a failing test that should go green
  • ~4 small features touching two or three files
  • ~3 refactors where behaviour must not change
  • ~3 "explain this" tasks over unfamiliar code
  • ~2 tasks you expect to be genuinely too hard
  • ~2 tasks with a deliberately ambiguous or under-specified brief

Those last two categories matter more than people expect. The hard ones show you whether a model fails loudly or fabricates confidently. The ambiguous ones show you whether it asks a clarifying question or barrels ahead on a guess — which, in an agent loop, is the difference between a useful teammate and a liability.

Scoring without kidding yourself

Resist a 1–10 quality score. You will not apply it consistently across sessions. Use three buckets:

  • Pass — tests green, diff is something you would merge.
  • Salvageable — right idea, needed edits.
  • Fail — wrong, or would have shipped a bug.

Record two more things per run, because they drive real cost:

  • Turns to completion in an agent loop. A model that gets there in 4 turns beats one that needs 18, even at equal quality.
  • Recovery behaviour. When a test fails, does it read the error and adjust, or repeat itself louder?

Running it fairly

A few controls stop you fooling yourself:

  1. Same harness. Same tool definitions, same system prompt, same repo state for every model. Reset with git stash between runs.
  2. Blind the results. Save outputs to files named by hash, score them later, reveal which model produced what afterwards. Expectation bias is real and large.
  3. Three runs on your ten most important tasks. Sampling is stochastic. A single run tells you very little about a close call.
  4. Fixed temperature. Whatever you use in production.

The part everyone skips

Re-run it. An eval set is only valuable if it is current — new model versions ship constantly, and your codebase drifts. Half an hour each quarter keeps it honest, and it turns "which model should we use" from an argument into a lookup.

This is also the only way to answer the question that actually matters: not "which model is best" but "is the cheaper one good enough for this". Usually, for most of your tasks, it is — and knowing precisely where it stops being good enough is worth more than any leaderboard.

A note on flat-rate access

Building an eval is exactly the workload that metered pricing punishes: dozens of full agent runs, most of them thrown away. It is worth doing regardless, but it is noticeably less painful when the meter is not running — which is one of the more legitimate arguments for flat-rate inference during development.

Common questions

How many tasks do I really need?

Twenty is a good target. Below about ten, run-to-run variance swamps the signal. Above thirty, most teams stop maintaining it, and a stale eval is worse than none.

Should I automate the scoring with another model?

For pass/fail on tasks with tests, use the tests. LLM-as-judge is reasonable for explanation quality but inherits the judge’s biases, so spot-check it against your own grading before trusting it.

Can I just use SWE-bench?

It is a useful filter for narrowing a shortlist, but it is public and therefore contaminable, and its task distribution is not your codebase. Use it to pick candidates, then use your own eval to pick the winner.

Similar articles

The Artificial Analysis Index Explained: What It Does Measure
Models
Models·9 min read

The Artificial Analysis Index Explained: What It Does Measure

A single number that ranks every model is convenient and easy to misread. What the AA Intelligence Index aggregates, and where it stops being useful.

Read
Blind Model Comparison: A Method That Survives Bias
Models
Models·9 min read

Blind Model Comparison: A Method That Survives Bias

Knowing which model wrote an answer changes how you judge it. A practical protocol for blind pairwise comparison that a small team can run in a day.

Read
Cost-Adjusted Model Scoring: Ranking by Value, Not Capability
Models
Models·9 min read

Cost-Adjusted Model Scoring: Ranking by Value, Not Capability

Capability leaderboards ignore price entirely. How to build a score that reflects what a model costs to run on your workload, and where it misleads.

Read