How to Choose a Coding Model Without Trusting Anyone’s Chart
Six properties decide whether a model is good at your codebase — and only one of them shows up on a leaderboard. A practical framework for picking.
Every model is "state of the art at coding". They cannot all be right, and the disagreement is not really about quality — it is that "good at coding" is at least six different properties, and different models are strong at different ones.
Here is the framework I would use, in the order that actually matters for day-to-day work.
1. Instruction adherence
Does it do what you asked, and only what you asked?
This is the most underrated property and the one that causes the most friction. You ask for a two-line fix and get a refactor of the whole module. The refactor might even be better — but you cannot review it quickly, and you did not ask for it.
Test it: give a deliberately narrow instruction ("change only the retry count, touch nothing else") and see whether the diff respects the boundary.
2. Recovery behaviour
What happens when it is wrong?
In an agent loop this dominates everything else. A model that writes mediocre code but reads the failing test and fixes itself will out-perform a brilliant model that repeats the same broken approach three times.
Test it: let it make a change that fails a test, and watch. Does it read the error, form a new hypothesis, and adjust? Or restate the same solution more confidently?
3. Codebase comprehension
Can it work with your conventions rather than generic ones?
Benchmarks favour self-contained problems. Real work is "add a field to this entity" across an ORM, a DTO, a validator, a migration and a test — following patterns already established in the repo.
Test it: ask for a change that should mirror an existing pattern. Does it find and follow the pattern, or invent its own?
4. Tool-calling reliability
Does it call the right tool with valid arguments, consistently?
A model that reasons beautifully but emits malformed tool calls is unusable in an agent. This is largely binary — models are either dependable here or they are not.
5. Latency
Not raw tokens per second — time to a useful result.
A fast model that needs four turns often beats a slow one that needs two. And in interactive work, time to first token matters more than throughput, because it determines whether the tool feels responsive.
6. Raw capability
The thing benchmarks actually measure: can it solve genuinely hard problems?
It matters least often, because most day-to-day work is not algorithmically hard. It matters enormously on the few tasks that are. Keep access to a strong model for those, but do not pay the premium on every request.
Matching model to task
The practical answer is usually "more than one":
- Boilerplate, tests, renames, docstrings — a fast, cheap model is genuinely sufficient. Do not overthink it.
- Feature work across several files — you want instruction adherence and codebase comprehension. This is the bulk of real work.
- Debugging — recovery behaviour dominates. Pick the model that reads errors well.
- Architecture and gnarly algorithms — reach for raw capability.
Teams that route by task usually land somewhere near "cheap model for most things, strong model for the hard 15%" — and spend far less than teams that send everything to the most expensive option.
How to actually decide
Pick three candidates from public benchmarks — that is what leaderboards are genuinely good for. Then evaluate them on your own work, because that is the only thing that predicts your experience.
Run the same ten tasks from your git history through each. Score pass / salvageable / fail. Count turns to completion. Blind the outputs so you are not scoring the brand.
The result is frequently that the cheap model handles most of your work, which is a far more useful finding than knowing which model tops a chart.
Two traps worth avoiding
Judging on first impressions. The first few responses from a new model feel different, and different reads as better or worse depending on your mood. Run the eval before forming a view; the gap between initial impression and measured result is routinely large.
Optimising for demos. Models that produce impressive one-shot output are not necessarily the ones that work well over a long session. The behaviours that matter in real use — staying in scope, recovering from errors, following existing patterns — do not show up in a screenshot.
Keep switching cheap
Whatever you pick will be superseded. Insulate yourself: use an OpenAI-compatible endpoint, keep model names in configuration rather than code, and keep the eval set current. Then "should we switch?" is a half-hour experiment instead of a migration.
Common questions
Should I just always use the most capable model?
Usually not. Most day-to-day coding does not need frontier capability, and the strongest models are often slower. Route the hard minority of tasks to the strong model and let a cheaper one handle the rest.
What single property predicts a good agent experience?
Recovery behaviour — what the model does after it is wrong. In a loop, a model that reads the error and adapts beats a stronger model that repeats itself.
How often should I re-evaluate my model choice?
Quarterly is enough for most teams, or whenever a model you use ships a major version. If your eval set is current, re-running it takes under an hour.