How to Choose a Coding Model Without Trusting Anyone’s Chart
Models

How to Choose a Coding Model Without Trusting Anyone’s Chart

Six properties decide whether a model is good at your codebase — and only one of them shows up on a leaderboard. A practical framework for picking.

Every model is "state of the art at coding". They cannot all be right, and the disagreement is not really about quality — it is that "good at coding" is at least six different properties, and different models are strong at different ones.

Here is the framework I would use, in the order that actually matters for day-to-day work.

1. Instruction adherence

Does it do what you asked, and only what you asked?

This is the most underrated property and the one that causes the most friction. You ask for a two-line fix and get a refactor of the whole module. The refactor might even be better — but you cannot review it quickly, and you did not ask for it.

Test it: give a deliberately narrow instruction ("change only the retry count, touch nothing else") and see whether the diff respects the boundary.

2. Recovery behaviour

What happens when it is wrong?

In an agent loop this dominates everything else. A model that writes mediocre code but reads the failing test and fixes itself will out-perform a brilliant model that repeats the same broken approach three times.

Test it: let it make a change that fails a test, and watch. Does it read the error, form a new hypothesis, and adjust? Or restate the same solution more confidently?

3. Codebase comprehension

Can it work with your conventions rather than generic ones?

Benchmarks favour self-contained problems. Real work is "add a field to this entity" across an ORM, a DTO, a validator, a migration and a test — following patterns already established in the repo.

Test it: ask for a change that should mirror an existing pattern. Does it find and follow the pattern, or invent its own?

4. Tool-calling reliability

Does it call the right tool with valid arguments, consistently?

A model that reasons beautifully but emits malformed tool calls is unusable in an agent. This is largely binary — models are either dependable here or they are not.

5. Latency

Not raw tokens per second — time to a useful result.

A fast model that needs four turns often beats a slow one that needs two. And in interactive work, time to first token matters more than throughput, because it determines whether the tool feels responsive.

6. Raw capability

The thing benchmarks actually measure: can it solve genuinely hard problems?

It matters least often, because most day-to-day work is not algorithmically hard. It matters enormously on the few tasks that are. Keep access to a strong model for those, but do not pay the premium on every request.

Matching model to task

The practical answer is usually "more than one":

  • Boilerplate, tests, renames, docstrings — a fast, cheap model is genuinely sufficient. Do not overthink it.
  • Feature work across several files — you want instruction adherence and codebase comprehension. This is the bulk of real work.
  • Debugging — recovery behaviour dominates. Pick the model that reads errors well.
  • Architecture and gnarly algorithms — reach for raw capability.

Teams that route by task usually land somewhere near "cheap model for most things, strong model for the hard 15%" — and spend far less than teams that send everything to the most expensive option.

How to actually decide

Pick three candidates from public benchmarks — that is what leaderboards are genuinely good for. Then evaluate them on your own work, because that is the only thing that predicts your experience.

Run the same ten tasks from your git history through each. Score pass / salvageable / fail. Count turns to completion. Blind the outputs so you are not scoring the brand.

The result is frequently that the cheap model handles most of your work, which is a far more useful finding than knowing which model tops a chart.

Two traps worth avoiding

Judging on first impressions. The first few responses from a new model feel different, and different reads as better or worse depending on your mood. Run the eval before forming a view; the gap between initial impression and measured result is routinely large.

Optimising for demos. Models that produce impressive one-shot output are not necessarily the ones that work well over a long session. The behaviours that matter in real use — staying in scope, recovering from errors, following existing patterns — do not show up in a screenshot.

Keep switching cheap

Whatever you pick will be superseded. Insulate yourself: use an OpenAI-compatible endpoint, keep model names in configuration rather than code, and keep the eval set current. Then "should we switch?" is a half-hour experiment instead of a migration.

Common questions

Should I just always use the most capable model?

Usually not. Most day-to-day coding does not need frontier capability, and the strongest models are often slower. Route the hard minority of tasks to the strong model and let a cheaper one handle the rest.

What single property predicts a good agent experience?

Recovery behaviour — what the model does after it is wrong. In a loop, a model that reads the error and adapts beats a stronger model that repeats itself.

How often should I re-evaluate my model choice?

Quarterly is enough for most teams, or whenever a model you use ships a major version. If your eval set is current, re-running it takes under an hour.

Similar articles

Best Model for Code Review: Optimise for Precision, Not Recall
Models
Models·9 min read

Best Model for Code Review: Optimise for Precision, Not Recall

An automated reviewer that flags everything gets muted within a week. How to pick and evaluate a model on false-positive rate, and why review has odd economics.

Read
The Best Model for Refactoring Is the One That Stays in Scope
Models
Models·9 min read

The Best Model for Refactoring Is the One That Stays in Scope

Refactoring is a constraint problem, not a generation problem. What the task actually demands from a model, and how to measure it on your own repo.

Read
A/B Testing Two Models Without Fooling Yourself
Models
Models·9 min read

A/B Testing Two Models Without Fooling Yourself

Comparing two models on live traffic sounds simple and usually is not. Sample sizes, paired designs, and the metrics that actually settle the question.

Read