The Best Model for Refactoring Is the One That Stays in Scope
Models

The Best Model for Refactoring Is the One That Stays in Scope

Refactoring is a constraint problem, not a generation problem. What the task actually demands from a model, and how to measure it on your own repo.

Refactoring is the one coding task where producing new ideas is a failure mode. The output is supposed to be behaviour-identical to the input. Every creative impulse the model has is, by definition, a bug.

That inverts most of what leaderboards reward. A benchmark asks "can it solve this problem?" A refactor asks "can it change fifteen call sites and change nothing else?" Those are different skills, and models are not equally good at both.

What the task actually demands

Break a real refactor down and you get four requirements, roughly in order of how often they are the thing that fails:

  • Scope discipline. You asked it to rename a method. It renamed the method, reformatted the file, "modernised" three loops into comprehensions, and added type hints. Now your diff is 400 lines and unreviewable.
  • Completeness. The opposite failure. It changed eleven of the fourteen call sites, and the three it missed are in a test helper, a string-based dynamic dispatch, and a file it never opened.
  • Consistency over distance. A large refactor runs long. The convention the model picks in file one must still hold in file twenty, forty minutes later.
  • Comprehension before editing. Knowing which call sites exist requires actually reading the codebase, not pattern-matching on the symbol name.

Note that raw reasoning capability barely appears on that list. Refactoring is rarely hard in the sense that a competitive programming problem is hard. It is hard in the sense that a merge is hard: lots of surface, low tolerance for error.

Why context window matters here more than elsewhere

For most tasks, a large context window is a convenience. For refactoring it is closer to a prerequisite, because the unit of work is the call graph rather than the file.

If the model cannot see every caller at once, it has to reconstruct them through search, and search results arrive as fragments without the surrounding invariants. That is exactly where the "missed three call sites" failure comes from.

The current generation has largely solved the raw capacity question. Kimi K3 and GLM 5.2 both advertise a one-million-token context window, as does DeepSeek V4 Pro. Whether a model uses that window well over a long session is a separate question, and one you should test rather than assume — attention quality at the far end of a long context degrades unevenly between models.

Long-horizon behaviour is the real differentiator

A refactor is a long-horizon task: dozens of edits, each of which must remain compatible with the ones before it. This is the property that separates models that feel good in a chat window from models that survive an agent loop.

GLM 5.2 is the interesting case here. Z.ai released it on 13 June 2026 as a mixture-of-experts model of roughly 744B total and 40B active parameters under an MIT licence, and its reported strengths cluster specifically around long-horizon coding evaluations — FrontierSWE, DeepSWE, Terminal-Bench 2.1 and SWE-bench Pro — rather than around single-shot puzzle solving. On the Artificial Analysis Intelligence Index it sits around 51, below Kimi K3 at roughly 57, and yet for sustained multi-file work it is a serious candidate. That gap between aggregate index and task-specific fit is the whole argument for evaluating by task.

How to measure it on your own code

You do not need a benchmark harness. You need your git history.

Find five refactors your team actually merged — a rename across modules, an interface extraction, a dependency swap, a signature change with call-site updates, a dead-code removal. Check out the parent commit, give the model the same instruction the ticket gave the human, and score four things:

  1. Does the test suite still pass? Binary, and non-negotiable. A refactor that changes behaviour is not a refactor.
  2. Blast radius ratio. Lines changed by the model divided by lines changed in the human commit. Around 1.0 is ideal. Anything above 2 means it is doing unrequested work, and you will pay for that in review time forever.
  3. Miss count. How many call sites did it fail to update? Your compiler or type checker finds most of these for free; dynamic languages will need a grep.
  4. Turns to green. If it broke something, how many iterations did it need to notice and fix it?

Blast radius ratio is the metric almost nobody tracks and the one that best predicts whether you will enjoy using a model for this. Two models can both reach a passing test suite while producing diffs that differ by an order of magnitude in reviewability.

Prompt structure changes the outcome more than model choice

Before concluding that a model is bad at refactoring, check that you asked properly. The same model behaves very differently under these two prompts:

Clean up the user service.

vs.

Rename UserService.fetch to UserService.load.
Update all call sites. Do not change any other code,
including formatting. If you find a call site you are
unsure about, list it instead of editing it.

The second prompt does three things: it names the exact transformation, it forbids collateral edits explicitly, and it gives the model a legal way to express uncertainty. That last part matters — without an escape hatch, a model that is unsure will guess, and guessing is how you get subtly broken code.

For large refactors, stage it. Ask for a plan and a list of affected files first, review that list, then let it execute file by file with the test suite running between steps. This converts one long-horizon task into several short ones, which is a much easier problem for any model.

Where a cheaper model is genuinely fine

Mechanical refactors — extract a constant, inline a variable, split a function, convert callbacks to async/await within one file — do not need a frontier model. They need a model that follows instructions and does not editorialise. Route them to whatever is fast and cheap, and save the strong model for refactors that cross module boundaries or touch a public interface.

Better still, remember that your IDE already does the safe subset perfectly. A language-server rename is deterministic and instant. Use the model for the refactors that require judgement about what the code should look like, not for the ones a parser can prove correct.

A decision rule

Pick two candidates: one strong long-horizon model for cross-module work, one fast model for single-file mechanical edits. Run the five-refactor eval above on both. Track blast radius ratio, not just pass rate. Keep the model name in configuration so that when the next release lands, re-running the eval is an afternoon rather than a migration.

Common questions

Does a bigger context window make a model better at refactoring?

It removes a hard ceiling but does not guarantee quality. A model needs to see every call site to update them all, so capacity matters — but how well it attends to the far end of a long context varies, and only testing on your own repo shows which models hold up.

Why does the model change more code than I asked for?

Usually because the instruction left room. Name the exact transformation, forbid collateral edits including formatting, and give it a way to flag uncertain call sites instead of guessing. Prompt structure moves this more than model choice does.

How should I score a refactoring eval?

Test suite passing is table stakes. The metric that actually discriminates is blast radius: lines the model changed divided by lines the human commit changed. Near 1.0 is good; above 2 means unrequested work you will pay for in review.

Similar articles

Best Model for Code Review: Optimise for Precision, Not Recall
Models
Models·9 min read

Best Model for Code Review: Optimise for Precision, Not Recall

An automated reviewer that flags everything gets muted within a week. How to pick and evaluate a model on false-positive rate, and why review has odd economics.

Read
Best Model for Legacy Code: Comprehension Beats Cleverness
Models
Models·9 min read

Best Model for Legacy Code: Comprehension Beats Cleverness

Legacy work is mostly reading, not writing. Why the model that scores highest on code generation is often the wrong one for a twenty-year-old codebase.

Read
How to Choose a Coding Model Without Trusting Anyone’s Chart
Models
Models·9 min read

How to Choose a Coding Model Without Trusting Anyone’s Chart

Six properties decide whether a model is good at your codebase — and only one of them shows up on a leaderboard. A practical framework for picking.

Read