Best Model for Code Review: Optimise for Precision, Not Recall
An automated reviewer that flags everything gets muted within a week. How to pick and evaluate a model on false-positive rate, and why review has odd economics.
Automated code review has an unusual failure mode: the tool that finds the most problems is usually the worst one. Every comment a human dismisses costs attention, and a reviewer that produces six speculative nits per pull request gets collapsed, filtered, then ignored — including on the pull request where it was right.
So the question is not "which model finds the most issues?" It is "which model can be trusted to stay quiet?" That reframing changes both the model you pick and the way you evaluate it.
Review is a judgement task, not a detection task
The genuinely mechanical parts of review are already solved and not by language models. Formatting is a formatter. Unused imports are a linter. Type errors are a type checker. Known vulnerable dependencies are a scanner. All of those are deterministic, instant and free, and a model that spends its output re-deriving them is wasting the slot.
What is left is the part that requires judgement: is this the right abstraction, does this change break an implicit contract elsewhere, is this error handling actually reachable, does this introduce a race, is the test asserting anything meaningful. Those are questions about intent and context, and they are the ones worth a model.
Which means the capability that matters is not raw problem-solving. It is calibration — knowing the difference between "this is wrong", "this is unusual but fine", and "I do not have enough context to say".
The economics are inverted, and that is useful
Most coding work is output-heavy: a short instruction, a long generated file. Review is the opposite. You feed in a large diff plus surrounding files plus conventions, and you want back a handful of specific comments.
Two consequences follow. First, review is cheap relative to how expensive it feels, because input tokens are the bulk of the cost and input is generally the cheaper side of the bill. Second, prompt caching is unusually effective here — if you prepend the same style guide, architecture notes and convention document to every review, that prefix is identical across every pull request in your repository.
Together these mean you can afford a stronger model for review than you can for generation, and you should. The bill scales with diff size, and diffs are small compared to the codebases they touch.
Context is the difference between a nit and a real finding
A diff on its own is close to unreviewable. It tells you what changed, not what the code around it guarantees. Almost every high-value review comment depends on something outside the diff: a caller that assumes the old behaviour, a convention established elsewhere, a test that will now be vacuous.
This is where large context windows earn their place. Kimi K3 and GLM 5.2 both offer a one-million-token window, which is enough to put an entire mid-sized service in the prompt alongside the diff. That is not merely convenient — it converts questions the model would otherwise have to guess at into questions it can answer by looking.
GLM 5.2 is worth a specific mention here. Released by Z.ai on 13 June 2026 as an MIT-licensed mixture-of-experts model of roughly 744B total and 40B active parameters, its reported strengths sit in long-horizon and repository-scale coding evaluations rather than in single-shot puzzles. Its Artificial Analysis Intelligence Index score of around 51 is below Kimi K3 at roughly 57, and for review that ordering may not hold, because review rewards codebase comprehension over raw problem difficulty. Test both; the aggregate index is not the metric for this job.
Evaluate with a clean-PR corpus, not just a bug corpus
Almost everyone builds a review eval by seeding bugs into pull requests and measuring how many the model catches. That measures recall, which is the less important half.
Build two corpora instead:
- Twenty pull requests with known defects. Real ones from your history, ideally the ones that caused incidents, checked out at the commit before the fix. Measure catch rate.
- Twenty pull requests that were merged clean and never needed a follow-up fix. Measure comment count. This is your false-positive corpus, and it is the one that predicts whether your team will keep the tool.
Report two numbers per model: defects caught out of twenty, and comments emitted on the clean set. A model that catches fourteen defects and emits three comments on clean pull requests is dramatically better than one that catches sixteen and emits forty, even though the second looks stronger on the metric most people report.
Rate every comment on the clean set as correct, harmless or noise. Noise includes anything stylistic your formatter already handles, anything phrased as "consider whether", and anything the model could not possibly know without context it was not given.
Prompt design controls the noise rate
Before switching models, tighten the instruction. Reviewers get quieter when you give them a budget and a bar:
Review this diff. Report only issues that would cause
incorrect behaviour, data loss, a security problem, or
a broken contract with existing callers.
Do not comment on style, naming, or formatting.
Maximum five comments. If there are no such issues,
reply exactly: No blocking issues found.
For each issue give: file and line, what breaks, and
the concrete input or sequence that triggers it.
Three parts do the work. The explicit severity bar removes the entire category of stylistic noise. The comment cap forces ranking, so you get the model's top findings rather than everything it noticed. And requiring a triggering input filters out speculative comments, because a model that cannot name the input usually does not have a real finding.
The permission to say nothing matters more than it looks. Without it, a model asked to review will find something, because finding nothing reads as unhelpful.
Where automated review genuinely does not belong
It will not tell you whether a feature should exist, whether an abstraction is worth its cost across the next year, or whether a change fits the direction the team agreed. Those are the parts of review that transfer knowledge between people, and outsourcing them is a net loss even when the output looks reasonable.
The honest positioning is a first pass that catches the mechanical-but-not-lintable class of defect — the unhandled error path, the off-by-one in a boundary condition, the missing await — so that human attention arrives at the pull request already spent on design rather than on scanning.
A shipping checklist
- Turn off everything your linter, formatter and type checker already cover.
- Give the model your conventions document as a cached prefix on every review.
- Cap comments and set an explicit severity bar in the prompt.
- Measure clean-PR comment count weekly. If it climbs, tighten the prompt before blaming the model.
- Keep human review for design, direction and anything that teaches somebody something.
Common questions
Should the review model be the same one that writes the code?
Not necessarily, and there is an argument for a different one. A model reviewing its own output tends to accept its own assumptions. A second model with a different training background flags things the first one considered settled.
How do I stop an AI reviewer from being noisy?
Set an explicit severity bar, cap the number of comments so it has to rank, require a concrete triggering input for each finding, and give it an exact phrase to use when it has nothing to say. Prompt changes usually fix this before a model swap does.
Is automated code review expensive to run on every pull request?
Less than expected. Review is input-heavy and output-light, and input is the cheaper side of the bill. Caching a fixed prefix of conventions and architecture notes across reviews cuts it further.