Automating Code Review With LLMs Without Drowning in Noise
Guides

Automating Code Review With LLMs Without Drowning in Noise

Automated review fails on precision, not capability. How to budget comments, give the model the context a diff omits, and measure whether anyone is acting on the output.

Automated code review has one failure mode that matters, and it is not missing bugs. It is posting fourteen comments on a three-line pull request, at which point developers stop reading any of them — including the one that was right.

Precision beats recall here by a wide margin. A reviewer that finds one real problem per ten pull requests and says nothing the rest of the time gets read. A reviewer that finds four real problems and ninety suggestions gets muted within a fortnight.

Do not review what a linter already reviews

The first thing an unconstrained model does is comment on formatting, naming, missing type annotations and the absence of docstrings. Every one of those is better handled by a deterministic tool that is faster, free, and produces identical output every run.

Run your formatter, linter and type checker first, and only invoke the model on code that already passes them. This removes most of the noise before it exists, and it stops the two systems contradicting each other in the same thread.

What is left is the category linters genuinely cannot reach: logic that contradicts its own comment, an error path that swallows the error, a new query without an index, a change that breaks an invariant established elsewhere in the codebase, a permission check that was moved rather than removed.

A diff is not enough context

The single biggest cause of confident wrong comments is reviewing a unified diff in isolation. The model sees a function call with two arguments and no definition, a variable that was validated forty lines above the hunk, an error that is handled by a decorator it cannot see.

Give it more. At minimum, request generous context lines rather than the default three:

gh pr diff 1234 --patch | head -c 200000
git diff -U25 "origin/main...HEAD" -- $(git diff --name-only origin/main...HEAD)

Better still, send the whole file for files under a few hundred lines, and the enclosing function or class for larger ones. Add the definitions of the symbols the diff calls. The cost difference is small relative to a wrong comment that a human has to spend ten minutes disproving.

Tell the model explicitly that it may be missing context, and instruct it to say nothing rather than speculate. Uncertainty phrased as a confident finding is the expensive failure.

Budget the comments

Hard-cap the number of comments per pull request. Three is a good starting point. Force ranking by asking for a severity on each finding and dropping everything below a threshold.

Ask for structured output so the cap is enforceable in code rather than by hoping the prompt was obeyed:

{
  "findings": [
    {
      "file": "src/billing/charge.ts",
      "line": 88,
      "severity": "high",
      "category": "correctness",
      "claim": "Retry loop re-charges on timeout because the idempotency key is regenerated per attempt.",
      "evidence": "key is computed inside the for-loop at line 84",
      "confidence": 0.8
    }
  ]
}

Then filter in your own code: drop anything below a confidence threshold, drop categories you have decided not to surface, sort by severity, take the top three. The prompt asks; your code decides.

Requiring an evidence field is a useful forcing function. A finding that cannot point at a specific line supporting it is usually a guess, and it is easy to discard automatically.

Summary comment or inline comments

Inline comments are more actionable and more annoying. They land in the review thread, they generate notifications, and a wrong one sits permanently next to the code it maligned.

Start with a single summary comment that the bot updates in place on each push, rather than posting a new one every time. Graduate to inline comments only for the categories where your measured precision is high — usually security and correctness, rarely style or design.

Updating one comment instead of appending is worth doing early. A bot that comments on every force-push turns a five-commit pull request into a wall of stale bot output.

Make the output easy to dismiss

Every finding should be trivially closeable. Mark the whole comment clearly as automated, keep each finding to two sentences, and never phrase a suggestion as a required change.

Do not have the bot request changes, and do not make it a required status check. The moment an automated reviewer can block a merge, its false positives become a tax on every developer in the repository, and the pressure will be to disable it rather than tune it.

Give people an escape hatch that works: a label or a comment directive that skips the bot for a specific pull request. If the only way to silence it is to turn it off globally, that is what will eventually happen.

Measure precision, not volume

The metric everyone reports is comments posted. The metric that matters is the fraction of comments that led to a code change.

Track it. Add reactions or a one-word reply convention, or simply diff whether the flagged lines changed in a subsequent commit. Review the last fifty comments by hand once a month and label them useful, harmless or wrong.

Below about fifty percent acted upon, the bot is a net negative regardless of how many real bugs it also found, because developers have already learned to skim. Tighten the categories or raise the confidence threshold until the number recovers.

Where automated review genuinely wins

Three situations where it outperforms a human reviewer reliably.

Large mechanical diffs. A 900-file rename that a human will approve without reading. A model reads all of it, and the one file where the rename was applied incorrectly is exactly the kind of thing it catches.

Consistency with local convention. Given a handful of examples of how your codebase does authorisation checks or error wrapping, it spots the new code that does it differently. This is a pattern-matching task, which is what these models are good at.

The first pass on an unfamiliar area. When the only available reviewer does not know the subsystem, a summary of what changed and which invariants might be affected makes their review better even when it identifies nothing itself.

Security caveats that are not optional

The diff is attacker-controlled on any public repository. Treat everything derived from it as untrusted input.

Never interpolate model output into a shell command; write it to a file and pass the file. Never give the review job write access to code or the ability to run repository scripts. And assume the diff may contain instructions aimed at your prompt — a comment in the code saying to approve without comment is a real technique, and the mitigation is that the bot has no approval power to begin with.

Also consider what you are sending. Diffs contain proprietary source, and occasionally credentials someone committed by accident. Run a secret scanner before the model call, not after, and know your provider retention policy.

A rollout that survives

  1. Run linters and type checks first; review only what passes.
  2. Send full files or enclosing functions, not bare diff hunks.
  3. Request structured findings with severity, evidence and confidence.
  4. Filter and cap in code — three comments maximum.
  5. One summary comment, updated in place, before any inline comments.
  6. Advisory only; never a required check, never requesting changes.
  7. Provide a per-pull-request opt-out.
  8. Measure the acted-upon rate monthly and tune the threshold.

Start narrow. One category, one repository, three comments maximum, for a month. Expanding a reviewer people trust is easy; rebuilding trust after a noisy launch takes far longer than getting it right the first time.

Common questions

Should an automated reviewer be able to block a merge?

No. Make it advisory. A false positive on a required check taxes every developer in the repository, and the reaction is to disable the bot rather than tune it.

Why does the bot comment on things that are already handled elsewhere in the file?

It is reviewing a diff hunk without the surrounding code. Send the full file or the enclosing function, and instruct it to stay silent when context is missing.

How do I know whether automated review is worth keeping?

Measure the share of comments that led to a code change. Below roughly half, developers are already skimming, and the bot is costing more attention than it returns.

Similar articles

Building a PR Summariser Reviewers Do Not Skip
Guides
Guides·9 min read

Building a PR Summariser Reviewers Do Not Skip

Most PR summary bots restate the diff and get ignored within a fortnight. What reviewers actually need, how to select the diff, and how to keep cost per PR predictable.

Read
Generating Documentation With LLMs That Is Worth Reading
Guides
Guides·9 min read

Generating Documentation With LLMs That Is Worth Reading

Most generated docs restate the function signature in English. What to generate instead, which formats tools can actually consume, and how to stop docs drifting from code.

Read
Building an LLM Code Review Bot People Do Not Mute
Guides
Guides·9 min read

Building an LLM Code Review Bot People Do Not Mute

Wiring a model into pull request review: what to send it, how to post comments that land, and the signal-to-noise threshold that decides whether the bot survives.

Read