Generating Tests With LLMs: Coverage Is Not the Goal
Guides

Generating Tests With LLMs: Coverage Is Not the Goal

Models are very good at writing tests that pass and prove nothing. How to get suites that actually catch regressions, and how mutation testing tells you which ones do.

Ask a model to write tests for a function and you will get tests. They will run, they will pass, and coverage will go up. Whether any of them would fail if the function were broken is a separate question, and it is the only one that matters.

This is the central problem with generated tests. The model reads the implementation and writes assertions describing what the implementation currently does. If the implementation is wrong, the test now enshrines the bug and reports success.

The tautology trap

Given an implementation, a model infers intent from the code rather than from the requirement. The result reads like a test and behaves like a snapshot:

def apply_discount(price, pct):
    return price - (price * pct)          # pct as a fraction, undocumented

# Generated test
def test_apply_discount():
    assert apply_discount(100, 0.2) == 80.0

That test passes. It also passes if the intended contract was a percentage from 0 to 100 and the function is off by a factor of a hundred. The test learned the bug.

Two mitigations, and you want both. Derive tests from the specification — the docstring, the ticket, the API contract — rather than only from the body. And review generated assertions specifically for the question "would this fail if the behaviour were wrong", which is a much faster review than reading the test as prose.

Coverage measures execution, not verification

Line coverage tells you a line ran. It says nothing about whether anything was checked. A test that calls a function inside a try block and asserts nothing produces full coverage of that function.

Generated suites inflate this metric more than handwritten ones, because producing calls is easy and producing meaningful assertions is not. Branch coverage is better than line coverage and still measures the wrong thing.

Mutation testing measures the right thing. It introduces small deliberate faults — flipping a conditional boundary, replacing a return with a null or a constant, negating a condition — reruns the suite, and reports how many of those faults a test caught. PIT, whose default operators include conditionals boundary, empty returns, true and false returns and null returns, does this for Java; Stryker covers JavaScript, TypeScript, C# and Scala; mutmut covers Python and has a CI flag that produces pipeline-appropriate exit codes.

The mutation score is a far more honest quality signal than line coverage. There is no universal target — what matters is which mutants survive, and a single fixed threshold across a whole codebase tends to mislead.

Use the model where it is genuinely strong

Three tasks where generation beats handwriting consistently.

Enumerating edge cases. Ask for the inputs that might break a function rather than for tests. Empty collections, unicode, negative zero, boundary values, timezone-crossing dates, concurrent access. You get a checklist you can triage, and you write the assertions yourself.

Characterisation tests for legacy code. When you need to refactor something nobody understands, snapshot-style tests that pin current behaviour are exactly right, and the tautology problem becomes a feature. Label them as characterisation tests so nobody mistakes them for a specification.

Filling out a table. Once you have written one good parametrised test, extending it to twenty cases is mechanical work a model does quickly and accurately.

@pytest.mark.parametrize("price,pct,expected", [
    (100, 0.0, 100.0),
    (100, 1.0, 0.0),
    (0, 0.5, 0.0),
])
def test_apply_discount(price, pct, expected):
    assert apply_discount(price, pct) == expected

Property-based tests survive refactoring

Example-based tests break when implementation details change. Properties do not, because they state invariants rather than outputs.

Models are good at proposing candidate properties, which is the hard part of property-based testing for most people. Round-trip properties, idempotence, ordering invariants, conservation of totals, monotonicity — ask for the invariants a function should satisfy, then implement the ones that are genuinely true with Hypothesis, fast-check or your ecosystem equivalent.

Be sceptical of proposed properties. A plausible-sounding invariant that is not actually required produces a test that fails for correct code, which is worse than no test.

Give the model your conventions

Generated tests that do not match your existing style get rewritten by hand, which erases the time saving. Include one representative existing test file in the prompt, and it will follow the fixtures, factories, naming scheme and assertion helpers you already use.

State the rules that are not visible in a single file: which fixtures exist, whether the database is real or mocked, that time must be frozen rather than read, that network access is forbidden. Otherwise you get tests that reach the network, sleep for real seconds, or depend on today being a weekday.

Reject the flaky ones automatically

Generated tests introduce nondeterminism at a higher rate than handwritten ones. Before any of them merge, run the new tests several times and in a randomised order:

pytest tests/generated -p no:randomly -q --count=5   # via pytest-repeat
pytest tests/generated -q -p randomly                 # random order

Anything that fails intermittently, or that only passes when run after another test, is deleted rather than debugged. A flaky test costs more attention over its lifetime than the bug it might have caught, and generated tests are cheap enough to discard without regret.

Also check runtime. A suite that grows by four minutes for marginal assertions makes every future change slower, and nobody attributes the slowdown to the day you bulk-generated tests.

A workflow that produces suites worth keeping

  1. Write or locate the specification — docstring, contract, ticket. Generate from that, not only from the implementation.
  2. Ask for edge cases first, triage them, then ask for tests covering the survivors.
  3. Include an existing test file so conventions carry over.
  4. Run the suite against deliberately broken code and delete every test that still passes.
  5. Run a mutation testing pass on the changed files; investigate surviving mutants.
  6. Repeat and shuffle to catch flakiness; delete rather than fix.
  7. Review assertions, not test bodies. The assertion is the test.

Step four is the cheap version of step five and takes about a minute. Break the function on purpose — invert a comparison, return a constant — and see what turns red. Tests that stay green are noise, and deleting them immediately is the highest-value thing you can do with generated output.

Where this does not help

Integration and end-to-end tests depend on system topology, fixtures, external services and timing that no model can infer from source. Generated versions of these tend to be plausible-looking and non-functional, and debugging them costs more than writing them would have.

Security tests are similar. A generated test for an authorisation check verifies the code path that exists, which is precisely the path that is not the vulnerability. Threat modelling is not a pattern-matching problem.

Unit tests on pure functions, parametrised expansions of an existing case, and characterisation tests around legacy code are the sweet spot. That is a narrower claim than the marketing, and it is still a lot of otherwise tedious work.

Common questions

Why do generated tests pass on buggy code?

Because they were written from the implementation, so they assert what the code currently does. Generate from the specification where one exists, and verify by breaking the code on purpose.

Is high coverage from generated tests worth anything?

On its own, little. Coverage records that a line executed, not that anything was checked. A mutation testing run tells you how many deliberate faults the suite actually catches.

What should I do with a generated test that fails intermittently?

Delete it. Flaky tests cost more attention over their lifetime than the bugs they catch, and regenerating is cheap enough that debugging one is rarely the right trade.

Similar articles

Evaluating Prompt Changes Without Fooling Yourself
Guides
Guides·9 min read

Evaluating Prompt Changes Without Fooling Yourself

You edited a prompt and the output looks better. Here is how to find out whether it actually is, with paired runs, enough samples and judges you can trust.

Read
Generating API Clients With an LLM Without Silent Drift
Guides
Guides·9 min read

Generating API Clients With an LLM Without Silent Drift

Where a model beats openapi-generator, where it quietly loses, and how to build a generate-compile-test loop that catches the hallucinated field before you ship it.

Read
Generating Regular Expressions With an LLM, Carefully
Guides
Guides·9 min read

Generating Regular Expressions With an LLM, Carefully

Models write regex fluently and confidently, which is the problem. How to specify the target, demand test cases, avoid catastrophic backtracking and spot dialect mismatches.

Read