Best Model for Test Generation: Coverage Is Not the Metric
Models are very good at producing tests that pass and prove nothing. Use mutation score to pick one, and treat test writing as your cheapest high-volume task.
Point any competent model at an untested module and you will get a test file in seconds. It will import the right things, cover most of the branches, and pass on the first run. It may also assert nothing that would ever fail.
This is the defining problem of generated tests. The model can see the implementation, so the cheapest way to write a passing test is to describe what the code currently does. That produces coverage without producing confidence, and the difference is invisible until an incident.
Tests that mirror the implementation
A test is supposed to encode the contract — what the function promises, independent of how it currently keeps that promise. A generated test frequently encodes the implementation instead.
The tell is a test that breaks whenever you refactor without changing behaviour. If renaming a private helper or reordering two independent operations turns the suite red, the suite is asserting mechanism rather than meaning. You have not gained a safety net; you have gained a second copy of the code that must be maintained alongside the first.
The sharper version of this problem: if the implementation has a bug, a model reading that implementation will happily write a test asserting the buggy behaviour, and now the bug is protected by a test. This is the strongest argument for generating tests from the specification, the docstring or the issue rather than from the code body.
Mutation score is the honest metric
Line coverage measures which lines were executed. It says nothing about whether an assertion would have caught a change to them. A test that calls a function and asserts it did not throw gives you full coverage and zero protection.
Mutation testing measures the thing you actually want. A mutation tool introduces small changes into your source — flipping a comparison, changing a boundary, removing a call, replacing a return value with a constant — and reruns the suite. Every mutant your tests fail to catch is a real change to behaviour your suite would have let through. The percentage caught is your mutation score.
Mature tooling exists across most ecosystems: Stryker for JavaScript, TypeScript, C# and Scala, PIT for the JVM, mutmut and Cosmic Ray for Python, cargo-mutants for Rust, go-mutesting for Go. Any of them will tell you within one run whether generated tests are worth their line count.
Run it on a module tested by a human and the same module tested by each candidate model. The ranking you get will not match the ranking by coverage percentage, and the mutation ranking is the one to trust.
What the task actually asks of a model
Test generation stresses a different set of properties than feature work:
- Adversarial imagination. Enumerating the inputs a developer did not consider — empty collections, boundary values, unicode in identifiers, concurrent calls, a clock that goes backwards. This is where models genuinely outperform tired humans, because it is recall over a large space of known failure patterns.
- Convention matching. Your suite has a shape: a fixture style, a naming pattern, a mocking approach. Tests that ignore it are technically correct and annoying forever. This is codebase comprehension, not capability.
- Restraint about mocking. The failure mode is a test that mocks every dependency until it verifies only that the mocks were called in the order the implementation calls them. Such tests pass permanently and detect nothing.
- Throughput. Test generation is high volume and individually low stakes. Latency and cost per file matter more here than on any other coding task.
Notice what is absent: frontier reasoning. Most unit tests are not intellectually demanding. This is the clearest case in the whole coding workflow for a fast, inexpensive model, and the money you save here is what funds a strong model on the tasks that need one.
Choosing between models
The open-weight field gives you real options at this price point. DeepSeek V4 Pro is a 1.6T mixture-of-experts model with 49B active parameters, MIT licensed, with a one-million-token context and an Artificial Analysis Intelligence Index around 44 — notably cheaper per token than the frontier tier, with reported strengths in algorithms and STEM reasoning that map well onto boundary-condition enumeration. MiniMax M3 sits at a comparable index score, roughly tied among open models.
At the top end, Kimi K3 — released by Moonshot AI on 16 July 2026, 2.8T parameters with open weights published on 27 July, one-million-token context, index around 57 — is a stronger model in general, and it is usually overqualified for writing a table-driven test for a validation function. Reach for it when generating tests for something genuinely intricate: a state machine, a concurrency primitive, a parser.
The routing rule writes itself. Cheap model for the bulk of the suite, strong model for the handful of modules where the specification itself is hard to reason about.
An evaluation you can run this afternoon
Pick five modules that already have good human-written tests. Delete the tests. For each candidate model:
- Generate a test file from the module's public interface and its documentation — not from the implementation body, if your harness lets you withhold it.
- Record whether it runs at all without editing. Broken imports and hallucinated fixture names are common and cost real time.
- Run mutation testing. Compare the score against the human-written baseline you deleted.
- Refactor the module without changing behaviour — rename a private helper, extract a function. Count how many generated tests break. Those were asserting implementation.
- Read the mocking. Count tests where every collaborator is mocked; those are near-worthless.
Steps three and four are the ones that discriminate. Every model passes step two, and coverage numbers will be similar across all of them.
Prompting for tests that assert something
Most of the quality gap closes with instruction changes:
Write tests for the public interface below.
Base assertions on the documented contract, not on how
the implementation happens to work. If the documented
behaviour and the implementation disagree, write the
test against the documentation and flag the conflict.
Include: happy path, empty and boundary inputs, and each
documented error condition.
Mock only network, filesystem and clock. Use real objects
for everything else.
Follow the conventions in the existing test file I have
attached as an example.
Attaching one exemplary existing test does more for convention matching than any amount of description. The instruction to flag contract conflicts is what turns test generation from a rubber stamp into something that occasionally finds a real bug.
The rule
Generate tests from the contract, not the code. Measure with mutation score, not coverage. Use a cheap fast model for the bulk of it and reserve the strong one for intricate modules. And treat any generated test that breaks on a behaviour-preserving refactor as a defect in the test, not in the refactor.
Common questions
Why do generated tests pass but not catch bugs?
Because the model can see the implementation, and describing what the code already does is the cheapest way to make a test pass. Generate from the documented contract instead, and verify with mutation testing rather than coverage.
Do I need an expensive model to generate unit tests?
Usually not. Most unit tests are not intellectually demanding, and this is a high-volume, low-stakes task where cost and latency dominate. Save the strong model for state machines, concurrency and parsers.
What is mutation testing and why use it here?
A tool makes small behavioural changes to your source and reruns the suite; every change your tests miss is a real gap. It is the only cheap way to tell whether generated tests assert anything, and mature tools exist for most languages.