Generating Documentation With LLMs That Is Worth Reading
Guides

Generating Documentation With LLMs That Is Worth Reading

Most generated docs restate the function signature in English. What to generate instead, which formats tools can actually consume, and how to stop docs drifting from code.

The default output of "document this codebase" is a docstring on every function that says, in a full sentence, exactly what the signature already said. Gets the user by ID. Args: user_id (str): The user ID. Returns: User: The user.

That is not documentation. It is noise with a maintenance cost, because it now has to be kept in sync with a signature that already communicated the same information more precisely.

Useful documentation records what the code cannot say for itself: why the approach was chosen, what the caller must not do, which invariants hold, what happens when things fail. Some of that a model can infer. Some of it only exists in someone head, and no amount of generation will conjure it.

Generate the parts that are recoverable from code

Sort documentation by whether the information exists in the repository at all.

Recoverable. What a module contains, how a request flows through the system, which exceptions a function can actually raise, which endpoints exist and what they accept, how to run the thing locally. All of this is derivable from source, tests, configuration and CI files, and a model reconstructs it faster than a human reads it.

Partly recoverable. Usage examples. A model can write a plausible one; whether it runs is a different matter, and an example that does not run is worse than none. Verify by executing them.

Not recoverable. Why this design rather than the obvious alternative. Which constraint made the ugly bit necessary. What broke last time somebody changed it. If a model produces confident text in this category, it invented it, and readers will believe it.

The practical rule: let the model draft everything recoverable, and leave a visible gap where the rationale belongs rather than accepting a fabricated one.

Pick a format your tooling already consumes

Free-form prose in a docstring is worth much less than the same content in a format your documentation generator, IDE and type checker understand.

In Python the mainstream choices are Google style, NumPy style and reStructuredText. Napoleon, the Sphinx extension, parses both Google and NumPy styles and converts them to reStructuredText before Sphinx processes them, so either works with autodoc. Google style uses indentation to separate sections and stays more compact; NumPy style uses underlines and more vertical space, and is the convention in NumPy, SciPy and pandas.

Whichever you pick, state it in the prompt along with an existing example, or you will get all three across one codebase.

def charge(customer_id: str, amount_cents: int) -> Charge:
    """Charge a customer and record the result.

    Args:
        customer_id: Stripe customer identifier.
        amount_cents: Amount in the smallest currency unit. Must be positive.

    Returns:
        The persisted Charge row.

    Raises:
        InsufficientFunds: Card declined for balance reasons.
        RateLimited: Upstream throttled the request; safe to retry.

    Example:
        >>> charge("cus_123", 500).status
        'succeeded'
    """

The Raises section is the one worth the most and the one humans skip most often. It is also directly recoverable — the exceptions a function can propagate are in the call graph — which makes it an ideal generation target.

Make examples executable

An example in a docstring that no longer works is the fastest way to lose reader trust. Written in doctest format, examples become tests:

pytest --doctest-modules src/

Now a generated example that hallucinated a method name fails in CI rather than misleading someone in six months. For TypeScript, the equivalent is extracting fenced code blocks from documentation and type-checking or executing them as part of the build.

This single practice changes the economics of generated documentation. Examples are the part models get subtly wrong most often, and they are also the part you can verify mechanically.

Generate at review time, not in a bulk pass

The one-off run that documents 4,000 functions produces a diff nobody can review, so it gets approved unread, and every mistake in it is now permanent and load-bearing.

Attach generation to the change instead. When a pull request adds or modifies a public function without a docstring, propose one in the review. The author is the person who knows whether it is right, and they are looking at the code at that exact moment.

The same applies to prose documentation. Rather than regenerating a guide, detect that a pull request changed something a document describes and flag the specific paragraph as possibly stale. Drift detection is more valuable than regeneration, because the failure mode of documentation is being wrong, not being absent.

Feed it more than the file

Documentation quality tracks context quality closely. A model looking at one file writes about one file.

Include the tests for the code being documented — they encode the intended contract better than the implementation does, including the edge cases someone deliberately handled. Include the callers, which show how the function is meant to be used. Include the module docstring and a neighbouring documented function so tone and format carry over.

For a README or architecture overview, the useful inputs are the directory tree, the dependency manifest, the CI workflow and the entry points. Those four together let a model describe how to build, run and test a project accurately, which is the part of a README people actually use.

Changelogs and release notes

This is where generation is least controversial, because the source material is complete and structured. Given the commits between two tags, a model produces a grouped, readable summary quickly.

git log --no-merges --pretty=format:'%s%n%b' v1.4.0..v1.5.0

Two rules make the output usable. Write for the consumer of the release rather than the author of the commit — users care that pagination stopped dropping the last page, not that a helper was refactored. And never let the model infer that something is a breaking change; take that from an explicit marker such as a conventional commit ! or a BREAKING CHANGE footer, because a missed breaking change in release notes is a real incident.

Keep the human review cheap

Generated documentation is only worth having if someone checks it, and checking is only sustainable if it is fast.

Ask for short output. Two or three sentences of summary, then the structured sections. Long generated prose does not get read, and it hides errors in the middle.

Require citations to code. If a docstring claims a function retries three times, it should be because a retry decorator says so. Findings you can trace are findings you can verify in seconds.

And keep generated content in the same file as the code it describes, not in a parallel documentation tree. Documentation stored next to the thing it documents is more likely to be updated in the same pull request; documentation in a separate repository is guaranteed to rot.

A practical policy

  1. Do not document what the signature already states.
  2. Prioritise failure modes, invariants and caller obligations.
  3. Fix one docstring format and supply an example of it in every prompt.
  4. Make examples executable and run them in CI.
  5. Generate per pull request, not in a bulk backfill.
  6. Include tests and callers in the context, not just the file.
  7. Take breaking-change flags from commit metadata, never from inference.
  8. Flag stale documentation rather than silently regenerating it.

The test for any generated document is simple: would a new engineer be better or worse off having read it. Text that restates the code makes them slower. Text that tells them what happens when the call fails makes them faster, and that is the material worth spending tokens on.

Common questions

Should I bulk-generate docstrings across an existing codebase?

Usually not. The diff is too large to review honestly, so errors get approved unread. Generate per pull request, where the author can verify the claim while looking at the code.

Which Python docstring format works best with generated output?

Any of Google, NumPy or reStructuredText, provided you fix one and say so in the prompt. Napoleon parses Google and NumPy styles for Sphinx autodoc, so both integrate cleanly.

How do I stop generated examples from being wrong?

Write them in doctest format and run them in CI. An example that references a method that does not exist then fails the build instead of misleading a reader months later.

Similar articles

Automating Code Review With LLMs Without Drowning in Noise
Guides
Guides·9 min read

Automating Code Review With LLMs Without Drowning in Noise

Automated review fails on precision, not capability. How to budget comments, give the model the context a diff omits, and measure whether anyone is acting on the output.

Read
Building a Changelog Generator People Actually Read
Guides
Guides·9 min read

Building a Changelog Generator People Actually Read

Restating commit subjects is not a changelog. How to pick the right input, separate classification from writing, handle reverts, and keep regeneration deterministic.

Read
Building a Code Search Tool That Beats grep
Guides
Guides·10 min read

Building a Code Search Tool That Beats grep

Semantic code search only pays off if you beat ripgrep on the queries it fails. Chunking at symbol boundaries, hybrid retrieval, incremental indexing and honest evaluation.

Read