Running LLMs in CI Without Flaky, Expensive Builds
Guides

Running LLMs in CI Without Flaky, Expensive Builds

Putting a model call inside CI adds a non-deterministic network dependency to your build. How to keep it cheap, secret-safe on forks, and incapable of blocking a merge.

Adding an LLM step to CI is easy to justify and easy to regret. You have introduced a paid, rate-limited, non-deterministic network call into the one system whose entire value comes from being fast, cheap and repeatable.

It can still be worth it. But the constraints are different from any other CI step, and the failure modes — leaked credentials on fork pull requests, builds blocked by an upstream outage, a bill that scales with commit frequency — are all avoidable if you design for them up front.

Decide what the step is allowed to do

Sort every proposed use into one of two categories before you write any YAML.

Advisory steps produce a comment, a summary, a suggestion. They must never fail the build. If the provider is down, the step is skipped and nobody notices. Review comments, release note drafts and changelog summaries belong here.

Gating steps can block a merge. The bar is much higher: the check must be deterministic enough that the same commit produces the same verdict, and a false positive must be cheap to override. Very few LLM outputs qualify.

The honest default is that model output is advisory and any real gate is a conventional deterministic check. If the model generates tests, the gate is that the tests pass — not that the model approved of something.

Fork pull requests are the security cliff

This is the part that has caused real repository compromises, so it is worth being precise.

Workflows triggered by pull_request from a fork do not receive your secrets. That is the safe default, and it means an LLM step in that workflow simply has no API key. Workflows triggered by pull_request_target run in the context of the base repository, with access to secrets and a privileged token — which is why people reach for it.

Combining pull_request_target with an explicit checkout of the pull request head is the dangerous pattern. You are executing untrusted code from a fork with maintainer credentials in the environment, and everything from a modified build script to a malicious dependency can exfiltrate your keys.

The safe structure is two workflows. One runs on pull_request, checks out fork code, and has no secrets. The other runs on pull_request_target, never checks out fork code, and does only the privileged thing — posting a comment or applying a label — using an artifact produced by the first.

name: pr-summary
on:
  workflow_run:
    workflows: ["pr-analysis"]
    types: [completed]

permissions:
  contents: read
  pull-requests: write

jobs:
  comment:
    runs-on: ubuntu-latest
    steps:
      - name: Download analysis artifact
        uses: actions/download-artifact@v4
        with:
          run-id: ${{ github.event.workflow_run.id }}
          name: analysis
      - name: Post comment
        env:
          GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
        run: gh pr comment "$PR_NUMBER" --body-file analysis.md

Set permissions explicitly at the top of every workflow and grant the narrowest set that works. contents: read plus pull-requests: write covers commenting; nothing about an LLM step needs write access to code.

Treat model output posted into a pull request as untrusted text. It was generated from a diff an outsider controls, so it can contain anything the diff author wanted it to contain. Never pipe it into a shell, never let it drive a subsequent command, and prefer --body-file over interpolating it into a command line.

Only run on what changed

The naive implementation sends the whole repository, or the whole diff, on every push. Cost then scales with commit frequency, which is the one variable you least want to tax.

Scope it. Run on pull request events rather than every push. Send only the changed files, and only the changed hunks where that is sufficient. Skip paths that do not benefit — lockfiles, generated code, vendored directories, anything matching your existing lint ignores.

CHANGED=$(git diff --name-only --diff-filter=ACMR "origin/$BASE_REF"...HEAD \
  | grep -E '\.(ts|py|go)$' \
  | grep -v -E '(dist/|vendor/|\.generated\.)' )
[ -z "$CHANGED" ] && { echo "nothing to analyse"; exit 0; }

Then cap the size. A pull request touching 800 files is exactly the one where a model summary is least useful and most expensive. Above a threshold, post a note saying the diff was too large and exit cleanly.

Cache by content hash

CI re-runs the same work constantly: a rerun after a flaky test, a rebase that changes nothing semantically, a merge queue evaluating the same tree twice.

Key a cache on the hash of the exact prompt input. If the hash matches, reuse the stored output and skip the call entirely. For file-level analysis, hash per file so a one-file change does not invalidate everything.

This is also what makes the step feel fast. A cached run finishes in seconds, and steps that finish in seconds do not get disabled by frustrated developers.

Make the step incapable of blocking

An advisory step should be structurally unable to fail the build. In GitHub Actions that means continue-on-error: true on the step and a short, explicit timeout on the job:

jobs:
  analyse:
    runs-on: ubuntu-latest
    timeout-minutes: 5
    steps:
      - id: llm
        continue-on-error: true
        run: python scripts/analyse.py > analysis.md

Set a client-side timeout too. The OpenAI Python SDK defaults to a ten-minute request timeout with two automatic retries, which means a single hung call can hold a runner for a long time and cost you money for nothing. Thirty to sixty seconds is plenty for a CI analysis call.

Watch concurrency as well. A merge queue or a monorepo fan-out can launch dozens of jobs at once, each hitting the same quota. Use a workflow concurrency group to serialise, and cancel superseded runs on the same branch so an abandoned push is not still burning tokens.

Pin the model and record what ran

A floating alias can change underneath you, and when it does, your CI output changes with no commit to blame. Pin an explicit model identifier in the workflow so upgrades are a reviewable pull request.

Write the model name, the prompt version and the token counts into the step output or a job summary. When someone asks why the review comments got worse last Tuesday, that log is the answer. It is also the only way to attribute spend to a specific workflow.

Track the cost per merged pull request

The metric that matters is not tokens per month, it is spend per merged pull request. That number tells you whether the step is proportionate, and it makes the trade-off legible when someone proposes adding a second one.

Sum token usage per workflow run, tag it with the repository, and put it somewhere visible. Automation whose cost nobody can see is automation that gets switched off in a panic during the next budget review rather than tuned.

A checklist before merging the workflow

  1. Advisory by default; only deterministic checks gate merges.
  2. No secrets in any workflow that checks out fork code.
  3. Explicit permissions block, narrowest scope that works.
  4. Model output treated as untrusted; never interpolated into shell.
  5. Changed files only, with a hard size cap.
  6. Content-hash cache so reruns are free.
  7. continue-on-error, job timeout, and a short client timeout.
  8. Concurrency group and cancellation of superseded runs.
  9. Pinned model, logged prompt version, logged token usage.

Common questions

Should an LLM check ever fail a build?

Rarely. Model output is advisory. If you want a gate, make it a deterministic check downstream of the generation — for example, that generated tests compile and pass.

How do I run an LLM step on pull requests from forks?

Split it in two. A no-secrets workflow analyses fork code and uploads an artifact; a separate privileged workflow that never checks out fork code posts the result.

How do I stop CI model calls from getting expensive?

Scope to changed files, cap diff size, cache on a content hash, run on pull requests rather than every push, and cancel superseded runs on the same branch.

Similar articles

GitHub Actions LLM Setup: A Workflow That Reviews PRs
Guides
Guides·9 min read

GitHub Actions LLM Setup: A Workflow That Reviews PRs

A working GitHub Actions job that calls an OpenAI-compatible model on pull requests — repository secrets, timeouts, rate limits, PR comments and cost control.

Read
Building a Changelog Generator People Actually Read
Guides
Guides·9 min read

Building a Changelog Generator People Actually Read

Restating commit subjects is not a changelog. How to pick the right input, separate classification from writing, handle reverts, and keep regeneration deterministic.

Read
Building a PR Summariser Reviewers Do Not Skip
Guides
Guides·9 min read

Building a PR Summariser Reviewers Do Not Skip

Most PR summary bots restate the diff and get ignored within a fortnight. What reviewers actually need, how to select the diff, and how to keep cost per PR predictable.

Read