Running LLMs in CI Without Flaky, Expensive Builds
Putting a model call inside CI adds a non-deterministic network dependency to your build. How to keep it cheap, secret-safe on forks, and incapable of blocking a merge.
Adding an LLM step to CI is easy to justify and easy to regret. You have introduced a paid, rate-limited, non-deterministic network call into the one system whose entire value comes from being fast, cheap and repeatable.
It can still be worth it. But the constraints are different from any other CI step, and the failure modes — leaked credentials on fork pull requests, builds blocked by an upstream outage, a bill that scales with commit frequency — are all avoidable if you design for them up front.
Decide what the step is allowed to do
Sort every proposed use into one of two categories before you write any YAML.
Advisory steps produce a comment, a summary, a suggestion. They must never fail the build. If the provider is down, the step is skipped and nobody notices. Review comments, release note drafts and changelog summaries belong here.
Gating steps can block a merge. The bar is much higher: the check must be deterministic enough that the same commit produces the same verdict, and a false positive must be cheap to override. Very few LLM outputs qualify.
The honest default is that model output is advisory and any real gate is a conventional deterministic check. If the model generates tests, the gate is that the tests pass — not that the model approved of something.
Fork pull requests are the security cliff
This is the part that has caused real repository compromises, so it is worth being precise.
Workflows triggered by pull_request from a fork do not receive your secrets. That is the safe default, and it means an LLM step in that workflow simply has no API key. Workflows triggered by pull_request_target run in the context of the base repository, with access to secrets and a privileged token — which is why people reach for it.
Combining pull_request_target with an explicit checkout of the pull request head is the dangerous pattern. You are executing untrusted code from a fork with maintainer credentials in the environment, and everything from a modified build script to a malicious dependency can exfiltrate your keys.
The safe structure is two workflows. One runs on pull_request, checks out fork code, and has no secrets. The other runs on pull_request_target, never checks out fork code, and does only the privileged thing — posting a comment or applying a label — using an artifact produced by the first.
name: pr-summary
on:
workflow_run:
workflows: ["pr-analysis"]
types: [completed]
permissions:
contents: read
pull-requests: write
jobs:
comment:
runs-on: ubuntu-latest
steps:
- name: Download analysis artifact
uses: actions/download-artifact@v4
with:
run-id: ${{ github.event.workflow_run.id }}
name: analysis
- name: Post comment
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: gh pr comment "$PR_NUMBER" --body-file analysis.md
Set permissions explicitly at the top of every workflow and grant the narrowest set that works. contents: read plus pull-requests: write covers commenting; nothing about an LLM step needs write access to code.
Treat model output posted into a pull request as untrusted text. It was generated from a diff an outsider controls, so it can contain anything the diff author wanted it to contain. Never pipe it into a shell, never let it drive a subsequent command, and prefer --body-file over interpolating it into a command line.
Only run on what changed
The naive implementation sends the whole repository, or the whole diff, on every push. Cost then scales with commit frequency, which is the one variable you least want to tax.
Scope it. Run on pull request events rather than every push. Send only the changed files, and only the changed hunks where that is sufficient. Skip paths that do not benefit — lockfiles, generated code, vendored directories, anything matching your existing lint ignores.
CHANGED=$(git diff --name-only --diff-filter=ACMR "origin/$BASE_REF"...HEAD \
| grep -E '\.(ts|py|go)$' \
| grep -v -E '(dist/|vendor/|\.generated\.)' )
[ -z "$CHANGED" ] && { echo "nothing to analyse"; exit 0; }
Then cap the size. A pull request touching 800 files is exactly the one where a model summary is least useful and most expensive. Above a threshold, post a note saying the diff was too large and exit cleanly.
Cache by content hash
CI re-runs the same work constantly: a rerun after a flaky test, a rebase that changes nothing semantically, a merge queue evaluating the same tree twice.
Key a cache on the hash of the exact prompt input. If the hash matches, reuse the stored output and skip the call entirely. For file-level analysis, hash per file so a one-file change does not invalidate everything.
This is also what makes the step feel fast. A cached run finishes in seconds, and steps that finish in seconds do not get disabled by frustrated developers.
Make the step incapable of blocking
An advisory step should be structurally unable to fail the build. In GitHub Actions that means continue-on-error: true on the step and a short, explicit timeout on the job:
jobs:
analyse:
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- id: llm
continue-on-error: true
run: python scripts/analyse.py > analysis.md
Set a client-side timeout too. The OpenAI Python SDK defaults to a ten-minute request timeout with two automatic retries, which means a single hung call can hold a runner for a long time and cost you money for nothing. Thirty to sixty seconds is plenty for a CI analysis call.
Watch concurrency as well. A merge queue or a monorepo fan-out can launch dozens of jobs at once, each hitting the same quota. Use a workflow concurrency group to serialise, and cancel superseded runs on the same branch so an abandoned push is not still burning tokens.
Pin the model and record what ran
A floating alias can change underneath you, and when it does, your CI output changes with no commit to blame. Pin an explicit model identifier in the workflow so upgrades are a reviewable pull request.
Write the model name, the prompt version and the token counts into the step output or a job summary. When someone asks why the review comments got worse last Tuesday, that log is the answer. It is also the only way to attribute spend to a specific workflow.
Track the cost per merged pull request
The metric that matters is not tokens per month, it is spend per merged pull request. That number tells you whether the step is proportionate, and it makes the trade-off legible when someone proposes adding a second one.
Sum token usage per workflow run, tag it with the repository, and put it somewhere visible. Automation whose cost nobody can see is automation that gets switched off in a panic during the next budget review rather than tuned.
A checklist before merging the workflow
- Advisory by default; only deterministic checks gate merges.
- No secrets in any workflow that checks out fork code.
- Explicit
permissionsblock, narrowest scope that works. - Model output treated as untrusted; never interpolated into shell.
- Changed files only, with a hard size cap.
- Content-hash cache so reruns are free.
continue-on-error, job timeout, and a short client timeout.- Concurrency group and cancellation of superseded runs.
- Pinned model, logged prompt version, logged token usage.
Common questions
Should an LLM check ever fail a build?
Rarely. Model output is advisory. If you want a gate, make it a deterministic check downstream of the generation — for example, that generated tests compile and pass.
How do I run an LLM step on pull requests from forks?
Split it in two. A no-secrets workflow analyses fork code and uploads an artifact; a separate privileged workflow that never checks out fork code posts the result.
How do I stop CI model calls from getting expensive?
Scope to changed files, cap diff size, cache on a content hash, run on pull requests rather than every push, and cancel superseded runs on the same branch.