LLMs for Log Analysis Without Sending Them Your Logs
Guides

LLMs for Log Analysis Without Sending Them Your Logs

You cannot fit a day of logs in a context window and should not try. Template mining, diffing and sampling first — then a model on the reduced set where it earns its keep.

The instinct during an incident is to paste logs into a model and ask what went wrong. It works on a hundred lines. It does not work on the four million lines your service emitted during the outage, and the naive fix — a bigger context window — is the expensive way to get a worse answer.

Logs are extraordinarily repetitive. The same twenty message shapes account for almost all the volume, with different timestamps and identifiers filled in. Compressing that redundancy with deterministic tooling before any model sees it is what makes the whole approach viable.

Reduce first, reason second

Log template mining is the standard technique. It separates the fixed skeleton of a message from its variable parts, so thousands of lines collapse into one template with a count.

The Drain algorithm is the well-known approach here: it clusters log lines online using a fixed-depth parse tree, which avoids building a deep unbalanced tree and keeps overhead low enough to run on a stream. Drain3 is the maintained production fork, adding configurable masking of variables and state persistence — it is stateful, so a restart needs the saved parse tree to avoid re-learning from scratch.

The output is what you send onward:

  482,193  GET /api/orders <*> 200 <*>ms
   31,004  cache miss for key <*>
    8,772  upstream timeout after <*>ms host=<*>
      219  failed to acquire connection from pool
        4  panic: nil map write in billing.Apply

Five lines instead of half a million, and the interesting one is already visible. This reduction is not a preprocessing detail — it is most of the value, and it works without any model at all.

Diff against a known-good window

The single most useful signal during an incident is not what the logs say, it is what changed. Mine templates over the incident window and over the same window from a healthy period, then compare the two distributions.

Templates that are new, or whose rate jumped by an order of magnitude, are the shortlist. Templates that vanished matter too — a heartbeat that stopped is often more diagnostic than an error that started.

Everything to this point is deterministic, cheap, reproducible and auditable. Reach for a model only for the step after: taking that shortlist plus the deploy history and describing what it plausibly means.

What a model is genuinely good at here

Three tasks, all downstream of reduction.

Narrating a sequence. Given fifty deduplicated events with timestamps, producing a readable timeline of what happened in what order. This is summarisation, which is the thing these models do best, and it turns a jumble into something a human on-call can act on.

Translating unfamiliar errors. A stack trace from a dependency nobody on the team has read, or a driver error code from a database somebody else operates. Explaining what a message means and what usually causes it is well-covered ground.

Proposing hypotheses to check. Not a root cause — a ranked list of things to look at, each with the specific query or command that would confirm or eliminate it. Framed that way, a wrong suggestion costs one command instead of an hour.

What it is bad at, and why it matters

Root cause attribution is the dangerous one. A model asked "what caused this" will always produce an answer, phrased with the same confidence whether it is derived from evidence or from what usually causes similar-looking symptoms elsewhere.

During an incident, a confident wrong hypothesis is worse than no hypothesis, because it sends the responder down a path and anchors everyone in the channel. Require every claim to cite the specific log lines supporting it, and treat an uncited claim as noise.

It is also bad at anything requiring exact counts, precise timing arithmetic, or correlation across many events. Those are aggregation queries. Run them in your log platform and give the model the answer rather than asking it to compute one.

And keep it out of alerting. Anomaly detection on log volume is a statistics problem with well-understood tooling, and a nondeterministic component in your alerting path produces pages you cannot reproduce or tune.

Structured logs make all of this dramatically better

If your services emit JSON with stable field names, most of the reduction problem solves itself — you can group, filter and count without inferring anything.

{"ts":"2026-08-09T10:14:22Z","level":"error","service":"billing",
 "trace_id":"6f2a...","event":"charge_failed","reason":"upstream_timeout",
 "customer_id":"cus_123","duration_ms":30012}

A stable event field is effectively a hand-written template, which is better than a mined one. A trace_id lets you pull one complete request path across services, which is the highest-value context you can hand to any analysis, human or otherwise.

When you cannot change the log format — a vendor appliance, a legacy service — template mining is the fallback that gets you most of the way there.

Redact before, not after

Logs are full of things that should not leave your infrastructure: email addresses, tokens that were meant to be masked and were not, internal hostnames, customer identifiers, occasionally a full request body.

Template mining helps here as a side effect, because the variable parts are exactly the parts that carry the identifiers, and they get replaced with placeholders. Do not rely on that alone. Run explicit pattern-based redaction for the categories you know about — bearer tokens, keys with recognisable prefixes, emails, card-shaped numbers — before anything crosses a network boundary.

Know the retention policy of whoever processes the data, and be able to state it. "We send our production logs to a third party" is a sentence that needs an owner and a documented answer, and finding that out during a security review is late.

Watch the cost shape

Log analysis has an unusual cost profile: near zero most of the time, then a spike exactly when things are broken and someone is running analysis repeatedly under pressure. That is the worst time to hit a rate limit or a spend cap.

Constrain it structurally. Cap the number of events sent per invocation. Cache the reduction so a second question about the same window does not reprocess anything. Set a per-incident ceiling and make exceeding it a deliberate action rather than an accident.

Predictable pricing helps here more than a low unit price, since the whole difficulty is that usage is bursty and unplanned — that is the scenario flat-rate access is well suited to, and it is worth comparing against your actual incident frequency rather than assuming.

A pipeline that works

  1. Ingest and normalise; prefer structured logs with a stable event field.
  2. Mine templates with Drain3 or equivalent; persist the parse tree state.
  3. Redact known-sensitive patterns explicitly.
  4. Diff the incident window against a healthy baseline window.
  5. Take the top new or spiking templates, plus a few raw examples of each.
  6. Add deploy and configuration change history for the same window.
  7. Ask for a timeline and a ranked list of hypotheses with checks, each citing specific lines.
  8. Verify with deterministic queries before acting.

Step four is the one people skip and the one that does the most work. A model reasoning over the delta between broken and healthy is doing a bounded comparison task. A model reasoning over an undifferentiated pile of logs is guessing, and it will sound exactly as sure of itself either way.

Common questions

Can I just paste logs into a model with a large context window?

For a small window, yes. At production volume the cost and the noise both scale badly. Mine templates first so you send a few hundred deduplicated events instead of millions of lines.

Should an LLM decide when to page someone?

No. Alerting needs to be deterministic and tunable. Use statistical anomaly detection for triggering, and use a model afterwards to summarise and suggest what to check.

How do I keep sensitive data out of the analysis?

Redact by pattern before anything leaves your network. Template mining masks most identifiers as a side effect, but add explicit rules for tokens, emails and customer IDs.

Similar articles

Building an Incident Assistant On-Call Engineers Trust
Guides
Guides·10 min read

Building an Incident Assistant On-Call Engineers Trust

Automate the first ten minutes of context gathering, not the diagnosis. Read-only tools, evidence-linked output, a latency budget, and why a wrong answer at 3am is expensive.

Read
curl Recipes for LLM APIs: Debug Before You Write Code
Guides
Guides·9 min read

curl Recipes for LLM APIs: Debug Before You Write Code

A working set of curl commands for OpenAI-compatible endpoints: streaming, timing, tool calls, error bodies, and building JSON safely with jq.

Read
Debugging LLM API Errors, Status Code by Status Code
Guides
Guides·9 min read

Debugging LLM API Errors, Status Code by Status Code

A field guide to the errors an LLM API actually returns: what each status means, which ones are worth retrying, and how to reproduce the failure in one curl command.

Read