Agent Observability: Tracing a Loop You Cannot Reproduce
Agent failures are rarely reproducible, so logs are not enough. What to record per step, how to span a tool loop, and which metrics predict a bad run.
A failing web request can be replayed. A failing agent run frequently cannot. Sampling is stochastic, the tools hit a world that has since changed, and the twelfth step depended on the exact wording of the fourth. If you did not record it while it happened, it is gone.
That single property is why agent observability is a different discipline from ordinary application logging, and why teams that instrument late end up rewriting their agent to find out why it misbehaved.
The unit of debugging is the run, not the call
Most LLM logging captures request and response pairs. That tells you what the model said, and almost nothing about why it said it, because the interesting causality is in the accumulated state.
The thing you actually need to reconstruct is: at step N, what was in the context, which tools were available, what did the model choose, what came back, and how did that change step N plus one. Individual call logs cannot answer that. A trace can.
So model a run as a tree. One root span for the task, child spans for each model call, sibling spans for each tool invocation, nested spans for subagents. The OpenTelemetry GenAI semantic conventions settled on this shape: an invoke_agent span with chat spans for model calls and execute_tool spans for tool invocations underneath.
Worth knowing before you standardise on it: those conventions are still marked as in development, and the gen_ai.* attribute names carry development stability. Adopt the span shape, expect attribute churn, and keep your own mapping layer thin enough to update.
What to record on every step
Be greedy here. Storage is cheaper than a reproduction attempt.
- The full tool call. Name and arguments, exactly as emitted. Truncated arguments hide the bug roughly half the time.
- The full tool result, or a deterministic reference to it. "Returned 4KB" is not a result.
- Token counts split into input, output, cached read and cached write. Cost attribution and context-growth debugging both depend on this.
- The model identifier actually served, not the alias you requested. Routing and fallbacks make these diverge.
- Finish reason. Stop, length, tool call, refusal, filter. Silent truncation looks identical to a bad answer unless you record this.
- Step index and parent span. So you can plot how the run branched.
- Latency split between time to first token and total generation, since they have different causes.
Prompt and completion content is the awkward one. It is the most useful field for debugging and the one most likely to contain customer data, which is why the conventions treat content capture as opt-in. Decide deliberately: capture in development, sample or redact in production, and never let it default on without a retention policy.
Metrics that predict a bad run
Traces tell you what happened once. Metrics tell you which runs to look at. A handful of aggregates catch most trouble:
Steps per task. The distribution matters more than the mean. A long tail is where loops live. Alert on the ninety-fifth percentile, not the average.
Tool error rate by tool. One tool failing 30% of the time will silently double your token spend as the model retries around it.
Repeated identical calls. The same tool with the same arguments twice in one run is a strong signal the model is stuck. It is trivial to detect and shockingly common.
Context size at each step. Plot it. If it grows monotonically to the ceiling, your compaction strategy is not working.
Cost per resolved task. Not cost per call. A cheaper model that needs three times the steps is not cheaper, and only this metric shows it.
Termination reason. Completed, step cap hit, error, human abort. A rising step-cap rate is the clearest early warning that quality is drifting.
Instrument the decision, not just the outcome
The most valuable annotation is usually not automatic. When your code makes a choice on the agent behalf — pruning context, selecting a model, rejecting a tool call, injecting a retry — record that as an event on the span.
Otherwise you get a trace where step 7 mysteriously lacks information that step 6 clearly had, and no indication that your own summariser removed it. Half of the confusing agent behaviour I have seen traced back to the harness, not the model.
Sampling without losing the failures
Full-fidelity traces on every run get expensive once volume is real. Head sampling — decide at the start — is the wrong tool here, because you cannot know at step one that step fourteen will be interesting.
Tail sampling fits agents much better: buffer the run, then keep it if it failed, exceeded a step or cost threshold, hit the step cap, took an unusual path, or was flagged by a user. Keep a small random baseline of healthy runs too, so you have something to compare against when a regression lands.
Traces become your evaluation set
The underrated payoff. Once runs are recorded with inputs, tool results and outcomes, you have a corpus of real tasks with real failure modes. That is a far better regression suite than anything you would write by hand, because it is drawn from what your users actually do.
Tag runs that went wrong, freeze them, and replay them against prompt or model changes. The tool results are already captured, so replay is deterministic even though the original run was not.
A minimum viable setup
- One trace per task, spans for every model and tool call, subagents nested under their parent.
- Record arguments, results, token counts, finish reasons and the served model identifier.
- Emit an event whenever your harness alters context or routing.
- Tail sample: keep failures, outliers and a small healthy baseline.
- Alert on step-count and cost percentiles, not means.
- Promote interesting failed runs into a replayable evaluation set.
You do not need a vendor to start. A structured log line per span, written to somewhere queryable, gets you most of the value on day one — and it is the difference between diagnosing a bad run and guessing at it.
Common questions
Why is logging prompts and responses not enough for agents?
Because the failure is usually in the accumulation, not in any single call. You need the whole run as a tree — which tools ran, what they returned, how context grew — to see why step twelve went wrong. Isolated call logs cannot show that.
Should I use the OpenTelemetry GenAI conventions?
Adopt the span shape, which is stable in practice: an agent span with model and tool spans nested under it. The gen_ai attribute names are still marked as in development, so keep a thin mapping layer you can update rather than hardcoding attributes everywhere.
How do I keep trace volume affordable?
Use tail sampling. Buffer each run and keep it only if it failed, hit a step or cost threshold, took an unusual path, or was flagged. Retain a small random sample of successful runs so you have a comparison baseline.