Measuring Agent Success Rate Without Fooling Yourself
Success rate is the easiest agent metric to report and the easiest to game. Here is how to define it so the number predicts production behaviour.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
Agents, tool calling, and the loop that makes them useful.
Success rate is the easiest agent metric to report and the easiest to game. Here is how to define it so the number predicts production behaviour.
ReadThe orchestrator-worker shape is the one multi-agent design that reliably pays for itself. Here is how to build the lead, the brief and the synthesis step.
ReadRunning tool calls concurrently cuts wall-clock time but not tokens, and it is only safe for some operations. How to decide what to parallelise.
ReadSix workers, four returned, two timed out. Whether to answer, retry or abort is a design decision — here is how to make it before the incident, not during.
ReadPrompt edits regress silently and every score is noisy. How to build a golden set, pick a scorer, handle variance and gate merges without a permanently red build.
ReadOne API key, many agents, one shared quota. How to build a fleet-wide limiter that respects token budgets, protects interactive work and survives bursts.
ReadAgent runs rarely reproduce, so the stored trace is the only evidence. What to capture, how to replay it, and where replay stops being faithful.
ReadAn agent stopped halfway through eleven changes. Checkpoints, compensating actions, staging areas and the undo plan you should require before it acts.
ReadA durable scratchpad survives compaction, crashes and handoffs. What belongs on it, what does not, and how to stop it becoming a second transcript.
ReadAgents that share a scratchpad drift, overwrite and confuse each other. Here are four shared-memory designs, what each one breaks, and how to pick.
ReadAgents spend most of their wall-clock time waiting. Speculative execution starts likely next steps early and discards the wrong guesses. Here is when it pays.
ReadLong sessions have to shed context. What to summarise, what must survive untouched, when to trigger it, and how to tell a good summary from a lossy one.
ReadShowing 37–48 of 74 articles