Detecting Agent Loops Before They Burn Your Budget
A stuck agent repeats the same failing action until something stops it. How to detect repetition cheaply and what to do once you have.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
A stuck agent repeats the same failing action until something stops it. How to detect repetition cheaply and what to do once you have.
ReadStreaming changes when tokens arrive, not how many are billed. Where it does affect spend, and the abandonment case that genuinely wastes money.
ReadDirect preference optimisation drops the reward model and the RL loop. What that simplification buys, what it costs, and why open-weight work leans on it.
ReadWire an OpenAI-compatible endpoint into Emacs — package choices, auth-source instead of hardcoded keys, buffer versus region workflows, streaming and keybindings.
ReadThe share of tasks your cheap model cannot finish decides whether a two-tier setup saves money or quietly doubles it. How to measure and act on it.
ReadPublic benchmarks rank a sample that is not your repository. How to choose tasks, source ground truth and size a set that actually predicts your work.
ReadYou edited a prompt and the output looks better. Here is how to find out whether it actually is, with paired runs, enough samples and judges you can trust.
ReadFive open-weight models now advertise 1M tokens. Advertised context and usable context are different things. How to tell which window actually holds up.
ReadThe router decides which experts see each token, and that one small network shapes quality, throughput and why identical prompts can behave differently.
ReadThe retry algorithm in detail: full versus decorrelated jitter, deadline propagation, idempotency, what to classify as retryable, and how to test it before production does.
ReadA fallback that behaves nothing like your primary turns an outage into a quality incident. How to pick a second model and prove it works.
ReadFew-shot examples decay, contradict each other and leak into output. How to choose them, order them, keep them in sync with your schema and know when to drop them.
ReadShowing 109–120 of 404 articles