Context Bloat: Paying Repeatedly for Tokens Nobody Reads
Every token added to a prompt is billed on that call and every call after it. Where bloat accumulates and what it actually costs to leave it there.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
Every token added to a prompt is billed on that call and every call after it. Where bloat accumulates and what it actually costs to leave it there.
ReadAgent transcripts grow until quality degrades and cost climbs. What to summarise, what to drop, and what must never be compacted away.
ReadA model advertising a one-million-token window does not reliably use one million tokens. The gap between the spec sheet and what actually works.
ReadFilling a million-token window costs between fourteen cents and three dollars depending on the model. The full comparison, including output ceilings.
ReadContinuous batching lets finished requests leave a batch and new ones join mid-flight. It is why modern inference servers hold high load without stalling.
ReadLines of code is the easiest denominator for AI spend and one of the worst. Where it misleads, where it genuinely works, and what to measure instead.
ReadAgent costs are driven by resent transcript, not generated output. Working out what one run actually costs and which lever moves it.
ReadAutomated review looks cheap per pull request until you count re-reviews, large diffs and false positives. Working out the real per-review figure.
ReadRefactors are the worst case for context bloat because every touched file must stay in view. A method for estimating the bill before you start the run.
ReadGenerating tests with a model is cheap and bounded. Making a red suite green is neither. How the two costs differ and how to put a ceiling on the expensive one.
ReadCapability leaderboards ignore price entirely. How to build a score that reflects what a model costs to run on your workload, and where it misleads.
ReadA critic agent reviews work another agent produced. The patterns that improve output, the ones that just add cost, and how to tell them apart.
ReadShowing 85–96 of 404 articles