What AI Actually Costs Per Developer Per Month
There is no single number, but there is a reliable way to derive yours. Work through the token arithmetic behind light, typical and heavy agentic usage.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
There is no single number, but there is a reliable way to derive yours. Work through the token arithmetic behind light, typical and heavy agentic usage.
ReadParameter counts are the most quoted and least understood model spec. What they measure, why total and active differ, and when the number predicts anything useful.
ReadFan-out is the dominant cost term once you run agents in parallel. Here is the arithmetic for coordinator and subagent spend, plus what actually reduces it.
ReadMost production LLM traffic does not need a frontier model. A practical framework for deciding which tasks can drop a tier without anyone noticing.
ReadMost tasks handed to agents are better served by a fixed workflow or a single model call. A decision rule, the failure modes, and what to build instead.
ReadAgent spend follows a heavy-tailed distribution, so the average is a poor planning number. Here is how to budget from percentiles and bound the tail instead.
ReadHallucination is not a bug that will be patched out. It follows from the training objective and from how we grade models. Here is the mechanism and the mitigations that work.
ReadTemperature zero is not determinism, and a seed is only best effort. Here is where LLM nondeterminism really comes from and how to build around it.
ReadShowing 397–404 of 404 articles