LLM API Error Codes: A Practical Reference
What each status code from an LLM API actually means, which are safe to retry, and how to handle the ones that look transient but are not. With real header names.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
What each status code from an LLM API actually means, which are safe to retry, and how to handle the ones that look transient but are not. With real header names.
ReadLLM latency is two different problems wearing one name. Understanding prefill and decode tells you which knob to turn when a request feels slow.
ReadYou cannot fit a day of logs in a context window and should not try. Template mining, diffing and sampling first — then a model on the reduced set where it earns its keep.
ReadLocal inference is free at the margin and expensive everywhere else. Work through the memory arithmetic, the real break-even, and where each option genuinely wins.
ReadAn agent running for hours hits limits a chat never does: context exhaustion, process restarts, stale state. The patterns that keep long jobs alive.
ReadThe Model Context Protocol standardises how agents connect to tools and data. What it solves, how it works, and when a plain function call is enough.
ReadToken counts are not returns. Here is how to convert engineer hours into money, pick a metric that survives scrutiny, and avoid the usual measurement traps.
ReadSwapping providers is two lines of config. Keeping quality, cost accounting and error handling intact is the actual work. A migration plan that survives contact with users.
ReadBoth land in the mid-forties on the Artificial Analysis index. They are not interchangeable, and the tie is a good lesson in why aggregate scores mislead.
ReadM3 is the smallest of the frontier-class open models and the only one that takes video natively. A look at sparse attention, the licence, and where it fits.
ReadMoE models decouple parameter count from compute per token, which is why a trillion-parameter model can be cheap to serve. Here is the mechanism and what it costs you.
ReadHow a small model inherits the behaviour of a much larger one, what gets lost along the way, and why distillation is the reason cheap models got good so fast.
ReadShowing 349–360 of 404 articles