LLM API Error Codes: A Practical Reference
Guides

LLM API Error Codes: A Practical Reference

What each status code from an LLM API actually means, which are safe to retry, and how to handle the ones that look transient but are not. With real header names.

LLM APIs mostly follow standard HTTP semantics, with a handful of quirks that cause real production incidents: a 429 that means two completely different things, a non-standard 5xx that some HTTP clients mishandle, and errors that arrive after a 200 because the response was streamed.

This is a reference for what each code means, what to do about it, and the specific things that catch people out. Exact error type strings vary by provider, so the codes below are grouped by behaviour rather than by vendor.

4xx: your request, your problem

400 — invalid request. Malformed JSON, an unknown parameter, a schema violation, a message array in an invalid order, or content rejected by policy. Never retry a 400 unchanged; the same request will fail identically.

The most common causes in practice are subtle rather than obvious. Sending a parameter a specific model does not support is a 400, and models within the same family often differ — Anthropic documents 400s for prefilling an assistant message on newer models, and for sending a thinking configuration a given model does not accept. If a request works on one model and 400s on another, check parameter support before checking your JSON.

401 — authentication. The key is missing, malformed, revoked or expired. Check for whitespace and newlines picked up from an environment file, and check you are sending the header the provider expects. Not retryable.

402 — billing or credits. Not every provider uses this, but some do. Anthropic documents 402 as a billing error; OpenRouter returns 402 when an account or key has insufficient credits. Where present it is unambiguous and adding money is the only fix.

403 — permission. The key is valid but not entitled: no access to that model, wrong workspace or organisation, an IP allowlist mismatch, or an unsupported country or region. Not retryable, and not fixable in code.

404 — not found. Usually a wrong model identifier or a wrong endpoint path. On OpenAI-compatible proxies this is also what you get when the base URL is missing or duplicating a version segment. Check whether your client appends a path suffix to the base URL you configured.

409 — conflict. A resource was modified concurrently or a unique value is already in use. Resolve the conflict, then retry.

413 — request too large. The payload exceeded a byte limit, which is separate from and stricter than the token limit. Anthropic publishes per-endpoint maximums — 32 MB for the Messages and token counting endpoints, 256 MB for batches, 500 MB for files. On the direct API this is returned by the edge before the request reaches the model, so a 413 tells you nothing about your token count. Base64 images and PDFs are the usual culprits.

422 — unprocessable. Well-formed but semantically rejected. Some SDKs surface this as a distinct exception type and suggest retrying; treat it as a request problem until proven otherwise.

429: the code that means two things

This is the one worth internalising, because the two cases require opposite responses.

  • Rate limited. You exceeded requests or tokens per minute. Transient. Back off and retry — this resolves on its own.
  • Out of money or over a cap. Credit balance exhausted, a project or organisation spend limit reached, or an account usage limit. Backing off does nothing. OpenAI distinguishes these with specific error codes including credit balance exhausted, organization spend limit exceeded and project spend limit exceeded.

A retry loop that treats every 429 as transient will hammer a quota error until it gives up, wasting minutes and producing an unhelpful log. Branch on the error code inside the body, not on the status alone.

There is a third, subtler case: acceleration limits. Anthropic notes that a sharp increase in an organisation traffic can produce 429s even below the nominal ceiling, and advises ramping traffic gradually. If you 429 immediately after deploying a scale-up, this is likely why.

Use the headers

Both major providers expose rate limit state on every response, successful ones included, which means you can monitor headroom instead of discovering the limit.

  • OpenAI-style: x-ratelimit-limit-requests, x-ratelimit-remaining-requests, x-ratelimit-reset-requests, and the token equivalents.
  • Anthropic-style: anthropic-ratelimit-requests-limit, -remaining and -reset, with separate series for input tokens and output tokens.
  • retry-after on a 429, indicating seconds to wait. Honour it rather than guessing.

Because Anthropic exposes requests, input tokens and output tokens as separate dimensions, the headers also tell you which limit you tripped, which determines which reset timestamp applies.

5xx: their problem, usually temporary

500 — internal error. Retry with exponential backoff. If it persists, capture the request ID and contact support.

503 — service unavailable. Overloaded. OpenAI also documents a slow-down condition triggered by a sudden traffic spike, where the advice is to drop back to your previous rate, hold it for around fifteen minutes, then increase gradually.

504 — timeout. The request took too long to process. The fix is architectural rather than a retry: stream the response, or use a batch endpoint for long jobs.

529 — overloaded. Anthropic uses this for temporary capacity exhaustion across all users. It is non-standard, so some HTTP libraries and proxies do not classify it as retryable by default. Check that your client handles it, because it is one of the more common transient failures in practice.

The errors that arrive after a 200

When you stream, the HTTP status is committed the moment headers are sent. An error occurring mid-generation cannot change it, so it arrives as an error event inside the event stream instead.

Code that only checks status codes will treat a truncated, failed stream as a successful empty response. Handle error events in the stream explicitly, and treat a stream that ends without its terminal event as a failure rather than a completion.

A retry policy that works

Never retry:  400, 401, 403, 404, 413, 422
Fix, retry:   409 (resolve conflict first)
Retry always: 500, 502, 503, 504, 529
Retry if:     429 AND error code is rate-limit,
              not quota or spend-limit

Backoff: exponential with full jitter, cap ~60s,
         3-5 attempts, honour retry-after when present

Two additions that pay for themselves. Log the provider request ID on every failure — both major providers return one, in a response header and in the error body, and support cannot help without it. And distinguish "retried and succeeded" from "succeeded first time" in your metrics, because a rising retry rate is the earliest warning of a capacity problem.

Most official SDKs already retry transient failures with backoff — commonly twice by default, honouring retry-after — so check what your client does before building your own layer on top and accidentally multiplying the attempts.

Common questions

Should I always retry a 429 from an LLM API?

No. A 429 can mean a per-minute rate limit, which is transient and worth retrying, or an exhausted credit balance or spend cap, which backoff will never resolve. Branch on the error code in the response body, not on the status alone.

What is a 529 error and why does my HTTP client ignore it?

It is Anthropic non-standard status for a temporarily overloaded API. Because it is outside the standard range, some HTTP libraries and proxies do not classify it as retryable by default, so it needs explicit handling in your retry policy.

Why did my streaming request fail after returning a 200?

The status is committed when headers are sent, so any error during generation arrives as an error event inside the stream. Handle those events explicitly, and treat a stream that ends without its terminal event as a failure.

Similar articles

Debugging LLM API Errors, Status Code by Status Code
Guides
Guides·9 min read

Debugging LLM API Errors, Status Code by Status Code

A field guide to the errors an LLM API actually returns: what each status means, which ones are worth retrying, and how to reproduce the failure in one curl command.

Read
Building a Chatbot From Scratch: The Parts Nobody Mentions
Guides
Guides·9 min read

Building a Chatbot From Scratch: The Parts Nobody Mentions

The model call is twenty lines. The other ninety percent is conversation state, idempotency, abuse limits and knowing when a reply was wrong. A build order that works.

Read
Choosing an AI Provider: The Questions to Ask
Guides
Guides·9 min read

Choosing an AI Provider: The Questions to Ask

A checklist of the questions that separate providers you can plan around from ones you cannot, covering model transparency, limits, compatibility and exit terms.

Read