Debugging LLM API Errors, Status Code by Status Code
A field guide to the errors an LLM API actually returns: what each status means, which ones are worth retrying, and how to reproduce the failure in one curl command.
Most LLM API failures are boring. A key is wrong, a message array is malformed, a context window overflowed. The reason they feel mysterious is that the useful detail is buried in a JSON body most client libraries throw away before you see it.
So the first rule of debugging these is: log the response body, not just the exception class. Everything below assumes you can see it.
The error envelope
OpenAI-shaped APIs return a consistent structure, and most compatible providers copy it:
{
"error": {
"message": "This model does not support the temperature parameter.",
"type": "invalid_request_error",
"param": "temperature",
"code": "unsupported_parameter"
}
}
The param field is the one people miss. When a 400 tells you which field it objected to, you are done — no bisecting required. Surface it in your logs and half your debugging time disappears.
Branch on HTTP status first and treat code as a refinement. Status codes are stable across providers; the string values of type and code are not.
4xx: your request, your fix
400 Bad Request. The body is malformed or a parameter is invalid for this model. In practice the recurring causes are a messages array in an order the API rejects, a tool_call_id in a tool result that does not match any preceding call, a content block the provider does not accept, or a parameter that exists on one model and not another.
401 Unauthorized. The key is missing, revoked, or belongs to a different organisation or project than the one the request targets. Check for whitespace and quote characters first — a key pasted from a chat window with a trailing newline produces exactly this.
403 Forbidden. Authentication worked, the request is not permitted. OpenAI documents unsupported country or region for this case, and also returns it when the calling IP does not match a configured project or organisation allowlist. If it started happening after an infrastructure change, suspect the allowlist.
404 Not Found. Two very different failures wearing the same number. Either the path is wrong — usually a doubled /v1, since most SDKs append it themselves — or the model name is not one this provider serves. Hit the models endpoint to tell them apart.
422 Unprocessable Entity. Less common, but some gateways use it for schema violations that OpenAI would return as 400. Read it as a 400.
429 is three different problems
A 429 can mean you exceeded requests per minute, exceeded tokens per minute, or ran out of prepaid quota. Only the first two are worth retrying, and the third will retry forever without ever succeeding.
Distinguish them by the error code. A quota exhaustion carries insufficient_quota, and no amount of backoff fixes a billing problem. Rate limiting carries a rate-limit code and typically a Retry-After header telling you the minimum seconds to wait.
OpenAI enforces several limits simultaneously — requests per minute and per day, tokens per minute and per day — and whichever you hit first triggers the 429. The response headers tell you which: x-ratelimit-remaining-requests and x-ratelimit-remaining-tokens are separate counters. If requests remaining is healthy and tokens remaining is zero, sending smaller prompts helps and sending fewer requests does not.
5xx: usually theirs, sometimes yours
500 and 502 mean something broke upstream. Retry with backoff; these normally clear in minutes.
503 means overload. Retry with longer delays than you would use for a 500, because hammering a capacity-constrained service is how a brownout becomes an outage. Anthropic uses 529 for the same idea, so if you handle multiple providers, treat 529 as a retryable overload rather than an unknown status.
The trap is that a 500 can be caused by your input. A request that reliably 500s on every attempt with the same body is not a transient fault — it is an input the provider fails to handle, and you will find it faster by bisecting the prompt than by retrying.
Failures that return 200
These are the expensive ones, because no exception is raised anywhere.
Truncated output. Check finish_reason on every response. A value of length means the model hit your token cap mid-sentence and the answer you are about to parse is incomplete. This is the single most common cause of "the JSON is invalid sometimes".
Empty content with a tool call. When the model decides to call a tool, content is often null and the payload is in tool_calls. Code that reads content unconditionally gets a null reference and blames the API.
Mid-stream failure. A streamed response returns 200 and starts sending, then dies. Your client has already rendered partial output. Handle the stream ending without a finish reason as an error path, not a success.
Content filtering. Some providers return a filtered or refused response as a normal 200 with a different finish reason. Log the finish reason so you can tell refusal apart from a bad answer.
Reproduce it in curl before you theorise
Frameworks add retries, middleware, prompt templating and message rewriting. Half the time the bug is in that layer, not the API. Strip it out:
curl -sS -i "$BASE_URL/chat/completions" \
-H "Authorization: Bearer $KEY" \
-H "Content-Type: application/json" \
-d '{"model":"MODEL","messages":[{"role":"user","content":"hi"}]}'
The -i flag prints response headers, which is where the rate-limit counters and Retry-After live. If curl succeeds and your app fails, the API is fine and the problem is between your code and the wire — usually a proxy, an SDK default, or a prompt template that produced something you did not intend.
To see exactly what your SDK sent, enable its debug logging rather than reconstructing the request by hand. The Python and Node OpenAI SDKs both expose request logging through an environment variable, and reading the real serialised body settles most arguments in seconds.
Context length errors deserve their own handling
A context_length_exceeded error is a 400, but retrying is not the fix and neither is a bigger model, necessarily. It means the sum of input tokens plus requested output tokens exceeded the window.
Two useful reflexes. First, check whether your requested output cap is the thing pushing you over — reserving 32,000 output tokens on a long conversation fails even when the input alone would fit. Second, count tokens before sending on any path that accumulates history, and truncate or summarise deliberately rather than discovering the limit in production.
A triage order that works
- Log the full error body, including
param. - Reproduce with curl and
-i. - Classify: 4xx means fix the request, 429 means check which limit, 5xx means retry with backoff.
- On a 200, check
finish_reasonbefore trusting the content. - If it only fails at scale, look at concurrency and retry amplification before blaming the provider.
The habit worth building is treating every LLM call as an untrusted network call with a structured failure mode, rather than as a function that returns a string. Almost every hard-to-debug incident in this space comes from code that assumed the second thing.
Common questions
Should I retry a 429 automatically?
Only when it is a rate limit. Check the error code first — a quota or billing exhaustion also returns 429 and will never succeed on retry, so it should surface as an alert instead.
Why does the same request work in curl but fail in my app?
Something between your code and the wire is changing the request. Enable SDK debug logging to see the serialised body, and check for middleware, proxies or prompt templates that rewrite messages.
What causes intermittent JSON parse failures on model output?
Usually truncation. Check finish_reason on the response — a value of length means the output was cut off at your token cap, so the JSON was never complete.