What "OpenAI-Compatible" Actually Means (And Where It Breaks)
Most AI tools speak one API shape. Understanding what compatibility covers — and the four places it usually leaks — saves hours of debugging a base URL swap.
OpenAI's /v1/chat/completions shape became the de facto standard for talking to language models. Enough tools assumed it that supporting it became mandatory for everyone else, and now "OpenAI-compatible" is a feature bullet on nearly every inference provider.
It is genuinely useful. It is also less complete than the phrase implies.
The part that always works
The core request shape is universal:
POST /v1/chat/completions
Authorization: Bearer YOUR_KEY
{
"model": "some-model",
"messages": [
{"role": "system", "content": "You are concise."},
{"role": "user", "content": "Explain tokenization."}
],
"temperature": 0.7,
"stream": false
}
Roles (system, user, assistant), the messages array, temperature, max_tokens, and the response envelope with choices[0].message.content and a usage block — all of that is reliable across providers.
This is why swapping providers is usually a two-line change:
OPENAI_API_KEY=your_key
OPENAI_BASE_URL=https://api.example.com/v1
Any SDK that reads those variables works unmodified. That is the whole value proposition, and it mostly holds.
The four places it leaks
1. Model names
There is no standard registry. Every provider names things differently, and passing an unknown name usually produces a 404 rather than something helpful. Check the provider's model list endpoint or docs rather than guessing.
2. Streaming details
Server-sent events are widely supported, but the fine print varies: whether a usage block appears in the final chunk, how tool calls are chunked, whether [DONE] is sent. Code that parses streams strictly tends to break across providers.
3. Tool calling
Broadly supported, unevenly implemented. Parallel tool calls, strict schema adherence, and streamed tool-call deltas are the usual gaps. If your agent depends on parallel calls, verify it explicitly rather than assuming.
4. Everything beyond chat
Compatibility almost always means chat completions. Embeddings, image generation, audio, files, assistants and fine-tuning are frequently absent. If your app uses more than chat, check each endpoint separately.
A five-minute compatibility check
Before committing, run these against any new provider:
# 1. Does it list models?
curl $BASE_URL/models -H "Authorization: Bearer $KEY"
# 2. Does a basic completion work?
curl $BASE_URL/chat/completions \
-H "Authorization: Bearer $KEY" \
-H "Content-Type: application/json" \
-d '{"model":"MODEL","messages":[{"role":"user","content":"hi"}]}'
# 3. Does streaming work?
curl -N $BASE_URL/chat/completions \
-H "Authorization: Bearer $KEY" \
-H "Content-Type: application/json" \
-d '{"model":"MODEL","messages":[{"role":"user","content":"count to 5"}],"stream":true}'
Then confirm a tool call round trip, and check whether usage is reported — you cannot track spend without it.
Headers and parameters that vary
Beyond the body, a few things differ quietly between providers. Some accept max_completion_tokens where others still expect max_tokens. Some ignore seed entirely rather than erroring, which means your "deterministic" runs are not. Others silently clamp temperature to a supported range.
Unknown parameters are usually ignored rather than rejected, which is convenient but means a typo in a parameter name fails silently — your carefully tuned top_p may never have been applied. Verify behaviour changes when you change a parameter, rather than assuming it took effect.
What to check on a gateway specifically
Routing gateways add a layer worth probing. Confirm whether streaming is genuinely passed through or buffered and replayed, since a buffered stream defeats the purpose. Check whether usage reflects the real upstream token counts. And confirm which model an alias actually resolves to — a name is a label, not a guarantee, and the only way to be certain is to ask the provider directly and get a straight answer.
Debugging a failed swap
- 401 — key format or header. Some gateways expect their own key, not the upstream provider's.
- 404 on the endpoint — usually a doubled or missing
/v1. Most SDKs append it themselves. - 404 on the model — wrong model name for this provider.
- Empty streamed output — a proxy buffering SSE. Confirm the response is not being collected before it is forwarded.
- Tool calls ignored — the model may not support them, or the harness is dropping the
toolsparameter.
Why this standard is worth caring about
The practical upshot is portability. Because Cline, Claude Code, OpenClaw, Roo, Aider, Continue and most SDKs all speak this shape, your tooling is not locked to whoever you started with. Changing providers costs a base URL and a model name.
That is unusual in infrastructure, and it is worth protecting: prefer providers that keep the standard shape rather than requiring a bespoke SDK, and you keep the ability to leave.
Common questions
Do I need a different SDK for an OpenAI-compatible provider?
No. The official OpenAI SDKs work by setting the base URL and key. That is the point of the standard, and it is the fastest way to test a new provider.
Why do I get a 404 after changing the base URL?
Usually a doubled /v1 — most SDKs append it, so the base should end at /v1 and not repeat it. If the path is right, the model name is probably not valid for that provider.
Does compatibility include embeddings?
Often not. "OpenAI-compatible" almost always means chat completions specifically. Verify any other endpoint your application depends on before migrating.