Migrating Off the OpenAI API Without Breaking Production
Swapping providers is two lines of config. Keeping quality, cost accounting and error handling intact is the actual work. A migration plan that survives contact with users.
The wiring part of a provider migration takes about ninety seconds. You change a base URL, change a key, change a model name, and requests start flowing somewhere else. That is not the risky part.
The risky part is that your prompts were tuned — deliberately or by accident — against one specific model, and nothing in your test suite will tell you when that tuning stops holding. Output gets slightly worse in ways that only show up as support tickets three weeks later.
Separate the two migrations
There are two independent changes hiding inside "we are moving off OpenAI", and doing both at once makes failures impossible to attribute.
The first is transport: a different host, different auth, different rate limits, different failure modes. The second is model behaviour: a different set of weights producing different text for the same prompt.
If you can, move transport first while keeping a model you already know, verify nothing broke, then change the model as a separate deploy. If the new provider does not serve anything you currently use, you cannot split it — in which case budget more time for evaluation, not less.
Build a golden set before you touch anything
Capture 50 to 200 real requests from production, with their current responses. Not synthetic examples — actual traffic, including the weird ones. Strip anything sensitive, store them as fixtures, and treat that file as the contract.
You do not need a scoring model to make this useful. Even a diff view where a human skims old versus new output catches the big regressions: a format that changed, a refusal that appeared, a response that got three times longer.
Where you can write a deterministic check, write one. Does the JSON parse. Does the SQL compile. Does the classification land in the allowed enum. Did the tool call name match. Those assertions are worth more than a vague quality score, because they fail loudly in CI.
The parameters that quietly differ
Request bodies are broadly portable, but a handful of fields have drifted.
max_tokensis deprecated on OpenAI chat completions in favour ofmax_completion_tokens, and the o-series reasoning models rejectmax_tokensoutright. Other providers still only acceptmax_tokens. Send whichever the target expects rather than assuming.seedis best-effort everywhere and ignored entirely by some backends. If your tests depend on reproducibility, verify it rather than trusting it.response_formatand strict schema enforcement vary a lot in how strictly they are honoured.- Unknown parameters are usually ignored rather than rejected, so a field the new provider does not support fails silently. Your carefully chosen
top_pmay simply not be applied.
The practical defence is to assert on behaviour, not on the request. Change a parameter, confirm the output actually changes, and only then believe it took effect.
Error handling needs remapping too
OpenAI returns a consistent error envelope with message, type, param and code fields, and a lot of production code branches on those strings — insufficient_quota, invalid_api_key, context_length_exceeded.
{
"error": {
"message": "...",
"type": "insufficient_quota",
"param": null,
"code": "insufficient_quota"
}
}
A new provider will use the same shape and different values, or the same values with different triggers. Grep your codebase for any string comparison against an error code before you migrate, and rewrite those branches to key off HTTP status first and the code only as a refinement.
Rate limit headers are the same story. OpenAI returns x-ratelimit-limit-requests, x-ratelimit-remaining-tokens, x-ratelimit-reset-requests and friends, plus Retry-After on eligible errors. Any dashboard or throttle you built on those header names needs checking against the new provider, which may return none of them.
Run both in shadow before you switch
The cheapest way to de-risk is to send a percentage of real traffic to both providers, serve the incumbent response to the user, and log the challenger for comparison.
import asyncio
from openai import AsyncOpenAI
incumbent = AsyncOpenAI(base_url="https://api.openai.com/v1", api_key=OLD_KEY)
challenger = AsyncOpenAI(base_url=NEW_BASE_URL, api_key=NEW_KEY)
async def handle(messages):
primary = incumbent.chat.completions.create(
model="MODEL_A", messages=messages
)
shadow = challenger.chat.completions.create(
model="MODEL_B", messages=messages
)
result, other = await asyncio.gather(primary, shadow, return_exceptions=True)
log_comparison(messages, result, other)
return result
Shadow traffic doubles your inference bill for the duration, so cap it — one percent of requests for a week usually surfaces the format regressions and the latency profile. Sample deliberately across your prompt types rather than taking the first N requests, which will over-represent whatever runs on a cron.
Timeouts and retries are not portable
The OpenAI Python SDK defaults to two retries and a ten-minute request timeout, retrying connection errors, 408, 429 and 5xx with exponential backoff. If you kept those defaults, you inherited a retry policy you did not choose, and it follows you to the new provider because it lives in the client, not the server.
Set both explicitly during a migration so that a slow new provider does not silently turn into tripled load:
client = OpenAI(base_url=NEW_BASE_URL, api_key=KEY, max_retries=0, timeout=30.0)
Then implement retries in one place you control, so you can see them in metrics. Two layers of automatic retry — SDK plus your own wrapper — multiply rather than add, and that is how a brief upstream blip becomes a self-inflicted outage.
Keep the accounting honest
Verify that the new provider reports a usage block, that it reports one while streaming if you rely on that, and that the counts are real upstream numbers rather than an estimate. Log token counts per request from day one of the shadow period, because "is this cheaper" is a question you want answered with your own traffic mix rather than a price-per-million comparison.
Pricing model matters here as much as unit price. If your workload is agentic — long contexts, many turns, frequent retries — per-token billing gets volatile in a way that flat-rate access does not, which is the case we built CozyAPI around. Either way, measure before you commit.
A migration checklist
- Freeze a golden set of real requests and current outputs.
- Write deterministic assertions for anything machine-checkable.
- Move transport and model in separate deploys where possible.
- Grep for hardcoded error codes and rate-limit header names.
- Set explicit timeouts and disable SDK-level retries.
- Shadow one percent of traffic for a week; diff the outputs.
- Confirm usage reporting matches your billing expectations.
- Keep the old credentials live and the switch behind a flag until you have a quiet week.
The rollback plan is the part people skip. If reverting means a code change and a deploy, you will hesitate at exactly the moment you should not. Put the provider choice in configuration and make the reversal a one-line change.
Common questions
How long should a provider migration take?
The code change is an afternoon. Plan a week of shadow traffic on top of it so you see the full range of production prompts, including whatever only runs on a weekly schedule.
Do I need to rewrite my prompts for a different model?
Sometimes. Prompts that rely on a specific formatting habit or refusal behaviour are the ones that break. Run your golden set first, then rewrite only what actually regressed.
Can I keep using the official OpenAI SDK?
Yes, for any provider that serves the chat completions shape. Set base_url and the key. Do check the SDK retry and timeout defaults, because they follow you to the new host.