Migrating Off the OpenAI API Without Breaking Production
Guides

Migrating Off the OpenAI API Without Breaking Production

Swapping providers is two lines of config. Keeping quality, cost accounting and error handling intact is the actual work. A migration plan that survives contact with users.

The wiring part of a provider migration takes about ninety seconds. You change a base URL, change a key, change a model name, and requests start flowing somewhere else. That is not the risky part.

The risky part is that your prompts were tuned — deliberately or by accident — against one specific model, and nothing in your test suite will tell you when that tuning stops holding. Output gets slightly worse in ways that only show up as support tickets three weeks later.

Separate the two migrations

There are two independent changes hiding inside "we are moving off OpenAI", and doing both at once makes failures impossible to attribute.

The first is transport: a different host, different auth, different rate limits, different failure modes. The second is model behaviour: a different set of weights producing different text for the same prompt.

If you can, move transport first while keeping a model you already know, verify nothing broke, then change the model as a separate deploy. If the new provider does not serve anything you currently use, you cannot split it — in which case budget more time for evaluation, not less.

Build a golden set before you touch anything

Capture 50 to 200 real requests from production, with their current responses. Not synthetic examples — actual traffic, including the weird ones. Strip anything sensitive, store them as fixtures, and treat that file as the contract.

You do not need a scoring model to make this useful. Even a diff view where a human skims old versus new output catches the big regressions: a format that changed, a refusal that appeared, a response that got three times longer.

Where you can write a deterministic check, write one. Does the JSON parse. Does the SQL compile. Does the classification land in the allowed enum. Did the tool call name match. Those assertions are worth more than a vague quality score, because they fail loudly in CI.

The parameters that quietly differ

Request bodies are broadly portable, but a handful of fields have drifted.

  • max_tokens is deprecated on OpenAI chat completions in favour of max_completion_tokens, and the o-series reasoning models reject max_tokens outright. Other providers still only accept max_tokens. Send whichever the target expects rather than assuming.
  • seed is best-effort everywhere and ignored entirely by some backends. If your tests depend on reproducibility, verify it rather than trusting it.
  • response_format and strict schema enforcement vary a lot in how strictly they are honoured.
  • Unknown parameters are usually ignored rather than rejected, so a field the new provider does not support fails silently. Your carefully chosen top_p may simply not be applied.

The practical defence is to assert on behaviour, not on the request. Change a parameter, confirm the output actually changes, and only then believe it took effect.

Error handling needs remapping too

OpenAI returns a consistent error envelope with message, type, param and code fields, and a lot of production code branches on those strings — insufficient_quota, invalid_api_key, context_length_exceeded.

{
  "error": {
    "message": "...",
    "type": "insufficient_quota",
    "param": null,
    "code": "insufficient_quota"
  }
}

A new provider will use the same shape and different values, or the same values with different triggers. Grep your codebase for any string comparison against an error code before you migrate, and rewrite those branches to key off HTTP status first and the code only as a refinement.

Rate limit headers are the same story. OpenAI returns x-ratelimit-limit-requests, x-ratelimit-remaining-tokens, x-ratelimit-reset-requests and friends, plus Retry-After on eligible errors. Any dashboard or throttle you built on those header names needs checking against the new provider, which may return none of them.

Run both in shadow before you switch

The cheapest way to de-risk is to send a percentage of real traffic to both providers, serve the incumbent response to the user, and log the challenger for comparison.

import asyncio
from openai import AsyncOpenAI

incumbent = AsyncOpenAI(base_url="https://api.openai.com/v1", api_key=OLD_KEY)
challenger = AsyncOpenAI(base_url=NEW_BASE_URL, api_key=NEW_KEY)

async def handle(messages):
    primary = incumbent.chat.completions.create(
        model="MODEL_A", messages=messages
    )
    shadow = challenger.chat.completions.create(
        model="MODEL_B", messages=messages
    )
    result, other = await asyncio.gather(primary, shadow, return_exceptions=True)
    log_comparison(messages, result, other)
    return result

Shadow traffic doubles your inference bill for the duration, so cap it — one percent of requests for a week usually surfaces the format regressions and the latency profile. Sample deliberately across your prompt types rather than taking the first N requests, which will over-represent whatever runs on a cron.

Timeouts and retries are not portable

The OpenAI Python SDK defaults to two retries and a ten-minute request timeout, retrying connection errors, 408, 429 and 5xx with exponential backoff. If you kept those defaults, you inherited a retry policy you did not choose, and it follows you to the new provider because it lives in the client, not the server.

Set both explicitly during a migration so that a slow new provider does not silently turn into tripled load:

client = OpenAI(base_url=NEW_BASE_URL, api_key=KEY, max_retries=0, timeout=30.0)

Then implement retries in one place you control, so you can see them in metrics. Two layers of automatic retry — SDK plus your own wrapper — multiply rather than add, and that is how a brief upstream blip becomes a self-inflicted outage.

Keep the accounting honest

Verify that the new provider reports a usage block, that it reports one while streaming if you rely on that, and that the counts are real upstream numbers rather than an estimate. Log token counts per request from day one of the shadow period, because "is this cheaper" is a question you want answered with your own traffic mix rather than a price-per-million comparison.

Pricing model matters here as much as unit price. If your workload is agentic — long contexts, many turns, frequent retries — per-token billing gets volatile in a way that flat-rate access does not, which is the case we built CozyAPI around. Either way, measure before you commit.

A migration checklist

  1. Freeze a golden set of real requests and current outputs.
  2. Write deterministic assertions for anything machine-checkable.
  3. Move transport and model in separate deploys where possible.
  4. Grep for hardcoded error codes and rate-limit header names.
  5. Set explicit timeouts and disable SDK-level retries.
  6. Shadow one percent of traffic for a week; diff the outputs.
  7. Confirm usage reporting matches your billing expectations.
  8. Keep the old credentials live and the switch behind a flag until you have a quiet week.

The rollback plan is the part people skip. If reverting means a code change and a deploy, you will hesitate at exactly the moment you should not. Put the provider choice in configuration and make the reversal a one-line change.

Common questions

How long should a provider migration take?

The code change is an afternoon. Plan a week of shadow traffic on top of it so you see the full range of production prompts, including whatever only runs on a weekly schedule.

Do I need to rewrite my prompts for a different model?

Sometimes. Prompts that rely on a specific formatting habit or refusal behaviour are the ones that break. Run your golden set first, then rewrite only what actually regressed.

Can I keep using the official OpenAI SDK?

Yes, for any provider that serves the chat completions shape. Set base_url and the key. Do check the SDK retry and timeout defaults, because they follow you to the new host.

Similar articles

What "OpenAI-Compatible" Actually Means (And Where It Breaks)
Guides
Guides·8 min read

What "OpenAI-Compatible" Actually Means (And Where It Breaks)

Most AI tools speak one API shape. Understanding what compatibility covers — and the four places it usually leaks — saves hours of debugging a base URL swap.

Read
Building a Chatbot From Scratch: The Parts Nobody Mentions
Guides
Guides·9 min read

Building a Chatbot From Scratch: The Parts Nobody Mentions

The model call is twenty lines. The other ninety percent is conversation state, idempotency, abuse limits and knowing when a reply was wrong. A build order that works.

Read
Choosing an AI Provider: The Questions to Ask
Guides
Guides·9 min read

Choosing an AI Provider: The Questions to Ask

A checklist of the questions that separate providers you can plan around from ones you cannot, covering model transparency, limits, compatibility and exit terms.

Read