Choosing an AI Provider: The Questions to Ask
Guides

Choosing an AI Provider: The Questions to Ask

A checklist of the questions that separate providers you can plan around from ones you cannot, covering model transparency, limits, compatibility and exit terms.

Most provider comparisons rank on price and a benchmark score. Both are the wrong axis, because both are trivially adjustable by the vendor and neither predicts the thing that actually hurts — finding out six weeks in that the model changed, the limit is lower than you assumed, or your tooling cannot talk to it.

What follows is a list of questions with a common property: a provider who answers them plainly is demonstrating something about how they will behave when something goes wrong. Vagueness on any of them is itself an answer.

Model transparency

You are buying a specific capability, and you need to know it is still what you bought.

  • Which model does each alias serve, named exactly? Including version. "Our fast model" is not a specification.
  • Is the served model identified in the response? If the response body does not echo a model identifier, a silent substitution is undetectable by you.
  • Can the model behind an alias change without notice? Many providers reserve this right. Knowing is fine; not knowing is not.
  • Is the model quantised relative to the reference weights? Aggressively quantised serving is a legitimate cost strategy and a material difference in output quality. Ask.
  • What is the deprecation policy? How much notice before a model is retired, and is there a pinned version you can stay on.

Silent substitution is the failure mode with the longest detection time, because nothing errors. Your evaluation scores just drift downward and someone eventually notices.

Limits, stated numerically

Every provider has limits. The question is whether they are published.

  • Requests per minute, tokens per minute, concurrent requests. Numbers, per model, at your tier.
  • Over what window is any quota measured? Per minute, per five hours, per week, per month. A generous monthly total does not help if one intense day is capped.
  • Do limits vary by time of day? Peak-hour shaping is common and rarely on the pricing page.
  • What happens at the limit? A 429 you can back off from, a hard stop until reset, a downgrade to a smaller model, or account review. These are different products sold under the same word.
  • Are limits exposed in response headers? If yes you can monitor headroom; if no you can only discover the ceiling by hitting it.
  • Is the quota per key, per user, or per organisation? Shared org-level quotas mean colleagues contend with each other.

The context window you actually get

Headline context numbers describe the model. What you can use is a property of the deployment.

  • What is the maximum input the endpoint accepts? Sometimes lower than the model specification.
  • What is the maximum output? Separate limit, frequently much smaller, and it silently truncates long generations.
  • Is output counted against the same window? It usually is, so filling the window leaves no room to answer.
  • Are there byte-level request size limits independent of tokens? These exist and are enforced before the model sees anything, which is why a large image or PDF payload can fail while being well within the token budget.
  • Is long context priced differently? Some providers charge a premium above a threshold.

Compatibility and exit

The cost of leaving determines how much leverage you have while you stay. Establish it before you commit, not after.

  • Is the endpoint OpenAI-compatible? If your existing tools work with a base URL and key change, switching cost is near zero and you keep your options.
  • Which extensions are supported? Tool calling, structured outputs, streaming, prompt caching, vision, reasoning parameters. Compatible does not mean feature-complete, and the gaps are usually in exactly these.
  • Are unknown parameters rejected or ignored? A gateway that 400s on any field it does not recognise will break clients that send provider-specific extras.
  • Any client restrictions? Some plans permit a vendor tool but not third-party harnesses or direct SDK use.
  • Can I export my usage data? You need it to forecast, and to price a competing offer.

What happens at the end of the term

The least glamorous section and the one that produces the most unpleasant surprises.

  • Does access stop, auto-renew, or degrade? Get this in writing.
  • Is unused allowance carried forward or forfeited?
  • What is the refund position if limits change materially mid-term, or if a model you depend on is retired?
  • How long is data retained after cancellation, and can you require deletion?
  • Do API keys keep working during a grace period, or fail closed at the exact moment of expiry? This determines whether expiry is an inconvenience or an outage.

Data, reliability and support

Three shorter sets that are nonetheless disqualifying if the answers are wrong.

Data. Are prompts or outputs used for training, by default or at all? How long is data retained, and where geographically? Is there human review of flagged content? Which sub-processors handle the traffic? For anything touching customer data these are not negotiable niceties.

Reliability. Is there a public status page with real historical incidents — an unbroken wall of green usually means unreported outages rather than perfect uptime. Is there a published uptime target, and does anything actually happen if it is missed? How are capacity constraints communicated?

Support. Is there a human, what is the response time, and what happens at 2am. Do error responses carry a request identifier you can quote, because without one no support conversation can go anywhere.

Evaluate before you commit

Every question above is about disclosure. This last part is about verification, and it takes an afternoon.

  1. Run twenty real tasks from your own workload through the provider. Not benchmarks. Yours.
  2. Run your heaviest realistic load at your normal peak hour, deliberately, while you still have a refund window.
  3. Log the model identifier on every response and alert on changes.
  4. Keep a fixed set of prompts with known-good outputs and re-run them weekly, forever. This is the only reliable detector of silent quality drift.
  5. Configure a second provider behind a flag before you need it. If switching costs a base URL and a key, you have an exit; if it costs a rewrite, you have a dependency.

Apply this to us as well. We sell a flat-rate pass, it is rate limited like everything else, and if we ever answer one of the questions above with an adjective where a number belongs, that is a reasonable basis for scepticism about the rest. The point of a checklist is that it is applied uniformly, including to the provider whose blog you happen to be reading.

Common questions

What is the most important question to ask an AI provider?

Which exact model each alias serves, and whether that can change without notice. Silent model substitution is the failure mode with the longest detection time, because nothing errors and quality simply drifts downward.

How do I check the context window a provider actually gives me?

Ask for the maximum accepted input and the maximum output separately, confirm whether output counts against the same window, and check for byte-level request size limits enforced before the model sees the request.

Why does OpenAI API compatibility matter when choosing a provider?

It sets your switching cost. If existing tools work after changing a base URL and key, you retain leverage and an exit path. Check which extensions are supported too, since compatible endpoints often differ on tool calling, caching and structured outputs.

Similar articles

Building a Chatbot From Scratch: The Parts Nobody Mentions
Guides
Guides·9 min read

Building a Chatbot From Scratch: The Parts Nobody Mentions

The model call is twenty lines. The other ninety percent is conversation state, idempotency, abuse limits and knowing when a reply was wrong. A build order that works.

Read
Debugging LLM API Errors, Status Code by Status Code
Guides
Guides·9 min read

Debugging LLM API Errors, Status Code by Status Code

A field guide to the errors an LLM API actually returns: what each status means, which ones are worth retrying, and how to reproduce the failure in one curl command.

Read
LLM API Error Codes: A Practical Reference
Guides
Guides·9 min read

LLM API Error Codes: A Practical Reference

What each status code from an LLM API actually means, which are safe to retry, and how to handle the ones that look transient but are not. With real header names.

Read