Is Unlimited AI Actually Unlimited?
Cost & Pricing

Is Unlimited AI Actually Unlimited?

No inference provider can sell genuinely uncapped compute. Here is what the word hides, why the limits exist, and exactly what to demand in writing before you buy.

No. Not from us, not from anyone, and any provider who tells you otherwise is either being loose with language or has not yet met a customer who tested it.

That is not a criticism of flat-rate pricing — it is a description of what flat-rate pricing is. The useful question is not whether limits exist but whether they are disclosed, where they sit, and what happens when you reach them. This piece is about how to find that out before you pay, including when the provider is us.

Why genuine unlimited is impossible

Inference is not a software licence. Every token generated occupies a GPU for a measurable slice of time, and that capacity is finite, expensive and slow to add — data centre buildouts run on multi-year timescales, not quarterly ones.

The economics follow directly. A flat-rate provider is making a bet on the distribution of usage across its customers: most will use a modest amount, some will use a lot, and the price is set so the average works. A single customer running maximum concurrency around the clock can consume more compute than dozens of ordinary subscriptions pay for.

So every flat-rate offer contains a mechanism to handle that case. The only variable is whether the mechanism is written down.

The seven places the limit actually lives

When a plan says unlimited, one or more of these is doing the real work. Learn to recognise them.

  • Fair use clauses. A terms-of-service paragraph reserving the right to throttle or suspend "excessive" or "abusive" usage, with excessive left undefined. This is the most common mechanism and the least measurable.
  • Undisclosed rate limits. Requests or tokens per minute enforced at the gateway. You discover them as 429s during your busiest hour.
  • Concurrency caps. A limit on simultaneous in-flight requests. Invisible to chat users, immediately binding for anyone running parallel agents.
  • Rolling windows. Quotas measured over a five-hour or weekly window rather than a month. The monthly total may be generous while a single intense day is not.
  • Time-of-day shaping. The same nominal allowance drains faster during peak hours. Your quota is unchanged; your ability to use it when you want is not.
  • Silent model substitution. Under load, requests route to a smaller or more heavily quantised model. Nothing errors. The output just gets worse, which is the hardest failure to detect.
  • Client restrictions. The plan permits the vendor's own tool but not third-party harnesses or SDK access. Unlimited within a client you did not choose.

None of these is inherently dishonest. Rate limits protect every other customer on the platform, and a provider with no abuse controls will be destroyed by the first person who points a scraper at it. The problem is not the limit. The problem is the gap between the word on the pricing page and the number in the gateway config.

What has actually happened in this market

The pattern is now well established across the coding-assistant category. Plans launch with generous or unstated limits, adoption outruns capacity, and limits get tightened, redistributed across rolling windows, or shifted between peak and off-peak. Third-party client access has been restricted and then reinstated on separate credit pools. Sessions that used to run for hours have ended in a fraction of that.

The recurring complaint in each case has been less about the tightening itself than about learning of it mid-task. A limit you planned around is a constraint. A limit you discover at 4pm on a deadline is an outage.

It is worth being clear that capacity constraints are genuine, not a pretext. When a provider says it is compute-constrained, that is usually true and usually not fixable on the timescale of your annoyance.

What to demand in writing

Ask these before you buy anything, from any provider. A vendor that answers all of them plainly is telling you something about how they will behave later.

  1. What are the numeric rate limits? Requests per minute, tokens per minute, concurrent requests. A number, not an adjective.
  2. Over what window is the quota measured? Per minute, per five hours, per week, per month. This determines whether one heavy day is possible.
  3. Does the limit vary by time of day? If yes, by how much and during which hours.
  4. What happens when I hit it? A 429 I can back off from, a hard stop until reset, a downgrade to a smaller model, or account review. These are very different products.
  5. Will I be told which model served my request? If the response does not identify the model, silent substitution is undetectable.
  6. Can I use my own client? An OpenAI-compatible endpoint with no client restrictions means your tools keep working.
  7. How much notice before limits change? And are changes announced, or discovered?
  8. What happens at the end of the term? Access stops, auto-renews, or degrades — and is unused allowance carried forward.

Get the answers in writing, in the documentation or an email, not in a chat window. If a provider will only describe limits qualitatively, treat the vagueness itself as the answer.

How to test a plan in the first week

Do not wait for the limit to find you.

  • Run your heaviest realistic workload deliberately, at your normal peak hour, on day two. Find the ceiling while you still have a refund window.
  • Log response headers. Rate limit headers exposing remaining requests and tokens are common; if they are present, you can monitor headroom instead of guessing.
  • Record the model identifier returned on every response, and alert if it changes.
  • Keep a small set of fixed prompts with known-good outputs and re-run them weekly. Silent quality changes show up here and nowhere else.
  • Keep a second provider configured behind a feature flag. The cost of switching should be a base URL and a key.

Our own position, plainly

We sell a flat-rate pass. It is not literally unlimited, because nothing is. It is rate limited, those limits exist to stop one account degrading service for everyone else, and the right thing for us to do is publish the numbers and tell you before we change them. Hold us to the same list above that you would apply to anyone else — if we ever answer one of those eight questions with an adjective instead of a number, that is a legitimate reason to distrust the rest.

The reasonable summary: flat rate is a good deal for steady heavy users because it removes the tax on curiosity and turns a variable line into a fixed one. It is not a promise of infinite compute, and a provider who describes it that way is setting you up for a discovery you would rather have made before paying.

Common questions

Can any AI provider offer genuinely unlimited usage?

No. Every token consumes GPU time, and capacity is finite and slow to expand. Flat-rate pricing is a bet on average usage across customers, so every such plan contains some mechanism to handle outliers. The only real question is whether it is disclosed.

What does a fair use policy mean on an AI plan?

It is a clause reserving the right to throttle or suspend usage the provider considers excessive, usually without defining excessive. It is the most common way an unlimited plan enforces a limit, and the hardest one to plan around.

How do I tell if a provider silently downgraded my model?

Log the model identifier returned in each response and alert on changes. Separately, keep a fixed set of prompts with known-good outputs and re-run them weekly. Silent substitution produces no error, so quality regression testing is the only reliable detector.

Similar articles

Free AI API Tiers: How to Compare Limits That Keep Moving
Cost & Pricing
Cost & Pricing·8 min read

Free AI API Tiers: How to Compare Limits That Keep Moving

Free tier numbers rot within weeks, and several vendors have stopped publishing them entirely. Here is a method for comparing them that survives the churn.

Read
Comparing Provider Pricing Models Without Getting Fooled
Cost & Pricing
Cost & Pricing·8 min read

Comparing Provider Pricing Models Without Getting Fooled

Per-token, per-seat, credits, tiers and flat rate all quote different units. Here is how to normalise them onto one number you can actually compare.

Read
The Real Cost of Open vs Closed Models
Cost & Pricing
Cost & Pricing·8 min read

The Real Cost of Open vs Closed Models

Open weights cost nothing to download and plenty to run. Here is how licence, hosting and switching costs actually compare, with current per-token rates.

Read