Free AI API Tiers: How to Compare Limits That Keep Moving
Cost & Pricing

Free AI API Tiers: How to Compare Limits That Keep Moving

Free tier numbers rot within weeks, and several vendors have stopped publishing them entirely. Here is a method for comparing them that survives the churn.

Every few months someone publishes a table of free AI API limits, and within a few weeks most of the cells are wrong. Quotas get cut, models get rotated out, and the tier you signed up for gets renamed.

There is a more durable version of this article: not the numbers, but the shape of the numbers — what dimensions a free tier is limited on, which of those actually bind you, and how to measure whether a free tier covers your workload before you build on it.

Several vendors have stopped publishing the numbers

This is the most important thing to know as of August 2026, and it is easy to verify yourself.

Google's Gemini API rate limits documentation (last updated 21 July 2026) no longer lists per-model free tier figures. It states that limits depend on your usage tier and directs you to view your active limits in AI Studio. GitHub's Copilot plan comparison page describes the Free plan qualitatively — limited access to a selection of features, auto model selection only — rather than committing to a monthly completion or chat allowance.

Read that as a signal, not an oversight. When a vendor moves quota from documentation into a console, the quota has become something they intend to adjust without a docs change. Any third-party table quoting those numbers is quoting a snapshot.

The counter-example is worth noting because it is the exception: OpenRouter still documents its free-variant limits plainly — 20 requests per minute, 50 requests per day, rising to 1,000 requests per day once your account has purchased at least 10 credits at any point. Published, specific, and checkable. That is what a free tier you can plan against looks like.

The five dimensions that matter

Whatever the provider, a free tier is constrained on some combination of these. Work out which one binds you.

  • Requests per minute (RPM). Binds interactive and parallel workloads. A batch job with concurrency 10 hits this instantly.
  • Tokens per minute (TPM). Binds long-context work. A single 200k-token request can exhaust a whole minute of budget.
  • Requests per day (RPD). Binds sustained use. This is the one that decides whether a free tier can carry a real project.
  • Model access. Free tiers usually route to the small or older models. Sometimes you do not get to choose at all.
  • Data terms. Free frequently means your prompts may be used for training or human review. This is the price, and for client work it is often disqualifying.

Most people compare on RPD because it looks like the headline. In practice TPM is what kills agentic use, because agent turns are enormous and the per-minute token budget is spent by the third tool call.

Work out whether a free tier fits, in five minutes

Do not eyeball it. Estimate your actual shape of demand.

requests_per_day  = tasks_per_day × turns_per_task
tokens_per_minute = avg_request_tokens × requests_per_minute_peak

A concrete example. Say you are building a small internal tool: 40 tasks a day, each a 6-turn agent loop, averaging 12,000 input tokens per turn.

requests_per_day  = 40 × 6           = 240
tokens_per_day    = 240 × 12,000     = 2.88M
peak burst        = 4 concurrent × 12,000 = 48,000 TPM

Now compare against the tier. 240 requests a day fits comfortably inside a 1,000 RPD allowance. But a 48,000 TPM burst will not fit inside a tier capped in the low tens of thousands of tokens per minute, and no amount of daily headroom rescues you. You would need to serialise the work, which turns a 20-second task into a two-minute one.

That asymmetry is the general lesson: free tiers are usually generous on request count and stingy on tokens, because tokens are what actually costs the provider money.

The failure modes to plan for

Free tiers do not fail politely. Budget engineering time for these:

  • 429 with no useful retry hint. Some free endpoints return a rate-limit error without a Retry-After header. You need your own backoff.
  • Silent model swaps. If the tier promises "auto model selection", the model behind your prompt can change without notice, and so can your output quality.
  • Shared org-level quotas. Limits often apply per organisation, not per key. Two colleagues on the same account contend with each other.
  • Cold-start latency. Free capacity is frequently the lowest priority queue. Median latency looks fine; the p99 does not.
  • Sudden withdrawal. Free model variants get retired with little notice. If your product depends on one specific free model ID, you have a single point of failure.

Where free tiers are genuinely the right answer

They are excellent for learning, for prototyping, for evaluating whether a model can do a task at all, and for low-volume personal scripts. There is no reason to pay for any of that.

They stop being the right answer at a fairly predictable point: when you need consistent latency, a model you chose deliberately, data terms you can show a client, or throughput that survives a burst. At that point the question is not free versus paid, it is metered versus flat-rate — which comes down to how spiky your usage is. Our own flat-rate pass exists for the steady-heavy end of that spectrum; if your usage is genuinely light and bursty, a free tier or a metered account is the cheaper answer and you should stay there.

A checklist before you build on a free tier

  1. Find the limits in the provider's own documentation or console. If you cannot find published numbers, assume they will change.
  2. Note the date you checked, in a comment next to the code that depends on them.
  3. Compute your RPM, TPM and RPD from your actual workload, not from intuition.
  4. Read the data-retention and training terms before any customer data touches it.
  5. Implement backoff and a fallback provider before you need them, not after the first outage.
  6. Keep the model ID in configuration. Free model IDs get retired; hardcoding one guarantees a bad afternoon.

The durable advice is simple: treat every free tier figure you read, including the ones in this article, as a measurement with a timestamp rather than a fact.

Common questions

Which free AI API tier is the most generous right now?

Any specific answer goes stale within weeks, and several vendors no longer publish per-model figures at all. Check the provider console or docs directly, note the date, and compare on tokens per minute rather than requests per day.

Why do I hit rate limits when I am nowhere near the daily quota?

Because the binding constraint is usually tokens per minute, not requests per day. One long-context agent turn can consume a whole minute of token budget on its own, so concurrency trips the limit long before the daily count does.

Can I run a production feature on a free tier?

Rarely a good idea. Free capacity is typically lowest priority, quotas can change without a documentation update, model variants get retired, and the data terms often permit training on your prompts. Use free tiers to evaluate, then pay for anything with a user waiting on it.

Similar articles

Is Unlimited AI Actually Unlimited?
Cost & Pricing
Cost & Pricing·9 min read

Is Unlimited AI Actually Unlimited?

No inference provider can sell genuinely uncapped compute. Here is what the word hides, why the limits exist, and exactly what to demand in writing before you buy.

Read
Comparing Provider Pricing Models Without Getting Fooled
Cost & Pricing
Cost & Pricing·8 min read

Comparing Provider Pricing Models Without Getting Fooled

Per-token, per-seat, credits, tiers and flat rate all quote different units. Here is how to normalise them onto one number you can actually compare.

Read
The Real Cost of Open vs Closed Models
Cost & Pricing
Cost & Pricing·8 min read

The Real Cost of Open vs Closed Models

Open weights cost nothing to download and plenty to run. Here is how licence, hosting and switching costs actually compare, with current per-token rates.

Read