Batching Requests to Save Money: When It Pays
Cost & Pricing

Batching Requests to Save Money: When It Pays

Batch endpoints trade latency for a large discount, and packing items into one prompt amortises overhead. Here is when each pays and when neither does.

Two entirely different things get called batching, they save money by different mechanisms, and they compose. Confusing them is why teams sometimes conclude batching did not help.

Asynchronous batch endpoints take a file of requests, process them on the provider's schedule, and return results later at a discount. Prompt-level packing puts several work items into one request so they share the system prompt and instructions. The first trades latency for a published discount. The second trades per-item isolation for amortised overhead.

Batch endpoints: the discount and its price

Both Anthropic and OpenAI publish a 50% discount on batch processing, as of August 2026. Anthropic's Message Batches API accepts up to 100,000 requests per batch, with most batches completing within an hour and a stated maximum of 24 hours.

The arithmetic is unusually simple. Halve your token cost; accept that results arrive when they arrive.

1,000,000 classification requests
  at 3,000 input and 60 output tokens each

sync:   3,000M input x $5/M  +  60M output x $25/M
      = $15,000 + $1,500 = $16,500

batch:  50% off = $8,250

$8,250 saved for accepting a delay. If nobody is waiting on the result, that is free money and there is no argument against taking it.

The cases where it does not apply are equally clear: anything user-facing, anything on a critical path, anything where the input depends on the previous output. Agent loops are structurally incompatible — each step needs the previous result — so batch endpoints do essentially nothing for agentic coding.

Where batch endpoints genuinely fit

  • Bulk classification and extraction over an existing corpus.
  • Backfills — summarising, tagging or embedding historical records.
  • Evaluation runs. An eval suite of 500 cases has no latency requirement at all, and halving its cost makes running it more often affordable.
  • Nightly reporting where the deadline is the morning.
  • Synthetic data generation for fine-tuning or testing.

One implementation detail that catches people: batch results come back in arbitrary order. Key them by the custom identifier you supplied, never by position. Results also carry a per-item status, so handle partial failures rather than assuming the whole batch succeeded.

Prompt-level packing: amortising the fixed cost

The second mechanism is unrelated to any endpoint, and it works on synchronous traffic too. Every request carries fixed overhead — system prompt, instructions, output schema, few-shot examples — that gets re-sent for each item. Packing several items into one request pays that overhead once.

Model it directly:

per_request_cost(k) = (fixed + k x per_item_input) x in_price
                    + (k x per_item_output) x out_price

cost_per_item(k)    = per_request_cost(k) / k

With a 2,000-token fixed preamble, 400 tokens per item, 50 output tokens per item, at $5 input and $25 output per million:

k = 1:  (2,000 + 400)  x 5e-6 + 50   x 25e-6 = $0.01325 per item
k = 5:  (2,000 + 2,000) x 5e-6 + 250 x 25e-6 = $0.02625 / 5 = $0.00525
k = 20: (2,000 + 8,000) x 5e-6 + 1000 x 25e-6 = $0.075 / 20 = $0.00375
k = 50: (2,000 + 20,000) x 5e-6 + 2500 x 25e-6 = $0.1725 / 50 = $0.00345

From $0.0133 to $0.0035 per item — a 74% reduction — with almost all of the gain arriving by k=20. The curve flattens because once the fixed cost is spread thin, the remaining per-item cost is irreducible.

That flattening is the useful insight. Packing 20 items captures most of the available saving; packing 200 adds little and introduces real risk.

The risks of packing, and how to size around them

Accuracy degrades with pack size. Models handle a handful of parallel items well and a hundred poorly, with errors concentrating in the middle of long lists. Measure this on your own data rather than assuming a safe k.

One failure spoils the batch. If the response is malformed or truncated, you lose all k items and retry all k. Expected cost including retries:

effective_cost_per_item = cost_per_item(k) / (1 - failure_rate(k))

Failure rate rises with k, so there is an optimum. If packing 50 items gives a 12% failure rate and packing 10 gives 2%, the smaller pack is often cheaper despite the higher nominal per-item cost.

Truncation risk. Output scales with k, so a pack sized against the input limit can still hit the output ceiling. Set max_tokens from k times expected per-item output, with headroom.

Attribution is lost. Per-item cost and latency become estimates once items share a request. If you need per-tenant billing accuracy, packing across tenants makes that harder.

A reasonable default: start at k=10, measure accuracy and failure rate, and increase only while both hold.

Composing the two, plus caching

They stack, and so does prompt caching. For a large offline job:

base:                          $16,500
pack 20 items per request:     roughly 25% of base  -> $4,150
submit via batch endpoint:     50% off              -> $2,075

Caching interacts differently. Packing already removes most of the repeated preamble, so the marginal value of caching that preamble drops — the two techniques address the same waste. Where they combine well is when a large shared reference document sits in the prefix: cache that, and pack the varying items after it.

Check whether your provider supports caching within batch requests before assuming both apply. This is a detail worth verifying rather than inferring.

A decision procedure

  1. Is anyone waiting? If no, use the batch endpoint. The 50% discount requires no engineering beyond a different submission path.
  2. Is the fixed overhead a large share of each request? Compute fixed / (fixed + per_item). Above about 40%, packing is worth building.
  3. Are items independent? Packing requires that item three does not depend on item two. Sequential dependency rules it out.
  4. Can you tolerate all-or-nothing failure? If a whole pack failing is unacceptable, keep k small and retry per item.

If the answer to the first question is "yes, a user is waiting" and to the third is "no, they are dependent" — which describes every agent loop — neither technique applies, and your levers are caching, context discipline and turn count instead.

What batching does not fix

It does not reduce the number of tokens you actually need. A batch endpoint discounts them and packing amortises the overhead, but a prompt carrying 8,000 tokens of unnecessary context still carries them, at half price or spread across twenty items.

Do the token reduction work first, then apply batching to what remains. Applied in the other order, batching hides the waste well enough that nobody goes looking for it.

Common questions

How much do batch API endpoints save?

Both Anthropic and OpenAI publish a 50% discount on batch processing as of August 2026, in exchange for asynchronous delivery. Anthropic states most batches finish within an hour, with a 24-hour maximum.

How many items should I pack into a single prompt?

Most of the saving arrives by around twenty items, because the fixed overhead is already spread thin by then. Start at ten, measure accuracy and malformed-response rate, and increase only while both hold steady.

Can I batch requests for an agent workflow?

No. Agent loops are sequentially dependent — each step needs the previous result — so neither asynchronous batching nor prompt packing applies. Use prompt caching, tool output truncation and turn caps instead.

Similar articles

Prompt Caching Economics: Break-Even, TTL and Hit Rate
Cost & Pricing
Cost & Pricing·8 min read

Prompt Caching Economics: Break-Even, TTL and Hit Rate

Cache writes cost more than normal input, so caching only pays above a read threshold. Here is the arithmetic for TTL choice, hit rate and breakpoint placement.

Read
Forecasting AI Spend Without Guessing
Cost & Pricing
Cost & Pricing·8 min read

Forecasting AI Spend Without Guessing

Most AI budget forecasts are a headcount multiplied by a hopeful number. Here is a model that decomposes spend into drivers you can actually measure and control.

Read
Input vs Output Token Pricing: Which One Actually Bills You
Cost & Pricing
Cost & Pricing·7 min read

Input vs Output Token Pricing: Which One Actually Bills You

Output tokens cost several times more per token, yet input usually dominates the bill. Here is how to compute your own blended rate and act on it.

Read