Batching Requests to Save Money: When It Pays
Batch endpoints trade latency for a large discount, and packing items into one prompt amortises overhead. Here is when each pays and when neither does.
Two entirely different things get called batching, they save money by different mechanisms, and they compose. Confusing them is why teams sometimes conclude batching did not help.
Asynchronous batch endpoints take a file of requests, process them on the provider's schedule, and return results later at a discount. Prompt-level packing puts several work items into one request so they share the system prompt and instructions. The first trades latency for a published discount. The second trades per-item isolation for amortised overhead.
Batch endpoints: the discount and its price
Both Anthropic and OpenAI publish a 50% discount on batch processing, as of August 2026. Anthropic's Message Batches API accepts up to 100,000 requests per batch, with most batches completing within an hour and a stated maximum of 24 hours.
The arithmetic is unusually simple. Halve your token cost; accept that results arrive when they arrive.
1,000,000 classification requests
at 3,000 input and 60 output tokens each
sync: 3,000M input x $5/M + 60M output x $25/M
= $15,000 + $1,500 = $16,500
batch: 50% off = $8,250
$8,250 saved for accepting a delay. If nobody is waiting on the result, that is free money and there is no argument against taking it.
The cases where it does not apply are equally clear: anything user-facing, anything on a critical path, anything where the input depends on the previous output. Agent loops are structurally incompatible — each step needs the previous result — so batch endpoints do essentially nothing for agentic coding.
Where batch endpoints genuinely fit
- Bulk classification and extraction over an existing corpus.
- Backfills — summarising, tagging or embedding historical records.
- Evaluation runs. An eval suite of 500 cases has no latency requirement at all, and halving its cost makes running it more often affordable.
- Nightly reporting where the deadline is the morning.
- Synthetic data generation for fine-tuning or testing.
One implementation detail that catches people: batch results come back in arbitrary order. Key them by the custom identifier you supplied, never by position. Results also carry a per-item status, so handle partial failures rather than assuming the whole batch succeeded.
Prompt-level packing: amortising the fixed cost
The second mechanism is unrelated to any endpoint, and it works on synchronous traffic too. Every request carries fixed overhead — system prompt, instructions, output schema, few-shot examples — that gets re-sent for each item. Packing several items into one request pays that overhead once.
Model it directly:
per_request_cost(k) = (fixed + k x per_item_input) x in_price
+ (k x per_item_output) x out_price
cost_per_item(k) = per_request_cost(k) / k
With a 2,000-token fixed preamble, 400 tokens per item, 50 output tokens per item, at $5 input and $25 output per million:
k = 1: (2,000 + 400) x 5e-6 + 50 x 25e-6 = $0.01325 per item
k = 5: (2,000 + 2,000) x 5e-6 + 250 x 25e-6 = $0.02625 / 5 = $0.00525
k = 20: (2,000 + 8,000) x 5e-6 + 1000 x 25e-6 = $0.075 / 20 = $0.00375
k = 50: (2,000 + 20,000) x 5e-6 + 2500 x 25e-6 = $0.1725 / 50 = $0.00345
From $0.0133 to $0.0035 per item — a 74% reduction — with almost all of the gain arriving by k=20. The curve flattens because once the fixed cost is spread thin, the remaining per-item cost is irreducible.
That flattening is the useful insight. Packing 20 items captures most of the available saving; packing 200 adds little and introduces real risk.
The risks of packing, and how to size around them
Accuracy degrades with pack size. Models handle a handful of parallel items well and a hundred poorly, with errors concentrating in the middle of long lists. Measure this on your own data rather than assuming a safe k.
One failure spoils the batch. If the response is malformed or truncated, you lose all k items and retry all k. Expected cost including retries:
effective_cost_per_item = cost_per_item(k) / (1 - failure_rate(k))
Failure rate rises with k, so there is an optimum. If packing 50 items gives a 12% failure rate and packing 10 gives 2%, the smaller pack is often cheaper despite the higher nominal per-item cost.
Truncation risk. Output scales with k, so a pack sized against the input limit can still hit the output ceiling. Set max_tokens from k times expected per-item output, with headroom.
Attribution is lost. Per-item cost and latency become estimates once items share a request. If you need per-tenant billing accuracy, packing across tenants makes that harder.
A reasonable default: start at k=10, measure accuracy and failure rate, and increase only while both hold.
Composing the two, plus caching
They stack, and so does prompt caching. For a large offline job:
base: $16,500
pack 20 items per request: roughly 25% of base -> $4,150
submit via batch endpoint: 50% off -> $2,075
Caching interacts differently. Packing already removes most of the repeated preamble, so the marginal value of caching that preamble drops — the two techniques address the same waste. Where they combine well is when a large shared reference document sits in the prefix: cache that, and pack the varying items after it.
Check whether your provider supports caching within batch requests before assuming both apply. This is a detail worth verifying rather than inferring.
A decision procedure
- Is anyone waiting? If no, use the batch endpoint. The 50% discount requires no engineering beyond a different submission path.
- Is the fixed overhead a large share of each request? Compute
fixed / (fixed + per_item). Above about 40%, packing is worth building. - Are items independent? Packing requires that item three does not depend on item two. Sequential dependency rules it out.
- Can you tolerate all-or-nothing failure? If a whole pack failing is unacceptable, keep k small and retry per item.
If the answer to the first question is "yes, a user is waiting" and to the third is "no, they are dependent" — which describes every agent loop — neither technique applies, and your levers are caching, context discipline and turn count instead.
What batching does not fix
It does not reduce the number of tokens you actually need. A batch endpoint discounts them and packing amortises the overhead, but a prompt carrying 8,000 tokens of unnecessary context still carries them, at half price or spread across twenty items.
Do the token reduction work first, then apply batching to what remains. Applied in the other order, batching hides the waste well enough that nobody goes looking for it.
Common questions
How much do batch API endpoints save?
Both Anthropic and OpenAI publish a 50% discount on batch processing as of August 2026, in exchange for asynchronous delivery. Anthropic states most batches finish within an hour, with a 24-hour maximum.
How many items should I pack into a single prompt?
Most of the saving arrives by around twenty items, because the fixed overhead is already spread thin by then. Start at ten, measure accuracy and malformed-response rate, and increase only while both hold steady.
Can I batch requests for an agent workflow?
No. Agent loops are sequentially dependent — each step needs the previous result — so neither asynchronous batching nor prompt packing applies. Use prompt caching, tool output truncation and turn caps instead.