If your team runs thousands of LLM calls that nobody is waiting on — nightly classification, evaluation suites, embedding backfills, document summarization — you are probably paying double. Every major provider now offers an LLM batch API that processes requests asynchronously at roughly half the price of real-time calls. In this guide we compare OpenAI’s Batch API, Anthropic’s Message Batches API for Claude, and Google’s Gemini Batch Mode, and show you when (and when not) to use each.
The short version: if a job can tolerate minutes-to-hours of latency, moving it to a batch endpoint is one of the easiest 50% cost cuts available in AI engineering in 2026 — no model change, no quality trade-off.
What Is an LLM Batch API and Why Use It?

A batch API accepts a large set of independent requests in one submission, queues them on the provider’s spare capacity, and returns the results later. Because providers can schedule this work around peak traffic, they pass the savings on to you.
The workflow is nearly identical across vendors:
- Build one request object per task (same body you would send to the normal chat/messages endpoint), each with a unique
custom_id. - Submit them — usually as a JSONL file or an inline array.
- Poll the batch status (or wait for a webhook/notification).
- Download results and join them back to your data by
custom_id, since output order is not guaranteed.
Key benefits beyond price:
- Separate rate limits — batch traffic draws from its own quota, so a million-row backfill will not throttle your production chatbot.
- Same models, same outputs — you get the identical model, just delivered asynchronously.
- Simpler retry logic — the provider handles queuing; you only re-submit the individual requests that errored.
OpenAI Batch API
OpenAI’s Batch API works with chat completions, the Responses API, and embeddings. You upload a JSONL file via the Files API with purpose="batch", then create a batch pointing at the endpoint you want to call.
- Discount: 50% off standard per-token pricing.
- Completion window: 24 hours (the only option today); many batches finish well before that.
- Limits: a single input file can hold tens of thousands of requests (check the current per-batch cap in OpenAI’s docs), and batch tokens count against a separate enqueued-token quota.
- Best for: embedding backfills, bulk classification, and teams already standardized on GPT models.
from openai import OpenAI
client = OpenAI()
f = client.files.create(file=open("requests.jsonl", "rb"), purpose="batch")
batch = client.batches.create(
input_file_id=f.id,
endpoint="/v1/chat/completions",
completion_window="24h",
)
print(batch.id, batch.status)Anthropic Message Batches API (Claude)
Anthropic’s Message Batches API lets you send Claude requests inline in a single API call — no separate file upload step.
- Discount: 50% off all token types, including prompt-cache writes and reads — so batching and prompt caching stack.
- Limits: up to 100,000 requests or 256 MB per batch, whichever comes first.
- Speed: most batches complete within an hour; anything unfinished after 24 hours expires.
- Best for: long-document processing and evals where a shared system prompt can be cached across every request.
import anthropic
client = anthropic.Anthropic()
batch = client.messages.batches.create(requests=[
{"custom_id": f"doc-{i}",
"params": {"model": "claude-sonnet-5-5", "max_tokens": 512,
"messages": [{"role": "user", "content": text}]}}
for i, text in enumerate(docs)
])
print(batch.id, batch.processing_status)Google Gemini Batch Mode
The Gemini API’s Batch Mode offers two submission styles, which makes it flexible for both small and very large jobs.
- Discount: 50% of standard cost.
- Turnaround: 24-hour target, usually much faster.
- Inline requests: pass a list of
GenerateContentRequestobjects directly, as long as the total stays under 20 MB. - File input: upload a JSONL file for larger jobs; results come back as a JSONL file.
- Best for: multimodal batches (images, PDFs, video frames) and teams on Google Cloud.
LLM Batch API Comparison Table
| Feature | OpenAI Batch API | Anthropic Message Batches | Gemini Batch Mode |
|---|---|---|---|
| Discount | 50% | 50% (all token types) | 50% |
| Submission | JSONL file upload | Inline JSON array | Inline (<20 MB) or JSONL file |
| Max window | 24 hours | 24 hours (most <1 hr) | 24-hour target |
| Batch size cap | Per-batch request + file-size cap | 100,000 requests / 256 MB | Inline 20 MB; larger via file |
| Stacks with caching | Cached-input pricing applies | Yes, cache discounts stack | Check current docs |
| Result order | Not guaranteed — use custom_id | Not guaranteed — use custom_id | Keyed per request |
When Should You Use a Batch API?
Great fits
- Nightly or weekly LLM evaluation runs over a fixed test set.
- Generating synthetic training data for fine-tuning.
- Tagging, moderation, or sentiment scoring of historical tickets, reviews, or logs.
- Re-embedding a document corpus after switching embedding models.
- Summarizing or extracting structured fields from large PDF archives.
Poor fits
- Anything user-facing: chat, autocomplete, agents waiting on tool results.
- Multi-step agent loops where step N depends on step N-1 (each step would wait in the queue).
- Tiny jobs — a few hundred requests rarely justify the extra orchestration.
Production Tips for Batch Workloads
- Always set a deterministic
custom_id(e.g., a database primary key) so you can join results and make re-runs idempotent. - Chunk huge jobs into several batches. A single malformed line or quota issue then affects only one chunk.
- Persist batch IDs in your job table and poll from a scheduler (cron, Temporal, Spring
@Scheduled) rather than holding a process open. - Combine with caching: put long, shared instructions at the start of the prompt so cached-token discounts apply on top of the batch discount.
- Route by urgency: an LLM router or LLM gateway can send latency-tolerant traffic to batch endpoints automatically.
- Handle partial failures: parse per-request error objects and re-queue only the failed IDs.

FAQ
Does an LLM batch API give lower-quality results?
No. Batch endpoints run the same models as the real-time API. Only delivery timing changes, not the model weights or decoding.
How much can I save with a batch API?
OpenAI, Anthropic, and Google all advertise a 50% discount versus standard pricing. On Claude, that discount also applies to prompt-cache reads and writes, so combined savings can be even higher.
How long does a batch job take?
All three providers target completion within 24 hours. In practice, Anthropic reports most batches finish within an hour, and Gemini says most jobs complete much faster than the target.
Can I stream responses from a batch API?
No. Batch APIs are asynchronous: you poll for status and download complete results. Use the standard endpoints if you need streaming.
Conclusion
An LLM batch API is the lowest-effort way to halve inference spend for offline workloads. Choose OpenAI’s Batch API if your stack is already GPT-centric, Anthropic’s Message Batches if you want caching and batching discounts to stack on long prompts, and Gemini Batch Mode for flexible inline-or-file submission and multimodal jobs. Start by auditing which of your current LLM calls have no user waiting on them — those are your first candidates.
Want more ways to cut AI costs? Subscribe to NewsifyAll for weekly hands-on guides to building cheaper, faster LLM applications.

