CodexHowSupport Us

Batch vs Realtime Savings Calculator

Standard, per run
$1.6000
Batch, per run
$0.8000
Saved per run
$0.8000 (50%)
Saved over 30 runs
$24.00

Batch commits to a fixed 24-hour completion window — not a maximum you can pay to shorten. A large computed saving is a reason to seriously consider Batch, not a verdict: it's only a fit if your workload can genuinely wait for that window.

Get an alert when a price on this page changes.

Flex is priced identically to Batch on every model that offers both — the distinction between them is the completion guarantee, not the price, so this calculator's numbers apply to Flex too.

Batch pricing is a flat, published discount against the Standard rate — that part of the decision is simple arithmetic. The part that actually determines whether Batch is right for a given workload is the fixed completion window, and that's a judgment call this calculator is built to make concrete rather than abstract.

The discount side is the easy half

Because the discount applies uniformly to every model that offers Batch, the dollar savings for a given workload are straightforward to compute once you know the token counts and the model — this calculator does that half the same way the token-cost estimator prices a Standard request, just against the Batch rate instead. That number on its own tends to look like an easy yes: a meaningful, guaranteed reduction for work that can tolerate a delay.

The half that actually needs a decision

What the discount number doesn't capture is whether your workload can actually tolerate Batch's fixed completion window in the first place — and "fixed" is doing real work in that sentence: it's not a maximum you can pay to shorten, it's the window, full stop. A workload where "done by tomorrow" is a genuinely acceptable promise is a clean fit. A workload where a result is needed within the hour isn't a fit at all, no matter how attractive the discount looks in isolation, because Batch simply doesn't offer a faster option at any price.

What the calculator actually shows

Enter your workload's token counts, model, and how many times you'd run it, and the calculator prices the same job twice — once at the Standard synchronous rate, once at Batch — and shows the gap in absolute terms, not just as a percentage. For a workload run repeatedly (a nightly job, a weekly report), it also projects that gap out over the period you specify, since a discount that looks modest per run can add up to something worth restructuring a pipeline around once it's multiplied by a month's worth of runs.

Reading the result honestly

A large computed saving is a reason to seriously consider Batch, not an automatic verdict — it's still worth asking, separately, whether the workload's actual deadline tolerates the fixed window, because the calculator has no way to know your operational constraints and isn't pretending to. Conversely, a workload where the computed saving looks small in absolute dollar terms might still be worth batching if the volume is large enough that the separate rate-limit pool matters more than the discount itself — Batch requests don't compete with your standard-tier throughput, which is a real operational benefit this calculator's dollar figure alone doesn't capture.

Flex, and why it isn't a separate row here

Flex is priced identically to Batch for every model that offers both, so a separate "Flex savings" calculation would just repeat this one's numbers under a different name. The distinction between the two is entirely about the completion guarantee, not the price — Batch commits to the fixed window, Flex doesn't carry that same guarantee — so this calculator treats Batch's numbers as Flex's numbers too, and the real decision between them is a reliability question this tool doesn't attempt to answer, because it isn't a cost question at all.

A note on volume and rate limits

Very large batch jobs interact with a separate rate-limit pool and, on the models that publish one, a genuinely large batch-queue ceiling — worth checking against the rate-limit reference if you're sizing a batch run large enough that even Batch's own generous queue might matter. Most individual jobs won't get near that ceiling, but a workload that does is worth planning against the published queue limit specifically, not assuming Batch has no ceiling at all.

A worked shape, not a worked number

Picture a job that reviews a stack of accumulated support tickets once a night rather than as they arrive — a textbook async fit, since nobody's waiting on any individual result the moment it's generated. Priced at Standard, that job costs whatever the token counts and model dictate; priced at Batch, the same job costs the discounted figure for the exact same tokens, and the only thing that changed is when the answers actually arrive. That's the entire decision in miniature: nothing about the quality of the answers changes, nothing about the tokens changes, only the price and the delivery guarantee — which is exactly why this calculator keeps the comparison this narrow instead of trying to model anything else.

When the discount isn't the deciding factor at all

Occasionally a workload's real constraint is neither price nor deadline but the separate rate-limit pool itself — a job so large that pushing it through Standard would blow past your account's per-minute ceiling long before the discount even mattered. For that shape specifically, Batch can be the only practical way to run the job at all, regardless of what this calculator's dollar comparison shows, since the alternative isn't "pay more," it's "can't actually send the requests fast enough."

Verified 2026-08-09 against https://developers.openai.com/api/docs/guides/batch.

Checked: https://developers.openai.com/api/docs/guides/batch · https://developers.openai.com/api/docs/pricing