Why Fast Mode Burns Your Codex Plan Faster
Fast mode looks like a single trade: pay more, wait less. Inside a ChatGPT plan it's actually two trades with two different prices, and the difference between them explains a complaint that comes up again and again — "I switched Fast on and my Codex allowance vanished." It didn't vanish by accident. It was billed at the steeper of the two rates, by design, and once you see how the rates differ, it's easy to keep Fast without letting it eat a plan.
The two multipliers
OpenAI publishes how speed modes count against a plan, relative to Standard mode on the same model. Against included subscription usage, Fast counts at 2.5x the Standard rate. Against purchased credits — and Enterprise pay-as-you-go usage — it counts at 2x. GPT-6 Astra Ultrafast follows the same pattern with a wider gap: 8x against included usage, 6x against credits.
OpenAI is careful to say these are billing multipliers, not speed figures, and that credit rates alone don't determine how quickly you use included limits. That second remark is the whole story: the allowance has its own, steeper exchange rate for speed.
Why the gap matters more than either number
Included usage is always spent first, and credits only after it runs out. So for anyone who uses Fast regularly, the steeper multiplier applies to the first, prepaid part of their usage — the part they think of as already paid for. A week of habitual Fast work drains the allowance noticeably faster than the speed-up alone would suggest, and the moment the allowance is gone, the same work starts drawing on credits at the gentler multiple.
That produces a counterintuitive pattern. Fast is relatively more expensive while you're inside your plan than after you've gone beyond it. The plan isn't punishing heavy users; it's pricing speed out of included usage at a premium of its own.
What "included" really costs
Whether that premium costs real money depends on one question: do you use your whole allowance?
If you rarely hit your limits, the steeper multiplier is close to free. Allowance you'd never have used anyway is being spent on speed, and nothing extra gets billed. Fast is a reasonable default in that situation, as long as you'd notice if your habits changed.
If you regularly hit your limits and top up with credits, it's a different calculation. Every unit of Fast work paid from the allowance uses up included usage that would otherwise have covered Standard work — which now has to be bought with credits. At that point the steeper included-usage multiplier is real money, paid at one remove, and it's the users who rely on Codex most who pay it.
The API-key comparison
With an API key, ChatGPT's multipliers don't apply; Codex bills at API token prices, where Fast is its own price row. For GPT-6.1 Sol, Fast costs 2x its Standard price — compare that with 2x for credits and 2.5x for included usage, and the plan allowance stands out as the place speed is billed most steeply. That doesn't make an API key the cheaper way to run Codex overall; a plan bundles usage, features and limits that a per-token bill doesn't. It does mean that if most of your Codex work is Fast, the subscription vs API calculator is worth re-running with that in mind.
Speed you feel versus speed you pay for
All of this would matter less if Fast always delivered value proportional to its price. It doesn't. The gain applies to the model's work, not the rest of a session: time spent running tests, reading files and waiting on tools doesn't shrink. An interactive session where you're watching each response arrive benefits a lot. A long autonomous run you'll read in the morning benefits not at all — and pays the same multiplier. Because /fast persists between sessions, the second case is how most accidental Fast spend happens: switched on for a tight debugging loop, left on for the overnight refactor.
A policy that keeps Fast cheap
- Keep Standard as the saved default. Turn Fast on with
/fastfor interactive stretches, and off before you hand a session a long task. Toggling Fast mode with/fastcovers the mechanics. - Make the state visible. Add the Fast item to the footer with
/statusline, so a session never runs on the premium tier without you seeing it. - Consider a cheaper model at Fast speed. If what you want is quick turnarounds, a cheaper model at Fast speed can cost less than a pricier model at Standard. The credits calculator compares every model on the same task at each speed.
- Watch the trend.
/usageshows daily and weekly activity; a jump with no change in the work is often a Fast toggle nobody switched back.
Questions to ask before switching it on
- Will I watch this output arrive? If not, the speed is wasted, and the multiplier is pure cost.
- Will the session outlive my attention? If it's going to keep working after you've stepped away, switch Fast off before that point, not after.
- Do I usually run out before my window resets? If you do, Fast is costing real credits at one remove. If you don't, it's spending allowance you'd otherwise leave unused.
- Would a cheaper model at Fast speed do the job? Often it would, at a lower combined cost than a more capable model at either speed.
Answering those honestly turns Fast from a setting into a decision, which is all the steeper multiplier asks of you.
The same logic, one tier up
Ultrafast, available in Codex on the top Pro tier and eligible Enterprise and Edu plans, widens every gap described here. It also uses included usage first and credits after, at steeper multiples than Fast on both. On the plans that have it, it's a tool for specific moments when you're blocked on GPT-6 Astra's output — not a setting to leave on. Choosing Standard, Fast or Ultrafast walks through when each tier pays for itself.
The takeaway
Fast isn't overpriced, but inside a plan it's priced twice: once in speed, once in how quickly it spends the allowance you've already paid for. Know which of those two prices you're paying, and switch it on when the wait is the thing you're buying.
Verified 2026-10-01 against CodexHow facts module (src/data/facts/) — see /about/#accuracy.