Choosing Standard, Fast or Ultrafast
OpenAI now sells the same model at more than one speed, and the price of going faster isn't a single number. It depends on how you pay: per token on an API key, in credits, or out of a ChatGPT plan's included usage — and each of those multiplies the cost of speed differently. Choosing well starts with knowing what each tier actually buys, then asking whether your workload is the kind that turns faster tokens into time you care about.
What each tier is
Standard is the default: no speed premium, and the baseline every other figure is measured against. Fast mode — the tier that was called Priority until July 30, 2026 — runs supported models faster for a per-token premium. In the API, OpenAI quotes up to 2.5x faster than Standard for GPT-5.6 Sol; inside Codex it states a smaller increase for GPT-5.6 and GPT-5.5, and "faster" without a figure for the GPT-6 models. Fast shares a model's Standard rate limit rather than adding headroom.
Ultrafast is the newest tier, added to the API on September 29, 2026, and it exists only for GPT-6 Astra (with GPT-5.6 Sol in a limited preview through OpenAI's account teams). OpenAI says it generates tokens up to 8x faster than Standard. It comes with its own token limits, separate from Astra's Standard ones, and it supports only US data residency and global processing. The Ultrafast mode reference lists its price and limits.
The price of speed depends on how you pay
On an API key, Fast and Ultrafast are separate rows on the pricing page, each a premium over Standard for the same model; the service tiers reference puts them side by side. Paying with ChatGPT credits, Fast multiplies the credits a task costs by 2x and GPT-6 Astra Ultrafast by 6x. Against a plan's included usage the multipliers are steeper: 2.5x for Fast and 8x for Ultrafast. OpenAI is explicit that these describe billing, not speed.
That asymmetry is the most useful single fact on this page. The same Fast session costs proportionally more of a plan's allowance than of a credit balance, so a plan user who reaches for Fast by habit runs out of included usage noticeably sooner than the speed-up alone would suggest. Why Fast mode burns your Codex plan faster works through the arithmetic, and the Codex credits calculator prices a specific task at each speed.
Faster tokens aren't the same as a faster task
OpenAI measures Ultrafast's gain in token generation speed, and says plainly that the comparison doesn't describe overall task completion time. An agentic Codex task spends much of its wall-clock time outside generation: running tests, waiting on tools, reading files, and making round trips between them. Speeding up generation shrinks only that slice. For tool-heavy agents OpenAI strongly recommends WebSockets, because without a persistent connection, network overhead between requests can eat into the latency gain. If your task is mostly waiting on a test suite, a faster tier mostly buys a more expensive wait.
When Fast earns its premium
Fast fits work where a person is watching the output arrive: an interactive Codex session you're steering turn by turn, or a user-facing application with steady traffic where latency is part of the product. It fits badly where nobody is waiting. OpenAI's own guidance is to avoid running large ETL or batch jobs in Fast mode, partly because a fast ramp in traffic can get some Fast requests downgraded to standard speed — billed at standard rates, and reported with service_tier: "default" in the response — which makes the bill harder to predict without making the job meaningfully faster.
When Ultrafast is worth it
Ultrafast is a narrow tool. In Codex it's available only on the top Pro tier and on eligible Enterprise and Edu plans, where it draws on included usage first and credits after; buying credits on another self-serve plan doesn't unlock it. In the API every customer can use it at low default rate limits. It makes sense where the latency of generation itself is the bottleneck and the work is valuable enough to absorb Astra's highest price row — a long, interactive Astra session where you're blocked on its output, say. For most Codex work, GPT-6.1 Sol at Standard or Fast speed is the cheaper way to get a quick answer, because it starts from a much lower Standard price.
When Standard is the right answer
Anything unattended belongs on Standard or below: CI runs, scheduled jobs, overnight refactors, anything you'll read the result of later. If the work is genuinely asynchronous, the cheaper tiers are worth a look too — Batch and Flex both cost less than Standard, for different reasons laid out in when Batch is worth the wait. Paying for speed on a job nobody is waiting for is the most common way speed tiers waste money.
Making it a per-session decision
In the CLI, /fast toggles Fast for the current model and remembers the choice, so the cheapest habit is to leave Standard as your default and switch Fast on only for sessions where you're actively waiting — see toggling Fast mode with /fast. On the API side, beware the project-level setting that makes Fast the default for every request that doesn't name a tier: it moves unrelated jobs onto the premium tier without a code change.
Checking what you actually paid for
The tier you requested isn't always the tier that ran. Read the service_tier field on responses, and group usage by service tier in the usage dashboard. On GPT-5.6 and earlier models, Fast requests report priority rather than fast, so a check that filters on fast alone will undercount them.
Verified 2026-10-01 against CodexHow facts module (src/data/facts/) — see /about/#accuracy.