CodexHowSupport Us

Ultrafast Mode: Price, Limits, Who Gets It

GPT-6 Astra by speed tier, per 1M tokens (short context unless stated)

Standard$10.00 in · $1.00 cached · $12.50 writes · $50.00 out
Fast$20.00 in · $2.00 cached · $25.00 writes · $100.00 out — 2.0× Standard
Ultrafast$60.00 in · $6.00 cached · $75.00 writes · $300.00 out — 6.0× Standard
Ultrafast, long context$120.00 in · $450.00 out (past the long-context threshold, whole request)

Ultrafast in the API

Added to the API2026-09-29
ModelsGPT-6 Astra; preview only for GPT-5.6 Sol
SpeedUp to 8× faster than Standard (token generation, not whole-task time)
Requestmodel "gpt-6-astra" with service_tier "ultrafast" — WebSockets recommended for tool-heavy agents
ResidencyUS data residency and global processing only
Default limit, Tiers 1–3500,000 TPM
Default limit, Tier 41,000,000 TPM
Default limit, Tier 55,000,000 TPM

Ultrafast inside Codex and ChatGPT Work

PlansThe top Pro tier and eligible Enterprise and Edu plans; off by default for Enterprise workspaces
Usage8× Standard against included usage, 6× against purchased credits — included usage first, then credits
Not unlocked byBuying credits on a lower Pro tier, or on any other self-serve plan

Ultrafast is OpenAI's fastest service tier, added to the API on September 29, 2026. It runs GPT-6 Astra — and only GPT-6 Astra, apart from a limited preview for GPT-5.6 Sol that goes through OpenAI's account teams — at up to 8x the token-generation speed of Standard. The tables above give its price against Standard and Fast, its default limits, and how it works inside Codex.

What the speed figure measures

OpenAI's comparison is about how quickly tokens are generated, not how long a task takes. An agentic task spends time on tool calls, tests and network round trips that a faster tier doesn't touch, so the end-to-end gain is smaller than the headline. For agents that make many tool calls in quick succession, OpenAI strongly recommends WebSockets: without a persistent connection, per-request network overhead can erase much of the latency gain. HTTP requests work too.

Price and limits

Every Ultrafast price cell is a multiple of Astra's Standard rate, and well above Fast — the first table computes the multiple. Ultrafast has its own token limits, separate from Astra's Standard table and lower at every usage tier above the first: OpenAI describes it as available to all API users at low rate limits, with higher limits by request through an account team. It supports US data residency and global processing only; EU and other regional processing endpoints aren't supported.

Inside Codex

In Codex and ChatGPT Work, Ultrafast is available on the top Pro tier and on eligible Enterprise and Edu plans. It uses included usage first, then credits, at 8x the Standard rate against included usage and 6x against credits. Enterprise workspaces have it off by default until an owner grants access, workspaces that require inference residency outside the United States can't use it, and with an API key Codex simply uses the API's Ultrafast prices.

When it's worth it

When you're actively waiting on Astra's output and generation time is the bottleneck. Choosing Standard, Fast or Ultrafast works through the decision, including the many cases where a cheaper model at Fast speed is the better buy.

Verified 2026-10-01 against https://developers.openai.com/api/docs/guides/ultrafast-mode.

Could not confirm: Any Ultrafast price or limit for GPT-5.6 Sol: it is preview-only through an account team, with no published row. Ultrafast for GPT-6.1 Sol is described only as "coming later".

Checked: https://developers.openai.com/api/docs/guides/ultrafast-mode · https://developers.openai.com/api/docs/pricing · https://learn.chatgpt.com/docs/agent-configuration/speed · https://developers.openai.com/api/docs/changelog