Ultrafast Mode: Price, Limits, Who Gets It
GPT-6 Astra by speed tier, per 1M tokens (short context unless stated)
| Standard | $10.00 in · $1.00 cached · $12.50 writes · $50.00 out |
|---|---|
| Fast | $20.00 in · $2.00 cached · $25.00 writes · $100.00 out — 2.0× Standard |
| Ultrafast | $60.00 in · $6.00 cached · $75.00 writes · $300.00 out — 6.0× Standard |
| Ultrafast, long context | $120.00 in · $450.00 out (past the long-context threshold, whole request) |
Ultrafast in the API
| Added to the API | 2026-09-29 |
|---|---|
| Models | GPT-6 Astra; preview only for GPT-5.6 Sol |
| Speed | Up to 8× faster than Standard (token generation, not whole-task time) |
| Request | model "gpt-6-astra" with service_tier "ultrafast" — WebSockets recommended for tool-heavy agents |
| Residency | US data residency and global processing only |
| Default limit, Tiers 1–3 | 500,000 TPM |
| Default limit, Tier 4 | 1,000,000 TPM |
| Default limit, Tier 5 | 5,000,000 TPM |
Ultrafast inside Codex and ChatGPT Work
| Plans | The top Pro tier and eligible Enterprise and Edu plans; off by default for Enterprise workspaces |
|---|---|
| Usage | 8× Standard against included usage, 6× against purchased credits — included usage first, then credits |
| Not unlocked by | Buying credits on a lower Pro tier, or on any other self-serve plan |
Ultrafast is OpenAI's fastest service tier, added to the API on September 29, 2026. It runs GPT-6 Astra — and only GPT-6 Astra, apart from a limited preview for GPT-5.6 Sol that goes through OpenAI's account teams — at up to 8x the token-generation speed of Standard. The tables above give its price against Standard and Fast, its default limits, and how it works inside Codex.
What the speed figure measures
OpenAI's comparison is about how quickly tokens are generated, not how long a task takes. An agentic task spends time on tool calls, tests and network round trips that a faster tier doesn't touch, so the end-to-end gain is smaller than the headline. For agents that make many tool calls in quick succession, OpenAI strongly recommends WebSockets: without a persistent connection, per-request network overhead can erase much of the latency gain. HTTP requests work too.
Price and limits
Every Ultrafast price cell is a multiple of Astra's Standard rate, and well above Fast — the first table computes the multiple. Ultrafast has its own token limits, separate from Astra's Standard table and lower at every usage tier above the first: OpenAI describes it as available to all API users at low rate limits, with higher limits by request through an account team. It supports US data residency and global processing only; EU and other regional processing endpoints aren't supported.
Inside Codex
In Codex and ChatGPT Work, Ultrafast is available on the top Pro tier and on eligible Enterprise and Edu plans. It uses included usage first, then credits, at 8x the Standard rate against included usage and 6x against credits. Enterprise workspaces have it off by default until an owner grants access, workspaces that require inference residency outside the United States can't use it, and with an API key Codex simply uses the API's Ultrafast prices.
When it's worth it
When you're actively waiting on Astra's output and generation time is the bottleneck. Choosing Standard, Fast or Ultrafast works through the decision, including the many cases where a cheaper model at Fast speed is the better buy.
Verified 2026-10-01 against https://developers.openai.com/api/docs/guides/ultrafast-mode.
Could not confirm: Any Ultrafast price or limit for GPT-5.6 Sol: it is preview-only through an account team, with no published row. Ultrafast for GPT-6.1 Sol is described only as "coming later".
Checked: https://developers.openai.com/api/docs/guides/ultrafast-mode · https://developers.openai.com/api/docs/pricing · https://learn.chatgpt.com/docs/agent-configuration/speed · https://developers.openai.com/api/docs/changelog