CodexHowSupport Us

GPT-6 Astra — Specs, Pricing & Limits

Specs

StatusPriced
Released (API changelog)2026-09-03
Context window1,050,000 tokens
Max output128,000 tokens
Knowledge cutoffApr 30, 2026
Reasoning effort levelslow, medium, high, xhigh, max
Long-context rule>272,000 input tokens: 2x input and cache rates, 1.5x output (full request)
Cache writesBilled at the cache-write rate instead of the input rate
ChatGPT credit rate (per 1M)250 in · 25 cached · 1,250 out

Pricing per 1M tokens, every published service tier

Standard — short contextin $10.00 · cached $1.00 · writes $12.50 · out $50.00
Standard — long contextin $20.00 · cached $2.00 · writes $25.00 · out $75.00
Batch — short contextin $5.00 · cached $0.50 · writes $6.25 · out $25.00
Batch — long contextin $10.00 · cached $1.00 · writes $12.50 · out $37.50
Flex — short contextin $5.00 · cached $0.50 · writes $6.25 · out $25.00
Flex — long contextin $10.00 · cached $1.00 · writes $12.50 · out $37.50
Fast — short contextin $20.00 · cached $2.00 · writes $25.00 · out $100.00
Fast — long contextin $40.00 · cached $4.00 · writes $50.00 · out $150.00
Ultrafast — short contextin $60.00 · cached $6.00 · writes $75.00 · out $300.00
Ultrafast — long contextin $120.00 · cached $12.00 · writes $150.00 · out $450.00

Rate limits by usage tier

FreeNot supported
Tier 1500 RPM · 500,000 TPM · 1,500,000 batch queue
Tier 25,000 RPM · 1,000,000 TPM · 3,000,000 batch queue
Tier 35,000 RPM · 2,000,000 TPM · 100,000,000 batch queue
Tier 410,000 RPM · 4,000,000 TPM · 200,000,000 batch queue
Tier 515,000 RPM · 40,000,000 TPM · 15,000,000,000 batch queue

Ultrafast token limits (separate from the table above)

Tiers 1–3500,000 TPM
Tier 41,000,000 TPM
Tier 55,000,000 TPM

GPT-6 Astra is OpenAI's most capable model and the only one on this site with an Ultrafast tier. It's built for the hardest end-to-end work — long coding tasks, computer use, research — and its per-token prices sit well above the Sol models. OpenAI's own argument for it is per task rather than per token: in several of its evaluations, Astra reached stronger results with substantially fewer output tokens, and its estimated cost per task came out lower than earlier models despite the higher rates. That's worth testing on your own work before accepting either way.

Moving a request to Astra takes more than a new model name. It doesn't accept the none reasoning effort, custom temperature or top_p, or log probabilities, and tool calling requires the Responses API. Its page lists effort levels from low to max but marks none of them as the default, so this site names no default; in Codex, OpenAI suggests starting Astra at its lightest setting. Behavior differs too: OpenAI says Astra is more likely than earlier models to stop and ask when more input could change the result, and more sensitive to instructions in AGENTS.md and skill files — see when Codex should ask instead of guessing.

Ultrafast has its own token limits, separate from the Standard table, and supports only US data residency and global processing. In Codex, Astra arrived switched off for Enterprise and Edu workspaces until an owner enables it, and Ultrafast is limited to the top Pro tier and eligible Enterprise and Edu plans. The Ultrafast reference covers both sides.

The tables above are computed from this site's facts module: every published service tier with both context bands, the effort levels, the Standard rate limits and Astra's separate Ultrafast limits.

Verified 2026-10-01 against https://developers.openai.com/api/docs/models/gpt-6-astra.

Could not confirm: The model page lists five effort levels but marks none of them as the API default, so this site names no default for GPT-6 Astra.

Checked: https://developers.openai.com/api/docs/models/gpt-6-astra · https://developers.openai.com/api/docs/pricing · https://developers.openai.com/api/docs/changelog · https://developers.openai.com/api/docs/guides/ultrafast-mode