GPT-5.4 Pro — Specs, Pricing & Limits
Specs
| Status | Priced |
|---|---|
| Context window | 1,050,000 tokens |
| Max output | 128,000 tokens |
| Knowledge cutoff | Aug 31, 2025 |
| Reasoning effort levels | Not published |
| Long-context rule | >272,000 input tokens: 2x input, 1.5x output (full request) |
Pricing (short context)
| Standard — input | $30.00/MTok |
|---|---|
| Standard — cached input | Not published |
| Standard — cache writes | Not published |
| Standard — output | $180.00/MTok |
| Fast mode | Not offered for this model |
Rate limits by usage tier
| Free | Not supported |
|---|---|
| Tier 1 | 500 RPM · 30,000 TPM · 90,000 batch queue |
| Tier 2 | 5,000 RPM · 450,000 TPM · 1,350,000 batch queue |
| Tier 3 | 5,000 RPM · 800,000 TPM · 50,000,000 batch queue |
| Tier 4 | 10,000 RPM · 2,000,000 TPM · 200,000,000 batch queue |
| Tier 5 | 10,000 RPM · 30,000,000 TPM · 5,000,000,000 batch queue |
GPT-5.4 Pro carries the largest sticker price on this site's roster, and the single most important thing to know about it isn't a price at all — it's that its entry-tier token-per-minute ceiling doesn't scale with its context window the way you'd assume from the price alone.
The window itself matches the current generation's frontier models. The entry-tier throughput ceiling doesn't: a single maximum-length prompt against this model can consume roughly a whole minute's entry-tier token budget more than thirty times over, a mismatch that isn't visible from the window or the price individually, only from checking both published numbers side by side. Anyone planning to use this model's window at anything close to its real size needs to plan the account tier alongside it, not after hitting a throttled request in production.
It publishes the long-context repricing rule the same way its GPT-5.6 counterparts do, and no Fast mode at all — at this price point, Fast's latency premium on top of an already-premium rate apparently isn't offered as an option.
The table below is computed directly from this site's facts module: pricing across every published service tier, the long-context rule, and the rate-limit table — read the rate-limit rows closely before assuming this model's price is the only planning constraint that matters.
Verified 2026-08-09 against https://developers.openai.com/api/docs/models/gpt-5.4-pro.
Checked: https://developers.openai.com/api/docs/models/gpt-5.4-pro · https://developers.openai.com/api/docs/pricing