GPT-5.4 Nano — Specs, Pricing & Limits
Specs
| Status | Priced |
|---|---|
| Context window | 400,000 tokens |
| Max output | 128,000 tokens |
| Knowledge cutoff | Aug 31, 2025 |
| Reasoning effort levels | Not published |
| Long-context rule | Not published |
Pricing (short context)
| Standard — input | $0.200/MTok |
|---|---|
| Standard — cached input | $0.020/MTok |
| Standard — cache writes | Not published |
| Standard — output | $1.25/MTok |
| Fast mode | Not offered for this model |
Rate limits by usage tier
| Free | Not supported |
|---|---|
| Tier 1 | 500 RPM · 200,000 TPM · 2,000,000 batch queue |
| Tier 2 | 5,000 RPM · 2,000,000 TPM · 20,000,000 batch queue |
| Tier 3 | 5,000 RPM · 4,000,000 TPM · 40,000,000 batch queue |
| Tier 4 | 10,000 RPM · 10,000,000 TPM · 1,000,000,000 batch queue |
| Tier 5 | 30,000 RPM · 180,000,000 TPM · 15,000,000,000 batch queue |
GPT-5.4 Nano is the cheapest model this site prices, on every tier it publishes — and cheapest doesn't mean "smallest version of a bigger number" here; it comes with real, specific constraints worth checking before defaulting to it purely on price.
No Fast mode is offered for this model at all — of the models that skip Fast entirely, this is the one at the opposite end of the price spectrum from the other members of that group, which is worth knowing if your mental model of "which models skip Fast" assumes it's only the largest, most expensive ones. Its entry-tier token-per-minute ceiling is also under half of what most other current models get at the same tier, which matters more for a high-volume batch-style workload than for occasional low-volume calls.
No long-context repricing rule is published for this model, consistent with its context window sitting well under the threshold that makes the rule relevant in the first place — checked directly against its own page rather than assumed.
For a workload that's genuinely simple per-call and running at real volume — classification, short extraction, lightweight formatting — this is worth checking first specifically because the per-token savings compound fast at volume, provided the entry-tier throughput ceiling doesn't become the actual bottleneck before the price does.
The table below is computed directly from this site's facts module: pricing across every published service tier and the rate-limit table.
Verified 2026-08-09 against https://developers.openai.com/api/docs/models/gpt-5.4-nano.
Could not confirm: No >272K long-context repricing rule is stated on this model's own page.
Checked: https://developers.openai.com/api/docs/models/gpt-5.4-nano · https://developers.openai.com/api/docs/pricing