GPT-5.6 Luna — Specs, Pricing & Limits
Specs
| Status | Priced |
|---|---|
| Context window | 1,050,000 tokens |
| Max output | 128,000 tokens |
| Knowledge cutoff | Feb 16, 2026 |
| Reasoning effort levels | Not published |
| Long-context rule | >272,000 input tokens: 2x input, 1.5x output (full request) |
Pricing (short context)
| Standard — input | $0.200/MTok |
|---|---|
| Standard — cached input | $0.020/MTok |
| Standard — cache writes | $0.250/MTok |
| Standard — output | $1.20/MTok |
| Fast — input | $0.400/MTok |
| Fast — output | $2.40/MTok |
Rate limits by usage tier
| Free | Not supported |
|---|---|
| Tier 1 | 500 RPM · 500,000 TPM · 5,000,000 batch queue |
| Tier 2 | 5,000 RPM · 2,000,000 TPM · 20,000,000 batch queue |
| Tier 3 | 5,000 RPM · 4,000,000 TPM · 40,000,000 batch queue |
| Tier 4 | 10,000 RPM · 10,000,000 TPM · 1,000,000,000 batch queue |
| Tier 5 | 30,000 RPM · 180,000,000 TPM · 15,000,000,000 batch queue |
GPT-5.6 Luna is the cheapest model in its family, and the one detail worth knowing before assuming "cheapest" means "most restricted" in every sense: Luna sits in a genuinely different rate-limit group from Sol and Terra, and the difference runs the opposite direction you'd guess.
Context window and maximum output length match the rest of the GPT-5.6 family exactly — Luna doesn't trade window size away for its lower price, the same way Terra doesn't. Where it actually differs from its siblings is throughput: Luna's batch-queue ceiling and its top-tier requests-per-minute and tokens-per-minute limits are substantially higher than Sol's or Terra's, not lower, at every tier where the two groups diverge. That's counterintuitive if you assume price and throughput move together — here, the cheapest model in the family is also the one with the most generous ceiling once you're spending enough to reach the higher usage tiers.
That makes Luna worth a serious look for any high-volume, latency-tolerant workload — bulk classification, large-scale extraction, anything running enough volume that the rate-limit ceiling itself becomes the binding constraint rather than the per-token price. It's a poor fit for anything that specifically needs Sol or Terra's reasoning quality on a hard task; Luna's position in the family is about price and throughput, not about matching the larger models' capability.
The published long-context repricing rule applies to Luna the same way it applies to the rest of the GPT-5.6 family — crossing the threshold reprices the full request, not just the portion above the line.
The table below is computed directly from this site's facts module: pricing across every published service tier, the long-context rule, and Luna's own rate-limit table.
Verified 2026-08-09 against https://developers.openai.com/api/docs/models/gpt-5.6-luna.
Checked: https://developers.openai.com/api/docs/models/gpt-5.6-luna · https://developers.openai.com/api/docs/pricing