GPT-6 Luna — Specs, Pricing & Limits
Specs
| Status | Priced |
|---|---|
| Released (API changelog) | 2026-09-22 |
| Context window | 1,050,000 tokens |
| Max output | 128,000 tokens |
| Knowledge cutoff | May 18, 2026 |
| Reasoning effort levels | none, low, medium (default), high, xhigh, max |
| Long-context rule | >272,000 input tokens: 2x input and cache rates, 1.5x output (full request) |
| Cache writes | Billed at the cache-write rate instead of the input rate |
| ChatGPT credit rate (per 1M) | 2.5 in · 0.25 cached · 12.5 out |
Pricing per 1M tokens, every published service tier
| Standard — short context | in $0.10 · cached $0.01 · writes $0.125 · out $0.50 |
|---|---|
| Standard — long context | in $0.20 · cached $0.02 · writes $0.25 · out $0.75 |
| Batch — short context | in $0.05 · cached $0.005 · writes $0.0625 · out $0.25 |
| Batch — long context | in $0.10 · cached $0.01 · writes $0.125 · out $0.375 |
| Flex — short context | in $0.05 · cached $0.005 · writes $0.0625 · out $0.25 |
| Flex — long context | in $0.10 · cached $0.01 · writes $0.125 · out $0.375 |
| Fast — short context | in $0.20 · cached $0.02 · writes $0.25 · out $1.00 |
| Fast — long context | in $0.40 · cached $0.04 · writes $0.50 · out $1.50 |
Rate limits by usage tier
| Free | Not supported |
|---|---|
| Tier 1 | 500 RPM · 500,000 TPM · 5,000,000 batch queue |
| Tier 2 | 5,000 RPM · 2,000,000 TPM · 20,000,000 batch queue |
| Tier 3 | 5,000 RPM · 4,000,000 TPM · 40,000,000 batch queue |
| Tier 4 | 10,000 RPM · 10,000,000 TPM · 1,000,000,000 batch queue |
| Tier 5 | 30,000 RPM · 180,000,000 TPM · 15,000,000,000 batch queue |
GPT-6 Luna is the cheapest priced model on this site and has the most recent knowledge cutoff of any of them — an unusual combination, since the low-cost tier usually trails on recency. OpenAI describes it as its most efficient model for focused, high-volume tasks: extraction, classification, transformation, structured summaries, and focused coding where you know what a good result looks like.
It keeps the full context window and output ceiling of the larger GPT-6 models, so choosing it is a capability decision rather than a capacity one. It accepts every effort level from none to max, with medium as the default; in Codex, OpenAI suggests starting it at High, and it supports Max but not Ultra there.
Its rate limits come from a different table than the Sol models'. Luna sits in the same high-throughput group as GPT-5.6 Luna and GPT-5.4 Mini, with much higher request and token ceilings at the top usage tiers — the reason high-volume pipelines land here, not just the price. Fast mode exists for it, though not with EU data residency.
In Codex it has a special role: it's the model Free and Go plans get, in the desktop app, and OpenAI's named replacement for GPT-5.4 Mini, which left Codex with ChatGPT sign-in on August 31, 2026. In Enterprise and Edu workspaces an administrator has to enable it first. GPT-6 Luna vs GPT-5.6 Luna compares it with its predecessor.
The tables above are computed from this site's facts module: every published service tier with both context bands, the effort levels, and Luna's own rate-limit table.
Verified 2026-10-01 against https://developers.openai.com/api/docs/models/gpt-6-luna.
Checked: https://developers.openai.com/api/docs/models/gpt-6-luna · https://developers.openai.com/api/docs/pricing · https://developers.openai.com/api/docs/changelog · https://learn.chatgpt.com/docs/pricing