CodexHowSupport Us

GPT-6 Luna — Specs, Pricing & Limits

Specs

StatusPriced
Released (API changelog)2026-09-22
Context window1,050,000 tokens
Max output128,000 tokens
Knowledge cutoffMay 18, 2026
Reasoning effort levelsnone, low, medium (default), high, xhigh, max
Long-context rule>272,000 input tokens: 2x input and cache rates, 1.5x output (full request)
Cache writesBilled at the cache-write rate instead of the input rate
ChatGPT credit rate (per 1M)2.5 in · 0.25 cached · 12.5 out

Pricing per 1M tokens, every published service tier

Standard — short contextin $0.10 · cached $0.01 · writes $0.125 · out $0.50
Standard — long contextin $0.20 · cached $0.02 · writes $0.25 · out $0.75
Batch — short contextin $0.05 · cached $0.005 · writes $0.0625 · out $0.25
Batch — long contextin $0.10 · cached $0.01 · writes $0.125 · out $0.375
Flex — short contextin $0.05 · cached $0.005 · writes $0.0625 · out $0.25
Flex — long contextin $0.10 · cached $0.01 · writes $0.125 · out $0.375
Fast — short contextin $0.20 · cached $0.02 · writes $0.25 · out $1.00
Fast — long contextin $0.40 · cached $0.04 · writes $0.50 · out $1.50

Rate limits by usage tier

FreeNot supported
Tier 1500 RPM · 500,000 TPM · 5,000,000 batch queue
Tier 25,000 RPM · 2,000,000 TPM · 20,000,000 batch queue
Tier 35,000 RPM · 4,000,000 TPM · 40,000,000 batch queue
Tier 410,000 RPM · 10,000,000 TPM · 1,000,000,000 batch queue
Tier 530,000 RPM · 180,000,000 TPM · 15,000,000,000 batch queue

GPT-6 Luna is the cheapest priced model on this site and has the most recent knowledge cutoff of any of them — an unusual combination, since the low-cost tier usually trails on recency. OpenAI describes it as its most efficient model for focused, high-volume tasks: extraction, classification, transformation, structured summaries, and focused coding where you know what a good result looks like.

It keeps the full context window and output ceiling of the larger GPT-6 models, so choosing it is a capability decision rather than a capacity one. It accepts every effort level from none to max, with medium as the default; in Codex, OpenAI suggests starting it at High, and it supports Max but not Ultra there.

Its rate limits come from a different table than the Sol models'. Luna sits in the same high-throughput group as GPT-5.6 Luna and GPT-5.4 Mini, with much higher request and token ceilings at the top usage tiers — the reason high-volume pipelines land here, not just the price. Fast mode exists for it, though not with EU data residency.

In Codex it has a special role: it's the model Free and Go plans get, in the desktop app, and OpenAI's named replacement for GPT-5.4 Mini, which left Codex with ChatGPT sign-in on August 31, 2026. In Enterprise and Edu workspaces an administrator has to enable it first. GPT-6 Luna vs GPT-5.6 Luna compares it with its predecessor.

The tables above are computed from this site's facts module: every published service tier with both context bands, the effort levels, and Luna's own rate-limit table.

Verified 2026-10-01 against https://developers.openai.com/api/docs/models/gpt-6-luna.

Checked: https://developers.openai.com/api/docs/models/gpt-6-luna · https://developers.openai.com/api/docs/pricing · https://developers.openai.com/api/docs/changelog · https://learn.chatgpt.com/docs/pricing