CodexHowSupport Us

Model Picker

How this shortlist was built

  • 10 priced model(s) on the roster to start.
  • Context window ≥ 200,000 tokens: removed 0.
  • Remaining 10 model(s) ordered by cost fit.
  1. 1. GPT-5.6 Luna1,050,000 tok window · $0.2/MTok in
  2. 2. GPT-5.4 Nano400,000 tok window · $0.2/MTok in
  3. 3. GPT-5.4 Mini400,000 tok window · $0.75/MTok in
  4. 4. GPT-5.3 Codex400,000 tok window · $1.75/MTok in
  5. 5. GPT-5.6 Terra1,050,000 tok window · $2/MTok in
  6. 6. GPT-5.41,050,000 tok window · $2.5/MTok in
  7. 7. GPT-5.6 Sol1,050,000 tok window · $5/MTok in
  8. 8. GPT-5.51,050,000 tok window · $5/MTok in
  9. 9. GPT-5.5 Pro1,050,000 tok window · $30/MTok in
  10. 10. GPT-5.4 Pro1,050,000 tok window · $30/MTok in

Two entries in the roster carry no published price or spec at all and are excluded from every shortlist, not guessed at.

Get an alert when a price on this page changes.

Every model comparison page on this site answers "how do these two differ." This tool answers a different question — "which one should I actually use" — by walking through what your task actually needs rather than asking you to already know which axes matter.

Why a picker instead of another table

A spec table is the right tool once you already know which two or three models are in contention. It's a poor tool for the more common starting position: staring at a lineup of ten priced models and not knowing which handful are even worth comparing closely. This picker exists for that earlier moment — a short set of questions about your actual task, translated into a shortlist, rather than a wall of numbers you have to cross-reference yourself.

The questions that actually move the answer

How large is the context your task needs to hold — a single file, a whole repository, a long multi-turn agentic session? That alone rules out the smaller-window models for anything genuinely large, before price even enters the conversation. Does the task benefit from very recent training data, or is it working against stable, well-established APIs and libraries where recency barely matters? That's the axis most people underweight, and it's a real, publishable difference between model generations, not a vague "newer is better" instinct. Is this latency-sensitive, cost-sensitive, or throughput-sensitive — because those three pull toward different tiers and sometimes different models entirely, and a task can only really optimize hard for one of them at a time. And does the task want explicit control over reasoning effort, a control published for exactly one model on this site's roster and not stated for any other?

How the recommendation is built

Nothing here is a hidden scoring formula you have to trust blindly — the picker filters the published model roster against your answers in a visible, checkable way: a context-window requirement removes anything smaller, a recency requirement removes anything with an older knowledge cutoff, a stated reasoning-effort requirement narrows to the one model that publishes those levels. What's left after filtering is your shortlist, typically two or three models, not a single forced answer — because for most real tasks, more than one model is a defensible choice, and the honest output of a filtering process is a shortlist, not a false certainty.

Why the CLI's own default isn't the automatic answer

The Codex CLI ships with a specific default model already selected, and it's a reasonable one for a lot of work — but "what the CLI does if you don't choose" and "what your specific task actually needs" are two different questions, and this site's whole premise is that no one page states the tradeoff between them plainly. If your answers point toward a different model than the CLI's default, that's not the tool disagreeing with OpenAI — it's your task's actual shape pointing somewhere specific, which is exactly the kind of judgment a generic default can't make for you.

What this tool won't tell you

It won't tell you which model produces better code for your specific codebase — that's a quality judgment this site has no reliable, sourced way to make, and any claim to the contrary would be exactly the kind of unverifiable assertion this site refuses to publish. What it will tell you, honestly and from published specs alone, is which models are even structurally suited to your task's context size, recency needs and cost profile — narrowing a ten-model lineup down to the few worth actually trying is most of the real work, and the quality call from there is yours to make with real output in front of you.

Revisiting your answer

Model lineups change — a new release, a repriced tier, a context window that grows on a future generation. Answers you get here are a snapshot against the currently published roster, and it's worth re-running the picker after any release this site's own weekly fact-check catches, rather than assuming a shortlist from months ago is still the right one.

A note on the models that don't answer at all

Two entries in the roster carry no published price and no published spec at all — their own detail pages return nothing, and every price cell in the pricing table is empty. The picker excludes them from every shortlist rather than guessing at where they'd land, because a recommendation built on an unpublished figure isn't really a recommendation, it's a placeholder dressed up as one, and this site draws a hard line against shipping those.

Verified 2026-08-09 against CodexHow facts module (src/data/facts/) — see /about/#accuracy.