CodexHowSupport Us

429 slow_down While Inside Your Rate Limits

The mistake

Seeing a 429 and reading it as "we've hit our rate limit" — then upgrading a usage tier, cutting requests per minute, or splitting traffic across keys, when the traffic was comfortably inside its RPM and TPM ceilings all along.

Why this happens

For years a 429 meant one thing to most developers: too many requests or tokens in a window. OpenAI now uses it for a second condition. A 429 carrying the code slow_down means your traffic increased too quickly — a ramp problem, not a ceiling problem — and it can fire even while you're within your published limits. The status code is the same; only the error code tells the two apart, and code that branches on the status alone treats them identically.

Why it matters

The responses to the two conditions point in opposite directions. A real rate-limit 429 is solved by more headroom: a higher tier, fewer tokens, spreading load. A slow_down 429 is solved by growing more gradually; a higher usage tier does nothing for it. A team that pays to reach the next tier to fix a ramp problem gets the same errors on the other side. On Fast mode, a ramp that's too steep has a quieter cost as well: some requests are downgraded to standard speed, billed at standard rates, and come back marked service_tier: "default".

The fix

Check error.code, not just the HTTP status — other errors share both status codes — and handle slow_down separately: follow Retry-After when it's present, reduce your request rate, then build back up gradually. OpenAI's rule of thumb: once traffic reaches 1,000,000 input tokens per minute, increase it by no more than 50% every 15 minutes. Shift traffic with feature flags over hours rather than switching instantly, and ramp gradually when moving to a new model or snapshot — a migration that flips all traffic at once is the classic trigger. When Retry-After is missing, increase the delay between retries and add a small random delay so workers don't retry in lockstep.

See also

Rate-limit tiers and headers covers the per-model ceilings and the headers that report them, and the 503 server_is_overloaded error is the other new error code that's easy to confuse with this one.

Verified 2026-10-01 against CodexHow facts module (src/data/facts/) — see /about/#accuracy.