503 server_is_overloaded: Retry, Don't Throttle
The mistake
Treating a 503 with the code server_is_overloaded as a signal that your own traffic is too heavy — and responding by throttling, shedding load or redesigning a pipeline that wasn't the problem.
Why this happens
Errors that arrive during a busy period look like they were caused by the busy period. OpenAI now separates the two situations explicitly: traffic that grows too fast gets a 429 with slow_down, while a 503 with server_is_overloaded means the model itself is temporarily overloaded. The second one isn't about your request rate at all, but it tends to show up at the same moments of peak demand, which makes the wrong diagnosis feel right. Older error handlers add to the confusion: on some endpoints a 503 with slow_down used to cover both conditions, so code written against that behavior may still be treating every 503 as a ramp problem.
Why it matters
Cutting your request rate in response to a 503 doesn't make the model less overloaded; it just makes your job slower. The opposite failure is worse: retrying immediately and repeatedly, from many workers at once, adds load to the exact moment the model is already short of capacity, and turns a brief overload into a long run of failures for your own pipeline. And a 503 isn't a quota or billing problem, so it won't be fixed by a tier upgrade or a top-up either.
The fix
Retry, but politely. If the response includes Retry-After, wait at least that long before the next attempt; if it doesn't, use exponential backoff, ideally with jitter so many workers don't retry in lockstep. Keep the retry budget bounded so a sustained outage surfaces as an alert rather than an endless loop. For streaming requests, these errors arrive before the stream starts; an error after output has begun can come as a stream event instead, and OpenAI advises against automatically replaying a request whose output you've already consumed. Log the error code separately from slow_down 429s, because the two call for different responses. If the work isn't latency-sensitive, it may belong on Batch, which completes within its window without your client having to manage retries call by call.
See also
The 429 slow_down error covers the ramp-rate case, and rate-limit tiers and headers explains Retry-After alongside the other response headers.
Verified 2026-10-01 against CodexHow facts module (src/data/facts/) — see /about/#accuracy.