SpyBara
Go Premium

Documentation 2026-09-04 23:59 UTC to 2026-09-05 17:01 UTC

2 files changed +22 −7. View all changes and history on the product overview
2026
Tue 29 22:57 Mon 28 22:57 Sat 26 23:59 Fri 25 23:58 Thu 24 23:58 Wed 23 23:58 Tue 22 23:57 Mon 21 23:00 Sat 19 23:00 Fri 18 22:59 Thu 17 10:04 Wed 16 20:58 Tue 15 22:59 Mon 14 22:58 Sun 13 15:02 Fri 11 20:00 Thu 10 18:01 Wed 9 23:59 Sat 5 17:01 Fri 4 23:59 Thu 3 23:00 Wed 2 22:59
Details

145To resolve this error, please follow these steps:145To resolve this error, please follow these steps:

146 146 

147- Pace your requests and avoid making unnecessary or redundant calls.147- Pace your requests and avoid making unnecessary or redundant calls.

148- If a `Retry-After` header is present, wait at least as long as it specifies before trying again. If it's missing, use exponential backoff with jitter and limit the number of retries. Each official SDK already honors this header for eligible retries. Read more in our [rate limit guide](https://developers.openai.com/api/docs/guides/rate-limits).148- If a `Retry-After` header is present, wait at least as long as it specifies before trying again. If it's missing, use exponential backoff with jitter and limit the number of retries. SDK support for long server delays varies by version and configuration. Read more in our [rate limit guide](https://developers.openai.com/api/docs/guides/rate-limits#retrying-with-exponential-backoff).

149- If you are sharing your organization with other users, note that limits are applied per organization and not per user. It is worth checking on the usage of the rest of your team as this will contribute to the limit.149- If you are sharing your organization with other users, note that limits are applied per organization and not per user. It is worth checking on the usage of the rest of your team as this will contribute to the limit.

150- If you are using a free or low-tier plan, consider upgrading to a pay-as-you-go plan that offers a higher rate limit. You can compare the restrictions of each plan in our [rate limit guide](https://developers.openai.com/api/docs/guides/rate-limits).150- If you are using a free or low-tier plan, consider upgrading to a pay-as-you-go plan that offers a higher rate limit. You can compare the restrictions of each plan in our [rate limit guide](https://developers.openai.com/api/docs/guides/rate-limits).

151- Reach out to your organization owner to increase the rate limits on your project151- Reach out to your organization owner to increase the rate limits on your project


229 229 

230## Python library error types230## Python library error types

231 231 

232Python raises `RateLimitError` for `429` responses and `InternalServerError` for `503` responses. If your handler previously caught only one of these classes for throttling and overload, handle both and inspect `error.code`. Video overload, for example, now returns `503` where it previously returned `429`. See [migration guidance](https://developers.openai.com/api/docs/guides/rate-limits#update-existing-error-handlers) for the endpoint-specific changes.

233 

232| Type | Overview |234| Type | Overview |

233| ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |235| ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |

234| APIConnectionError | **Cause:** Issue connecting to our services. <br /> **Solution:** Check your network settings, proxy configuration, SSL certificates, or firewall rules. |236| APIConnectionError | **Cause:** Issue connecting to our services. <br /> **Solution:** Check your network settings, proxy configuration, SSL certificates, or firewall rules. |


239| InternalServerError | **Cause:** Issue on our side. <br /> **Solution:** Retry your request after a brief wait and contact us if the issue persists. |241| InternalServerError | **Cause:** Issue on our side. <br /> **Solution:** Retry your request after a brief wait and contact us if the issue persists. |

240| NotFoundError | **Cause:** Requested resource does not exist. <br /> **Solution:** Ensure you are the correct resource identifier. |242| NotFoundError | **Cause:** Requested resource does not exist. <br /> **Solution:** Ensure you are the correct resource identifier. |

241| PermissionDeniedError | **Cause:** You don't have access to the requested resource. <br /> **Solution:** Ensure you are using the correct API key, organization ID, and resource ID. |243| PermissionDeniedError | **Cause:** You don't have access to the requested resource. <br /> **Solution:** Ensure you are using the correct API key, organization ID, and resource ID. |

242| RateLimitError | **Cause:** You have hit your assigned rate limit. <br /> **Solution:** Pace your requests and follow `Retry-After` when it's present. Each official SDK already honors this header for eligible retries. Read more in our [Rate limit guide](https://developers.openai.com/api/docs/guides/rate-limits). |244| RateLimitError | **Cause:** You have hit your assigned rate limit or increased traffic too quickly. <br /> **Solution:** Pace your requests and follow `Retry-After` when it's present, subject to your retry limits. Read more in our [Rate limit guide](https://developers.openai.com/api/docs/guides/rate-limits#retrying-with-exponential-backoff). |

243| UnprocessableEntityError | **Cause:** Unable to process the request despite the format being correct. <br /> **Solution:** Please try the request again. |245| UnprocessableEntityError | **Cause:** Unable to process the request despite the format being correct. <br /> **Solution:** Please try the request again. |

244 246 

245 247 


348If you encounter a `RateLimitError`, please try the following steps:350If you encounter a `RateLimitError`, please try the following steps:

349 351 

350- Send fewer tokens or requests or slow down. You may need to reduce the frequency or volume of your requests, batch your tokens, or use exponential backoff when `Retry-After` isn't present. You can read our [Rate limit guide](https://developers.openai.com/api/docs/guides/rate-limits) for more details.352- Send fewer tokens or requests or slow down. You may need to reduce the frequency or volume of your requests, batch your tokens, or use exponential backoff when `Retry-After` isn't present. You can read our [Rate limit guide](https://developers.openai.com/api/docs/guides/rate-limits) for more details.

351- When `Retry-After` is present, wait at least as long as it specifies before retrying. The official Python library already honors this header for eligible retries.353- When `Retry-After` is present, wait at least as long as it specifies before retrying. The Python library can stop automatic retries when a server delay exceeds its supported limit. If you retry at the application level, respect the original delay and account for SDK retries.

352- You can also check your API usage statistics from your account dashboard.354- You can also check your API usage statistics from your account dashboard.

353 355 

354 356 

Details

67| x-ratelimit-remaining-project-tokens | 57000 | The remaining number of tokens permitted before exhausting the project-scoped token rate limit. |67| x-ratelimit-remaining-project-tokens | 57000 | The remaining number of tokens permitted before exhausting the project-scoped token rate limit. |

68| x-ratelimit-reset-project-tokens | 3s | The time until the project-scoped token rate limit resets to its initial state. |68| x-ratelimit-reset-project-tokens | 3s | The time until the project-scoped token rate limit resets to its initial state. |

69 69 

70Project-token headers may be present when a project-scoped token limit applies. `Retry-After` may be present on `429` responses caused by a temporary rate limit. It does not mean that quota, billing, or other errors that require user action can be resolved by retrying.70Project-token headers may be present when a project-scoped token limit applies. `Retry-After` may be present on `429` responses caused by a temporary rate limit and `503` responses caused by temporary model overload. It does not mean that quota, billing, or other errors that require user action can be resolved by retrying.

71 71 

72### Fine-tuning rate limits72### Fine-tuning rate limits

73 73 


98 98 

99Enterprise customers whose pay-as-you-go traffic routinely hits ramp-rate limits can consider [Scale Tier](https://openai.com/api-scale-tier/) for more predictable capacity on eligible models. For GPT-5.6 and later models, see [Reserved Tier](https://openai.com/api-reserved-tier/). Capacity tiers don't change how you should handle a `slow_down` response: follow `Retry-After` when it's present, reduce traffic, and ramp gradually.99Enterprise customers whose pay-as-you-go traffic routinely hits ramp-rate limits can consider [Scale Tier](https://openai.com/api-scale-tier/) for more predictable capacity on eligible models. For GPT-5.6 and later models, see [Reserved Tier](https://openai.com/api-reserved-tier/). Capacity tiers don't change how you should handle a `slow_down` response: follow `Retry-After` when it's present, reduce traffic, and ramp gradually.

100 100 

101#### Update existing error handlers

102 

103If your application handled earlier throttling and overload responses, check both the HTTP status and `error.code`:

104 

105- On endpoints that previously returned `503` with the `slow_down` code for both conditions, rapid traffic increases now return `429` with `slow_down`. Model overload remains `503` but uses `server_is_overloaded`.

106- Video requests rejected before creating a job previously returned `429` with the `invalid_request_error` type and `rate_limit_exceeded` code for these conditions. Rapid traffic increases now return `429` with `rate_limit_error` and `slow_down`; model overload returns `503` with `service_unavailable_error` and `server_is_overloaded`. Errors reported in a video job's status are a separate case.

107 

108Handle both `429` and `503` in your SDK error handlers. For example, Python, TypeScript, and Ruby use `RateLimitError` for `429` and `InternalServerError` for `503`; Java uses `RateLimitException` and `InternalServerException`. Keep support for earlier response codes while your application can still receive them. Other errors can use the same HTTP statuses, so inspect the error body before choosing a recovery action.

109 

110For streaming requests, these HTTP error responses apply before the stream starts. An error after streaming begins can arrive as a stream event; don't automatically replay a request after consuming output.

111 

101### What are some steps I can take to mitigate this?112### What are some steps I can take to mitigate this?

102 113 

103The OpenAI Cookbook has a [Python notebook](https://developers.openai.com/cookbook/examples/how_to_handle_rate_limits) that explains how to avoid rate limit errors, as well an example [Python script](https://github.com/openai/openai-cookbook/blob/main/examples/api_request_parallel_processor.py) for staying under rate limits while batch processing API requests.114The OpenAI Cookbook has a [Python notebook](https://developers.openai.com/cookbook/examples/how_to_handle_rate_limits) that explains how to avoid rate limit errors, as well an example [Python script](https://github.com/openai/openai-cookbook/blob/main/examples/api_request_parallel_processor.py) for staying under rate limits while batch processing API requests.


110 121 

111When a request exceeds a temporary rate limit, the API returns a `429` error. The response can include a `Retry-After` header that tells you how many seconds to wait before trying again. Treat this value as a minimum: wait at least that long and add a small random delay so multiple clients don't retry at the same time.122When a request exceeds a temporary rate limit, the API returns a `429` error. The response can include a `Retry-After` header that tells you how many seconds to wait before trying again. Treat this value as a minimum: wait at least that long and add a small random delay so multiple clients don't retry at the same time.

112 123 

113Each [official OpenAI SDK](https://developers.openai.com/api/docs/libraries#install-an-official-sdk) automatically retries eligible rate-limit errors and honors `Retry-After` when it's present. You don't need to parse the header or add another retry loop for standard API calls.124Each [official OpenAI SDK](https://developers.openai.com/api/docs/libraries#install-an-official-sdk) automatically retries eligible `429` and `503` responses, subject to its retry settings. Handling of `Retry-After`, especially long delays, varies by SDK version and configuration. Check the retry behavior of your installed version rather than assuming every server delay is supported.

125 

126If a valid server delay exceeds the supported or configured maximum retry delay, stop retrying and defer the request rather than retrying sooner. An SDK can return the original HTTP error when it declines a delay above its limit. Continue to handle cancellation and timeout errors separately: a canceled request or expired deadline can stop retries without returning that HTTP error. A timeout for each attempt isn't necessarily a deadline for the entire operation.

114 127 

115If you're using your own HTTP client, follow `Retry-After` when the header is present and contains a valid value. If it's missing or invalid, fall back to exponential backoff with jitter. Limit both the number of attempts and the total time spent retrying. If you add application-level retries, account for the retries your SDK already performs. Don't retry quota, billing, or other errors that require you to take action.128If you're using your own HTTP client, follow `Retry-After` when the header is present and contains a valid value. If it's missing or invalid, fall back to exponential backoff with jitter. Limit both the number of attempts and the total time spent retrying. If you manage retries in your application, disable SDK retries or account for them in those limits so nested retry loops don't multiply requests. Don't retry quota, billing, or other errors that require you to take action.

116 129 

117Exponential backoff means waiting briefly after an unsuccessful request, then increasing the delay after each unsuccessful retry. This continues until the request succeeds or reaches a configured retry limit.130Exponential backoff means waiting briefly after an unsuccessful request, then increasing the delay after each unsuccessful retry. This continues until the request succeeds or reaches a configured retry limit.

118 131 


124 137 

125Note that unsuccessful requests contribute to your per-minute limit, so continuously resending a request won’t work.138Note that unsuccessful requests contribute to your per-minute limit, so continuously resending a request won’t work.

126 139 

127Below are a few example solutions **for Python** that use exponential backoff.140The Python examples below demonstrate fallback backoff. They don't inspect `Retry-After`: before using them, add handling for valid server hints so the wrappers don't retry sooner than requested. Disable SDK retries or account for them in your application's retry limits.

128 141 

129 142 

130 143