Rate limits
Per-key request limits, a per-account cap on in-flight generations, and the headers that tell you where you stand.
| Limit | Value | Scope |
|---|---|---|
| Submissions | 30 / minute | per key |
| Reads (poll, list, models, credits) | 300 / minute | per key |
| In-flight generations | 10 queued at once | per account (all queued tasks) |
| Failed authentications | 20 / minute | per IP |
Every authenticated response carries x-ratelimit-limit, x-ratelimit-remaining and x-ratelimit-reset (Unix seconds). A 429 adds retry-after in seconds.
Two kinds of 429
rate_limitedmeans the per-key request rate for this window is used up. Wait forretry-afterand continue.concurrency_limitedmeans 10 of your generations are stillqueued. Poll and let them finish before submitting more. All queued tasks count, including older tasks. Concurrent submissions reserve capacity before billable work starts.
Neither is charged. Need more? Contact us with your use case and expected volume.