Budgets and limits
Monthly budget
Each plan includes a monthly usage budget. Every request counts against it based on the model used and the tokens processed. Track it on the Usage page.
When the budget runs out, requests return 402 with code budget_exhausted. The x-sc-budget-renews-at response header tells you when it resets. There are no automatic overage charges.
To keep working before the reset, you can upgrade your plan or add budget.
Smart routing uses efficient models for simple work, so your budget lasts longer than if every request is pinned to a premium model.
Smart pacing
Smart pacing is optional and off by default. Turn it on from the Pacing page if you'd rather slow down than stop.
With Smart pacing on, heavier use adds a short delay before each response, so an always-on agent can't spend a month's budget in a few days. Speed usually recovers within about an hour. With Smart pacing off, requests run at full speed until the budget is used.
Rate limits
If you send requests too quickly, you get 429 with code rate_limit_exceeded and a retry-after header. Wait that many seconds before retrying.
Request limits
| Limit | Paid plans | Free trial |
|---|---|---|
| Input per request | about 850,000 tokens | about 500,000 tokens |
| Output per request | 50,000 tokens | 32,768 tokens |
| Images per request | 32, up to 10 MB each | 32, up to 10 MB each |
If you ask for more output than the limit, it is reduced to the limit. Oversized input is trimmed where possible; if it still doesn't fit, the request returns 400.
Fair use
Plans are intended for individual developers and their agents. Occasional bursts are expected. See the fair use policy for details.