Skip to main content

Budgets and limits

Monthly budget​

Each plan includes a monthly usage budget. Every request counts against it based on the model used and the tokens processed. Track it on the Usage page.

When the budget runs out, requests return 402 with code budget_exhausted. The x-sc-budget-renews-at response header tells you when it resets. There are no automatic overage charges.

To keep working before the reset, you can upgrade your plan or add budget.

Stretch your budget

Smart routing uses efficient models for simple work, so your budget lasts longer than if every request is pinned to a premium model.

Smart pacing​

Smart pacing is optional and off by default. Turn it on from the Pacing page if you'd rather slow down than stop.

With Smart pacing on, heavier use adds a short delay before each response, so an always-on agent can't spend a month's budget in a few days. Speed usually recovers within about an hour. With Smart pacing off, requests run at full speed until the budget is used.

Rate limits​

If you send requests too quickly, you get 429 with code rate_limit_exceeded and a retry-after header. Wait that many seconds before retrying.

Request limits​

LimitPaid plansFree trial
Input per requestabout 850,000 tokensabout 500,000 tokens
Output per request50,000 tokens32,768 tokens
Images per request32, up to 10 MB each32, up to 10 MB each

If you ask for more output than the limit, it is reduced to the limit. Oversized input is trimmed where possible; if it still doesn't fit, the request returns 400.

Fair use​

Plans are intended for individual developers and their agents. Occasional bursts are expected. See the fair use policy for details.