Rate Limits and Quotas
Requests per minute, the daily decision budget, request and input size caps, and the headers that tell you where you stand.
Three separate ceilings apply to a decision, and they fail differently.
Requests per Minute
Each key has its own limit, counted in a rolling one-minute window. The default is 60; you can raise it to 300 in Settings → API keys.
Every response carries where you stand:
X-RateLimit-Limit-Requests: 60
X-RateLimit-Remaining-Requests: 59
X-RateLimit-Reset-Requests: 2026-09-24T19:29:00.000Z
Exceeding it answers 429 with code: rate_limit_exceeded and a Retry-After
header in seconds. Nothing is charged.
Daily Decisions per Key
Each key also carries a daily budget, counted in decisions rather than in a
currency, and reset on the schedule the key is configured with (daily by
default). The floor for this endpoint is 10,000 decisions per window; if
your key's spend limit is set higher than that, the higher number applies, and
a key with its spend limit set to 0 has no daily ceiling at all.
Why this is not the chat spend limit
An API key's spend limit is denominated in chat credits and defaults to 300 a day — a number that makes sense for chat completions and would throttle a decision workload to a fraction of what a token plan buys. This endpoint counts decisions instead and never applies a ceiling below 10,000.
Exhausting it answers 429 with code: api_key_spend_limit_exceeded, the
window's reset_at, and a link to the key's settings. Nothing is charged.
Neither ceiling is a spending limit. Your token and credit balances are what actually caps spend; see Billing.
Size Caps
| Cap | Value | What happens past it |
|---|---|---|
| Request body | 4 MB | 413 request_too_large |
state | 100,000 characters | 400, before the request leaves our server |
| Input tokens | 32,000 (state + questions) | 422 input_too_large, not charged |
| Questions | 20 per call | 400 |
| Choice labels | 2–24 per question | 400 |
| Score tiers | 2–10 per question | 400 |
instructions | 1,000 characters | 400 |
The 100,000-character cap on state is a cheap guard, not the real limit: it
stops an obviously oversized payload from costing a round trip. The real limit
is the 32,000-token context, and exceeding it returns 422 without a
charge — input is never silently truncated.
Staying Inside Them
- Batch questions, not requests. Twenty questions about one state is one request against your rate limit and one state charged once.
- Run decisions in parallel up to your rate limit; they are independent.
- Back off on
429usingRetry-Afterrather than a fixed sleep. - Use one key per workload so a runaway job cannot exhaust the budget your production path depends on.