Rate Limits and Quotas

Requests per minute, the daily decision budget, request and input size caps, and the headers that tell you where you stand.

Three separate ceilings apply to a decision, and they fail differently.

Rate limit60 req / minPer key, configurable up to 300
Daily decisions10,000Per key, per UTC day
Request body4 MBPer request
Input32,000 tokensState plus questions

Requests per Minute

Each key has its own limit, counted in a rolling one-minute window. The default is 60; you can raise it to 300 in Settings → API keys.

Every response carries where you stand:

X-RateLimit-Limit-Requests: 60
X-RateLimit-Remaining-Requests: 59
X-RateLimit-Reset-Requests: 2026-09-24T19:29:00.000Z

Exceeding it answers 429 with code: rate_limit_exceeded and a Retry-After header in seconds. Nothing is charged.

Daily Decisions per Key

Each key also carries a daily budget, counted in decisions rather than in a currency, and reset on the schedule the key is configured with (daily by default). The floor for this endpoint is 10,000 decisions per window; if your key's spend limit is set higher than that, the higher number applies, and a key with its spend limit set to 0 has no daily ceiling at all.

Why this is not the chat spend limit

An API key's spend limit is denominated in chat credits and defaults to 300 a day — a number that makes sense for chat completions and would throttle a decision workload to a fraction of what a token plan buys. This endpoint counts decisions instead and never applies a ceiling below 10,000.

Exhausting it answers 429 with code: api_key_spend_limit_exceeded, the window's reset_at, and a link to the key's settings. Nothing is charged.

Neither ceiling is a spending limit. Your token and credit balances are what actually caps spend; see Billing.

Size Caps

CapValueWhat happens past it
Request body4 MB413 request_too_large
state100,000 characters400, before the request leaves our server
Input tokens32,000 (state + questions)422 input_too_large, not charged
Questions20 per call400
Choice labels2–24 per question400
Score tiers2–10 per question400
instructions1,000 characters400

The 100,000-character cap on state is a cheap guard, not the real limit: it stops an obviously oversized payload from costing a round trip. The real limit is the 32,000-token context, and exceeding it returns 422 without a charge — input is never silently truncated.

Staying Inside Them

  • Batch questions, not requests. Twenty questions about one state is one request against your rate limit and one state charged once.
  • Run decisions in parallel up to your rate limit; they are independent.
  • Back off on 429 using Retry-After rather than a fixed sleep.
  • Use one key per workload so a runaway job cannot exhaust the budget your production path depends on.