Billing

Two wallets, which one the API spends first, why output tokens are free, and what is never charged at all.

Your account holds two balances, and the API spends them in a fixed order.

API spendsTokens firstThen 1 credit per call
Billed quantityInput tokensOutput tokens are free
Failed requestsNot chargedHolds are released automatically
Purchased tokensRoll overNo expiry while the account is active

The Two Wallets

WalletWhere it comes fromWhere it is spent first
TokensBought with a plan, denominated in input tokensThe API
CreditsWelcome gift and daily check-inThe web tools (playground, batch)

The API spends the token wallet first. If the token balance cannot cover the whole call, the entire call costs 1 credit instead — never a partial drain of both. Which one paid is on every response:

"usage": {
  "input_tokens": 503,
  "output_tokens": 70,
  "charged_tokens": 503,
  "charged_credits": 0,
  "wallet": "tokens"
}

and on every response header pair X-Jev-Tokens-Used / X-Jev-Credits-Used.

The web tools do the opposite: they spend credits first (1 credit per playground run, 1 per batch row) and fall back to input tokens when credits run out. One balance, two entry points, no model-specific sub-wallets.

What a Decision Costs

A decision is billed on the input tokens it read — the state, the questions and the model's own scaffolding. Output tokens are counted in usage and never billed.

That makes short states genuinely cheap and makes extra questions nearly free: a state read once answers up to 20 questions for the same input tokens. If you have four things to decide about one ticket, ask them in one call.

A worked example

The three-question ticket triage in the introduction reported 503 input tokens and 70 output tokens. It cost 503 tokens from the token wallet. The 70 output tokens cost nothing.

Two-Phase Charging

Every call reserves an estimate before it runs and settles against the real usage afterwards:

  1. Reserve. We estimate the input tokens and hold that amount. The estimate is deliberately conservative, so it usually holds slightly more than the call ends up using.
  2. Settle. The response comes back with real usage, we charge exactly that, and the difference goes straight back to the same balance it came from.
  3. Release. If the call fails at any point — upstream error, timeout, network, your own cancellation — the entire hold is released and nothing is charged.

Failed requests are never charged

Every non-2xx answer from this endpoint leaves your balance where it started. A 422 for oversized input, a 503 while the model is unavailable, a 504 timeout: all released in full. The message on those errors says so too.

When Neither Wallet Can Pay

If the token balance cannot cover the call and there is no credit either, the call is refused with a 402 before anything runs. The body tells you exactly what was needed and what you have:

{
  "error": {
    "message": "Insufficient balance. This decision needs about 503 input tokens, or 1 credit if the token wallet cannot cover it. Your account has 0 tokens and 0 credits.",
    "type": "insufficient_quota",
    "code": "insufficient_quota",
    "required_tokens": 503,
    "required_credits": 1,
    "current_tokens": 0,
    "current_credits": 0,
    "recharge_url": "https://jev-ai.org/pricing"
  }
}

What Is Never Charged

  • Creating, listing, renaming, rotating or revoking an API key.
  • GET /api/v1/models.
  • Reading a request's status, or cancelling it.
  • Any request that did not return 200.

Where to See the Ledgers

Your account shows the token ledger and the credit ledger side by side, and API usage breaks spend down per key and per request. Purchased tokens roll over for as long as the account is active.