Billing
Two wallets, which one the API spends first, why output tokens are free, and what is never charged at all.
Your account holds two balances, and the API spends them in a fixed order.
The Two Wallets
| Wallet | Where it comes from | Where it is spent first |
|---|---|---|
| Tokens | Bought with a plan, denominated in input tokens | The API |
| Credits | Welcome gift and daily check-in | The web tools (playground, batch) |
The API spends the token wallet first. If the token balance cannot cover the whole call, the entire call costs 1 credit instead — never a partial drain of both. Which one paid is on every response:
"usage": {
"input_tokens": 503,
"output_tokens": 70,
"charged_tokens": 503,
"charged_credits": 0,
"wallet": "tokens"
}
and on every response header pair X-Jev-Tokens-Used / X-Jev-Credits-Used.
The web tools do the opposite: they spend credits first (1 credit per playground run, 1 per batch row) and fall back to input tokens when credits run out. One balance, two entry points, no model-specific sub-wallets.
What a Decision Costs
A decision is billed on the input tokens it read — the state, the questions and
the model's own scaffolding. Output tokens are counted in usage and never
billed.
That makes short states genuinely cheap and makes extra questions nearly free: a state read once answers up to 20 questions for the same input tokens. If you have four things to decide about one ticket, ask them in one call.
A worked example
The three-question ticket triage in the introduction reported 503 input tokens and 70 output tokens. It cost 503 tokens from the token wallet. The 70 output tokens cost nothing.
Two-Phase Charging
Every call reserves an estimate before it runs and settles against the real usage afterwards:
- Reserve. We estimate the input tokens and hold that amount. The estimate is deliberately conservative, so it usually holds slightly more than the call ends up using.
- Settle. The response comes back with real
usage, we charge exactly that, and the difference goes straight back to the same balance it came from. - Release. If the call fails at any point — upstream error, timeout, network, your own cancellation — the entire hold is released and nothing is charged.
Failed requests are never charged
Every non-2xx answer from this endpoint leaves your balance where it started.
A 422 for oversized input, a 503 while the model is unavailable, a 504
timeout: all released in full. The message on those errors says so too.
When Neither Wallet Can Pay
If the token balance cannot cover the call and there is no credit either, the
call is refused with a 402 before anything runs. The body tells you exactly
what was needed and what you have:
{
"error": {
"message": "Insufficient balance. This decision needs about 503 input tokens, or 1 credit if the token wallet cannot cover it. Your account has 0 tokens and 0 credits.",
"type": "insufficient_quota",
"code": "insufficient_quota",
"required_tokens": 503,
"required_credits": 1,
"current_tokens": 0,
"current_credits": 0,
"recharge_url": "https://jev-ai.org/pricing"
}
}
What Is Never Charged
- Creating, listing, renaming, rotating or revoking an API key.
GET /api/v1/models.- Reading a request's status, or cancelling it.
- Any request that did not return
200.
Where to See the Ledgers
Your account shows the token ledger and the credit ledger side by side, and API usage breaks spend down per key and per request. Purchased tokens roll over for as long as the account is active.