Routing

Route the Request Before You Pay a Large Model to Read It

Most requests that reach a frontier model do not need one. Order lookups, password resets, "where is my refund" — they need a branch in your application, and the model call is just an expensive way to find out which branch.

A router is a classification problem wearing a generation costume. Give it to a model that classifies: one call decides the intent, whether a record lookup is needed, and whether the whole thing can be handled without a human or a large model at all.

  • P50 ~0.22s upstream
  • One label plus every label’s probability
  • Up to 20 questions per call
  • Input tokens only

The problem

The Job: A Cheap, Fast Branch at the Front of the Pipeline

A router sits on the hot path. Every request pays for it, so its latency is added to every response and its cost is multiplied by your whole volume. That budget is what rules out "send it to the big model and ask what to do" — you end up paying frontier prices to decide not to use the frontier model.

It also has to be honest about ambiguity. Real inbound messages carry two intents as often as one ("cancel it if it cannot arrive by the 14th"), and a router that returns a single confident label for those will silently send half of them to the wrong place. What you want is the split: when the top two labels are 0.52 and 0.48, that is a case for the fallback path, not a routing decision.

And the label set changes. Product adds a queue, support splits a team, a new workflow appears. Re-training a classifier for that is a sprint; editing a list of labels and their descriptions should be a pull request.

Why a decision model

Why a Decision Model Beats Asking an LLM to Pick a Label

A chat model can pick a label. The trouble starts at the second decimal place and at the second million requests.

Labels are declared, so the model cannot invent one

A choice question takes an object of label to description. The answer is always one of your labels — never a synonym, never a new category the model liked better, never "billing_or_technical". You do not need a fallback branch for outputs that are not in your enum.

The distribution is the ambiguity signal you were missing

Every choice answer carries a probability for every label and a confidence. Route on the winner when it is clear, and send the near-ties to a human or a bigger model. That one number turns a routing table into a system with a sane failure mode.

Latency small enough to sit in front of everything

Upstream reports a P50 around 0.22s. That is a budget you can spend on every request, including the ones you are about to answer from a cache.

The label set is configuration, not a training run

Adding a queue means adding a key and a sentence describing it. There is no dataset to label, no fine-tune to schedule, and no drift between the classifier and the routing table it feeds.

Ask the follow-up questions in the same breath

Intent, "does this need a record lookup", and "can this be fully automated" are three decisions your dispatcher needs. Asked together they read the message once and cost one call.

Playground

Route a Real Inbound Message

A chat message that is half an order-status question and half a conditional cancellation. Run it and look at how the probability is split before you decide what your dispatcher should do with it.

Pick the handler before you spend a model call on the reply

Model

1Text

Inbound message

205 / 100,000

2Questions

3 in this request

Editing is free. Sign in and you come straight back here with your text and questions — no card needed.

3Answers

Sample · real Jev output
Choice98% confidence

Select the single intent that best describes what the sender wants the business to do next.

cancel_or_refund
  • cancel_or_refund98%
  • order_status2%
  • account_help0%
  • product_question0%
  • returns_exchange0%
  • other0%
Yes / No

Does answering this message require looking up a specific order or account record?

96%
Yes
Yes / No

Could an automated flow resolve this end to end, without a human agent reading the message?

66%
Yes

Recorded 2026-09-30 on typesafe/jev-1.13-20260917539 input tokensEdit the text or questions, then run for a live answer.

In code

What You Would Write Next

The escalation rule is two lines because the model hands you the numbers it is unsure about instead of hiding them behind a confident sentence.

Route, then dispatchtypescript
const { answers } = await jev.decide({
  state: message.body,
  questions: {
    intent: {
      type: "choice",
      instructions:
        "Select the single intent that best describes what the sender wants next.",
      // criteria is an OBJECT for choice questions: label -> description.
      criteria: {
        order_status:     "Where is my order, when will it arrive, tracking.",
        cancel_or_refund: "Cancel an order or subscription, or ask for money back.",
        product_question: "Pre-purchase questions about specs, sizing or stock.",
        account_help:     "Login, address, payment method or profile changes.",
      },
    },
    needs_order_lookup: {
      type: "noul",
      instructions:
        "Does answering this require looking up a specific order or account?",
    },
  },
});

const { choice, probabilities, confidence } = answers.intent;

// Act on the distribution, not only the winner. A 0.52/0.48 split is a
// case to escalate - it is not a routing decision.
const [, runnerUp] = Object.values(probabilities).sort((a, b) => b - a);
if (confidence < 0.7 || runnerUp > 0.3) {
  return handoff.toGeneralQueue(message, probabilities);
}

return dispatch[choice](message, {
  needsLookup: answers.needs_order_lookup.noul > 0.5,
});

The endpoint, both criteria container shapes and the full error contract are in the developer docs, and the API page has a brief you can paste straight into a coding agent.

Real cost

One Run Costs $0.000023

Measured, not estimated. On 2026-09-25 we sent the exact payload the playground above loads to jev-1.13-20260917 and read the numbers below straight out of the response’s usage block — measured on the inbound message in the playground above, with all three routing questions in one call.

It checks out against the list rate of $0.042 per million input tokens: 539 ÷ 1,000,000 × 0.042 = $0.000023. The 106 output tokens were counted and not billed, which is why asking four questions about one state costs barely more than asking one.

Input tokens
539

The only thing charged

One run
$0.000023

Round trip 0.32s from a laptop, network included

1,000 runs
$0.023

Same questions, same length of input

1,000,000 runs
$22.64

At the model list rate, before our margin

Those are model costs. On this site a playground run costs 1 credit from your credit balance, and an API call debits exactly those input tokens from your token balance — whichever balance applies, output stays free and a failed request is never charged. The pricing page has the per-plan rates, and your own runs will differ in length from this example, so treat this as a worked figure rather than a quote.

Questions

LLM Router: The Questions People Actually Ask

How many labels can a choice question have?

Enough for a realistic routing table. Keep each label’s description specific about its boundaries — most routing mistakes come from two labels whose descriptions overlap, not from the model being unable to read.

Should I route on the winning label or on the probabilities?

On the probabilities. Take the winner when it clears your threshold, and send everything else down a fallback path. A router without an "I am not sure" branch will make the same mistake a thousand times before anyone notices.

Can I use this to pick between two LLMs rather than two queues?

Yes — that is the same question with different labels. A common shape is one choice question for the destination model and one score question for difficulty, so cheap requests go to a small model and the hard tail goes to a large one.

What about multi-intent messages?

Model them as several yes/no questions rather than one choice. "Does this ask about order status?" and "Does this ask to cancel?" can both be true, and each comes back with its own probability. A choice question is for picking exactly one.

Does the router add much latency?

Upstream reports a P50 around 0.22s and we measured the routing call in the playground above at roughly a third of a second end to end from a laptop, network included. It is small relative to the model call it is deciding whether to make.

What happens when the model is unavailable?

You get a 503 before anything is reserved or charged, never a wrong label. Treat it as a retryable status with backoff, and keep a default route so an outage degrades to "send it to the general queue" instead of failing the request.

Run It on Your Own Data

Editing is free and browsing is free. Signing in brings you back to this page with your text and your questions, no card required.