Routing

Decide Where a Request Should Go Before You Spend Money on It

Routing sits on the hot path: every request pays for it, so latency and unit cost are the whole design constraint. That is exactly what rules out asking a frontier model which branch to take.

These decisions are all "pick one of my labels, and tell me how sure you were" — with the second half doing most of the work, because the near-ties are where routers quietly go wrong.

Pages

Routing in Detail

Each page states the problem, explains why a typed decision suits it better than a generated review, loads its scenario into the playground, and prints what one run measured.

The shape

What These Decisions Have in Common

The label set is yours and it changes. A choice question declares it inline, so adding a queue is a pull request rather than a retraining job, and the model can never answer with a category you did not define.

Ambiguity is normal and must be visible. Real messages carry two intents; the probability split is what lets your dispatcher send those to a fallback instead of guessing.

They are cheap per call and enormous in aggregate. The only price that matters is the one multiplied by your whole request volume.

Ready to run

Scenarios You Can Load in One Click

These ship with the playground — pick one from the scenario menu and it arrives with its state and its questions already written. Editing is free; only running uses your account.

Support ticket triage

Decide which team owns an inbound ticket, how fast it has to be answered, and whether it can be closed by an automated reply — before a human reads it.

queueurgencyneeds_humanrefund_requested

Intent routing

Classify what an incoming message is asking for so the application can dispatch it to the right tool, workflow or agent instead of sending everything to a large model.

intentneeds_order_lookupself_service_possible

Questions

Routing FAQ

Is a decision model faster than a small chat model?

Upstream reports a P50 around 0.22s, and it returns no tokens to stream, so there is no time-to-first-token to wait out. Measure it against your own candidate on your own traffic before committing — that is the only comparison that counts.

How do I handle a request that matches two labels?

Model it as several yes/no questions rather than one choice. A choice question is for picking exactly one; overlapping conditions each want their own probability.

What should the router do when the model is unavailable?

You get a 503 before anything is reserved or charged. Keep a default route so an outage degrades to a general queue rather than failing the request, and retry with backoff.

Try It on Your Own Text

Browsing and editing cost nothing. Sign in only when you want to run a decision, and you come straight back with your work intact.