Use cases

Where a Decision Beats a Paragraph

Every page here is a job people currently do by asking a large language model to read something and write about it, then parsing the writing. A decision model skips the writing: you declare the answer shape, and you get a label, a score or a probability with the full distribution behind it.

Each page states the problem honestly, loads its own scenario into the playground so you can run it, shows the code you would write next, and prints what one run actually cost — measured from a live call, starting at $0.000023.

  • 6 tool pages
  • 4 categories
  • Measured costs, not estimates
  • No sign-in to browse

Evaluation

Judge Something Against a Rubric and Get a Number You Can Sort By

Evaluation is the category where the shape of the output matters most, because the output is not the product — it is a column in a table you are going to sort, trend and alert on.

All evaluation use cases

Routing

Decide Where a Request Should Go Before You Spend Money on It

Routing sits on the hot path: every request pays for it, so latency and unit cost are the whole design constraint. That is exactly what rules out asking a frontier model which branch to take.

All routing use cases

Retrieval

Decide What Is Worth Reading, and Whether It Says What You Think

Retrieval has two failure modes that look identical from the outside: you fetched the wrong thing, or you fetched the right thing and asserted something it does not say. Telling them apart is what makes a RAG system debuggable.

All retrieval use cases

Data matching

Compare Two Records and Return a Probability You Can Threshold

String similarity solves the easy pairs and then stops. What is left needs judgement about which evidence is decisive: a shared registration number outweighs a changed address; a similar name outweighs nothing at all.

All data matching use cases

Questions

Before You Pick One

What kinds of problem is a decision model actually good at?

Ones where your application has to branch and the branch depends on reading something unstructured. If the output of the model call is a label, a number or a boolean that goes into an if statement, this is the right shape of tool. If the output is text a person will read, it is not.

How is this different from using GPT or Claude with structured output?

Structured output guarantees the JSON parses. It does not give you a calibrated probability for every alternative, and you still pay for the reasoning tokens that produced the answer. Here the answer shape is declared in the request, every alternative comes back with its probability, and output tokens are counted but never billed.

Can I combine several of these in one call?

If they are about the same state, yes, and you should. The state is read and charged once no matter how many questions ride along, so routing plus a difficulty score plus a needs-human flag costs barely more than any one of them alone.

Where do the per-run costs on these pages come from?

From live calls. Each tool page prints the usage block returned by an actual request sending exactly the payload its playground loads, along with the arithmetic that reproduces it from the published input-token rate. None of the figures are estimates.

Do I have to sign in to try these?

Only to run one. Browsing every page, editing the sample text and rewriting the questions are all free and need no account. Signing in returns you to the page you were on with your work intact, and needs no card.

Start from Whichever One Is Your Problem

Or open the playground and paste something you already classify by hand. The scenario menu carries every example on this site.