Evaluation

Judge Something Against a Rubric and Get a Number You Can Sort By

Evaluation is the category where the shape of the output matters most, because the output is not the product — it is a column in a table you are going to sort, trend and alert on.

Everything here is the same mechanical job: read one artefact, apply criteria you wrote down, return a position on a scale or a label from a set, and say how sure you were.

Pages

Evaluation in Detail

Each page states the problem, explains why a typed decision suits it better than a generated review, loads its scenario into the playground, and prints what one run measured.

The shape

What These Decisions Have in Common

They all need comparability over time. A grade is only useful if this week’s number means what last week’s number meant, which is why calibration matters more here than raw accuracy on any single case.

They all produce a review queue. No evaluation is fully automatic; the useful thing is knowing which 2% a human should read. The confidence on each answer picks that set far better than random sampling does.

They are all volume problems. An eval that costs enough to think about is an eval you run before releases and not on every request — which is the same as not having it when something regresses on a Tuesday.

Ready to run

Scenarios You Can Load in One Click

These ship with the playground — pick one from the scenario menu and it arrives with its state and its questions already written. Editing is free; only running uses your account.

Content moderation

Apply a written policy to user-generated content and return the enforcement action, how severe the violation is, and whether the post exposes personal data.

actionseverityexposes_personal_data

Review scoring

Turn free-text reviews into structured signal: how positive they are, what they are about, and whether they describe a defect the product team should see.

sentimentprimary_topicreports_defectwould_repurchase

Lead scoring

Grade inbound leads the moment the form is submitted, so sales sees the ready-to-buy ones first and nobody hand-sorts a queue.

fitbuying_stagebudget_signalroute_to_sales

LLM guardrail

Screen a model-generated draft for policy risk, hallucination and leaked personal data in a single fast call, so a slow model does not have to review itself.

safe_to_sendrisk_typehallucination_risk

Questions

Evaluation FAQ

Can a model grade another model fairly?

It can apply criteria you wrote, consistently, at a price that lets you apply them to everything. That is a narrower claim than fairness, and it is the one worth relying on: pair it with a human review of the low-confidence tail rather than treating any automated grade as final.

How do I stop my eval drifting when the grader changes?

Pin the model. `jev-1.13` is a fixed build; `jev-latest` is a rolling alias. Use the pinned id for anything you trend over time, and re-baseline deliberately when you move.

What is the cheapest way to add more criteria?

Put them in the same call. The state is read and charged once no matter how many questions ride along, so a fifth criterion costs a few dozen input tokens rather than a fifth request.

Try It on Your Own Text

Browsing and editing cost nothing. Sign in only when you want to run a decision, and you come straight back with your work intact.