Judge Something Against a Rubric and Get a Number You Can Sort By
Evaluation is the category where the shape of the output matters most, because the output is not the product — it is a column in a table you are going to sort, trend and alert on.
Everything here is the same mechanical job: read one artefact, apply criteria you wrote down, return a position on a scale or a label from a set, and say how sure you were.
Pages
Evaluation in Detail
Each page states the problem, explains why a typed decision suits it better than a generated review, loads its scenario into the playground, and prints what one run measured.
The shape
What These Decisions Have in Common
They all need comparability over time. A grade is only useful if this week’s number means what last week’s number meant, which is why calibration matters more here than raw accuracy on any single case.
They all produce a review queue. No evaluation is fully automatic; the useful thing is knowing which 2% a human should read. The confidence on each answer picks that set far better than random sampling does.
They are all volume problems. An eval that costs enough to think about is an eval you run before releases and not on every request — which is the same as not having it when something regresses on a Tuesday.
Ready to run
Scenarios You Can Load in One Click
These ship with the playground — pick one from the scenario menu and it arrives with its state and its questions already written. Editing is free; only running uses your account.
Content moderation
Apply a written policy to user-generated content and return the enforcement action, how severe the violation is, and whether the post exposes personal data.
actionseverityexposes_personal_data
Review scoring
Turn free-text reviews into structured signal: how positive they are, what they are about, and whether they describe a defect the product team should see.
sentimentprimary_topicreports_defectwould_repurchase
Lead scoring
Grade inbound leads the moment the form is submitted, so sales sees the ready-to-buy ones first and nobody hand-sorts a queue.
fitbuying_stagebudget_signalroute_to_sales
LLM guardrail
Screen a model-generated draft for policy risk, hallucination and leaked personal data in a single fast call, so a slow model does not have to review itself.
safe_to_sendrisk_typehallucination_risk
Questions
Evaluation FAQ
Can a model grade another model fairly?
It can apply criteria you wrote, consistently, at a price that lets you apply them to everything. That is a narrower claim than fairness, and it is the one worth relying on: pair it with a human review of the low-confidence tail rather than treating any automated grade as final.
How do I stop my eval drifting when the grader changes?
Pin the model. `jev-1.13` is a fixed build; `jev-latest` is a rolling alias. Use the pinned id for anything you trend over time, and re-baseline deliberately when you move.
What is the cheapest way to add more criteria?
Put them in the same call. The state is read and charged once no matter how many questions ride along, so a fifth criterion costs a few dozen input tokens rather than a fifth request.
Try It on Your Own Text
Browsing and editing cost nothing. Sign in only when you want to run a decision, and you come straight back with your work intact.
