LLM as a Judge, Without the Judge Writing an Essay
The standard recipe is to ask a large model to grade another model’s answer and explain its reasoning. It works well enough to demo and badly enough to be a problem: the grade arrives inside prose you have to parse, two runs of the same input disagree, and the score clusters on whatever number the grader likes that week.
A decision model removes the essay. You hand it the question, the answer and the reference, and you get back a number on the scale you defined, plus the probability it assigned to every tier.
- Ordered score scales
- Full probability distribution
- No prose to parse
- Free output tokens
The problem
The Job: Turn an Answer into a Row in a Table
Evaluation is only useful when it is comparable. You want a column per criterion, a number per answer, and the ability to sort by it, diff last week against this week, and alert when a release moves a metric. That means the grader has to emit the same shape every time, for every answer, including the weird ones.
A chat model asked to grade will emit that shape most of the time. The failures are the expensive part: a refusal on a sensitive answer, a 4.5 where the rubric has four tiers, a JSON block wrapped in an apology, a score that silently drifts after a model update. Each one is a row you either drop or fix by hand, and dropping rows is how an eval quietly starts measuring the easy cases.
The other half of the job is disagreement. A grade of "3" tells you nothing about whether the grader was sure. When you are choosing which two hundred of ten thousand answers a human should look at, "the model split 0.48 / 0.52 between tier 2 and tier 3" is the most useful thing on the page.
Why a decision model
Why a Decision Model Beats Asking an LLM to Write a Review
Both approaches read the same text. The difference is what comes back and what it costs to trust it.
The output shape is the contract, not a request
A score question is declared with an ordered list of tiers, and the answer is a position on that scale. There is no format to ask for, no schema to validate after the fact, and no retry budget for the calls that came back as prose. Bad input shapes are rejected as a 400 before the model runs.
You get the distribution, not just the verdict
Every answer carries the probability of each tier or label, plus a confidence. That is what makes an eval actionable: sort by confidence to find the cases the grader could not decide, and route those to a human instead of sampling at random.
Scores are calibrated, so they are comparable across runs
The model is trained to report probabilities that mean what they say. A grade of 3 with 0.9 probability means something stable enough to trend over time — which is the only reason to keep an eval suite at all.
Reasoning tokens are the expensive part, and you are not buying them
A judge that writes 300 words of justification per answer bills you for 300 words per answer. Here output tokens are counted and free, and only the input is charged — so grading ten thousand answers is an arithmetic problem, not a budget conversation.
Four criteria cost barely more than one
The state is read once per call no matter how many questions ride along. Accuracy, helpfulness, failure mode and a ship/no-ship flag in one request cost a few hundred extra input tokens, not four times the price.
Grade an Answer Against Its Reference
A support answer that contradicts the billing policy it was supposed to follow. Run it and watch the failure-mode distribution, not just the winning label.
Grade a generated answer against a rubric, one call per answer
1Text
Question, answer and reference621 / 100,000
2Questions
4 in this requestEditing is free. Sign in and you come straight back here with your text and questions — no card needed.
3Answers
Answers appear here
Each answer returns a probability for every option and a confidence score.
In code
What You Would Write Next
One call per answer, two criteria, and a human-review queue driven by the model’s own confidence rather than a random sample.
const res = await fetch("https://jev-ai.org/api/v1/systemone/", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.JEV_API_KEY}`,
"Content-Type": "application/json",
"Idempotency-Key": `eval-${runId}-${caseId}`,
},
body: JSON.stringify({
model: "jev-1.13",
state: [
`QUESTION\n${testCase.question}`,
`REFERENCE\n${testCase.reference}`,
`ASSISTANT ANSWER\n${candidate.answer}`,
].join("\n\n"),
questions: {
factual_accuracy: {
type: "score",
instructions:
"Rate how well the answer matches the reference. Judge only against the reference.",
// criteria is an ARRAY for score questions, lowest tier first.
criteria: [
"Contradicts the reference on the central point",
"Partly wrong - one material claim is unsupported",
"Broadly right with a minor inaccuracy",
"Fully supported by the reference",
],
},
safe_to_send: {
type: "noul",
instructions: "Is this answer safe to send with no human review?",
},
},
}),
});
const { answers } = await res.json();
// score is a DECIMAL - the expected value of the distribution, which is
// exactly what makes it sortable. Never type it as an integer.
const accuracy: number = answers.factual_accuracy.score; // e.g. 0.41
const spread = answers.factual_accuracy.probabilities; // { "0": .., "3": .. }
const confident = answers.factual_accuracy.confidence > 0.8;
if (!confident) queueForHumanReview(caseId, spread);The endpoint, both criteria container shapes and the full error contract are in the developer docs, and the API page has a brief you can paste straight into a coding agent.
Real cost
One Run Costs $0.000033
Measured, not estimated. On 2026-09-25 we sent the exact payload the playground above loads to jev-1.13-20260917 and read the numbers below straight out of the response’s usage block — measured on the question, reference and answer in the playground above, with all four criteria in one call.
It checks out against the list rate of $0.042 per million input tokens: 795 ÷ 1,000,000 × 0.042 = $0.000033. The 125 output tokens were counted and not billed, which is why asking four questions about one state costs barely more than asking one.
- Input tokens
- 795
- One run
- $0.000033
- 1,000 runs
- $0.033
- 1,000,000 runs
- $33.39
The only thing charged
Round trip 0.97s from a laptop, network included
Same questions, same length of input
At the model list rate, before our margin
Those are model costs. On this site a playground run costs 1 credit from your credit balance, and an API call debits exactly those input tokens from your token balance — whichever balance applies, output stays free and a failed request is never charged. The pricing page has the per-plan rates, and your own runs will differ in length from this example, so treat this as a worked figure rather than a quote.
Questions
LLM as a Judge: The Questions People Actually Ask
Is this a replacement for human evaluation?
No. It is a replacement for the sampling step. Grade every answer cheaply, then spend your human attention on the cases where the model was not confident — which is a far better use of a reviewer than reading a random 2%.
How is this different from asking GPT or Claude to return JSON?
Structured output makes a chat model emit valid JSON; it does not make the number inside it calibrated, and it still bills you for the reasoning tokens that produced it. Here the scale is part of the request, the probability of every tier comes back with the answer, and output tokens are free.
Can I use my existing rubric?
Yes, if it is an ordered scale. Write your tier labels into the criteria array lowest-first and put the boundary cases in the tier text rather than in the instructions. A rubric that is really several independent judgements should become several questions in the same call.
Why is the score a decimal when my rubric has four tiers?
Because it is the expected value of the distribution, not the argmax. A 1.89 on a four-tier scale means the model is mostly on tier 2 with weight on tier 1 — which sorts correctly, where a rounded integer would throw away the part you need. Type it as a float.
What happens if the answer I am grading is very long?
The context window is 32,000 tokens. Oversized input comes back as a 422 and is never silently truncated or charged, so a long transcript is a chunking decision you make deliberately rather than a quality regression you discover later.
Does a failed grading call cost anything?
No. A decision is billed only when it returns 200; any other status releases its hold in full. Retries with a stable Idempotency-Key are free and cannot produce a duplicate grade.
Related
Other Decisions in This Shape
LLM router
Decide which model, tool or workflow should handle a request, with a probability you can threshold.
RetrievalRAG evaluation
Check retrieval quality and answer grounding chunk by chunk, cheaply enough to run it on everything.
Data matchingEntity matching
Link records that fuzzy matching gets wrong, with a probability you can set a merge threshold on.
Run It on Your Own Data
Editing is free and browsing is free. Signing in brings you back to this page with your text and your questions, no card required.
