Evaluation

LLM as a Judge, Without the Judge Writing an Essay

The standard recipe is to ask a large model to grade another model’s answer and explain its reasoning. It works well enough to demo and badly enough to be a problem: the grade arrives inside prose you have to parse, two runs of the same input disagree, and the score clusters on whatever number the grader likes that week.

A decision model removes the essay. You hand it the question, the answer and the reference, and you get back a number on the scale you defined, plus the probability it assigned to every tier.

  • Ordered score scales
  • Full probability distribution
  • No prose to parse
  • Free output tokens

The problem

The Job: Turn an Answer into a Row in a Table

Evaluation is only useful when it is comparable. You want a column per criterion, a number per answer, and the ability to sort by it, diff last week against this week, and alert when a release moves a metric. That means the grader has to emit the same shape every time, for every answer, including the weird ones.

A chat model asked to grade will emit that shape most of the time. The failures are the expensive part: a refusal on a sensitive answer, a 4.5 where the rubric has four tiers, a JSON block wrapped in an apology, a score that silently drifts after a model update. Each one is a row you either drop or fix by hand, and dropping rows is how an eval quietly starts measuring the easy cases.

The other half of the job is disagreement. A grade of "3" tells you nothing about whether the grader was sure. When you are choosing which two hundred of ten thousand answers a human should look at, "the model split 0.48 / 0.52 between tier 2 and tier 3" is the most useful thing on the page.

Why a decision model

Why a Decision Model Beats Asking an LLM to Write a Review

Both approaches read the same text. The difference is what comes back and what it costs to trust it.

The output shape is the contract, not a request

A score question is declared with an ordered list of tiers, and the answer is a position on that scale. There is no format to ask for, no schema to validate after the fact, and no retry budget for the calls that came back as prose. Bad input shapes are rejected as a 400 before the model runs.

You get the distribution, not just the verdict

Every answer carries the probability of each tier or label, plus a confidence. That is what makes an eval actionable: sort by confidence to find the cases the grader could not decide, and route those to a human instead of sampling at random.

Scores are calibrated, so they are comparable across runs

The model is trained to report probabilities that mean what they say. A grade of 3 with 0.9 probability means something stable enough to trend over time — which is the only reason to keep an eval suite at all.

Reasoning tokens are the expensive part, and you are not buying them

A judge that writes 300 words of justification per answer bills you for 300 words per answer. Here output tokens are counted and free, and only the input is charged — so grading ten thousand answers is an arithmetic problem, not a budget conversation.

Four criteria cost barely more than one

The state is read once per call no matter how many questions ride along. Accuracy, helpfulness, failure mode and a ship/no-ship flag in one request cost a few hundred extra input tokens, not four times the price.

Playground

Grade an Answer Against Its Reference

A support answer that contradicts the billing policy it was supposed to follow. Run it and watch the failure-mode distribution, not just the winning label.

Grade a generated answer against a rubric, one call per answer

Model

1Text

Question, answer and reference

621 / 100,000

2Questions

4 in this request

Editing is free. Sign in and you come straight back here with your text and questions — no card needed.

3Answers

Answers appear here

Each answer returns a probability for every option and a confidence score.

In code

What You Would Write Next

One call per answer, two criteria, and a human-review queue driven by the model’s own confidence rather than a random sample.

Grade one answertypescript
const res = await fetch("https://jev-ai.org/api/v1/systemone/", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.JEV_API_KEY}`,
    "Content-Type": "application/json",
    "Idempotency-Key": `eval-${runId}-${caseId}`,
  },
  body: JSON.stringify({
    model: "jev-1.13",
    state: [
      `QUESTION\n${testCase.question}`,
      `REFERENCE\n${testCase.reference}`,
      `ASSISTANT ANSWER\n${candidate.answer}`,
    ].join("\n\n"),
    questions: {
      factual_accuracy: {
        type: "score",
        instructions:
          "Rate how well the answer matches the reference. Judge only against the reference.",
        // criteria is an ARRAY for score questions, lowest tier first.
        criteria: [
          "Contradicts the reference on the central point",
          "Partly wrong - one material claim is unsupported",
          "Broadly right with a minor inaccuracy",
          "Fully supported by the reference",
        ],
      },
      safe_to_send: {
        type: "noul",
        instructions: "Is this answer safe to send with no human review?",
      },
    },
  }),
});

const { answers } = await res.json();

// score is a DECIMAL - the expected value of the distribution, which is
// exactly what makes it sortable. Never type it as an integer.
const accuracy: number = answers.factual_accuracy.score;      // e.g. 0.41
const spread = answers.factual_accuracy.probabilities;        // { "0": .., "3": .. }
const confident = answers.factual_accuracy.confidence > 0.8;

if (!confident) queueForHumanReview(caseId, spread);

The endpoint, both criteria container shapes and the full error contract are in the developer docs, and the API page has a brief you can paste straight into a coding agent.

Real cost

One Run Costs $0.000033

Measured, not estimated. On 2026-09-25 we sent the exact payload the playground above loads to jev-1.13-20260917 and read the numbers below straight out of the response’s usage block — measured on the question, reference and answer in the playground above, with all four criteria in one call.

It checks out against the list rate of $0.042 per million input tokens: 795 ÷ 1,000,000 × 0.042 = $0.000033. The 125 output tokens were counted and not billed, which is why asking four questions about one state costs barely more than asking one.

Input tokens
795

The only thing charged

One run
$0.000033

Round trip 0.97s from a laptop, network included

1,000 runs
$0.033

Same questions, same length of input

1,000,000 runs
$33.39

At the model list rate, before our margin

Those are model costs. On this site a playground run costs 1 credit from your credit balance, and an API call debits exactly those input tokens from your token balance — whichever balance applies, output stays free and a failed request is never charged. The pricing page has the per-plan rates, and your own runs will differ in length from this example, so treat this as a worked figure rather than a quote.

Questions

LLM as a Judge: The Questions People Actually Ask

Is this a replacement for human evaluation?

No. It is a replacement for the sampling step. Grade every answer cheaply, then spend your human attention on the cases where the model was not confident — which is a far better use of a reviewer than reading a random 2%.

How is this different from asking GPT or Claude to return JSON?

Structured output makes a chat model emit valid JSON; it does not make the number inside it calibrated, and it still bills you for the reasoning tokens that produced it. Here the scale is part of the request, the probability of every tier comes back with the answer, and output tokens are free.

Can I use my existing rubric?

Yes, if it is an ordered scale. Write your tier labels into the criteria array lowest-first and put the boundary cases in the tier text rather than in the instructions. A rubric that is really several independent judgements should become several questions in the same call.

Why is the score a decimal when my rubric has four tiers?

Because it is the expected value of the distribution, not the argmax. A 1.89 on a four-tier scale means the model is mostly on tier 2 with weight on tier 1 — which sorts correctly, where a rounded integer would throw away the part you need. Type it as a float.

What happens if the answer I am grading is very long?

The context window is 32,000 tokens. Oversized input comes back as a 422 and is never silently truncated or charged, so a long transcript is a chunking decision you make deliberately rather than a quality regression you discover later.

Does a failed grading call cost anything?

No. A decision is billed only when it returns 200; any other status releases its hold in full. Retries with a stable Idempotency-Key are free and cannot produce a duplicate grade.

Run It on Your Own Data

Editing is free and browsing is free. Signing in brings you back to this page with your text and your questions, no card required.