Retrieval

RAG Evaluation on Every Chunk, Not on a Sample of Fifty

RAG fails in two places, and they need different fixes. Retrieval can bring back the wrong passage, or retrieval can be fine and the generator can assert something the passage never said. A single "is this answer good?" score cannot tell you which one happened.

Both are yes/no questions about a pair of texts. A decision model answers them at a price that makes running the check on every chunk, on every request, an ordinary thing to do rather than a quarterly exercise.

  • Grounded / contradicted / unrelated as declared labels
  • Per-tier probabilities
  • 32K context
  • 422 on oversize, never silent truncation

The problem

The Job: Separate a Retrieval Bug from a Generation Bug

The standard offline metrics — context precision, context recall, faithfulness, answer relevance — are all pairwise judgements underneath: does this passage support this claim, is this claim traceable to this context. Frameworks implement them by asking a large model, in prose, one pair at a time. That is why most teams run them on a curated set of fifty questions and hope it generalises.

Fifty questions is enough to notice a catastrophe and not enough to notice a regression. The failures that matter in production are narrow: one document format that chunks badly, one date range that retrieves the superseded version of a policy, one phrasing that pulls the FAQ instead of the contract. A sample almost never contains them.

The other thing a sample cannot do is run inline. If a grounding check is cheap and fast, it stops being an offline metric and becomes a guardrail: score the citation before the answer goes out, and fall back to "I could not verify this" when it fails.

Why a decision model

Why a Decision Model Beats Asking an LLM to Critique the Retrieval

Grounding is not a writing task. It is a comparison with three or four possible outcomes, and you want the same answer twice.

Entailment, contradiction and irrelevance are different answers

A choice question distinguishes "the source supports this", "the source says the opposite" and "the source is about something else". A prose critique blurs the last two into "this does not really support the claim", which sends you looking for a generation bug when you have a retrieval bug.

Cheap enough to run on every chunk you retrieve

At the cost measured below, checking ten retrieved passages per request costs a fraction of a cent. That is the difference between an eval you run before a release and a filter you run in the request path.

The grounding score is a threshold you can tune

An ordered strength scale plus its probability distribution gives you a dial: pass strong grounding straight through, hold the middle for a second retrieval pass, and refuse to answer at the bottom. All three thresholds come from one number.

Oversized input fails loudly

A passage that exceeds the 32,000-token context returns a 422 and is not charged. It is never silently truncated — which matters enormously here, since a truncated source is exactly how a grounding check starts reporting false negatives.

The same call answers the metric and the guardrail

Supported-or-not for your dashboard, relationship for your debugging, strength for your threshold: three questions, one read of the passage, one charge.

Playground

Check a Citation That Is Nearly Right

A claim that says 30 days against a source that says ninety. This is the failure mode a keyword overlap score will never catch and a prose critique will hedge about.

Does the source actually support the claim?

Model

1Text

Claim and cited source

435 / 100,000

2Questions

3 in this request

Editing is free. Sign in and you come straight back here with your text and questions — no card needed.

3Answers

Sample · real Jev output
Yes / No

Does the cited source support the claim as stated, without the reader having to assume anything the source does not say?

98%
No
Choice79% confidence

Select how the cited source relates to the claim.

contradicted
  • contradicted84%
  • partially_supported16%
  • unrelated0%
  • fully_supported0%
Score90% confidence

Rate how strongly the cited source grounds the claim.

0.10No grounding — the source is irrelevant or contradictory
No grounding — the source is irrelevant or contradictoryWeak — requires a generous reading of the sourceAdequate — supports the substance with small gapsStrong — the source states the claim explicitly

Recorded 2026-09-30 on typesafe/jev-1.13-20260917594 input tokensEdit the text or questions, then run for a live answer.

In code

What You Would Write Next

Ten chunks in parallel, one verdict each. Contradicted chunks are a content problem; a pile of unrelated ones is a retriever problem.

Score every retrieved chunktypescript
const checks = await Promise.all(
  retrieved.map((chunk) =>
    jev.decide({
      state: `CLAIM\n${claim}\n\nCITED SOURCE\n${chunk.text}`,
      questions: {
        supported: {
          type: "noul",
          instructions:
            "Does the cited source support the claim as stated, without the reader assuming anything the source does not say?",
        },
        relationship: {
          type: "choice",
          instructions: "Select how the cited source relates to the claim.",
          criteria: {
            fully_supported:     "The source states the claim, or entails it directly.",
            partially_supported: "The source supports part of the claim but not all of it.",
            contradicted:        "The source states something incompatible with the claim.",
            unrelated:           "A different subject; neither supports nor contradicts.",
          },
        },
      },
    })
  )
);

// A retrieval bug and a generation bug look different here, which is the
// whole point of asking for the relationship instead of a single score.
const contradicted = checks.filter(
  (c) => c.answers.relationship.choice === "contradicted"
);
const unrelated = checks.filter(
  (c) => c.answers.relationship.choice === "unrelated"
);

if (contradicted.length) return refuseAndFlag(claim, contradicted);
if (unrelated.length / checks.length > 0.5) return rerank(claim);

The endpoint, both criteria container shapes and the full error contract are in the developer docs, and the API page has a brief you can paste straight into a coding agent.

Real cost

One Run Costs $0.000025

Measured, not estimated. On 2026-09-25 we sent the exact payload the playground above loads to jev-1.13-20260917 and read the numbers below straight out of the response’s usage block — measured on the claim and cited source in the playground above, with all three grounding questions in one call.

It checks out against the list rate of $0.042 per million input tokens: 594 ÷ 1,000,000 × 0.042 = $0.000025. The 91 output tokens were counted and not billed, which is why asking four questions about one state costs barely more than asking one.

Input tokens
594

The only thing charged

One run
$0.000025

Round trip 0.30s from a laptop, network included

1,000 runs
$0.025

Same questions, same length of input

1,000,000 runs
$24.95

At the model list rate, before our margin

Those are model costs. On this site a playground run costs 1 credit from your credit balance, and an API call debits exactly those input tokens from your token balance — whichever balance applies, output stays free and a failed request is never charged. The pricing page has the per-plan rates, and your own runs will differ in length from this example, so treat this as a worked figure rather than a quote.

Questions

RAG Evaluation: The Questions People Actually Ask

Does this replace Ragas, DeepEval or my existing eval harness?

It replaces the model call inside them. The metric definitions, the datasets and the reporting are still yours; what changes is that the pairwise judgement underneath becomes typed, calibrated and cheap enough to run on the full corpus instead of a sample.

Can I run this inline, on the request path?

That is the more valuable use. Check the citations before the answer is returned, and degrade to "I could not verify that from the documents" when grounding is weak. The latency is small enough to sit inside a response that is already waiting on a generation call.

How do I handle a claim that spans several retrieved chunks?

Split the claim, not the context. Ask one question per atomic claim against the chunk that was cited for it. A composite claim graded against a bundle of context is exactly the measurement that hides the one sentence that was invented.

What if my documents are longer than the context window?

You already chunk them for retrieval; check the chunk that was cited rather than the document. If a single chunk exceeds 32,000 tokens the call returns a 422, unbilled, so you find out at integration time rather than through a quiet accuracy drop.

Is the model reading the source, or pattern-matching the words?

Try the playground example above: the claim and the source share almost all of their vocabulary and differ on one number. Lexical overlap scores that pair as a match; the grounding question does not.

Does this work for languages other than English?

Write your instructions and criteria in the language your corpus uses and test it on your own passages before you rely on it. We publish measurements for English examples only, so treat anything else as something you verify rather than something we claim.

Run It on Your Own Data

Editing is free and browsing is free. Signing in brings you back to this page with your text and your questions, no card required.