Retrieval

Decide What Is Worth Reading, and Whether It Says What You Think

Retrieval has two failure modes that look identical from the outside: you fetched the wrong thing, or you fetched the right thing and asserted something it does not say. Telling them apart is what makes a RAG system debuggable.

Both questions compare two texts and return one of a few outcomes. Priced per chunk rather than per sample, they stop being an offline metric and become a filter you run in the request path.

Pages

Retrieval in Detail

Each page states the problem, explains why a typed decision suits it better than a generated review, loads its scenario into the playground, and prints what one run measured.

The shape

What These Decisions Have in Common

They run per chunk, not per answer. That multiplies volume by ten or more, which is why unit cost decides whether the check is feasible at all.

They need a distinction, not a score. "Contradicted" and "unrelated" are different bugs with different fixes, and a single relevance number collapses them.

They protect a context window. Every passage you let through costs frontier prices downstream and competes with the passage that actually had the answer.

Ready to run

Scenarios You Can Load in One Click

These ship with the playground — pick one from the scenario menu and it arrives with its state and its questions already written. Editing is free; only running uses your account.

Citation check

Verify a retrieved passage against the claim it is cited for — the grounding check that makes a RAG pipeline trustworthy, run on every chunk instead of a sample.

supportedrelationshipgrounding_strength

Questions

Retrieval FAQ

Can I run a grounding check inline instead of offline?

That is the more valuable use. Verify the citations before the answer leaves your server and degrade to "I could not verify that from the documents" when grounding is weak.

How big a passage can I check at once?

Up to the 32,000-token context. Oversize input returns a 422 and is never charged or silently truncated — which matters here, because a truncated source is how a grounding check starts producing false negatives.

Does this replace an embedding reranker?

No, they answer different questions. A reranker orders passages by similarity; these decisions say whether a passage entails a claim and what kind of page it is. Most pipelines want both.

Try It on Your Own Text

Browsing and editing cost nothing. Sign in only when you want to run a decision, and you come straight back with your work intact.