Decide What Is Worth Reading, and Whether It Says What You Think
Retrieval has two failure modes that look identical from the outside: you fetched the wrong thing, or you fetched the right thing and asserted something it does not say. Telling them apart is what makes a RAG system debuggable.
Both questions compare two texts and return one of a few outcomes. Priced per chunk rather than per sample, they stop being an offline metric and become a filter you run in the request path.
Pages
Retrieval in Detail
Each page states the problem, explains why a typed decision suits it better than a generated review, loads its scenario into the playground, and prints what one run measured.
The shape
What These Decisions Have in Common
They run per chunk, not per answer. That multiplies volume by ten or more, which is why unit cost decides whether the check is feasible at all.
They need a distinction, not a score. "Contradicted" and "unrelated" are different bugs with different fixes, and a single relevance number collapses them.
They protect a context window. Every passage you let through costs frontier prices downstream and competes with the passage that actually had the answer.
Ready to run
Scenarios You Can Load in One Click
These ship with the playground — pick one from the scenario menu and it arrives with its state and its questions already written. Editing is free; only running uses your account.
Citation check
Verify a retrieved passage against the claim it is cited for — the grounding check that makes a RAG pipeline trustworthy, run on every chunk instead of a sample.
supportedrelationshipgrounding_strength
Questions
Retrieval FAQ
Can I run a grounding check inline instead of offline?
That is the more valuable use. Verify the citations before the answer leaves your server and degrade to "I could not verify that from the documents" when grounding is weak.
How big a passage can I check at once?
Up to the 32,000-token context. Oversize input returns a 422 and is never charged or silently truncated — which matters here, because a truncated source is how a grounding check starts producing false negatives.
Does this replace an embedding reranker?
No, they answer different questions. A reranker orders passages by similarity; these decisions say whether a passage entails a claim and what kind of page it is. Most pipelines want both.
Try It on Your Own Text
Browsing and editing cost nothing. Sign in only when you want to run a decision, and you come straight back with your work intact.
