Web Context: Decide What a Fetched Page Is Before You Pay to Read It
A browsing agent fetches ten pages to answer one question. Two of them are consent walls, three are category listings that only link elsewhere, one is a blog post with the answer buried under a fitting guide, and one is the product page it wanted. Feeding all ten into a context window is how a cheap question becomes an expensive one — and how the relevant sentence gets lost among forty thousand tokens of navigation.
The question "is this page worth reading" is a decision, and it is much cheaper than reading the page with a frontier model. Answer it first, and the expensive model only ever sees pages that earned their place.
- Page-type labels you declare
- Relevance to the query that fetched it
- Keep / extract / drop as an ordered score
- Output tokens free
The problem
The Job: Keep the Context Window Honest
Fetched web pages are mostly not content. Navigation, cookie banners, newsletter interstitials, related-product rails and footers routinely make up most of the tokens on a page, and boilerplate strippers only get the obvious ones. What survives still has to be judged against the question that sent you there, which no heuristic can do.
The cost of getting this wrong compounds twice. You pay for the junk tokens on the way in, and you pay again in quality: a long context full of near-misses makes the generation step worse, not better, because the relevant passage now competes with six plausible distractors.
It also decides how an agent behaves. "This page is a consent wall" means fetch a different URL. "This is a listing page" means follow a link rather than read. "One useful section" means extract instead of quote. Those are branches, and a branch wants a label, not a summary.
Why a decision model
Why a Decision Model Beats Asking an LLM to Summarise the Page First
Summarising to decide whether to read is the expensive way round. Decide first.
The filter must cost a fraction of what it saves
A gate that costs as much as the call it is protecting is not a gate. At the price measured below, triaging ten pages costs a small fraction of a cent while removing most of what would have gone into the context window.
Page type is a branch, and branches need labels
Consent wall, listing, documentation, product page, article. Declared as labels, they drive what the crawler does next. Extracted from prose with a regex, they drive an incident.
Relevance is judged against the query, not in the abstract
Put the query in the state alongside the page. "Does this contain what is needed to answer this question, even if it is not the page’s main subject" is precisely the judgement that lets you keep a fitting guide that happens to list the dimensions.
An ordered keep/extract/drop scale drives the pipeline
One score with a tier distribution gives you three actions and their thresholds: drop it, extract the useful section, or keep the page whole.
Blocked and empty pages are a first-class answer
A cookie wall is not a low-relevance page, it is a fetch that failed. Giving it its own label means your crawler can retry or route around it instead of recording a bad result.
Triage a Page a Crawler Just Fetched
A kitchen fitting guide that buries the dimension the query asked for, under a subscribe prompt and a cookie notice. Run it and see the keep/extract/drop score.
Decide what a fetched page is before you pay to reason about it
1Text
Fetched page802 / 100,000
2Questions
4 in this requestEditing is free. Sign in and you come straight back here with your text and questions — no card needed.
3Answers
Answers appear here
Each answer returns a probability for every option and a confidence score.
In code
What You Would Write Next
Truncate the page before you send it: you are deciding whether the page is worth reading, and the first few thousand tokens almost always settle that.
const triaged = await Promise.all(
fetched.map(async (page) => {
const { answers } = await jev.decide({
state: [
`QUERY THAT SENT US HERE\n${query}`,
`URL\n${page.url}`,
`FETCHED TEXT (truncated)\n${page.text.slice(0, 8000)}`,
].join("\n\n"),
questions: {
answers_query: {
type: "noul",
instructions:
"Does this page contain the information needed to answer the query, even if it is not the main subject of the page?",
},
page_type: {
type: "choice",
instructions: "Select what kind of page this is.",
criteria: {
documentation: "Reference material, specifications or manuals.",
product_page: "One purchasable product, with its details.",
editorial: "A guide, blog post or review.",
listing: "A category or search page that mostly links elsewhere.",
blocked_or_empty:"Consent wall, paywall, error, or no usable content.",
},
},
context_value: {
type: "score",
instructions:
"Rate how much of this page is worth putting into a model context window for this query.",
criteria: [
"None - drop it",
"One or two useful sentences buried in boilerplate",
"A useful section worth extracting",
"Substantially on-topic - keep the page",
],
},
},
});
return { page, ...answers };
})
);
// Three branches, one score.
const keep = triaged.filter((t) => t.context_value.score >= 2.5);
const extract = triaged.filter((t) => t.context_value.score >= 1 && t.context_value.score < 2.5);
const retry = triaged.filter((t) => t.page_type.choice === "blocked_or_empty");The endpoint, both criteria container shapes and the full error contract are in the developer docs, and the API page has a brief you can paste straight into a coding agent.
Real cost
One Run Costs $0.000033
Measured, not estimated. On 2026-09-25 we sent the exact payload the playground above loads to jev-1.13-20260917 and read the numbers below straight out of the response’s usage block — measured on the fetched page in the playground above, with all four questions in one call.
It checks out against the list rate of $0.042 per million input tokens: 794 ÷ 1,000,000 × 0.042 = $0.000033. The 120 output tokens were counted and not billed, which is why asking four questions about one state costs barely more than asking one.
- Input tokens
- 794
- One run
- $0.000033
- 1,000 runs
- $0.033
- 1,000,000 runs
- $33.35
The only thing charged
Round trip 0.28s from a laptop, network included
Same questions, same length of input
At the model list rate, before our margin
Those are model costs. On this site a playground run costs 1 credit from your credit balance, and an API call debits exactly those input tokens from your token balance — whichever balance applies, output stays free and a failed request is never charged. The pricing page has the per-plan rates, and your own runs will differ in length from this example, so treat this as a worked figure rather than a quote.
Questions
Web Context: The Questions People Actually Ask
Should I send the whole page or a truncated version?
Truncate. You are deciding whether the page deserves a full read, and the first few thousand tokens nearly always settle it. Sending the whole page to decide whether to send the whole page defeats the purpose.
Does this fetch pages for me?
No. This is a decision endpoint, not a crawler. You fetch with whatever you already use — your own client, a scraping API, a headless browser — and send the text you got back.
How does it compare to a reranker?
A reranker orders passages by relevance to a query and is excellent at that. This answers different questions in the same call: what kind of page this is, whether it is a consent wall, whether it is primarily commercial. Those drive crawler behaviour, which an ordering cannot.
Can I run this on every page an agent fetches?
That is the intended use, and the cost is why it is reasonable. Triaging ten pages per query costs a small fraction of a cent, against a context window you would otherwise pay frontier prices to fill with navigation.
What about HTML — should I strip it first?
Strip it. Tags and inline styles are tokens you pay for and signal the model does not need. Extracted text, with the URL and the query alongside it, is the cheapest input that still supports the judgement.
Will it tell me whether a page is trustworthy?
It will tell you what kind of page it is and whether it is primarily selling something, which are observable properties of the text. Do not ask it to adjudicate truth — for that, retrieve the authoritative source and use a grounding check.
Related
Other Decisions in This Shape
LLM as a judge
Grade generated answers against a rubric and get sortable scores instead of a paragraph of praise.
RoutingLLM router
Decide which model, tool or workflow should handle a request, with a probability you can threshold.
RetrievalRAG evaluation
Check retrieval quality and answer grounding chunk by chunk, cheaply enough to run it on everything.
Run It on Your Own Data
Editing is free and browsing is free. Signing in brings you back to this page with your text and your questions, no card required.
