Retrieval

Web Context: Decide What a Fetched Page Is Before You Pay to Read It

A browsing agent fetches ten pages to answer one question. Two of them are consent walls, three are category listings that only link elsewhere, one is a blog post with the answer buried under a fitting guide, and one is the product page it wanted. Feeding all ten into a context window is how a cheap question becomes an expensive one — and how the relevant sentence gets lost among forty thousand tokens of navigation.

The question "is this page worth reading" is a decision, and it is much cheaper than reading the page with a frontier model. Answer it first, and the expensive model only ever sees pages that earned their place.

  • Page-type labels you declare
  • Relevance to the query that fetched it
  • Keep / extract / drop as an ordered score
  • Output tokens free

The problem

The Job: Keep the Context Window Honest

Fetched web pages are mostly not content. Navigation, cookie banners, newsletter interstitials, related-product rails and footers routinely make up most of the tokens on a page, and boilerplate strippers only get the obvious ones. What survives still has to be judged against the question that sent you there, which no heuristic can do.

The cost of getting this wrong compounds twice. You pay for the junk tokens on the way in, and you pay again in quality: a long context full of near-misses makes the generation step worse, not better, because the relevant passage now competes with six plausible distractors.

It also decides how an agent behaves. "This page is a consent wall" means fetch a different URL. "This is a listing page" means follow a link rather than read. "One useful section" means extract instead of quote. Those are branches, and a branch wants a label, not a summary.

Why a decision model

Why a Decision Model Beats Asking an LLM to Summarise the Page First

Summarising to decide whether to read is the expensive way round. Decide first.

The filter must cost a fraction of what it saves

A gate that costs as much as the call it is protecting is not a gate. At the price measured below, triaging ten pages costs a small fraction of a cent while removing most of what would have gone into the context window.

Page type is a branch, and branches need labels

Consent wall, listing, documentation, product page, article. Declared as labels, they drive what the crawler does next. Extracted from prose with a regex, they drive an incident.

Relevance is judged against the query, not in the abstract

Put the query in the state alongside the page. "Does this contain what is needed to answer this question, even if it is not the page’s main subject" is precisely the judgement that lets you keep a fitting guide that happens to list the dimensions.

An ordered keep/extract/drop scale drives the pipeline

One score with a tier distribution gives you three actions and their thresholds: drop it, extract the useful section, or keep the page whole.

Blocked and empty pages are a first-class answer

A cookie wall is not a low-relevance page, it is a fetch that failed. Giving it its own label means your crawler can retry or route around it instead of recording a bad result.

Playground

Triage a Page a Crawler Just Fetched

A kitchen fitting guide that buries the dimension the query asked for, under a subscribe prompt and a cookie notice. Run it and see the keep/extract/drop score.

Decide what a fetched page is before you pay to reason about it

Model

1Text

Fetched page

802 / 100,000

2Questions

4 in this request

Editing is free. Sign in and you come straight back here with your text and questions — no card needed.

3Answers

Answers appear here

Each answer returns a probability for every option and a confidence score.

In code

What You Would Write Next

Truncate the page before you send it: you are deciding whether the page is worth reading, and the first few thousand tokens almost always settle that.

Filter before the context windowtypescript
const triaged = await Promise.all(
  fetched.map(async (page) => {
    const { answers } = await jev.decide({
      state: [
        `QUERY THAT SENT US HERE\n${query}`,
        `URL\n${page.url}`,
        `FETCHED TEXT (truncated)\n${page.text.slice(0, 8000)}`,
      ].join("\n\n"),
      questions: {
        answers_query: {
          type: "noul",
          instructions:
            "Does this page contain the information needed to answer the query, even if it is not the main subject of the page?",
        },
        page_type: {
          type: "choice",
          instructions: "Select what kind of page this is.",
          criteria: {
            documentation:   "Reference material, specifications or manuals.",
            product_page:    "One purchasable product, with its details.",
            editorial:       "A guide, blog post or review.",
            listing:         "A category or search page that mostly links elsewhere.",
            blocked_or_empty:"Consent wall, paywall, error, or no usable content.",
          },
        },
        context_value: {
          type: "score",
          instructions:
            "Rate how much of this page is worth putting into a model context window for this query.",
          criteria: [
            "None - drop it",
            "One or two useful sentences buried in boilerplate",
            "A useful section worth extracting",
            "Substantially on-topic - keep the page",
          ],
        },
      },
    });
    return { page, ...answers };
  })
);

// Three branches, one score.
const keep    = triaged.filter((t) => t.context_value.score >= 2.5);
const extract = triaged.filter((t) => t.context_value.score >= 1 && t.context_value.score < 2.5);
const retry   = triaged.filter((t) => t.page_type.choice === "blocked_or_empty");

The endpoint, both criteria container shapes and the full error contract are in the developer docs, and the API page has a brief you can paste straight into a coding agent.

Real cost

One Run Costs $0.000033

Measured, not estimated. On 2026-09-25 we sent the exact payload the playground above loads to jev-1.13-20260917 and read the numbers below straight out of the response’s usage block — measured on the fetched page in the playground above, with all four questions in one call.

It checks out against the list rate of $0.042 per million input tokens: 794 ÷ 1,000,000 × 0.042 = $0.000033. The 120 output tokens were counted and not billed, which is why asking four questions about one state costs barely more than asking one.

Input tokens
794

The only thing charged

One run
$0.000033

Round trip 0.28s from a laptop, network included

1,000 runs
$0.033

Same questions, same length of input

1,000,000 runs
$33.35

At the model list rate, before our margin

Those are model costs. On this site a playground run costs 1 credit from your credit balance, and an API call debits exactly those input tokens from your token balance — whichever balance applies, output stays free and a failed request is never charged. The pricing page has the per-plan rates, and your own runs will differ in length from this example, so treat this as a worked figure rather than a quote.

Questions

Web Context: The Questions People Actually Ask

Should I send the whole page or a truncated version?

Truncate. You are deciding whether the page deserves a full read, and the first few thousand tokens nearly always settle it. Sending the whole page to decide whether to send the whole page defeats the purpose.

Does this fetch pages for me?

No. This is a decision endpoint, not a crawler. You fetch with whatever you already use — your own client, a scraping API, a headless browser — and send the text you got back.

How does it compare to a reranker?

A reranker orders passages by relevance to a query and is excellent at that. This answers different questions in the same call: what kind of page this is, whether it is a consent wall, whether it is primarily commercial. Those drive crawler behaviour, which an ordering cannot.

Can I run this on every page an agent fetches?

That is the intended use, and the cost is why it is reasonable. Triaging ten pages per query costs a small fraction of a cent, against a context window you would otherwise pay frontier prices to fill with navigation.

What about HTML — should I strip it first?

Strip it. Tags and inline styles are tokens you pay for and signal the model does not need. Extracted text, with the URL and the query alongside it, is the cheapest input that still supports the judgement.

Will it tell me whether a page is trustworthy?

It will tell you what kind of page it is and whether it is primarily selling something, which are observable properties of the text. Do not ask it to adjudicate truth — for that, retrieve the authoritative source and use a grounding check.

Run It on Your Own Data

Editing is free and browsing is free. Signing in brings you back to this page with your text and your questions, no card required.