Jev Agent: How to Use Jev AI Inside an AI Agent, Not as One
Sep 30, 2026

Jev Agent: How to Use Jev AI Inside an AI Agent, Not as One

Jev agent, explained: Jev AI is not an agent but the fast decision step inside one. Where it fits in the loop, five patterns with real numbers, and code.

A week after Jev launched, someone asked on Cursor's forum for TypeSafe's model to be added to Cursor Agent. The first reply asked how you would even chat with it. The second put it more bluntly: Jev is not an orchestrator and not a reasoner, it is a very fast, very cheap classifier — use it as a tool your agent calls, not as another agent you talk to.

That exchange is the whole story behind the search Jev agent. There is no agent inside Jev. It does not write text, call tools or run a loop. What it does, in a few hundred milliseconds and for a few thousandths of a cent, is make the small decisions an agent loop is full of: which tool next, whether this call is safe, whether the run is finished. If you are new to the model, start with our explainer on what Jev AI is; this guide is about where it goes in an agent.

A quick note if your search results look odd: Google also completes "jev agent" to Japanese encephalitis agent, because JEV is the abbreviation for Japanese encephalitis virus. This article is about the AI model.

Everything below comes from TypeSafe's own cookbooks, integration guides published by OpenRouter, LangChain and Vercel, and one agent-review call we measured ourselves. By the end you will know which steps of your loop to hand to Jev, how each maps to a question type, what it costs, and where it should stay out.

Last updated September 30, 2026. The integrations below are days or weeks old and changing quickly — the linked pages are the source of truth.

Is There a "Jev Agent"? The Short Answer

No, not from TypeSafe. TypeSafe ships one model family, Jev, served over an API. The sites and repositories named "Jev Agent" are independent projects built around that API.

The distinction matters because the two jobs are different:

An agent's main model (an LLM)Jev
OutputText, code, tool callsTyped answers with probabilities
Holds a conversationYesNo
Plans multi-step workYesNo
Calls tools itselfYesNo — your code acts on its answer
Typical job in a loopReason, generate, actDecide, route, gate, check
Cost profilePays for long outputsPays for input only; output tokens are free

So a working "Jev agent" is an ordinary agent — an LLM, tools and a loop — with Jev called at the points where the loop needs a judgement rather than a paragraph. In that role it is best thought of as a Jev AI tool: a function your code or your agent calls, which returns a number instead of prose.

Where Jev Fits in an Agent Loop

Vercel's announcement of Jev on AI Gateway lists the agent jobs it is suited to: choosing the next tool or subagent, deciding whether to continue, retry, ask the user or stop, scoring urgency or risk before an action, and verifying outputs as a guardrail. Map those onto a loop and each one turns into a specific question type.

Step in the loopQuestion for JevTypeWho acts on it
A request arrivesWhich handler or model should take this?choiceYour router
Before actingWhich tool (or skill) fits this turn — or none?choice + noulThe agent harness
Filling a tool callWhich of these allowed values does the user mean?choice per argumentYour dispatcher
Before a risky callIs this call supported by the request and within policy?noulA gate: approve, block or ask a human
After a stepDid that step succeed? Is the answer grounded in the source?noulThe loop
End of a turnContinue, retry, ask the user or stop?choiceThe loop
After the runDid the run complete the task, and where did it break?noul, choice, scoreYour evaluation pipeline

Three properties make Jev a good fit for these slots specifically.

Many questions, one pass. Jev reads the state once and answers every question in the request in parallel, so asking for the tool and its arguments and a safety check costs one round trip, not three.

Output is free. Agent state is long — conversation history, tool results, policies. With Jev you pay to read it once; there is no long generated answer on top.

Confidence is a second axis. A choice comes back with a probability for every option. A split of 0.52 against 0.48 is not a decision, it is a case to escalate — and in an agent, knowing when not to act is most of what a gate is for.

Five Jev Agent Patterns, With Real Numbers

Most write-ups stop at listing these patterns. Here is what each looks like in practice, with the figures the people who built them published.

1. Pick the tool or skill — and allow "none"

Agents with large skill rosters choose from a truncated index, and on turns where nothing fits they often load something anyway. TypeSafe's skill suggestion cookbook puts two Jev requests in front of that choice: one ranks all 182 skills in Nous Research's Hermes catalog and asks whether the turn needs a skill at all, the second re-reads the top three in full and may reject them all. The winner goes into one line of the system prompt.

Over 488 requests against Claude Haiku 4.5, the agent loaded the wrong skill 7.3% of the time with the suggestion, against 16.8% without it, and loaded a skill when none fitted 4.0% of the time against 9.8%.

The same idea scales down to a browser. Browser Use's open-source Jev Ultrafast agent has Jev choose both the operation (click, type, select, scroll, done) and the page element in a single request, and only calls a small LLM when there is text to write.

2. Fill tool arguments from closed sets

When a tool's arguments come from fixed lists, Jev can fill them without the model inventing a value. TypeSafe's function calling cookbook does this for a market-data assistant with 10 functions and 28 fillable arguments: each argument becomes a choice over the values the function accepts, plus a "stated" noul that asks whether the user said anything about that argument at all. If not, the argument is left out and the function's default applies.

The detail worth copying is how it reports confidence: the call's confidence is its weakest argument, not the product of all of them, because one wrong argument is enough to spoil the call.

3. Gate risky tool calls: approve, block or review

A static human-in-the-loop rule sees a tool name and its arguments. It cannot tell a refund the customer asked for from one they did not. OpenRouter's cookbook for gating agent tool calls with Jev sends the proposed call, the ticket and the policy to Jev as a few yes-or-no questions, then maps the probabilities to three outcomes: run it, refuse it with a reason, or pause for a person. In their runs a check cost under $0.0001.

Their advice on where this stops is the right one: keep the static rule for tools that are always dangerous. The gate is for tools whose safety depends on the situation.

4. Route each turn to the right model

Not every turn needs your most expensive model. LangChain's Jev harness post ships a model-routing middleware that lets Jev pick between a fast and a powerful model from criteria you write, and an auto-mode middleware that uses Jev to check tool calls before they execute. OpenRouter goes one step further with a hosted Jev Router model that picks the model and reasoning effort for every request. We cover the LangChain side in detail in Jev with LangChain, and OpenRouter's in Jev on OpenRouter.

5. Grade every run, not a sample

An agent's final message is the least reliable part of its trace: "Done — refunded and emailed the customer" reads the same whether the refund succeeded or failed three times. Grading the trace itself — did the task complete, which step broke, did it claim success for something that never happened — is a job for noul and choice questions over the transcript.

We measured exactly that on September 25, 2026 against jev-1.13 (build jev-1.13-20260917): a 901-input-token agent transcript with a failure taxonomy of labelled options cost $0.00003784 and came back in 498 ms, round trip from our machine. At that price you can grade agent runs straight from the trace on every run in CI and production, not a weekly sample.

These patterns combine. A support agent can route the ticket, gate the refund and grade the run with three calls. Our Jev AI use cases cover routing, evaluation, retrieval and matching, each with a runnable example and the measured cost of one run.

A Minimal Jev Gate You Can Copy

Here is pattern 3 reduced to one function, calling the jev-ai.org API exactly as our API docs describe. It asks two noul questions about a proposed tool call and returns approve, block or review.

const APPROVE_AT = 0.9;  // both checks must clear this to run unattended
const BLOCK_BELOW = 0.1; // either check below this is a confident "no"

type Verdict = 'approve' | 'block' | 'review';

export async function gateToolCall(
  proposedCall: { tool: string; args: unknown },
  ticket: string,
  policy: string,
): Promise<Verdict> {
  let scores: number[];
  try {
    const response = await fetch('https://jev-ai.org/api/v1/systemone/', {
      method: 'POST',
      headers: {
        Authorization: `Bearer ${process.env.JEV_API_KEY}`,
        'Content-Type': 'application/json',
      },
      body: JSON.stringify({
        model: 'jev-1.13',
        state: { proposed_call: proposedCall, ticket, policy },
        questions: {
          requested: {
            type: 'noul',
            instructions: 'The ticket clearly asks for the action in proposed_call.',
          },
          allowed: {
            type: 'noul',
            instructions: 'The action in proposed_call is permitted by the policy.',
          },
        },
      }),
    });
    if (!response.ok) return 'review';
    const { answers } = await response.json();
    scores = [answers.requested.noul, answers.allowed.noul];
  } catch {
    // Network error or malformed body: fail closed.
    return 'review';
  }

  // Anything other than a clean answer goes to a person, never to "approve".
  if (!scores.every((p) => typeof p === 'number' && p >= 0 && p <= 1)) return 'review';

  if (scores.every((p) => p >= APPROVE_AT)) return 'approve';
  if (scores.some((p) => p <= BLOCK_BELOW)) return 'block';
  return 'review';
}

Three choices in that code are deliberate:

  • Two-sided thresholds. A noul of 0.51 is a coin flip that happens to lean yes. Acting above 0.9, refusing below 0.1 and sending the middle to a person uses the probability instead of throwing it away.
  • Fail closed. A network error, a 429 or a malformed body returns review. A broken check must never become an approval.
  • The state carries the evidence. Jev only sees what you send. If the policy is not in the state, it cannot judge against the policy.

Start with thresholds like these, then look at what lands in review for a week and adjust. If you would rather see the questions work before writing code, paste a ticket and a proposed call into the playground first.

What Jev Can't Do in an Agent

Knowing the edges keeps Jev in the slots where it helps.

  • It cannot plan or write. It will not decompose a task, draft a reply or produce the text a tool needs. Pair it with an LLM for those.
  • It judges only the evidence you give it. OpenRouter's coding-agent recipe makes the point well: Jev cannot know that a command reads a credential file unless the command text says so. Keep irreversible and credential-touching actions on a static deny list that runs before Jev.
  • Confidence is not correctness. Calibration is a property of many predictions, not a promise about one. Route the uncertain middle to a person or a stronger model.
  • Every question must be declared. Jev cannot tell you something you did not ask. Open-ended "what went wrong here?" stays with an LLM.
  • Text in, only. Screenshots, audio and files need turning into text or structured fields before they become state.

Can I Use Jev AI to Trade?

You can use Jev inside trading software, but not the way many search results suggest. Jev does not forecast prices. It answers questions about the text you send it, with calibrated probabilities — so what it can judge is language, not markets.

TypeSafe's own example shows the responsible version. In its function calling cookbook, a market-data assistant turns "compare nvda amd and msft over the past three months" into compare_returns(symbols=['NVDA', 'AMD', 'MSFT'], window='3mo') with a confidence of 0.94. Jev's job there is understanding the request and filling the call — the data, the charts and the numbers come from ordinary code.

That suggests where Jev can sit in a trading tool:

  • Parsing instructions into typed function calls, with confidence you can threshold.
  • Checking an order against what the user actually said before it is sent, with low-confidence cases held for confirmation.
  • Routing or triaging news, alerts and support messages.

What it should not do is decide on its own to buy or sell. Keep execution behind hard limits and human confirmation, and treat any "Jev trading bot" that promises returns with suspicion — especially downloadable apps that ask for exchange keys, which we cover in our Jev AI download guide. Nothing here is financial advice.

How to Add Jev to Your Agent This Week

  1. List the decisions your loop makes — routing, tool choice, approvals, stop conditions. Pick the one that costs the most today, in tokens or in human attention.
  2. Write it as questions. One decision per question; boundary cases in the option descriptions, not in the instructions.
  3. Test on real examples in the Jev AI playground until confident answers match what you would decide.
  4. Wire it in behind thresholds through the Jev AI API, failing closed, and log every answer with its probabilities.
  5. Review the uncertain middle weekly and move thresholds based on what you find.

If your agent is a coding agent — Claude Code, Cursor or Codex — the hooks are different; our guide to using Jev with Claude Code and Cursor covers permission-prompt gating there.

FAQ

Is Jev an AI agent?

No. Jev is a decision model: it returns typed answers with probabilities and never writes text or calls tools. It is used inside agents, for routing, gating and checking.

Is "Jev Agent" an official TypeSafe product?

No. TypeSafe offers the Jev model and its API. Sites and projects called "Jev Agent" are independent.

What is a Jev AI tool?

A function that sends state and typed questions to Jev and returns the answers — usually exposed to an agent as a tool for routing, approval or verification. Every example in our use-case library is one such call, with the payload you can reuse.

How much does one agent decision cost?

It depends on how much state you send, since only input is billed. Our measured agent-transcript review read 901 tokens and cost $0.00003784. OpenRouter reports its tool-call gate at under $0.0001 per check.

Can Jev replace the LLM that runs my agent?

No. The agent still needs a generative model to plan, write and call tools. Jev replaces the judgement calls you might otherwise make with that LLM — which is usually where it saves the most time and money.

The Bottom Line

A Jev agent is an ordinary agent that stops asking its LLM to make every small decision.

  • Jev is not an agent and cannot be one: no text, no tools, no loop.
  • It fits the judgement points — tool choice, argument filling, approvals, stop conditions and run grading — each mapping to a choice, noul or score question.
  • Published results are concrete: wrong skill loads cut from 16.8% to 7.3% in TypeSafe's test, and a full agent-run review cost us under four thousandths of a cent.
  • Keep static rules for always-dangerous actions, and send the uncertain middle to a person.

Pick one decision your agent makes a thousand times a day and try it. Start from the closest match in our Jev AI use cases — routing, grading and grounding checks, each runnable in the browser with its measured cost.

Sources

Cookbook and integration figures are those published by TypeSafe, OpenRouter and LangChain as of September 30, 2026. The agent-review cost and latency are our own measurement from September 25, 2026 and reflect our payload and network location.

Try Jev AI Free in the Playground

Wondering how to try Jev AI? Sign in, take the five welcome credits and run it — no card required.