TypeSafe AI Jev Benchmarks: What the Numbers Actually Show
Sep 30, 2026

TypeSafe AI Jev Benchmarks: What the Numbers Actually Show

TypeSafe AI Jev benchmarks, sorted: vendor workflow evals, independent tests and our own calls. How smart Jev is, where it wins, and every caveat.

"193.6× faster, 444.6× cheaper." That line on TypeSafe AI's homepage is the most-repeated number about Jev, and it is usually followed by a question nobody answers cleanly: compared with what, measured how, and by whom? Search for Jev benchmarks and you get a pile of pages quoting the same multipliers, a few quoting different ones, and almost no one separating what TypeSafe measured from what anyone else did.

There is a reason the picture is murky. TypeSafe has deliberately not published scores on public benchmarks like MMLU or LMArena, and it has said it does not plan to. What exists instead is one vendor-run evaluation with unusual rules, a homepage demo, a handful of independent tests on GitHub, and customer anecdotes.

We run jev-ai.org, a playground and API built on the Jev model, so we have a practical interest in knowing how good it really is. For this article we pulled every data point from TypeSafe's evaluation site workflow by workflow, read the code it published to reproduce them, collected the independent tests we could verify, and added the six calls we measured ourselves. Every number below says who produced it. If you need the basics first, read what Jev AI is and come back.

Last updated September 30, 2026. Scores are as published on that date; TypeSafe says it will only publish one-off evals when the product changes, so check the sources below for anything newer.

In medical search results, JEV is the Japanese encephalitis virus. This article is about Jev, the decision model from TypeSafe AI.

TypeSafe AI Jev Benchmarks at a Glance

ClaimNumberWho measured itThe catch
Speed and cost vs LLMs193.6× faster, 444.6× cheaperTypeSafeFrom its own workflow evals; TypeSafe calls these "the higher end of real world gains"
Workflow accuracy, four-task average67.8%, at about $0.0004 and 0.4 s per caseTypeSafe"Accuracy" means agreement with answers from two frontier LLMs, not ground truth
Best LLM on the same evals74.1%, at about $0.084 and 23 s per caseTypeSafeThe top model beat Jev on accuracy in all four workflows
Side-by-side demo$0.000081 in 0.114 s vs $0.013880 in 8.566 sTypeSafeOne short, dense query that TypeSafe says flatters Jev
Zero-shot classification pilot0.910 and 0.870 accuracy on two datasets; 0.480 on a thirdIndependent developer300 examples, compared with one open classifier, not an LLM
Real customer workload5–18× faster, with greater accuracyVercel engineer, via TechCrunchOne safety classifier; no published dataset
Our six production-shaped calls$0.000023–$0.000038 each; median round trip about 310 msjev-ai.orgOur payloads, our network location

The short reading: the price is real, the speed is real, and the intelligence claim is "roughly as good as a mid-tier frontier LLM on decision workflows", with real weak spots. The rest of this article shows where each piece comes from.

Why There Is No MMLU Score for Jev

Most model announcements lead with a leaderboard. TypeSafe went the other way. In the FAQ of its launch post, it says it "deliberately chose not to publish performance against public benchmarks" and plans only one-off evals when it ships product updates. Its stated reasons: put no weight on public benchmarks, and encourage users to build evals for their own use cases, which it argues are easy to write for decision tasks.

There is also a practical reason. Jev cannot write text, so benchmarks built around free-form answers — essays, code, chain-of-thought math — do not apply to it in their usual form. You can only compare Jev with an LLM on tasks where both are forced to pick from the same set of answers.

So when a page quotes a "Jev MMLU score" or an arena rank, treat it as someone else's experiment, not something TypeSafe published.

TypeSafe's Workflow Evals, Explained

The numbers behind the homepage multipliers come from evals.typesafe.ai, which TypeSafe calls "workflow evals". The method is worth understanding, because it is not how most benchmarks work.

  • Four workflows, written in code. Security incident triage (240 cases), agent trace review (111), invoice processing (150) and customer service (204). Each breaks a real process into narrow questions — yes/no, pick-one, rate-on-a-scale — and code combines the answers into actions.
  • The harness is assumed correct. Every model gets the same workflow and the same questions. Only the model changes.
  • The "right answer" is what two frontier LLMs said. Reference labels are the average of GPT-6 Astra and Claude Fable 5.1, both at high reasoning effort, answering every question in the harness. A model's accuracy is how often its final actions match that consensus exactly.
  • Everyone else runs at default settings. All other models use their provider's default reasoning setting. LLMs are driven through TypeSafe's own wrapper that constrains them to the same typed answers.

Here is the four-workflow average for each model in workflow mode, as plotted on TypeSafe's site. Cost and time are per case.

Model (as labelled on the site)Agreement with referenceCost per caseTime per case
OpenAI "sol"74.1%$0.083623.3 s
Claude Opus 573.1%$0.176137.8 s
OpenAI "terra"67.9%$0.030410.1 s
Jev67.8%$0.00040.4 s
Claude Sonnet 567.8%$0.117478.1 s
OpenAI "luna"66.8%$0.003312.9 s
DeepSeek V4 Pro (via Fireworks)65.5%$0.041386.5 s
DeepSeek V4 Flash (via Fireworks)64.4%$0.005951.9 s
Claude Haiku 4.553.6%$0.019512.5 s

The same page also scores each LLM with the whole policy written as a single prompt instead of a workflow. Every model does worse that way; the best prompt-mode score is Claude Opus 5 at 64.8%.

Jev workflow by workflow

The average hides a lot. Split by task, Jev is never the most accurate, and its rank moves from third to eighth.

WorkflowJevBest modelJev's rank of 9
Security incidents61.7%, $0.0001, 0.3 sOpus 5: 66.2%, $0.0574, 15.1 s3rd
Agent trace observability71.6%, $0.0003, 0.5 s"sol": 76.6%, $0.0575, 40.3 sTied 6th
Invoice processing61.8%, $0.0011, 0.5 s"sol": 79.1%, $0.2152, 34.3 s8th
Customer service76.0%, $0.0001, 0.4 s"sol": 78.3%, $0.0323, 10.1 s4th

Invoice processing is Jev's weakest result by a wide margin. It is tempting to blame numbers, but TypeSafe's invoice page says sums, dates, account numbers and statuses are computed in code rather than asked, so the gap is not simply arithmetic, and TypeSafe does not explain it. Treat it as a reminder to test on your own document-heavy workflows.

What the multipliers look like against a fair opponent

TypeSafe says the headline 193.6× and 444.6× figures come from these workflows, but not which model they are measured against. So compare like with like: against the models that land at similar accuracy, the gap is smaller than the headline but still large. These are our ratios from the rounded figures above, so treat them as approximate:

  • vs OpenAI "terra" (67.9%) — about 76× cheaper and 25× faster for the same accuracy.
  • vs OpenAI "luna" (66.8%), the cheapest LLM near Jev's accuracy — about 8× cheaper and 32× faster.
  • vs the most accurate model, "sol" (74.1%) — about 200× cheaper and 58× faster, for 6.3 points less agreement.

That is the honest shape of the claim: at equal accuracy, Jev is one to two orders of magnitude cheaper and faster than the LLMs TypeSafe tested. Whether 67.8% is good enough depends entirely on your task and your fallback.

The caveats TypeSafe itself lists

Credit where due: TypeSafe publishes most of the reasons to be careful. From the launch post and eval site:

  • The workflows were written by TypeSafe's own model-capabilities team, so "some bias could exist", although TypeSafe says they were not chosen to flatter Jev and are not in its training data.
  • Using OpenAI and Anthropic models as the reference biases results toward those models.
  • LLMs were run through TypeSafe's structured wrapper, which TypeSafe says is the most accurate way to get decisions from them but slower and more expensive than asking for plain answers.
  • Timing was generally measured from laptops on the US West Coast, close to TypeSafe's service.
  • TypeSafe expects the 193.6× and 444.6× figures to be "on the higher end of real world gains".

The code to rerun everything is public: WorkflowEvals on TypeSafe's GitHub, Apache-2.0 licensed, with inputs and reference labels downloaded from Hugging Face. You need your own API keys for every provider you want to test.

The Side-by-Side Demo Number

TypeSafe's homepage shows the same query sent to Jev and to an LLM: Jev finishes in 0.114 s for $0.000081, the LLM in 8.566 s for $0.013880. The launch post names the LLM as GPT-5.6 Terra at default reasoning — chosen, TypeSafe says, because it is the most comparable to Jev in intelligence on average. That is about 75× faster and 170× cheaper.

TypeSafe is upfront that this demo is favourable: the questions were simplified with readable keys, and the state is a short, dense paragraph chosen to highlight the difference in sampling. The one disagreement between the two models in the recorded run was on a churn-likelihood question TypeSafe calls genuinely ambiguous. Read it as an illustration of why Jev is fast — all answers in parallel instead of token by token — not as a typical ratio.

Independent Jev Benchmarks

Outside TypeSafe, the useful evidence so far comes from developers who published their code and raw results, plus one customer quoted on the record. These are self-reported by their authors and small, but they are reproducible and none of them is affiliated with TypeSafe.

TestWhat was comparedResultCaveat
jev-benchmarks (AbdelStark)Jev vs the open GLiNER2.5 classifier, zero-shot text classificationJev 0.910 vs 0.700 on AG News; 0.870 vs 0.610 on 72-label Banking77; 0.480 vs 0.440 on emotion, where Jev was also worse calibrated300-example pilot; public datasets may be in training data
jev-benchmark (wondertwins)Jev on chess and on "which NPC is the player talking to?"NPC addressee F1 0.96 on clean text, 0.93 on noisy transcripts; chess no better than random from a raw board, about 950 Elo with move facts computed in codeChess is deliberately outside Jev's design
jev-plays-doom (tirukovelamanoj)Jev vs a hand-coded aiming script in a Doom scenarioTied at 6.55 kills per episode; 212 ms latency across 2,642 callsOne scenario; results swung with option wording
Vercel, quoted by TechCrunchJev vs an LLM-based safety classifier5–18× faster, with greater accuracyAnecdote; no data published

The pattern across all four is consistent. Jev is strong when the task is a judgment over text with well-defined options — topic and intent classification, who is being addressed, is this unsafe. It gets weak when the task needs calculation, search or subtle reading of emotion. The Doom result is covered in detail in our explainer on why Jev AI can play Doom.

Community leaderboards such as JevBench have also appeared. They are not run by TypeSafe; check who wrote the tasks and how the reference answers were made before comparing their scores with anything above.

Our Own Measurements

We cannot rerun TypeSafe's accuracy evals without every provider's keys, but we can report what real calls cost and how long they take. On September 25, 2026 we sent six production-shaped requests to jev-1.13 (build jev-1.13-20260917) through OpenRouter — answer grading, intent routing, citation checks, entity matching, agent-run review and web-page triage — each asking several typed questions at once.

Every call cost between $0.000023 and $0.000038, exactly in line with TypeSafe's list price of $0.042 per million input tokens, with output tokens counted and billed at zero. The median round trip from our machine was about 310 ms, with one outlier near a second — inside TypeSafe's published 70–500 ms range once you add the network hop from our location. The per-call breakdown, workload by workload, is in the Jev explainer linked above. On price and speed, the vendor numbers held up for us. Accuracy is the claim you have to test yourself.

How Smart Is Jev AI?

About as smart as a capable mid-tier LLM on decisions, and not smart at all at things it was never built to do. On TypeSafe's workflows it matched Claude Sonnet 5 and OpenAI's "terra" on average and trailed the best models by six points. On independent classification tests it clearly beat a dedicated open classifier on two of three datasets.

Where it falls down is documented by TypeSafe itself. Its Jev 1.13 jaggedness page lists the failure modes it knows about:

  • Literal reading. It answers the question you wrote, not the one you meant.
  • Math and numbers. It does not count reliably and is "not a calculator"; hex colours and raw numbers do worse than words.
  • Dates and times. Ordering dates or checking windows is unreliable; extract the parts and compare in code.
  • Indirection. Double negatives and multi-hop questions cost accuracy.
  • Large, noisy state. Accuracy falls as irrelevant detail grows.
  • Adversarial content. Text written to steer the answer can move it.
  • Generation. It is not trained to produce text at all.

That list lines up with much of the independent spread: chess needs search, emotion is subtle — and routing, triage and addressee detection are simple judgments where Jev does well. TypeSafe's homepage FAQ says the same in one line: tasks like complex mathematics or chess-like planning "may be better suited to large reasoning models".

How Is Jev AI Unique?

Benchmarks score one thing — did it pick the right answer — and miss most of what makes Jev different:

  • Calibrated confidence on every answer. Each decision comes with probabilities and a confidence your code can threshold. In the independent classification pilot, that let Jev automatically handle 83–86% of examples at an error rate of 5% or less on two datasets, against 24–27% for the open classifier. On the emotion dataset, the same approach covered almost nothing — calibration is only as good as the task fit.
  • Questions answered in parallel. Jev reads the state once and answers many questions about it at the same time, so the fourth question costs almost nothing extra.
  • Answers that cannot break your types. A choice question can only return one of your labels. It can pick the wrong one, but it cannot invent one or return malformed output.
  • Input-only pricing. $0.042 per million input tokens, output free. TypeSafe puts that at 238× below the input price of Claude Fable 5.1.
  • Consistency over determinism. TypeSafe designs Jev to give similar answers to inputs that mean the same thing, even when the wording changes.

If you are deciding whether that trade suits your stack, our comparison of Jev vs LLM structured output walks through it, and we are rounding up Jev AI alternatives for the cases where it does not.

How to Read Any Jev Benchmark

Most pages quoting Jev numbers skip the questions that decide whether a number means anything. Before you trust one, check:

  1. Who wrote the tasks? Vendor-written workflows and tester-written datasets both carry bias; say which.
  2. What counts as correct? Human labels, a public dataset, or another model's answers? TypeSafe's "accuracy" is agreement with two LLMs.
  3. What reasoning setting did the LLMs use? Default, high, or none changes both accuracy and cost dramatically.
  4. Where was latency measured from? A laptop next to the data centre and a server on another continent give different numbers for the same model.
  5. Was cost list price or a router's price? Routing through a third party can change both.
  6. Was calibration measured? Accuracy alone ignores Jev's main selling point: knowing when it is unsure.

Run Your Own Jev Benchmark in an Afternoon

TypeSafe's advice is to build your own eval, and for decision tasks it genuinely is quick.

  1. Collect 50–300 real examples from your own traffic, with the answer you would want.
  2. Write the decision as narrow questions — one choice, score or noul per judgment — and keep arithmetic in code.
  3. Pin the version. Use jev-1.13 rather than a moving alias so reruns are comparable.
  4. Record four things: accuracy, how many cases clear your confidence threshold, latency from where you will actually run it, and cost per case.
  5. Compare against what you use today. TypeSafe's MIT-licensed System One adapter sends the same questions to OpenAI, Anthropic or Gemini models for a like-for-like comparison. Our guide to Jev on GitHub covers the official repositories.

The fastest way to run step 4 without writing a harness: upload your rows as a CSV to batch processing on jev-ai.org, run all of them in one pass, and export the answers and probabilities to compare with your labels, or try a single example first in the Jev AI playground.

FAQ

Is Jev better than GPT or Claude?

Not on raw accuracy. On TypeSafe's own workflow evals, the best OpenAI and Anthropic models agreed with the reference answers more often than Jev did in every workflow. Jev's case is that it gets close — level with Claude Sonnet 5 on average — at a small fraction of the cost and time.

Has TypeSafe published MMLU, GPQA or arena scores for Jev?

No. TypeSafe says it deliberately chose not to publish results on public benchmarks and plans only one-off evals at product updates. Most public LLM benchmarks also assume text generation, which Jev does not do.

Can I reproduce TypeSafe's Jev benchmarks?

Yes. TypeSafe published the code as WorkflowEvals on GitHub, with datasets on Hugging Face. You will need API keys for Jev and for each LLM you want to compare, and scores are measured as agreement with the published reference answers.

How fast is Jev in practice?

TypeSafe publishes 70–500 ms end to end. Independent developers reported medians of about 212 ms and 236–256 ms from their locations, and our own calls from a different region had a median round trip of about 310 ms.

The Bottom Line

The TypeSafe AI Jev benchmarks tell a narrower story than the homepage multipliers, and a more useful one.

  • On TypeSafe's workflow evals, Jev scored 67.8% agreement — level with mid-tier frontier LLMs, six points behind the best — at about $0.0004 and 0.4 seconds per case.
  • Against LLMs of similar accuracy, it is roughly 8–76× cheaper and 25–32× faster; TypeSafe's own 444.6× sits at what it calls the higher end.
  • Independent tests agree on the shape: strong on classification and routing, weak on numbers, planning and subtle judgment.
  • The price and speed are easy to verify on your own bill. The accuracy is not — measure it on your own data.

The quickest way to find out whether Jev is smart enough for your decision is to run it. Open the playground with a real example, then score a few hundred rows through the Jev AI API or batch upload and compare with what you use today.

Sources

All TypeSafe figures are vendor-reported and were read from its site on September 30, 2026; costs on the eval site are rounded to four decimal places, so our ratios are approximate. Independent results are self-reported by their authors. The six measured calls are our own and reflect our payloads and network location.

Try Jev AI Free in the Playground

Wondering how to try Jev AI? Sign in, take the five welcome credits and run it — no card required.