The clip is everywhere: a green-on-black dashboard, a Doom marine strafing and firing, and a caption saying an AI model is making the calls. Then the obvious question. Jev is the model that famously cannot write a sentence. How does something that never generates text play a first-person shooter, and in real time?
The short version is that Jev does not really "play" Doom the way a person does. It never sees the screen, it does not aim, and it does not plan a route through the level. What it does is answer a handful of small, typed questions about the game, about ten times a second, and some ordinary code turns those answers into button presses. That combination is the whole trick, and it is the same pattern you would use to put Jev inside any piece of software.
We run jev-ai.org, a playground and API built on the Jev model, so we read the demo the way a developer would. This article is built from TypeSafe AI's own launch post and the Doom video's captions (timestamps quoted below), The Register's launch coverage, TypeSafe's published model limits, and three open-source Doom agents on GitHub whose READMEs and published results we read. If you are new to the model itself, start with what Jev AI is and how it differs from an LLM.
Last updated September 30, 2026. TypeSafe has promised a full walkthrough of the demo but had not published it when we checked. Limits and prices below are as published on that date.
A quick note if you arrived from a plain search for "jev": in medical results, JEV is the Japanese encephalitis virus. This article is about Jev, the AI model from TypeSafe AI.
Why Can Jev AI Play Doom? The Short Answer
Jev can play Doom because a shooter, broken down the right way, is a stream of small decisions, and small decisions are exactly what Jev is built for. Four things have to be true at once, and in TypeSafe's demo they are:
- The game is turned into text. Jev never sees pixels. It is fed a structured description of the situation — health, nearby enemies, incoming projectiles, available pickups.
- Each moment becomes a few typed questions. "Should the trigger be held down right now?" "What is the player's top priority?" "Should the player be dodging?" Each has a fixed set of allowed answers.
- The answers come back fast enough. TypeSafe says a whole batch of questions returns in about 100 ms, which works out to about 10 decisions a second.
- The answers are cheap enough to ask constantly. The engineer behind the demo put the cost of 10 queries a second at roughly $7 an hour.
Code does everything else: it builds the state, lists the legal options, turns the chosen answers into movement and fire, and loops. TypeSafe calls this "a composition of AI primitives."
What TypeSafe Actually Showed
The first-hand source is the Doom section of TypeSafe's launch post from September 15, 2026, which embeds a 2 minute 42 second video. We pulled the video's auto-generated English captions rather than paraphrasing a summary. Here is what the sources actually say, with where each claim comes from.
| Claim | Source | Where |
|---|---|---|
| Jev makes the decisions "inside the software loop" — where to move, where to aim, whether to hold the trigger | Demo video | 0:07–0:12 |
| About 100 milliseconds "per battery of questions", about 10 decisions a second | Demo video | 0:11–0:20 |
| Input is "a structured description of the situation": health, nearby enemies, incoming projectiles, available pickups | Demo video | 0:29–0:39 |
| Answers are composed into the player's next action, repeated as the game changes | Demo video | 0:52–0:58 |
| The default strategy is written as text and can be changed live; told "do not fire, simply dodge", the player's behaviour changes immediately | Demo video | 1:05–1:23 |
| A second composition deciding where to explore next is needed to get through a level | Demo video | 1:32–1:40 |
| 10 queries a second costs about $7 an hour | Launch post | Doom section |
| The demo runs on structured state as text, "not on images (yet…)" | Launch post | Doom "Nuance" note |
| "A non-AI doom bot could play better" | Launch post | Doom "Nuance" note |
The video's cover frame shows the decision panel itself. Under a heading called Judgments, each question is listed with a bar for every allowed answer:
- FIRING — "Should the player's trigger be held down right now?" with the options
fireandhold_fire;hold_firewins at a confidence of 0.95. - GOAL — "Considering
player,enemies, anditems, what is the player's highest-priority goal right now?" with options includingscout,stock_ammo,add_armorand a restore-health option, which wins at 0.70 confidence. - DODGE — a follow-up that takes the current top priority as given ("…trying to restore health with medikit B. What does this exact moment call for?") and chooses between
carry_on,dodge_left, a dodge-right option and others.
Above the questions sits a "standing order" box — the text strategy the model reads on every tick — and a small counter reading 10 enemies spawned, 4 killed. The Register, which covered the launch the next day, summed it up the same way: Jev can play Doom "when fed structured data describing the player's game state."
Note what is not in any of this: no win rate, no level completed, no comparison with a human player. TypeSafe presented Doom as a fun demo of real-time decisions, and said so.
How a Decision Model Drives a Game Loop
A Jev call has one shape: some state goes in, a set of typed questions goes with it, and a typed answer with probabilities comes back for each question. A game loop simply calls that shape over and over.
- Observe. Code reads the engine — player health, ammo, enemy types, distances and bearings, projectiles, pickups — and writes it into a small JSON object.
- Ask. The same request carries every question for this tick: fire or hold, top goal, dodge direction, where to move. Each is a
choiceover options the code already knows are legal. - Compose. Code combines the answers into one action. If the goal is "restore health" and the dodge answer is
dodge_left, the motor command is "strafe left toward the medikit". - Act and repeat. The engine keeps running while the next request is in flight, then the loop starts again with fresh state.
Here is what one tick could look like as a request to the jev-ai.org API. The question wording is taken from the demo's dashboard; the state fields and labels are ours, for illustration.
POST /api/v1/systemone/
{
"model": "jev-1.13",
"state": {
"standing_order": "Stay alive first. Kill enemies that block the way to the exit.",
"player": { "health": 34, "armor": 0, "ammo": 18 },
"enemies": [
{ "kind": "imp", "distance": 310, "bearing_deg": -20, "throwing_fireball": true }
],
"items": [{ "kind": "medikit", "id": "B", "distance": 140, "bearing_deg": 35 }]
},
"questions": {
"firing": {
"type": "choice",
"instructions": "Should the player's trigger be held down right now?",
"criteria": {
"fire": "Hold the trigger down this tick.",
"hold_fire": "Do not fire this tick."
}
},
"goal": {
"type": "choice",
"instructions": "Considering player, enemies and items, what is the player's highest-priority goal right now?",
"criteria": {
"kill_enemy": "Engage the nearest dangerous enemy.",
"restore_health": "Reach a health pickup.",
"stock_ammo": "Collect ammunition.",
"scout": "Explore toward unseen areas."
}
},
"incoming_danger": {
"type": "noul",
"instructions": "Will a projectile hit the player within the next second if they keep their current course?"
}
}
}
Three details in that request carry most of the weight:
- All the questions ride in one call. Jev reads the state once and answers every question in parallel, so asking four things costs about the same time as asking one. Sending them as separate requests would multiply the latency per tick.
- Every answer is one of the options you listed. Jev cannot return a move that does not exist, a malformed JSON object or an apology. That is what lets code act on the answer without a parser in between.
- Each answer carries a confidence. Code can act on a 0.95
hold_fireand treat a 0.52 split differently — fall back to a scripted rule, or keep the previous action for one more tick.
The standing order is plain text in the state, which is why the demo can switch strategy mid-game: change the sentence and the very next batch of answers is judged against it. There is no retraining and no prompt engineering in the chat sense, only a different input.
The Numbers That Make It Work: Latency, Frequency and Cost per Frame
Whether a model can drive a game loop is mostly arithmetic. Here are the figures that matter, each with its source.
| Figure | Value | Source |
|---|---|---|
| Doom engine speed | 35 tics per second | ViZDoom, as used by the open-source agents below |
| Jev decisions in the demo | About 10 per second | TypeSafe demo video |
| Time per batch of questions | About 100 ms | TypeSafe demo video |
| Jev end-to-end response time | 70–500 ms | TypeSafe launch post |
| Independent median, one Doom agent | 212 ms across 2,642 calls | tirukovelamanoj/jev-plays-doom on GitHub |
| Cost of the demo loop | About $7 per hour at 10 queries/second | TypeSafe launch post |
| TypeSafe list price | $0.042 per million input tokens; output free | TypeSafe models page |
| TypeSafe rate limits | 40 requests/second and 100K tokens/second | TypeSafe models page (marked as adjusting) |
Frequency. At 10 decisions a second, the model makes a fresh call every three to four engine tics. That is slower than a human's reflexes but fast enough to track an enemy, fire and sidestep a fireball, which is what the video claims.
Cost per frame. Ten queries a second is 36,000 calls an hour. By our arithmetic, at about $7 an hour each call costs roughly $0.0002, and at TypeSafe's list price that is about 4,600 input tokens per call. TypeSafe has not published a per-call figure; this is our back-of-envelope estimate, assuming the $7 was at list price, but a plausible size for a JSON game state plus four or five questions. Output is free, so asking more questions about the same state adds very little.
Headroom. Those same numbers put the demo at about a quarter of TypeSafe's published per-second request limit and a little under half of its token-per-second limit. A Doom bot fits inside one account's limits; a server running a hundred of them would not, without a custom plan.
Why a chat model could not do this. TypeSafe's launch table puts end-to-end response time for frontier LLMs at 3 to 329 seconds. In its own side-by-side demo, GPT-5.6 Terra took 8.566 seconds against Jev's 0.114 seconds on the same short query. At 8.6 seconds a decision, a Doom player would get one instruction every 300 engine tics — the fireball has landed long before the answer does. Speed is not a nice-to-have here; it is the precondition.
Our own measurements point the same way. In six real jev-1.13 calls we timed on September 25, 2026 (build jev-1.13-20260917, through OpenRouter, network hop included), the median round trip was about 310 ms — slower than the demo's 100 ms because of distance to TypeSafe's West Coast service, but still enough for three to five decisions a second.
What Jev Is Not Doing in Doom
Most write-ups stop at "an AI played Doom", which leaves readers imagining a model watching the screen and outplaying people. The sources describe something narrower, and the gap matters if you want to build anything similar.
| Common assumption | What the sources show |
|---|---|
| Jev watches the screen | No. The input is text: a structured description of the game. TypeSafe says images are not supported "yet", and Jev's model page lists text-only input. |
| Jev aims the gun | Not precisely. The open-source agents compute enemy bearings in code and leave the model the judgment calls — fire or hold, engage or flee. |
| Jev plans a route through the level | No. TypeSafe says a second composition deciding where to explore is needed to get through a level, and its own FAQ says "chess-like planning" is better left to reasoning models. |
| Jev plays better than a bot | No. TypeSafe itself says a non-AI Doom bot could play better. In one independent test, Jev tied a hand-coded aiming script. |
| The model learned to play Doom | No. The same weights serve every customer. The Doom behaviour comes from the state, the questions and the standing order you send. |
What is impressive is the part a scripted bot cannot do: you can change the strategy with a sentence, and the bot can react sensibly to states nobody wrote a rule for. TypeSafe's own framing is that a bot "reactive to different representations of game state" and able to follow instructions was the point.
What Independent Developers Found
TypeSafe has not released its Doom code, but within days of launch several developers published their own Jev Doom agents on GitHub. None is affiliated with TypeSafe, and each says so or makes no such claim. Their READMEs are the best public evidence of how the pattern behaves outside a promotional video.
- jev-plays-doom by tirukovelamanoj runs ViZDoom's
defend_the_centerscenario, where the player cannot move and must turn and shoot. Over 20 seeded episodes, Jev averaged 6.55 kills — exactly the same as a hand-coded aiming script — with 212 ms latency across 2,642 calls. The more useful finding is about wording: when the fire threshold was stated in the option descriptions, Jev averaged 6.20 kills; when the options only described intent, it never chose to attack and finished at −0.60. Confidence fell from 0.88 to 0.41 as the phrasing got vaguer, which is the calibration doing its job. - jev-doom by olivier-motium runs Freedoom in a local dashboard with the choices, probabilities, latency and cost visible, and lets you edit standing orders mid-run. It sends structured observations, not pixels, runs 90-second sessions by default behind a $15 spending cap, and states plainly that it makes no level-completion or performance claim.
- doom-jev by AmoghCreator decouples the 35-tic engine loop from an asynchronous inference loop at about 10 Hz, so the game never freezes while a request is in flight. Jev picks the high-level goal; trigonometry in code steers the crosshair.
You will also find a "JevDoom" simulation on jev.works. That site belongs to a GitHub organization called TypeSafe Community, whose profile describes it as unofficial. Our guide to the Jev GitHub landscape covers how to tell official TypeSafe repositories from lookalikes.
The shared lesson: code does the geometry, Jev does the judgment, and the questions decide the quality of play.
Try the Pattern on Your Own Decisions
You do not need a game engine to see why this works. The mechanics are identical for any loop that makes a small decision about changing state — a queue, a moderation feed, an agent choosing its next tool. Open the Jev AI playground, paste a snapshot of some state you already have, and ask two or three choice or noul questions about it. You will see the probabilities and confidence that a game loop would branch on, and editing a scenario needs no account.
Is Jev Good for Video Game Development?
For some jobs, yes; for the ones the Doom clip suggests, mostly no. The dividing line is the same as everywhere else: Jev is strong at common-sense judgments over text and weak at anything that needs precise numbers or long-range planning.
| Game job | Fit | Why |
|---|---|---|
| Deciding which NPC the player is talking to (voice or chat) | Strong | An independent test reported F1 of 0.96 on clean text and 0.93 on noisy speech-to-text transcripts |
| NPC reactions to player text, tone or intent | Strong | A choice or score over declared reactions, returned in a few hundred milliseconds |
| Moderating player chat and names | Strong | A noul or choice you can threshold, with a confidence for escalation |
| High-level bot behaviour: engage, retreat, heal, explore | Works | What the Doom demo shows, with code handling execution |
| Aiming, collision, physics, pathfinding | Wrong tool | Numeric precision belongs in code; TypeSafe's own notes say Jev "is not a calculator" |
| Deep tactics and search, chess-style | Weak | The same independent developer estimated Jev at roughly 950 Elo in chess, even with move facts computed in code |
| Reading the screen directly | Not supported | Text input only for now |
Two practical limits to plan around. First, Jev's documented weak spots — numbers, dates, indirect questions, very large states — show up quickly in games, so compute distances, bearings and cooldowns in code and pass the results in. Second, per-key rate limits cap how fast any one loop can run. On jev-ai.org, a key defaults to 60 requests a minute and can be raised to 300 — five decisions a second — which suits event-driven game logic such as dialogue, moderation and NPC reactions better than a 10 Hz combat loop.
Build Your Own Real-Time Jev Loop: A Checklist
If the demo made you want to wire Jev into something that runs continuously, these are the rules the Doom agents converge on.
- Serialize only what the decision needs. TypeSafe notes accuracy falls as the state fills with irrelevant detail. Nearest enemies, not the whole map.
- Keep numbers in code. Convert raw coordinates into bearings, distances and named buckets before they reach the model.
- Batch every question for a tick into one request. Parallel questions are nearly free; parallel requests are not.
- Write options with explicit boundaries. The
defend_the_centerresults show a threshold written into the option descriptions is the difference between firing and never firing. - Decouple the loops. Let the world keep running while a request is in flight, and hold the last action until the next answer lands.
- Gate on confidence. Below a threshold you choose, fall back to a scripted rule instead of trusting a coin flip.
- Cap the spend. At 10 calls a second, a forgotten loop is 36,000 requests an hour.
FAQ
Does Jev see the Doom screen?
No. Jev only accepts text: a string, a JSON object or an array of text. The demo feeds it a structured description of the game state, and TypeSafe says the Doom demo is "not on images (yet…)".
Can I run TypeSafe's Doom demo myself?
Not the official one. TypeSafe said it intends to publish an in-depth walkthrough and host hacking events, but as of September 30, 2026 we found no Doom code in its official GitHub organization or its docs. Independent reimplementations exist on GitHub; they are community projects and need your own API key.
How much does it cost to let Jev play Doom?
TypeSafe's figure is about $7 an hour at 10 queries a second. Actual cost depends on how big your state is and how often you call, because Jev bills input tokens only.
Why Doom?
Doom is the classic "can it run Doom?" test, and it is a good stress test for a decision model: the state changes every fraction of a second, the options are small and well defined, and a slow answer is visibly useless. It shows speed and instruction-following in a way a support-ticket demo cannot.
Does playing Doom mean Jev is smart?
It means Jev is fast and follows instructions. For how Jev actually scores against large language models on decision tasks — and how much of that comes from TypeSafe itself — see our breakdown of TypeSafe AI's Jev benchmarks.
The Bottom Line
Jev can play Doom because TypeSafe turned the game into the kind of problem Jev is built for: text state in, a handful of typed decisions out, ten times a second, for a few dollars an hour.
- It never sees the screen; code describes the game and executes the moves.
- About 100 ms per batch of questions is what makes a real-time loop possible; a chat model at several seconds per answer cannot keep up.
- It is not a better player than a scripted bot, and TypeSafe does not claim it is. Its edge is following new instructions without new code.
- The same loop fits any software that makes repeated small decisions.
If you have a decision your software makes over and over, that is your Doom. Try it in the playground with your own state, then send the identical request through the Jev AI API when the answers look right.
Sources
- Introducing System One Models & Jev — TypeSafe AI — The Doom demo video, the $7/hour and 10 queries/second figures, the text-not-images note, and response-time claims.
- TypeSafe AI debuts model for machines that plays Doom — The Register — Independent launch coverage of the demo and the 0.114 s vs 8.566 s side-by-side.
- Models — TypeSafe AI docs — List price, text-only input, and the current rate limits.
- jev-plays-doom — GitHub — Independent Doom agent with seeded scores, latency over 2,642 calls and the option-wording experiment.
- jev-doom — GitHub — Independent Freedoom dashboard using structured observations, standing orders and a spending cap.
The video timestamps come from the auto-generated captions on TypeSafe's embedded video, read on September 30, 2026. Community repositories report their authors' own results and are not affiliated with TypeSafe. Rate limits are marked by TypeSafe as adjusting and may have changed since.




