In the last few days, I’ve seen a lot of people using Jev, some of them claiming complete bullshit while others claiming more realistic gains. I tried Jev Playground and reviewed some projects to see how real developers are using it. I noticed a pattern: Jev isn’t the system. It’s one layer inside a much larger architecture.
One project that really stood out and demonstrates Jev’s ability to offer genuine value is Christian Mathiesen’s “Jev Plays Pokémon Red”. The headline reads, “Jev plays the game autonomously with no scripts or cheats” and it’s spot on, which is impressive. However, after diving into the codebase, I think the most important takeaway isn’t necessarily about Jev.
Jev didn’t make the project possible on its own. The harness did, and Jev was the last piece of the puzzle.
What Jev is (and isn’t)?
Jev is a decision model. You give it a state and a set of questions along with the structured output you expect from it, either choice, score, or true/false. It isn’t a chatbot and doesn’t generate text. One key takeaway is that it provides probability distributions.
Here is a simple example of how it looks like. You provide a state:
{
"state": {
"battle": "Gym Leader Brock",
"myPokemon": {
"species": "SQUIRTLE",
"level": 14,
"hp": "28/40",
"types": "WATER"
},
"enemy": {
"species": "ONIX",
"level": 14,
"hp": "35/35",
"types": "ROCK/GROUND"
}
}
}Then, you define the question you want Jev to answer and the criteria:
{
"decision": {
"type": "choice",
"instructions": "You are in a Pokémon battle. Choose the best action.",
"criteria": {
"Use TACKLE": "NORMAL move, power 35. Not very effective against ROCK/GROUND.",
"Use BUBBLE": "WATER move, power 20. Super effective x4 against ROCK/GROUND.",
"Use POTION": "Heal SQUIRTLE by up to 20 HP. Uses the turn."
}
}
}And Jev provide the following structured answer:
{
"answers": {
"decision": {
"type": "choice",
"choice": "Use BUBBLE",
"confidence": 0.99,
"probabilities": {
"Use BUBBLE": 0.99,
"Use POTION": 0.01,
"Use TACKLE": 0
}
}
}
}It doesn’t see the screen, never presses a button and doesn’t remember your last request. Each call is fresh and stateless over the list of options you provide.
All these constraints, make Jev fast and cheap enough to run on every single decision while playing a video game like Pokémon Red. A continuous 24 hours of gameplay could cost $1~1.70. A frontier lab LLM like GPT-5.6 Terra could cost $60~110 while also being multiple times slower as well. That’s the edge a decision model provide.
But how Jev actually play the game?
That’s the most interesting thing from analyzing this project (and any other serious one about Jev). From around 4,400 lines of code, only 200 of them talk to Jev. It is specifically used for choosing a focus, define where to go (using grid like coordinates), choose battle action, picking an answer from a menu and spelling nickname.
Everything else, is a harness that directly read the Game Boy’s RAM in order to capture the map, your position, your party’s HP, bag, badges, story event flags and battle stats. The harness is capable of capturing all the information needed and structure it in a way the decision model can provide a legal/valid action to perform.
To me that’s amazing, especially because Jev’s decision are only as good as the information you provide to it. That’s why a great harness, a well structured architecture, a frontier lab model used to reasoning and finally Jev to make decisions, is the chain that makes emerging projects to be useful and not simply AI slop trying to hype any new thing that gets released.
The layer Jev replaced
While I definitely believe that Jev is one piece of the puzzle, its important to understand how the same scenario looks like without it.
Write the decisions in code.
if (hpFraction < 0.25) goHeal();
else if (teamLevel < objective.level - 5) goTrain();
else if (balls > 0 && party.length < 6 && isNewType(wild)) throwBall();
else if (bestMove.effectiveness >= 2) useMove(bestMove);
...This works for about a day, then break miserably and require you to keep adjusting a blob of code that grows unbounded. Every new situation is a new branch, every branch interact with others and the rule tree becomes the project.
Train a new model.
When seeing Jev classification/extraction capabilities for the first time, I immediately related to custom trained models that were capable of doing similar tasks very fast and provide score as well. What separate Jev from this is its ability to generalize. NER and classifier models requires a labeled dataset and a training step in order to provide “decisions”. If new constraints or scenarios surface, more data and a new training must be done.
Its powerful, but slow and expansive. ML operations becomes the bottleneck.
Use a frontier lab LLM.
Provide the context and the structured output schema to the frontier lab and you get Jev’s equivalent output. The problem becomes latency and cost. These models are slow and too expensive to be used on every decision.
A decision model closes the gap, and it’s a real one. Evaluating 100,000 scenarios with sub second latency while being cost effective is a real game-changer.
It simply isn’t the gap that hype posts and attention-seekers describe.
The core is around the decisions
Jev Plays Pokémon Red is a great project, and it’s a good showcase of what a decision model makes cheap and scalable. A few years ago, getting this quality of in-game judgment across an entire RPG meant either a rule engine big enough to be a job on its own or a training pipeline. Here it’s five function calls.
But look at what those five calls sit on: RAM decoding, a ROM-derived world graph, counterfactual pathfinding, failure memory, milestone verification and deterministic execution. That’s the core. Jev removed the last blocker.
Jev isn’t the next big thing. It’s what finally lets the systems you were already building make decisions well, quickly, and cheaply enough to scale.
Make sure to review Christian Mathiesen project. I found it quite interesting and cool.
Reverse Pitch is one post a week on software engineering and the career around it, written from inside the work.



