
JEV from TypeSafe AI opens a new category: a System One model that returns typed decisions with probabilities and calibrated confidence, in a tenth of a second and at a fraction of LLM cost. What that means for decision trees, routing, lead scoring and guardrails, and how the architecture differs from a regular language model.
Most AI systems I build for clients do not need the model to talk. They need it to decide. Which team gets this ticket, is this lead worth a phone call today, is the answer the bot is about to send safe enough to go out without a human in the loop. And yet at every one of those junctions sits an LLM generating text, while we write code that parses the text back into values software can understand. It works, but it is slow, expensive, and fragile exactly where it hurts most: the format.
JEV, the first model from TypeSafe AI, takes a different approach. It receives your system state and returns a typed decision, with probabilities and calibrated confidence. No free text, no parsing, no hoping the JSON comes back valid this time. In this article I break down what this model actually is, how it differs from a regular LLM in architecture and performance, and where I would plug it in today, for example in the decision trees behind bots and automations.
The name comes from Daniel Kahneman. System 1 thinking is the fast, intuitive kind that recognizes an angry face in a split second. System 2 is the slow, deliberate kind that works through long multiplication. TypeSafe argues that a very large share of what software needs from AI is System 1 work: classification, routing, scoring, verifying a claim. Decisions a skilled person makes in seconds given the right context, which need an answer rather than a well written paragraph.
Their definition of JEV is short and sharp: a frontier-intelligence function call. Unstructured state goes in, typed probabilistic decisions come out. The model is named after the economist William Stanley Jevons, who observed that when a resource gets cheaper and more efficient, demand for it explodes. The hint is clear: when an AI decision costs close to nothing and takes a tenth of a second, it suddenly pays to put one everywhere.
A regular LLM generates text, and even with structured output it still generates tokens that you hope will match the schema. Anyone who has run such a system in production knows the moment the model returns a slightly different field name, or politely wraps the JSON in an explanation. JEV returns values from the type you defined, and according to TypeSafe an out-of-type output is mathematically impossible rather than filtered after the fact. To be precise: this eliminates format hallucinations, not content mistakes. The model can still pick the wrong option, but it is wrong within the type, and with a confidence number you can act on.
An autoregressive language model generates a token, waits, generates the next. That is the structural reason an answer takes seconds and output tokens cost several times more than input. JEV samples all its answers in parallel, with no sequential generation. This is a different game rather than an incremental improvement: response times of 70 to 500 milliseconds, and pricing where output is simply free, because there are no output tokens.
Ask a regular LLM for a confidence score and you get a number that sounds confident almost always. These models are known to be overconfident, and inconsistently so. TypeSafe's central claim is that JEV is trained so its confidence tracks its real accuracy. From an engineering point of view this may be the most important property of all: you can write code in the style of act if confidence is above 0.9, otherwise hand off to a human, and the threshold actually splits the cases correctly. Deciding when not to decide becomes part of the system.
Chat models are trained with RLHF, reinforcement learning from human feedback, which teaches them to please the reader. Newer approaches like RLVR reward verifiably correct answers. TypeSafe developed a third algorithm for JEV, RLCD, Reinforcement Learning for Calibrated Decisions: the model is rewarded not only for the right decision but for reporting probabilities that reflect its real chance of being right.
JEV's interface is built from three primitives, and every request is a state plus one or more questions:
Choice: pick an option from a closed list. This is the tool for routing and classification, and the answer includes the choice, a probability distribution over all the options, and a confidence level.
Score: rate the state against a rubric you define. Lead scoring, conversation quality, ticket priority. Again with probabilities and confidence.
Noul: is this statement true, as a continuous value between 0 and 1. To me this is the most interesting primitive, because it turns fact checking against a source into a function call: is the sentence the bot is about to send supported by the document it quoted.
The guiding principle in their documentation is strict: atomic questions. A good JEV question is one a knowledgeable person could answer in a few seconds given the right context. Composing small decisions into a large process happens in your code, not inside the model. Anyone who has built AI agents will recognize a refreshing inversion here: instead of pushing ever more logic into a giant prompt, the logic returns to software, and the model answers small, precise questions.
TypeSafe publishes a benchmark of a full operational workflow: JEV finished in 0.114 seconds what took a leading language model 8.5 seconds, 193 times faster, and 444 times cheaper. Official pricing: 42 dollars per billion input tokens, roughly 200 times less than frontier models, with output free. Response times run 70 to 500 milliseconds, which puts the model in territory that until now belonged to ordinary code: inside an HTTP request, inside a loop over a dataset, inside every incoming message.
An obvious disclosure: these are vendor numbers, on benchmarks the vendor chose. I would apply a healthy discount before any purchasing decision. But even after the discount, the advantage is structural rather than promotional: a model that does not generate tokens one after another really is playing a different game, and no future version of a chat model will close that particular gap.
This is the part that interests me most as someone who builds bots and automations. The classic decision tree is the most battle tested tool in business software: clear conditions, predictable paths, easy to debug and easy to explain to a client. Its one weakness is that the world arrives unstructured. Is this request urgent is not an if statement, and that is exactly where trees break down or fill up with keyword rules that age badly.
The pattern a System One model enables: the tree stays in code, visible, testable and version controlled. But every soft node, every place that until now needed a human or a guess, becomes a Choice or Noul question. And the probabilities give every node a third branch that was always missing: not just yes or no, but also not sure enough, hand this to a person. A tree like that handles the confident cases automatically and escalates exactly the cases that deserve human eyes.
A few concrete places I would put this today: intent routing in a support bot, the node that decides whether a message is an information question, a problem report or a hot lead. Lead scoring before the CRM, with a threshold above which a rep calls the same day. A guardrail over another LLM's answers, a Noul check that every reply is supported by its source before it reaches the customer. And bulk classification of records, thousands of historical tickets or documents, which at 42 dollars per billion tokens turns from a project into a query.
It is important to understand that JEV does not replace the LLM, it frees the LLM from the work it is bad at. In a hybrid architecture the fast decisions move to the decision model and the writing stays with the language model. In a real bot it looks like this: JEV routes the request and screens for urgency, the LLM writes the actual answer, and JEV checks it against the sources before it goes out. Three calls, two of them in a tenth of a second at near zero cost, and the whole system becomes both faster and more trustworthy.
When you need text, there is nothing to discuss, this model does not generate it. When you need a long chain of reasoning or multi step planning, that is System 2 work for reasoning models. When hard rules are enough, do not bring a model to a place an if statement handles perfectly and for free. And finally, a sober vendor consideration: this is a young model from a young company, in a category that is itself new. I would wrap the calls in a thin abstraction layer that allows swapping vendors or falling back to an LLM with structured output, and run a trial period where decisions are logged but do not act on their own.
Most of the AI a business actually needs day to day is not impressive conversation but small decisions at high frequency: where does this go, how urgent is it, is it safe. Until now the cost of each such decision, in time and money, made us ration them. If the System One promise holds, the direction flips: a cheap, fast decision layer embedded in every process, with human escalation exactly where it belongs. This is the kind of system I build for clients, a combination of AI agents, automations and internal tools, and models like JEV are about to become another core tool in that kit. If you want to think through where a decision layer like this meets your business, talk to me.
Want a consultant? Click here to schedule a call.
Schedule a call
Everyone says "train the AI on your data". RAG does nothing of the sort: no weights change, your data stays in your database, and answers are grounded at question time. What actually happens, why it is better for business, and when fine-tuning does make sense.

AI FOMO makes companies buy platforms nobody uses. The adoption path that works is boring: map processes, pilot with measurable ROI, RAG on company knowledge, evals and guardrails, team enablement. What it saves, and how I help companies get there in weeks.

How I ship AI chat features safely: Clerk-gated access, OpenAI ChatKit sessions, prompt/response guardrails, and performance-minded client loading in a Next.js 16 App Router codebase.
Wherever you are in the world, let's work together on your next project.
Tel Aviv, Israel
Prefer to talk directly? Schedule a call and we can discuss your project live.
Schedule a call