Jev AI: The Model That Makes Decisions, Not Text


Jev is a new AI model that generates no text, only typed decisions with calibrated confidence. What it does, what it costs, and where it actually falls down.
On 15 September 2026, a new lab called TypeSafe AI released Jev, a model that generates no text at all. You send it some information and a typed question, and it returns a category, a score or a yes-or-no probability. Nothing else. TypeSafe claims it runs up to 193.6 times faster and costs up to 444.6 times less than frontier models on this kind of work. The idea underneath is simple and worth understanding even if you never write code: most business automation is not writing, it is deciding, and we have been paying chat models to do it.
What Jev actually is
Jev is the first of what TypeSafe calls System One models. In the launch post, founder Diogo Almeida describes the question that led to it: models have been superhuman at chat for years, so where is all the automation? He helped build the instruction-following methods behind ChatGPT at OpenAI, and concluded that something big was missing.
The name comes from Daniel Kahneman's Thinking, Fast and Slow, and its split between fast intuitive System 1 thinking and slow deliberate System 2 reasoning. Reading a stop sign is System 1. Working through a proof is System 2. Chat models are built for the second kind and get used for the first constantly.
TypeSafe describes Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out. In plain terms, it is a smart if-statement. It reads a messy situation and returns a clean answer your software can act on, without writing a sentence about it.
The other name is a clue to the thinking. Jev is named after William Stanley Jevons, the economist known for observing that when steam engines got more efficient, coal use went up rather than down. TypeSafe expects the same with machine intelligence: every order of magnitude drop in cost unlocks far more uses.
The three questions it can answer
Every request has the same shape. You send a state, which is any text or data your system already holds, plus a set of questions. Each question is one of three types, per TypeSafe's documentation.
Choice: pick one option from a list you defined. Support tickets into Billing, Technical or Sales, for example. A choice supports up to 255 options.
Score: rate something on a scale you defined, such as urgency or quality, returning a position on that scale with probabilities for each level.
Noul: a yes-or-no question, returning the probability that the answer is yes, from 0 to 1.
Questions are evaluated in parallel and in isolation against the same state, so asking ten questions about a ticket costs barely more time than asking one. Every answer comes with a calibrated probability, which is the part that matters most and gets discussed least. Ask a chat model how confident it is and you get a number that sounds convincing and means very little. Models trained on human feedback learn to sound sure of themselves, because people prefer confident answers. TypeSafe puts the problem well in its comparison table: if a model can do a task 95 percent of the time but cannot tell you when it is in the other 5 percent, you cannot automate that task. Jev was trained with a method TypeSafe calls Reinforcement Learning for Calibrated Decisions, or RLCD, which optimises for honest probabilities rather than for answers people like. Speaking on the Latent Space podcast, Almeida argued that training models to be helpful to humans actively impaired them for automation. If that calibration holds up, it changes how you build. High confidence goes straight through. Medium confidence takes a cautious path. Low confidence goes to a person. The automation rate becomes a dial you set rather than a gamble you take.Why calibration is the real story
Price: $0.042 per million input tokens, with output free, against $0.20 to $10 per million input tokens for LLMs by TypeSafe's own comparison. Output is free because there is almost nothing to meter. Speed: 70 to 500 milliseconds end to end, against seconds to minutes for frontier models. Headline claim: 193.6 times faster and 444.6 times cheaper, from TypeSafe's workflow evaluations.The numbers, and how to read them
Now the caveats, which are worth reading because TypeSafe published them itself rather than waiting to be caught out.
The reference answer in those evaluations is the average of GPT-6 Astra and Claude Fable 5.1, which TypeSafe notes biases results towards OpenAI's and Anthropic's models.
The workflows were written by TypeSafe's own capabilities team, and the company acknowledges some bias could exist.
TypeSafe says it expects those gains to sit at the higher end of real-world results.
The company says it cannot prove the price is not subsidised, and that proving sustainability will take time.
Its published speed evaluations were generally run from laptops on the US West Coast, where the service is based.
All four workflows are published at evals.typesafe.ai, including the disagreements. A vendor that publishes where its model loses is easier to trust than one that does not, but these are still vendor-run numbers. Test on your own data before you believe any of them.
What "cannot hallucinate" really means
TypeSafe says Jev cannot hallucinate, and this is the claim most likely to be misread. What it means is narrow and real: the possible outputs are defined in advance, so the model cannot return a value outside your schema. It cannot invent a category you did not list or produce a malformed response. Type errors are mathematically impossible. TypeSafe is clear that the zero figure in its chart is a guarantee of schema matching, not an empirical measurement.
It does not mean the answer is right. Jev can pick the wrong option from your list, and it can do so with high confidence. Critics on the launch thread made exactly this point: a valid type can still hold a wrong value. That criticism is fair, and it does not cancel the benefit. Removing an entire class of failure, the malformed or invented output, is genuinely useful. It just is not the same as being correct.
Where it fails
TypeSafe publishes a page it calls jaggedness, listing the known weaknesses of the current version. It is the most useful page in the documentation, and the limitations are specific. Drawn from that documentation:
It reads questions literally, answering what you wrote rather than what you meant.
It is not a calculator and does not count reliably.
It reads dates as text rather than as ordered quantities, so date comparison belongs in your code.
Accuracy falls as the state fills with information irrelevant to the decision, so filtering before the call matters.
Multi-hop reasoning and layers of indirection degrade it.
It does not treat input as potentially hostile by default, so adversarial content needs explicit handling.
It will not reconcile contradictory criteria.
It is not trained to generate text, so it cannot draft a reply, explain its reasoning or write code.
There is also a subtler failure worth knowing. If you force a choice between options and the right answer is not among them, the model picks something anyway. Independent testing found that asking an unrelated question with only Billing, Technical and Sales available produced a confident-looking answer with a low confidence score. Give it an "other" option and it used that correctly. Your answer list is part of the design, not an afterthought.
Who should actually care about this
If you build software
Any place you are currently spending an LLM call to get back a single word is a candidate. Ticket routing, content moderation, lead scoring, tagging, guardrails around agent tool calls. Jev is available through TypeSafe's API on an early-access waitlist, through OpenRouter, and through the Vercel AI SDK.
If you run a business
You do not need to use this directly, but you should know it exists, because it changes what is affordable. Work that was too expensive to run on every record becomes cheap enough to run on all of them. TypeSafe's own Doom demo makes the point: an engineer worried about making ten queries a second, which worked out at roughly seven dollars an hour. Classifying every support ticket, every review, every inbound lead now costs less than the electricity to run the server.
If you teach or study AI
This is a useful correction to the idea that AI means chatbots. The interesting frontier is not only bigger models but better-shaped ones. Edapt covers this kind of practical distinction in its AI workshops for working professionals, with dates on the events page.
Frequently asked questions
What is Jev AI?
Jev is a model from TypeSafe AI, released in early access on 15 September 2026. It does not generate text. You send it a state and typed questions, and it returns a choice from your options, a score on your scale, or a yes-or-no probability, each with a calibrated confidence value, in roughly 70 to 500 milliseconds.
What is a System One model?
System One is TypeSafe's name for a class of models built for fast structured decisions rather than text generation. The name comes from Daniel Kahneman's distinction between fast intuitive System 1 thinking and slow deliberate System 2 reasoning.
Is Jev a replacement for ChatGPT or Claude?
No, and TypeSafe does not claim it is. Jev cannot write, explain, summarise or code. It handles the decision points inside software while generative models handle anything where language is the product. The intended pattern is to use both, with Jev around the LLM calls rather than instead of them.
How much does Jev cost?
$0.042 per million input tokens, with output tokens free. There is no published plan tier and no self-serve free tier. TypeSafe says the price may be subsidised and expects it to fall over time.
Can Jev hallucinate?
It cannot return a value outside the schema you defined, which TypeSafe describes as mathematically guaranteed rather than measured. It can still choose the wrong option from your list, including with high confidence, so schema safety is not the same as correctness.
What is Jev not good at?
TypeSafe documents the weak spots: arithmetic and counting, date comparison, literal reading of questions, multi-hop reasoning, noisy input containing irrelevant detail, adversarial content, contradictory criteria, and anything requiring generated text.
How do I get access to Jev?
TypeSafe runs an early-access waitlist for its own API. The model is also reachable through OpenRouter and through the Vercel AI SDK. The current public version is Jev 1.13.
Who built Jev?
TypeSafe AI, founded in 2024 and out of stealth on launch day. Chief executive Diogo Almeida previously worked at OpenAI on the instruction-following research behind ChatGPT. Reporting puts the company's seed funding at $40 million led by DCVC.
Where this leaves things
It is early. Jev is behind a waitlist, the benchmarks are vendor-run, and the pricing may not survive contact with real costs. Anyone building something important on it this month is taking a risk.
The idea behind it is more durable than the product. For four years the assumption has been that a more capable chat model is the answer to every problem, and a lot of production systems now pay for a paragraph when the code only needed a category. Whether or not TypeSafe is the company that fixes that, the observation is correct, and more models shaped like this are likely to follow. The company's manifesto argues AI needs an interface software can depend on. That part is hard to disagree with. More analysis like this is on the Edapt blog.
TypeSafe AI, Introducing System One Models and Jev, 15 September 2026, including the company's own nuance and caveat sections. TypeSafe AI documentation and its System One concepts page, for the question types and the jaggedness limitations. TypeSafe workflow evaluations, for the benchmark workflows and disagreements. Latent Space, Jev: System One models for Prod, not God, interview with Diogo Almeida. OpenRouter, Jev 1.13 and Vercel, How to classify, route and score with Jev and AI SDK, for access and pricing.Sources