Claude Fable 5.1 costs $10 per million input tokens. That is $10,000 per billion.
Jev, from TypeSafe, costs $42 per billion. Output is free, described by the company as "too cheap to meter."
That is a 238x gap, and the obvious next question is what the catch is. The catch is that there is no chat box. You cannot ask Jev anything. It does not write, does not explain itself, and will never produce a sentence. Your software calls it. You never will.
The bet is against three years of industry direction
Every frontier lab has spent three years making models think for longer. Reasoning effort settings, chain of thought, test-time compute, more deliberation for a better answer. TypeSafe went the other way and named the bet after Daniel Kahneman. From their own launch post, the model class name "draws on the distinction between fast, intuitive System 1 thinking and slow, deliberate System 2 reasoning."
The claim underneath is worth sitting with. Most decisions inside working software never needed deliberation. Is this spam. Which queue does this go in. Does this need a human. Those are multiple-choice questions, and the industry has been paying frontier prices to have a System 2 machine answer them in prose that your code then has to parse back into a value.
Almeida
"Human-in-the-loop tasks: chatbots, copilots, coding agents. General and powerful, but requires human oversight because their freedom also means they might go off the rails."
Why the missing text box is the point
Three answer types, and that is the whole model.
Choice picks one of a defined set and returns a probability for each option. Score places something on described, ordered levels. Noul returns the probability that a statement is true. You send state, which is the context the judgment needs, and questions about that state. Many independent questions ride in one request and run in parallel.
Because the output shape is guaranteed rather than generated, TypeSafe states a schema hallucination rate of 0%. That claim is narrower than it sounds and worth stating precisely: a guaranteed schema removes one failure mode and leaves the other one completely intact. Typed output guarantees the interface, not the truth. The model can be confidently, correctly-shaped wrong.
What it buys you is that every answer arrives with a number attached, and the number was trained to mean something. TypeSafe's training method is called RLCD, Reinforcement Learning for Calibrated Decisions, and the stated property is simply "higher confidence means higher accuracy." Their diagnosis of the alternative is that conventional models "tend to be overconfident and inconsistent."
That is the part that matters if you have ever tried to decide whether an agent can act alone. You cannot set a threshold on a confidence score that does not track accuracy. See What Is a Level-3 AI Agent?
The operator move is a gate, not a swap
Nobody is replacing their frontier model with this. The move practitioners landed on in the first week is architectural: put the cheap judgment in front of the expensive one.
One email tool reported swapping two AI calls per email for a single Jev call and keeping its existing pipeline behind it. The cheap thing makes the small choice, and the expensive model only runs on what survives the gate.
Which gives a CEO one question to ask about an agent pipeline they already pay for: how many of these frontier calls are answering a multiple-choice question? Every one of those is a gate candidate, and the rest still cost $10,000 per billion because the rest are actually generation.
The numbers, and how much to trust each one
The strongest figure in the launch window is the one where somebody ran both systems on the same feed at the same time. Elvis Sun classified 384 news headlines across 15 brands in 24.9 seconds for $0.19. Claude Opus 5, on the same 384, got through 4 of them for $0.77.
Others posted their own runs in the same days: 1,891 competitor ads in 19 seconds for $0.12; three million session replay events triaged in 40 seconds for $2.17; 500 emails for three and a half cents. An SEO agency reported a client audit falling from roughly $250 to roughly $25 for the same output.
Every one of those is an unaudited run posted by an enthusiastic user inside a launch window. Treat them as reports, not measurements. The 384-versus-4 comparison earns more weight than the rest only because both systems ran the same input at the same moment and both results were published.
On adoption, Vercel reported Jev reached roughly 13% of AI Gateway teams on day one, which it called the fastest in the gateway's history, at twice the GPT-5.6 family and six times Fable 5.1. That is real and it is also a launch-week number from the gateway that benefits from routing volume.
What this does to the bill
Not what you would expect. OpenRouter's Alex Atallah, on the same launch, predicted the opposite of a saving:
Atallah
"100x more AI usage."
A price cut at the bottom tier does not reduce the bill. It moves where the bill is spent, and it makes judgments economical that nobody was making at all. Every event in a stream can now carry a decision. That is the same pattern as Token Prices Fell and the Bills Went Up, arriving one tier lower.
Before you plan around it
Jev was waitlist-gated as of September 18, so this is not something a team does on Monday without getting access first.
The founder is Diogo Almeida, reported as a co-inventor of ChatGPT, after roughly two years in stealth. He leads with the constraint rather than burying it, which is the correct read of the product: it cannot generate text, and the gains are not free.
And the durable part is not Jev. It is the pattern. A cheap calibrated gate in front of expensive generation survives this particular model being replaced by a better one, and it is available to anyone whose pipeline is quietly asking a frontier model to pick between three options.