Ask Jev whether a job candidate is strong in Python and it returns 0.81. Ask it how much Python experience that same candidate has and it returns 2.05. Same resume, same request. The two numbers are not on the same scale, neither is a confidence score, and confusing them is the most expensive mistake a team makes in its first week.

Both figures come straight from TypeSafe's docs. The gap between them is the product. Jev is the first model in a class TypeSafe calls System One. The price and the missing text box got the attention at launch. The three answer types decide whether it works in your building.

Figure 1. The three answer types
Two return a confidence. One does not.
CHOICE SCORE NOUL Pick one of a defined set Place on ordered levels Is this statement true? RETURNSRETURNSRETURNS the chosen option a probability per option a continuous value the full distribution one number, 0 to 1 a confidence value a confidence value no confidence field 0.81 means the modelpicked that option 2.05 sits betweenlevel 2 and level 3 0.5 means yes and noare equally likely A Noul near 0.5 is not medium intensity. It is the model saying it does not know.
The missing confidence row on the right is the single most common source of misread output in the first week.

Three answer types, and that is the whole surface

Choice picks one option from a set you define. It returns the winning option, a probability for every option you listed, and a confidence value. Which team owns this ticket. Which of your nine product categories this is. You write the list, so the model cannot invent a tenth.

Score places the material on ordered levels you describe in words. It returns a continuous value that can land between two levels, the probability on each level, and a confidence. The docs work a bug report scoring 1.43 on a three-level severity scale, with 0.57 of the probability on "workaround exists" and 0.43 on "no workaround". Write the levels as situations, not degrees. "Broken or degraded feature, but workaround exists" gives the model something to match. "Moderately severe" does not.

Noul returns one number between 0 and 1: the probability that a statement is true. Does this message request a refund. Does this contain personal data. And unlike the other two, a Noul carries no confidence value at all. The docs explain why: "A Noul's probability distribution has only two outcomes, yes and no, so the single noul value describes it completely."

The 0.5 that fools everyone

A Noul of 0.5 does not mean "moderately". It means the model gives yes and no similar probability. It describes uncertainty about a binary question, not the quantity of anything. The docs say it flatly: a Noul value "runs from 0 to 1, but it's not a scale of the thing you asked about."

Back to the candidates. Asked "Is the candidate strong in Python?", four resumes returned 0.03, 0.14, 0.81 and 0.92. Those are not experience levels. They are four answers to one proposition, and the spacing was chosen by nobody. The same four, run through a Score with four written levels, landed at 0.0, 1.0, 2.05 and 2.89, each near a level a human wrote.

So a dashboard that buckets Noul values into low, medium and high has invented levels the model never saw. Want degree, ask a Score and write the levels. Want a yes or no, ask a Noul and set a threshold.

You send state and questions. You do not send a prompt.

A request has two halves. State is the material being judged: a support message, an order record, your refund policy, packaged as one string or JSON object. Questions are the judgments you want made about it, each carrying instructions, the judgment, and criteria, the possible answers.

The part that changes how you budget: every question in a request sees the same state and runs in parallel. TypeSafe's line is that "adding questions barely changes the response time and costs only the tokens for the extra questions." Their cookbook batched 13 questions into one call instead of 13 and came out 11.5x cheaper and 9.6x faster, answers unchanged. The argument about whether a feature justifies another AI call stops applying.

Confidence is a shape, not a grade

Choice and Score answers both carry a confidence from 0 to 1, and it is not what most people assume. It is computed from how the probability is spread. All of it on one option gives 1.0. The flatter it spreads, the lower the number. That is the entire definition.

Which means confidence is not a correctness score, and it is not permission to act. That bug report scored 1.43 at a confidence of 0.35, and the low number is not a failure: the model split 0.57 and 0.43 across neighbouring levels because the bug sits between them. Low confidence on a Choice, in TypeSafe's words, "often means none of the options are a clear winner over the others." Several acceptable answers spread probability too, so a low number on a harmless preference invalidates nothing.

The trap is treating every low number as a defect and rewriting the question until it climbs. The docs model the better move in code: different thresholds for different stakes, because showing the wrong screen is recoverable and moving the wrong money is not. Your risk tolerance stops being a paragraph in a prompt and becomes a number in a file with an owner. Same gate as Level-3 autonomy, cheap enough now for every event.

Why Kahneman's name is on it

Diogo Almeida
Diogo
Almeida

"After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I've spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today."

TypeSafe launch post  ·  September 16, 2026

RLCD stands for Reinforcement Learning for Calibrated Decisions, and TypeSafe positions it as a third training path beside RLHF, which produced chatbots, and RLVR, which produced reasoning models. The objective is different: the model returns decisions and probabilities instead of text, and "higher probability should correspond to a greater chance that the answer is correct."

The class name follows from that. TypeSafe's launch post says it draws on "the distinction between fast, intuitive System 1 thinking and slow, deliberate System 2 reasoning", and the docs add the qualifier that matters: "Here, the emphasis is on fast, focused judgments."

Typed output guarantees the interface, not the truth

Calibration means that across many predictions, outcomes assigned 0.8 should happen about 80% of the time. The docs then say the quiet part in their own voice: "Calibration is measured across groups of predictions; it does not guarantee that an individual answer is correct."

A guaranteed schema removes one failure mode and leaves the other completely intact. Your code will never again fail to parse the answer. The answer can still be wrong, and now it arrives wrong in a well-formed envelope with a number attached. The model can be confidently, correctly-shaped wrong.

So a pilot does not prove "it returned valid JSON", because it always will. Take 500 cases your team already labelled, run them, and check whether the answers that came back at 0.8 were right about 80% of the time on your data. The only evidence worth signing against.

Decomposition is the part that pays

TypeSafe flags one idea above every other in its build guide: "This is probably the most important concept in this guide." Stop asking one fuzzy question. Instead of "Is this message spam?", their example asks three narrow ones over the same material: does the body ask for a password, does the sender identity mismatch, does it claim an unexpected reward. Combining happens in your code, not the model.

spam_risk = 0.45 * requests_credentials + 0.30 * sender_identity_mismatch + 0.25 * unexpected_reward

Simplified from TypeSafe's example, not verbatim API syntax. The three numbers in front are the entire policy. "The weights are yours," the docs say. "When the combined result doesn't match what your team would decide, change them in code and run again."

Broad questions hide several judgments behind one answer. Atomic questions expose them, so the rule your company operates by stops living inside a prompt nobody can diff and turns into three numbers somebody signed off on. It costs nothing in latency, because the questions run in parallel anyway.

The durable skill is not Jev. It is knowing which of your judgments are genuinely one judgment and which are five stacked into one. Ask smaller questions.

Which of your judgments are five judgments?

Book a free Diagnostic: 30 to 45 minutes, no deck, no pitch. We take one fuzzy decision your pipeline currently asks a frontier model to make and break it into the narrow questions it was hiding.

Book the Diagnostic →
Sources
1TypeSafe documentation, docs.typesafe.ai, September 2026: the Choice, Score and Noul primitive pages and the confidence page. Choice returns the chosen option, a probability per option and a confidence. Score returns a continuous value across described ordered levels, the distribution and a confidence. Noul returns a single probability that a statement is true and carries no confidence field.
2TypeSafe on confidence: it is "a statistic computed from the probability distribution", and "Calibration is measured across groups of predictions; it does not guarantee that an individual answer is correct." A low confidence "often means none of the options are a clear winner over the others."
3TypeSafe on the name: "The System One name comes from the concept Daniel Kahneman popularized in his book Thinking, Fast and Slow... Here, the emphasis is on fast, focused judgments."
4RLCD, Reinforcement Learning for Calibrated Decisions, is TypeSafe's stated training method. Diogo Almeida, founder, on the launch.
5The batching figures (11.5x cheaper, 9.6x faster, answers unchanged across 13 questions in one request) are from TypeSafe's own cookbook and are the vendor's measurements.
6The composition example is simplified from TypeSafe's documentation and is not verbatim API syntax.
John Tan
John Tan

Founder and CEO of nativefirst.ai. Embeds with scaling founders and CEOs to ship Level-3 agents and AI workflows in production.