Independent, unofficial explainer. Not affiliated with or endorsed by TypeSafe AI. Content checked against docs.typesafe.ai on 5 October 2026; interactive values are simulated unless marked as docs examples. Checked against real Jev (typesafe/jev-1.13-20260917) the same day: field names and confidence formulas match, values within 0.05 of the docs. Rerun it with tools/validate_jev.py.
Jev doesn't write text. Your code sends it a piece of text and a typed question; it returns a typed answer with probabilities. This page explains the three question types, how each one measures its confidence, and how to turn that into decisions.
state: a ticket, a message, a documentinstructionsconfidenceEach tab explains one type with fixed cases, an interactive tool and the real API call. A good reading order is left to right.
Is the customer asking for a human agent?
noul = 0.84the probability of yesConfidence: not returned. The distance from 0.5 is the certainty; your threshold makes the decision.
Which team should handle this?
choice = "returns"confidence 0.42Confidence: how far the top probability climbs from the even split 1/n towards certainty.
How severe is the reported issue?
score = 1.11confidence 0.84Confidence: how far the probability sits from the peak level, compared with total ignorance.
Each type has its own picture of total ignorance. Confidence measures how far the answer is from it, rescaled so ignorance is 0 and certainty is 1.
Probability and confidence answer different questions. The probability says which answer; confidence says how sure. A top probability of 0.61 sounds like "61% sure", but with 3 options the floor is already 0.33, so confidence is only 0.42. Whether a 0.8 really comes true 80% of the time is a third question, calibration, which only your own labelled data can answer.
A side-by-side table, three quick tests, what goes wrong with the wrong type, and a worked example: "Is the customer angry and asking for a refund?"
The Low / Medium / High bands on every tab use illustrative cut-offs of 0.5 and 0.7. TypeSafe's docs set thresholds per action, by the cost of being wrong: a balance check can act on less certainty than a money transfer.
Sources: Confidence · Noul · Choice · Score · Confidence-gated routing
Ask a yes/no question, or give a statement to judge: does this message ask for a refund, does this resume mention distributed systems, does this comment contain personal data. Jev returns one number, the probability that the answer is yes. That single number carries both the answer and the certainty, so there is no separate confidence field.
One answer, one field. Everything else is computed by your code. Value from the docs' human-escalation example.
noulconfidence|2p − 1|. Your code works it out.decisionnoul with a threshold you choose, based on the cost of being wrong.noulA Noul question asks the TypeSafe model to evaluate a yes/no question and return the probability that the answer is yes.
A value near 1 is a strong yes. A value near 0 is a strong no. A value near 0.5 means the model gives yes and no similar probability.
confidenceA Noul answer is a single probability p that the answer is yes, and TypeSafe returns no separate confidence for it.the docs' optional formula. It is the Choice formula with two options:
noul value Jev returns, from 0 to 1Which side of the middle is the needle on, and how far from the middle?
With two possible answers, a model with no idea gives each the same probability, . That is the ignorance point. Unlike Choice and Score, it sits in the middle of the scale, not at an end: moving away from it in either direction adds certainty.
So 0.10 and 0.90 are equally certain, just in opposite directions. Both have confidence 0.80.
The needle's side gives the answer. Its distance from 0.5 gives the certainty.
"Is the customer asking for a human agent?"
TypeSafe's docs show what Jev returned for six different messages. Read them as a dial: the dashed line is the coin-flip point, 0.5.
Thanks, that fixed it!
How do I reset my password?
I need this sorted today, whatever it takes.
Are you a bot?
Is there any way to speak to someone about my invoice?
I have asked three times now. Can I please just talk to a real person?
The threshold is a business decision, not a model setting
Jev's answers stay the same. Only the threshold changes, and with it which conversations go to a person.
The docs' rule: "Where to set the threshold depends on the cost of being wrong."
"Use 0.5 when yes and no are equally easy to act on."
"Raise it when acting on a false yes is expensive, such as paging someone or issuing a refund."
"Lower it when missing a true yes is expensive, such as failing to flag a safety issue."
The invoice message (0.84) is escalated at 0.5 and 0.2 but not at 0.9. Nothing about the model changed: you decided a false escalation costs more there. "Are you a bot?" (0.40) only reaches a person when missing a real request is the bigger risk.
Pick a message or drag the dial. Move the threshold to see the decision change.
instructionsstatecriteriatrue and false descriptions of what a yes and a no mean (see the API card, view 1b).noul
A Noul gives back one number. Put two conditions in one question, such as "Is the customer angry and asking for a refund?", and that number has to summarise both: you can no longer tell which part is true. Ask one Noul per condition and combine them in code. And phrase every question so that a high value means yes.
angry_and_refund0.45Which part is uncertain? Angry without asking for a refund, a calm refund request, or both half-true? You can't tell.
is_angry0.92wants_refund0.08Angry, but not asking for a refund. A clear answer per condition; your code decides what the pair means.
Pick one option from a set (support department, product category, language). Options have no order, so the only thing that matters is how much the top option beats a fair share.
One answer, three fields. This tab is about confidence, but you read it together with the other two. Values from the docs' ticket-routing example.
choiceconfidenceprobabilitiesratioprobabilities.choiceThe option with the highest probability.
confidenceA number from 0 to 1 computed from how probabilities is spread. A flat shape, with probability spread across several options, means low confidence. A single peak on one option means high confidence.the option whose probability is largest; that probability is
Where does the top probability sit between total ignorance and certainty?
A model with no idea which option is right gives every option the same probability, . That is the ignorance floor: the lowest the top probability can ever be, because n probabilities below could not add up to 1.
More options means a lower floor. A top probability of 0.6 is a weak lead between 2 options and a strong one among 5.
Same top probability, different number of options
In all 3 examples the top option, Returns, gets 0.6. Only the number of options changes, and with it the ignorance floor .
The higher the floor, the less a 0.6 lead is worth. Confidence measures the climb from the floor to certainty, not the raw 0.6.
"Returns, or maybe Billing." A coin-flip with a slight lean.
"Returns, others far behind." A clear lead.
"Returns, the rest scattered." A strong lead.
The raw top probability is 0.60 every time, yet confidence goes from 0.20 to 0.50. A threshold on pmax alone would treat a near coin-flip and a strong lead the same way.
Pick an example or drag the bars. Every number below updates live.
Shoes arrived two weeks late and in the wrong size. Also I see two charges of $120 on my card.
Jev picks returns, but the double charge pulls weight toward billing: confidence = (0.61 − 0.333) / 0.667 ≈ 0.42, a moderate answer.
and both give 0.40. The formula looks only at , so a close two-way race scores the same as one leader with scattered runners-up. TypeSafe lists the top-to-second ratio as an alternative when you care about that race.
Rate something on an ordered scale of 2 to 10 described levels: bug severity, formality, satisfaction. Order changes the question. Being split between neighbours (1 vs 2) is a small doubt; being split between opposite ends (0 vs 2) means the model has no idea. So Score confidence measures how far the probability sits from the peak, not just how tall the peak is.
One answer, four fields. This tab is about confidence, but you read it together with the other three. Values from the docs' PDF-export example.
scoreconfidenceprobabilitieslegendscoreThe position on the level number line, from 0 to the top level number. It's each level number multiplied by its probability, added up.
confidenceA number from 0 to 1 computed from how probabilities is spread. A single peak on one level means high confidence.
Probability on a neighboring level lowers confidence less than the same probability on a level further away.
How far is the probability from the peak, compared with how far it would be under total ignorance?
A model with no idea which level is right gives every level the same probability, . That flat answer has no peak, so its spread is measured from the middle of the scale. The average distance it ends up with is the ignorance spread, : the yardstick every real answer is compared against.
More levels means a bigger yardstick, because a flat answer over a long scale is spread further.
Same top probability, different confidence
In all 3 examples the peak is fixed: 0.6 sits on the highest level of the severity scale, Blocker (level 4, the 5th level, because levels are numbered from 0).
The remaining 0.4 moves. The further it sits from the peak, the more it lowers score confidence.
In financial terms, each unit of distance from the peak costs confidence, and the budget is the ignorance spread. For 5 levels that budget is .
"Blocker, maybe only critical." A small doubt.
"Blocker, or only major." A real doubt.
"Blocker, or cosmetic?" Not understood.
The Choice formula sees only the 0.6 and gives 0.50 every time. Score confidence falls from 0.67 to 0 as the runner-up moves away. Watch the score field too: in the third case it averages to 2.4, "Major", a level the model gave no probability at all.
Pick an example or drag the bars. Every number below updates live.
Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too.
The peak is level 1, "workaround exists". The 0.11 on "blocking" is only one level away, so it costs little: confidence = 1 − 0.11 / 0.667 ≈ 0.84, and the score lands at 1.11, just past the peak.
The score field is an average: each level times its probability, added up. Very different answers can share it. Below, a model that is sure the bug has a workaround and a model split between "cosmetic" and "blocking" both return score 1.00. Only confidence tells them apart, which is why you read the two fields together. The Choice formula would not help either: the docs' neighbours case (0, 0.5, 0.5) and opposite-ends case (0.5, 0, 0.5) have the same top probability, but Score confidence gives them 0.25 and 0.
Which level is the peak when two levels tie? The docs define as "the most likely level" but don't say how ties are broken, and it changes the answer. For (0.4, 0.4, 0.2), choosing level 0 as the peak gives confidence 0; choosing level 1 gives 0.10. This page takes the lowest tied level. Try the "Tie" example, and check against a real Jev response before relying on it.
Start with the shape of the answer. Yes or no? Use Noul. One of several options with no order? Use Choice. A position on an ordered scale? Use Score. For the last two, one question decides it: is a near miss less wrong than a far miss?
| Noul | Choice | Score | |
|---|---|---|---|
| The answer is | Yes or no | One of several unordered categories | A level on one ordered scale |
| Near miss vs far miss | Only two answers, so no distance | All mistakes are equally wrong | One step off is less wrong than three steps off |
| Main output | noul: probability of yes | choice: which option, plus probabilities | score: where on the scale, plus probabilities |
| Confidence | Not returned. Optional: |2p − 1|, distance from 0.5 | How far the top option rises above an even split | How far the probability sits from the peak |
| Examples | Asks for a refund? Contains personal data? Wants a human? | Department, product category, language, intent | Severity, urgency, satisfaction, formality, fit |
| Your code typically | Compares with a threshold; combines several Nouls with AND / OR | Routes to one handler | Sorts, prioritises, applies thresholds, feeds a weighted total |
| Watch out for | Two conditions in one question; questions phrased so that yes means "absent" | Using it on an ordered scale: distance information is lost | Using it on unordered categories: distances are made up |
TypeSafe's own guidance agrees: Score is not for discrete unordered categories, and plain yes/no questions belong to Noul. Score docs · Choice docs · Noul docs
For Choice versus Score. If any of them says "order matters", use Score.
Meaningful average → Score.
Order carries meaning → Score.
Needs a position → Score.
Both mistakes give you a confidence number that looks fine and means the wrong thing.
Choice gives both 0.50. It can't tell a close call from a model that hasn't understood the ticket. Score gives them 0.67 and 0.
The same billing/shipping split scores 0.25 or 0 depending only on the order you happened to list the options. The number means nothing.
One question with two conditions, and how to fix it
A support team wants Jev to spot customers who are angry and want their money back, and send them to a senior agent. The natural first attempt is to ask exactly that, in one Noul.
"angry_refund": {
"type": "noul",
"instructions": "Is the customer angry and asking for a refund?"
}
Three messages come in. The compound question squeezes each into one number, and the numbers look alike. Values are illustrative.
| Message | What is really going on | angry_refund |
|---|---|---|
Three weeks late and I'm furious. I don't want my money back, I want the order! | Angry, no refund | 0.35 |
Hello, could you please refund order #4411? It doesn't fit. | Calm, wants a refund | 0.40 |
This is getting annoying. Maybe I should just ask for my money back. | Somewhat annoyed, maybe a refund | 0.45 |
All three land around 0.4. At any threshold they are routed the same way, yet the first needs the order chased, the second needs the refund flow, and the third needs a person. The number cannot tell you which condition is uncertain, and no code can get that back.
Ask one Noul per condition, both phrased so that a high value means yes, in the same request. They run in parallel, so the second costs almost no time. Values are illustrative.
| Message | is_angry | wants_refund | Your code routes to |
|---|---|---|---|
| Three weeks late, furious | 0.92 | 0.06 | Chase the order, apologise |
| Polite refund request | 0.05 | 0.95 | Standard refund flow |
| Annoyed, maybe a refund | 0.62 | 0.55 | A person: both are unclear |
# yes / no / None (unclear), with cut-offs set by the cost of each mistake
def read(p, yes_at, no_at):
return True if p >= yes_at else False if p <= no_at else None
angry = read(a["is_angry"].noul, 0.7, 0.3) # missing an angry customer is costly
refund = read(a["wants_refund"].noul, 0.8, 0.2) # a wrong refund is costly
if angry is None or refund is None: send_to_person(ticket) # a condition is unclear
elif angry and refund: send_to_senior_agent(ticket)
elif refund: start_refund_flow(ticket)
elif angry: chase_order_and_apologise(ticket)
else: handle_normally(ticket)
Example: Free / Pro / Enterprise customers routed to different teams. If a near miss really is less harmful, use Score and act on the peak level. If any mistake is equally bad, use Choice.
Example: a severity scale that blends impact and reach. Split it into separate Score questions, as the docs recommend, and combine them in your code.
If the two options are really "yes" and "no", use Noul. One probability carries both the answer and the certainty, and the threshold stays in your code.