Blog / Calibration / FIG. 114
Per-Choice Probabilities, Explained
LLM probability output, explained: how to read per-choice scores, margins and the unclear band, and when calibrated probabilities are worth trusting.
Most models answer with a word. Jev, the decision model from TypeSafe AI, answers with a distribution: per ecosystem documentation, its Choice API returns a probability for each choice in the set you supplied. That's a different kind of LLM probability output than the token logprobs some chat APIs expose, and it changes what your code can do with an answer.
This page owns the output shape: what the numbers are, how to read them, and the mistakes that turn a rich signal back into a coin flip. The API surface itself (auth, model id, SDKs) belongs to the API page, and the case for closed answers over parsed JSON belongs to the structured outputs guide. If you're new to the model, what Jev is comes first.
A label throws information away; a distribution keeps it
Picture two verdicts on "Is this ticket a billing dispute?", both of which a label-only API would report as YES:
- YES 0.97, NO 0.02, UNCLEAR 0.01
- YES 0.51, NO 0.12, UNCLEAR 0.37
The first is a ruling. The second is the model telling you, in numbers, that the ticket is ambiguous and it's leaning rather than knowing. A label-only interface flattens both into the same word, and your pipeline treats them identically. Per-choice probabilities let you treat them differently, which is the whole point of having them.
Three readings matter in practice:
The top probability. How strongly the model favors its best answer. This is what most confidence gates check.
The margin. The gap between the top two choices. A 0.48 vs 0.45 split between two labels is a tie, whatever the argmax says. Margins catch the case where the top probability looks moderate but the real story is two overlapping labels (usually a choice set design problem, not a model problem).
Mass on the escape hatch. Probability sitting on UNCLEAR (or your custom "insufficient evidence" option) is the model declining to commit. That's information, not failure. Route it; don't round it away.
Gating on the distribution as an LLM confidence score
The standard move is a gate: act on confident verdicts, escalate the rest. Here is the shape as pseudocode, not real SDK syntax:
// pseudocode: illustrative only, not a real client API
probs = verdict.probabilities // one score per choice
top, second = two highest entries of probs
if probs["UNCLEAR"] is high: route to human or bigger model
else if top.score - second.score is small: escalate (it's a tie)
else if top.score is above threshold: act on top.label
else: escalate
Where those thresholds sit is a business decision wearing a math costume. A wrong auto-approval and a wrong escalation have different costs, and the threshold should reflect that ratio. The jeval-confidence build does exactly this, as reported: it calibrates against labeled data and sets the hand-off line from the cost of a mistake. One builder gating a coding agent's shell commands reports measuring 18 commands to decide where the thresholds should sit; small sample, right instinct. The full method lives in the confidence thresholds guide.
Probabilities are not promises: calibrated probabilities
A score of 0.9 is only useful if things scored 0.9 are right about 90 percent of the time. That property is calibration, and no model gets it for free on your data. Calibrated probabilities are ones you have checked: take a few hundred labeled cases, bucket verdicts by score, and compare each bucket's score to its actual accuracy. If the 0.9 bucket is right 75 percent of the time, your gate at 0.9 is looser than you think.
Two honest notes. First, there are no official Jev calibration benchmarks; builders publish their own checks, and yours is the one that counts. Second, calibration is per question: a well-calibrated spam question tells you nothing about your urgency question. How to calibrate walks the bucket-and-compare ritual.
The mistakes that waste the signal
Logging only the winner. Store the full distribution and the question version with every verdict. The winner alone can't tell you later whether a bad call was a confident mistake or a coin flip you acted on.
Averaging across questions. A 0.8 on "is this spam" and a 0.8 on "is this urgent" are not the same unit. Combine decisions in code with explicit rules, not by blending scores.
Comparing across question versions. Reword a question and its score distribution shifts. Old thresholds on new wording are guesses.
Re-asking until the number looks better. If a verdict lands in the uncertain band, escalate. Re-rolling for a higher score is p-hacking with an API key.
Assuming normalization. Whether scores are guaranteed to sum to one, and how UNCLEAR is treated, are wire-format details; docs.typesafe.ai is the source, not this page. Write gates that don't depend on it.
Frequently asked questions
What does a per-choice probability mean?
It's the model's score for each option in your closed answer set, per ecosystem documentation of the Choice API. Read the top score, the gap to the runner-up, and the mass on UNCLEAR together; any one alone hides something.
Are Jev's probabilities calibrated?
Treat them as uncalibrated until you've checked on your own labeled data, per question. Builders report doing exactly that, and the calibration guide covers the method.
How is this different from logprobs on a chat model?
Logprobs score generated tokens; per-choice probabilities score your answer options directly. You get a distribution over decisions instead of reconstructing one from text.
Is a per-choice probability the same as an AI confidence score?
It's the raw material for one. The top probability and the margin are what most teams mean by an AI confidence score, but the number only means what a calibration check on your own labeled data says it means.
What threshold should I use?
The one your error costs justify: stricter where a wrong action is expensive, looser where escalation is the expensive part. The thresholds guide shows how to derive it rather than guess.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.