shipwithjev

Blog / Calibration / FIG. 107

Confidence Thresholds: The Complete Guide

Confidence threshold LLM guide: how to set probability cutoffs from labeled data and error costs, with two-sided bands and builder receipts.

A confidence threshold is the line where a model's verdict stops being a suggestion and starts being an action. Set it by instinct ("0.8 sounds safe") and you get either a pipeline that escalates everything or one that confidently misroutes the cases that mattered. This confidence threshold LLM guide is about setting that line on evidence. It applies to any model that returns probabilities; the examples use Jev, TypeSafe AI's decision model, because per ecosystem documentation it returns a probability for each choice natively.

Custody note: what happens after the line (which tier the uncertain slice escalates to) is the cascade, covered in LLM routing. This page is the craft of drawing the line itself.

What the number actually is

A per-choice probability is the model's score for each answer in your set. A threshold turns that score into a decision: accept the top choice if its probability clears the bar, otherwise do something else.

Two warnings before touching any cutoff. First, a probability is only as meaningful as its calibration: a verdict at 0.9 should be right about nine times in ten on your data, and whether it is can only be measured, never assumed. Second, the number depends on the question. Reword the question and the distribution shifts, so thresholds belong to a specific question version, not to the model.

How to set probability cutoffs from data

The method is short and non-negotiable.

  1. Label a few hundred real cases with the answer you'd accept as correct. The 100-case calibration method is the floor; high-stakes routes deserve more.
  2. Run your question and keep every probability, not just the top choice.
  3. Sweep the threshold. For each candidate cutoff, compute two numbers: accuracy on the cases above it, and the share of cases falling below it (your escalation rate).
  4. Price the errors. A false "not fraud" and a false "fraud" rarely cost the same. Weight each error type by what it costs you, then pick the cutoff that minimizes total cost, including the cost of escalating.
  5. Write it down with a version. Threshold, question version, date, and the labeled set it was fit on.

Tooling is catching up to this ritual. jeval calibrates against labeled data and sets the hand-off line from the cost of a mistake, per its entry, so low-confidence cases go to a person "on the numbers rather than on a guess." jevcal fits and drift-checks thresholds against labeled data. And one builder put a probability gate in front of a coding agent's shell, write, and edit calls, then measured 18 commands to decide where the thresholds should sit (build). Eighteen is small, and he measured anyway; that's the right instinct.

Patterns that beat a single cutoff

Two-sided bands. For a yes/no gate, use two lines: auto-accept above the high one, auto-reject below the low one, and send the middle band to review. One cutoff forces every uncertain case into a wrong bucket.

Per-class thresholds. In a multi-label set, rare and costly classes ("security incident") deserve lower bars for flagging than common, cheap ones ("general question"). A single global cutoff over-serves the easy classes and under-serves the dangerous ones. Moderation tooling in the directory already works this way: jevmod returns a probability per category with thresholds you set.

Respect "unclear." If your choice set includes an UNCLEAR option (the default set does, per ecosystem documentation), a high probability on UNCLEAR is itself a routing signal. Don't threshold it away.

User-owned thresholds. For personal tools, let the user move the line and give them an escape hatch. Jev Slop Guard blurs posts scoring above a user-set threshold, with a reveal button for mistakes: a cutoff that admits it can be wrong.

The fraud gate, read as a threshold story

The canonical receipt is Hassan's fraud cascade: Jev judged 100 emails in 1.42 seconds, the cases it was unsure about went to Kimi K3, and the full pipeline got 96 of 100 right for about $0.07, as reported. The routing gets the headlines, but the threshold did the work. Draw the "unsure" line too high and the expensive model sees everything, wiping out the cost win. Too low and the uncertain errors sail through unescalated. The reported result is what a well-placed line looks like; the builder's exact cutoff isn't in our entry, so find yours on your data. How accurate the base model is in general is a separate question, with no official benchmarks, covered honestly in Jev accuracy.

Keeping confidence thresholds honest after launch

Thresholds drift because inputs drift. Watch two signals: the escalation rate (rising means reality and your question are diverging; falling to zero usually means the gate broke) and agreement on a rotating human-audited sample. Re-fit on a schedule and whenever the question changes. And one rule no threshold overrides: irreversible actions (payments, deletions, account changes) never execute on a lone verdict, however confident.

Frequently asked questions

What is a good confidence threshold for an LLM classifier?

There is no universal number; it depends on your question, your data, and what each kind of error costs. Fit it by sweeping cutoffs over a few hundred labeled cases and picking the lowest-total-cost point.

Should I use one threshold for every class?

Usually not. Rare, costly classes deserve lower flagging bars than common, cheap ones, and yes/no gates work better with a two-sided review band.

Are Jev's probabilities calibrated?

That has to be measured on your task; no official benchmarks exist. Tools like jeval and jevcal exist to check calibration against labeled data before you trust a cutoff.

What happens to cases below the threshold?

They escalate to a bigger model, a human, or a safe default, depending on stakes. The escalation catalog is in escalation design patterns, and the cascade core is LLM routing.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.