Blog / Calibration / FIG. 108
Calibrating a Judge: The 100-Case Method
LLM calibration in one afternoon: the 100-case method for verdict agreement testing, reading disagreements, and checking probability buckets.
Every judge pipeline gets calibrated eventually. The only choice is whether you do it deliberately in an afternoon or a customer does it for you in production. LLM calibration, in the sense builders mean it, is two checks: does the model agree with a careful human on your task, and do its confidence numbers mean what they claim? This page is the ritual for both, in depth, sized to one question and 100 cases. The examples use Jev, TypeSafe AI's decision model, but the method works for any judge.
Where this sits: the getting-started guide introduces the "score 100 known cases" step in a paragraph, and the evals guide turns it into permanent suite infrastructure. This page is the step itself, done properly.
Why 100, and why one question at a time
A hundred cases is small enough to label by hand in an hour or two and large enough that the big failure modes show up. It is not large enough for precise numbers: each case is a full percentage point, so 91 and 94 percent agreement are roughly the same result. Treat the output as a diagnosis, not a benchmark.
One question at a time, because a battery of ten questions scored together hides which one is broken. Calibrate the question that drives the most consequential action first.
The LLM calibration method, step by step
1. Draw the 100. Pull real inputs from production or history, not ones you wrote. Mirror the real mix, but make sure every answer in your choice set appears several times, and deliberately include the edge cases you already know about. Write down how you sampled, so later rounds are comparable.
2. Label blind, before running anything. Record your answer for each case without seeing the model's. If you can, have a second person label the same 100 independently. Human-to-human agreement on your question is your realistic ceiling; if two careful people agree only 85 percent of the time, the question is the problem, and no model will fix it.
3. Run the question and keep everything. Top choice, the full per-choice probabilities (native per ecosystem documentation), the exact question text, and a version tag. In pseudocode (illustrative, not API syntax; the real request format is at docs.typesafe.ai):
# PSEUDOCODE: the shape of a calibration run, not Jev syntax
for case in sample_100:
result = ask(case.input, QUESTION_V3, choices=["yes", "no", "unclear"])
log(case.id, case.human_label, result.top_choice, result.probabilities, "v3")
4. Score agreement three ways. Overall agreement, agreement per answer class, and a confusion table (which labels get mistaken for which). Overall agreement flatters you when one class dominates; per-class numbers are where problems hide.
5. Read every disagreement. This is the step people skip and the one that matters most. Sort each miss into one of four bins: question ambiguity (two readings were reasonable; fix the wording), label error (you were wrong; fix the label), genuinely hard (reasonable people would hesitate; route it to escalation), or bad input (truncated, garbled, off-topic; add an input-quality check). In most first rounds, the first bin is the biggest.
6. Revise, then confirm on fresh cases. Rewrite the question using the operational rewrite rules, rerun, and check the same 100. Then draw a new 100 and label them before you look. Improvements that only appear on the original sample are overfitting to cases you've memorized.
7. Check the confidence numbers. Bucket verdicts by top-choice probability and compute agreement per bucket. If the 0.9-and-up bucket agrees with humans far less than nine times in ten, the probabilities are overconfident on your task, and any cutoff you set needs to account for that. That bucket table is the direct input to setting confidence thresholds.
8. Record the baseline. Question version, sample description, agreement overall and per class, bucket table, date. This is the number every future change gets compared against.
Verdict agreement testing as a standing job
Calibration is not a launch ritual. Inputs drift, products change, and a question that agreed 95 percent of the time in September can quietly slide. Keep a small rotating sample under human review every week or month, compare against the baseline, and re-run the full method whenever the question text changes. Building the labeled set that feeds all of this is its own craft, covered in building gold sets.
The ecosystem is automating parts of the loop. jev-align builds calibrated classifiers from human feedback: you label examples, GEPA optimizes the question, and out comes a Jev function with a known error rate, per the entry. jeval calibrates against labeled data and sets the hand-off line from the cost of a mistake. Tools speed up the rewrite and the math; they don't replace the human labels, and they don't replace reading the misses.
Frequently asked questions
What is LLM calibration?
Checking that a model's verdicts agree with careful human judgment on your task, and that its confidence scores match how often it's actually right. Both are measured on labeled cases, never assumed.
Are 100 cases enough to calibrate a judge?
Enough to find the big problems and set a first baseline, not enough for precise accuracy claims. Confirm on a fresh sample, and grow the set for high-stakes questions.
What agreement rate should I aim for?
As close to your human-to-human agreement as possible, since that's the realistic ceiling. The acceptable floor depends on what happens after the verdict; see the evals guide for turning this into ongoing measurement.
How often should I recalibrate?
Whenever the question text changes, and on a regular rotating sample in production. A falling agreement number is the earliest warning you'll get.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.