Blog / Recipes / FIG. 113
Choice Set Design: Labels That Judge Well
Classification label design for decision models: build answer sets that are exclusive, exhaustive, and action-shaped, so verdicts stop wobbling.
Most classification label design happens in about four seconds: someone types "billing, technical, account, other" into a config file and moves on. Then the pipeline ships, a third of tickets land in "other", and the team spends a month blaming the model. The model was fine. The labels were a shrug with commas.
A decision model like Jev, TypeSafe AI's model for closed-set verdicts, can only answer from the set you hand it. That makes the answer set half the question. Question wording gets most of the attention (and deserves it); this page owns the other half: the labels themselves, how they relate to each other, and how to build an answer set taxonomy that a judge can actually use.
A choice set is a set, not a list
The single most useful mental shift: stop thinking of labels as a list of things that exist, and start thinking of them as a partition of every input you will ever see. Two properties follow.
Mutually exclusive. For any input, at most one label should be clearly right. If "billing" and "refund request" can both be true of the same ticket, the model will split its confidence between them, and your downstream code will read that split as uncertainty when it's actually overlap. Either merge them, nest them, or turn one into its own yes/no question.
Exhaustive, honestly. Every input needs somewhere legitimate to go. Per ecosystem documentation, Jev's Choice API defaults to YES / NO / UNCLEAR when you don't customize the set, and that third option is doing real work. When you write a custom set, keep an escape hatch, and be precise about which kind you mean: "unclear" (the evidence is ambiguous), "none of these" (the input is clear but out of scope), and "other" (a real category you haven't named yet) are three different facts. Collapsing them into one bucket turns your most informative verdicts into noise.
Shape labels around actions, not ontology
The test that fixes most taxonomies: if two labels trigger the same downstream action, they're one label. If one label triggers two different actions depending on something the label doesn't capture, it's two labels.
Taxonomies built from ontology ("what kinds of tickets exist in the universe?") sprawl. Taxonomies built from actions ("which queue does this go to?") stay small, because the number of things your system can do is small. A support router with nine queues needs nine labels plus an escape hatch, not the 40-leaf tree someone drew in a planning doc.
The corollary for multi-label problems: don't fake them with a single choice set. "Which of these issues does the review mention?" is really five yes/no questions wearing a trench coat. Ask them separately. The SuperX scorer asks 61 questions per draft for a reported $0.0004, so at decision-model prices decomposition is effectively free, and each answer becomes independently debuggable.
Names are instructions
The model reads your labels. A label called "P2" tells it nothing; a label called "billing dispute: customer contests a charge" tells it what the category means. Treat each label name as a miniature definition:
- Use plain descriptive names, not internal codes. Map to codes in your own code afterwards.
- Put the distinguishing feature in the name when two labels sit close together ("refund requested" vs "refund status question").
- Keep grammatical shape parallel. A set mixing nouns, questions, and sentence fragments reads like a set assembled by committee, because it was.
- Avoid loaded or overlapping vocabulary. If "urgent" appears in one label, don't let "critical" appear in another unless you've legislated the difference.
Position is worth a test, too. Chat-model judges are known to drift toward the first option in comparisons, per the judge guide's bias list. Nobody has published whether choice order moves Jev's probabilities, so shuffle order on a calibration batch and check whether verdicts move. If they do, you've learned something about your labels (usually that two of them overlap).
Size: small sets, staged when they grow
Per ecosystem documentation, choice sets cap around 255 options per question; the jev-tree page covers that limit and the recursive workaround (build), so we won't repeat it. The craft point is different: long before the cap, big flat sets get worse. Twenty near-neighbor labels means twenty chances to split probability across overlapping options. When a set passes a dozen or so labels, ask whether a two-stage question (coarse category, then fine label within it) would be clearer. Usually it is.
Taxonomies evolve; version them
Your labels will change. A new product launches, a queue splits, legal invents a category. Every change to a choice set is a change to the instrument: adding a label steals probability from its neighbors, renaming one shifts what the model thinks it means. Tag every verdict with the set version, and never compare label distributions across versions without saying so out loud. The question versioning guide covers the mechanics; the short version is that a choice set is config, and config lives in version control.
Frequently asked questions
How many labels should a classification choice set have?
As many as your downstream actions require, plus an honest escape hatch. Past a dozen or so near-neighbor options, a staged two-question design is usually clearer, and the ~255 cap (per ecosystem documentation) is covered on the jev-tree page.
Should I include an "other" label?
Include an escape hatch, but decide which one you mean: "unclear" for ambiguous evidence, "none of these" for out-of-scope inputs, "other" for a category you haven't named. Mixing them destroys the signal each one carries.
Is multi-label classification possible with a single choice set?
Not cleanly. Split it into one yes/no question per label; at reported decision-model prices the extra questions cost almost nothing, and each answer can be calibrated on its own.
Do label names affect verdict quality?
Yes, because the model reads them as part of the question. Descriptive names with the distinguishing feature built in beat internal codes every time; map to codes in your own code afterwards.
How do label probabilities relate to choice set design?
Overlapping labels split probability between themselves, which looks like uncertainty but is really a taxonomy bug. The probability outputs explainer shows how to read those splits.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.