shipwithjev

Blog / Classification / FIG. 154

Zero-Shot Classification: The Practical Guide

Zero-shot classification labels text into categories a model never trained on. How the three generations work, where each breaks, how to calibrate.

Zero-shot classification sorts text into categories the model never trained on. You supply the labels at inference time, the model decides which fits, and changing the labels changes the classifier instantly. No labeled training set, no fine-tuning run.

That property is why decision models like Jev, from TypeSafe AI, are functionally zero-shot classifiers with probabilities attached. It's also why the hard part of classification moved from training to wording.

How zero-shot classification works: three generations

NLI-based. Each label becomes a hypothesis, like "This text is about billing," and a natural-language-inference model scores how strongly the text entails it. This was the classic open-source route and made zero-shot a one-liner in popular libraries. Solid for coarse labels, sensitive to how hypotheses are phrased.

Embedding-based. Embed the text and a description of each label, and the nearest label wins. Cheap and scalable, weak at fine distinctions.

Instruction-following. Chat models and decision models read the labels as instructions, including boundary clauses. Decision models return a probability for every choice rather than a single pick, per ecosystem documentation of the Choice API. That distribution is what makes thresholds and escalation possible.

Why zero-shot doesn't mean zero work

The label wording is the model. "Billing" and "billing (charges, refunds, invoices; not login problems)" are different classifiers. Writing judge questions covers boundary clauses, the single biggest accuracy lever.

Give the model a way out. A forced choice between four labels guarantees confident nonsense on the fifth kind of input. An unclear option routes those to a person, per choice-set design.

And calibrate anyway. Zero training data doesn't mean zero evaluation data: run a hundred real items, compare with a person's labels, fix the labels that disagree, and repeat. That hour is the difference between a demo and a system, and the calibration guide walks through it.

Zero-shot vs trained classifiers

Zero-shot wins on cold starts, labels that change, moderate volume, and long-tail categories you'd never collect enough examples for.

Trained classifiers win when labels are frozen, volume is enormous, and the lowest possible per-item cost matters. The honest pattern is to start zero-shot, then shadow-test a trained model once you have the labeled history, and graduate if it matches. The old-versus-new comparison covers when that switch pays.

What you don't do is fine-tune the decision model itself. The tuning happens in your questions, as the fine-tuning answer explains.

A zero-shot pipeline in one afternoon

  1. Write each label with a boundary clause, and include an unclear option.
  2. Run 100 real items. Have a person label the same items.
  3. Read every disagreement. Fix the wording that caused it.
  4. Re-run, then save the wording as a version.
  5. Set a confidence threshold; below it, a person decides.

For cost, one builder classified 500 emails for 3.5 cents, as reported in the build entry. The calibration runs cost less than the coffee you drink during them.

Frequently asked questions

What's the difference between zero-shot and few-shot classification?

Few-shot adds a handful of labeled examples, in the prompt or through light fine-tuning. Zero-shot uses only the label names and descriptions.

How many labels can zero-shot classification handle?

Label overlap limits you before label count does. For decision models, the Choice API caps around 255 choices per question per ecosystem documentation, and larger taxonomies split into stages.

Is zero-shot classification accurate enough for production?

For routing, triage, and screening, often yes after calibration. Measure agreement on your own data before trusting any number.

Does zero-shot classification work in other languages?

It depends on the model. For Jev specifically, multilingual evidence is still thin, so test on your own language samples first.

When should I switch to a trained classifier?

When labels have frozen, volume is huge, and a shadow test shows the trained model matching your zero-shot agreement rate.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.