shipwithjev

Blog / Classification / FIG. 166

Multi-Label Text Classification, the Verdict Way

Multi-label text classification assigns every label that applies. Why N independent yes/no questions often beat one multi-hot model, and how to measure it.

Multi-label text classification assigns every label that applies, not just the best one. A support ticket can be billing, urgent, and a churn risk all at once, and a classifier that must pick one is wrong twice.

The traditional approach trains one model with a multi-hot output. The verdict approach asks N independent closed questions, one per label, each with its own threshold. Here's how both work and when the second wins.

Multi-label vs multi-class

Multi-class means exactly one label per item: mutually exclusive, like which queue a ticket goes to. Multi-label means zero or more.

The most common mistake is forcing a multi-label problem into a multi-class choice set. Ask "is this billing or urgent?" and the model has to pretend one of them is false. Choice-set design covers how to spot this before it ships. For where multi-label sits among the other shapes, see the text classification guide.

The traditional approaches

Binary relevance: one binary classifier per label, trained independently.

Classifier chains: each label's prediction feeds the next classifier, capturing some correlation between labels.

One network, many outputs: a single model with a sigmoid output per label.

All three need labeled data for every label, and each rare label suffers its own class imbalance. Adding a label means relabeling and retraining.

The verdict way: N independent questions

Each label becomes its own yes/no/unclear question, with boundary clauses written the way the question-writing guide teaches. Then:

  • Each label gets its own threshold. "Urgent" wants high recall, since missing one is expensive. "Churn risk" might want high precision to avoid crying wolf.
  • Adding a label is adding a question. No retrain.
  • Evidence is per label, so an auditor can see why each label was applied.

Decomposed batteries are cheap. The SuperX build scores each draft with 61 questions for $0.0004, as reported in the build entry. Cost grows linearly with the number of labels, so a 20-label battery costs roughly twenty single verdicts per item.

Handling label correlation

Independent questions assume independent labels, and sometimes they aren't. "Refund request" implies "billing." Two fixes:

  • Stage the questions: ask the parent first, and ask children only when it's yes. Fewer calls, consistent output.
  • Post-hoc consistency rules: if a child is yes and its parent is no, flag the item for review.

Rules like these are cheap and catch the contradictions that make multi-label output look careless.

Measuring multi-label performance

  • Per-label precision and recall. Always, first.
  • Micro average: pools every decision, so frequent labels dominate.
  • Macro average: averages per-label scores, so every label counts equally.
  • Hamming loss: the fraction of label decisions that are wrong.
  • Exact match ratio: items where every label is right. Strict and usually low.

Calibrate each label separately on the same hundred items, per the calibration guide.

Frequently asked questions

Is multi-label the same as multi-class classification?

No. Multi-class picks exactly one label per item; multi-label picks any subset, including none.

How many labels is too many for the question approach?

Cost and latency grow linearly with labels. Past a few dozen per item, add a cheap first stage that decides which questions are worth asking.

Can one question return multiple labels?

Keep one label per question. Combined questions blur boundaries and make per-label thresholds impossible.

How do I handle a label that's rarely true?

Use a high-recall threshold, review its positives, and deliberately include positives in your calibration set so you can measure it at all.

What metric should I report for multi-label classification?

Per-label precision and recall, plus macro F1 when every label matters equally. More on the model side in what Jev is.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.