shipwithjev

Blog / Classification / FIG. 153

Text Classification in 2026: The Working Guide

Text classification assigns labels to text. The four ways to do it in 2026, how to pick one, how to measure it, and what it costs per real receipts.

Text classification assigns a label from a fixed set to a piece of text: spam or not, which department, urgent or routine, positive or negative. It's the most common NLP job in production, and in 2026 the question isn't whether you can do it. It's which of four approaches fits your volume, your labels, and how often those labels change.

What counts as text classification

Four shapes cover nearly everything:

  • Binary: spam or not, relevant or not.
  • Multi-class: exactly one of several labels, like routing a ticket to one queue.
  • Multi-label: any number of labels at once, covered on its own page.
  • Hierarchical: a taxonomy where the first label decides which labels come next.

Sentiment analysis, intent detection, moderation, and lead scoring are all NLP text classification wearing different hats. The useful output is always a label plus a confidence. A label alone can't tell you when to ask a human.

The four approaches

Rules and keywords. Fast, transparent, and brittle. Great for compliance hard-stops, bad at paraphrase.

Trained classifiers. Classical models or fine-tuned transformers. Cheap per item at volume, stable, but they need labeled data and a retrain whenever labels change.

Zero-shot and few-shot models. Labels supplied at inference time, no training run. The zero-shot guide covers the three generations.

Decision models. Closed questions in, a probability per choice out. That's the shape of Jev, the decision model from TypeSafe AI, explained here. Priced like classification, worded like instructions.

The techniques field manual breaks these into nine specific methods, ranked by when each wins.

How to pick

No labeled data and labels that change monthly: zero-shot or a decision model, then calibrate. Tens of thousands of labeled examples and frozen labels: a trained classifier probably wins on cost. Somewhere in between, and most teams are: run both in shadow mode and let the numbers decide, the approach the old-versus-new comparison lays out.

If the decisions need to be explained later to an auditor or a customer, favor approaches that log evidence and question versions with every label.

What it costs now

One builder classified 500 emails for 3.5 cents, as reported in the build entry. That's one question per email; batteries with more questions scale linearly. Trained models run nearly free per item, but the labeling, training, and maintenance bill arrives up front.

The fastest way to get your own number is classifying a CSV: about twenty minutes from export to labeled file.

How to measure it

Accuracy lies on imbalanced data. If 95% of mail is fine, a classifier that labels everything fine scores 95% and catches nothing. Use per-class precision and recall and look at the confusion matrix.

Then measure what matters most: agreement with a person on a hundred of your own examples. And check whether confidence means anything. If items scored 0.9 are right about 90% of the time, the thresholds you set will behave.

Frequently asked questions

What is text classification?

Assigning a label from a fixed set to a piece of text, such as spam or not, or which team should handle a ticket. Production systems also return a confidence, so uncertain items can go to a person.

What's the difference between text classification and sentiment analysis?

Sentiment analysis is one kind of text classification, where the labels are feelings like positive, negative, or neutral.

Do I need labeled data for text classification?

For trained classifiers, yes, often thousands of examples. For zero-shot and decision models, only a small evaluation set of about a hundred.

How many categories can one classifier handle?

It depends on the approach. Jev's Choice API caps around 255 choices per question per ecosystem documentation, and taxonomies past that recurse in stages.

Is an LLM overkill for text classification?

Often, yes. Classification-shaped models priced for classification exist precisely because chat models are expensive for picking labels.

What's the best first text classification project?

Export a few hundred real items, classify them, and compare against a human's labels on a sample.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.