Blog / Classification / FIG. 160
Text Classification Techniques, Ranked by When
Nine text classification techniques from keyword rules to decision models, each with when it wins, when it fails, and what it costs to run.
Most lists of text classification techniques rank by sophistication. This one ranks by when: which technique wins given your data, your volume, and how often your labels change. The working guide covers the concepts. This is the field manual.
Techniques that need no training
1. Keyword and regex rules. Wins on tiny fixed vocabularies and compliance hard-stops ("contains a card number"). Fails on paraphrase, sarcasm, and anything a thesaurus can defeat. Rules versus verdicts covers the handoff point.
2. Zero-shot NLI. Each label becomes a hypothesis scored by an inference model. Wins for quick prototypes with a handful of coarse labels. Fails on nuanced domain labels and is sensitive to phrasing.
3. Embedding similarity. Embed the text and short label descriptions, and the nearest label wins. Wins with many labels at huge scale and low cost. Fails on fine distinctions, since similar isn't the same as correct. See embeddings versus verdicts.
4. LLM prompting. Ask a chat model for the label. Wins on nuance and when you want an explanation too. Fails on cost and consistency at volume.
5. Decision models. Closed questions with a probability per choice. Wins when labels change often, nuance matters, and you need cheap per-call pricing. Fails when the task actually needs generated text.
Techniques that need labeled data
6. Classical ML. TF-IDF features with Naive Bayes, logistic regression, or an SVM. Wins with lots of labeled data, stable labels, and tight compute budgets, and it makes a strong baseline. Fails on context and on any label you didn't train.
7. Fine-tuned transformers. BERT-family encoders trained on your labels. Wins at high volume with frozen labels, where it's often the accuracy leader. Fails when labels change, because every change is a retrain. The BERT comparison has the shadow-test method.
8. Few-shot fine-tuning. Contrastive fine-tuning of sentence transformers on a handful of examples per label (SetFit is the best-known recipe). Wins when you have tens of examples, not thousands. Fails when the few examples don't cover the edge cases.
Techniques that combine
9. Cascades. A cheap first pass decides the easy items and escalates the uncertain ones to something smarter or to a person. Wins almost everywhere at production scale. Five cascade architectures maps the variants.
The "when" in one pass
Labels change monthly: techniques 2 to 5. Ten thousand labeled examples and frozen labels: 6 or 7. Fifty labeled examples: 8 or 5. Mistakes are expensive: 9, with a human in the uncertainty band. Millions of simple items a day: 3 or 6, possibly as the first stage of a cascade.
Frequently asked questions
Which text classification technique is most accurate?
The one that scores best on your own labeled examples. Rankings from other people's data don't transfer reliably.
Can I combine classification techniques?
Yes, and at production scale you usually should. Cascades that route easy cases cheaply and hard cases carefully are the default.
Is Naive Bayes still worth using?
As a baseline, yes. For spam-like tasks with plenty of labeled data, it's fast, cheap, and surprisingly hard to beat.
Where does zero-shot NLI fall short?
On subtle domain labels and phrasing sensitivity. Instruction-following models handle nuance better, covered in what Jev is.
What's the fastest way to compare two techniques?
Run both on the same 100 to 200 labeled items and compare per-class precision, recall, and cost.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.