shipwithjev

Blog / Comparisons / FIG. 101

Jev vs Embeddings: Ruling vs Resembling

Embeddings vs LLM classification: embeddings measure likeness, Jev rules on a question. When each wins, and why the best stacks use both.

The embeddings vs LLM classification debate usually gets framed as old versus new. It's actually about two different questions. An embedding tells you how much one piece of text resembles another. A decision model like Jev, from TypeSafe AI, rules on a question you asked about the text. "Is this ticket like the refund tickets?" and "Is this ticket asking for a refund?" sound identical, and the gap between them is the whole page.

The one-paragraph summary of embeddings among all the alternatives lives in Jev alternatives. This page owns the line between the two.

Semantic similarity vs judgment

Embeddings turn text into vectors, and nearby vectors mean similar meaning. Classification by embedding works by proximity: embed labeled examples, embed the new item, take the label of whatever it lands closest to. It's cheap, fast, and genuinely good when your categories are topics.

It struggles when the category depends on a condition rather than a topic. "A refund request unless the order is under 30 days old" is a rule. "A complaint that mentions a safety issue" is a rule. "A review that praises the product but reports a defect" is two facts in tension. Embeddings compress all of that into one point in space, and the "unless" gets lost in the average.

Jev answers a question instead. Per ecosystem documentation, you give it a question and candidate answers and it returns a probability for each. The question can carry the condition, the exception, and the boundary, because the model reads it every time. It doesn't generate text; it just rules.

A quick test for which one you need: if you could explain the category to a new hire by showing five examples, embeddings will probably do. If you'd have to explain it with a sentence containing "unless", "but only if", or "even though", you want a ruling.

Where embeddings still win

  • Retrieval. Finding the 50 documents out of a million that might matter. Nearest-neighbor search is built for this, and a ruling on a million items per query isn't.
  • Clustering and discovery. Grouping thousands of open-ended responses to see what themes exist before you know what questions to ask.
  • Stable topic routing at huge volume. When categories are topics that don't change, proximity is fast and free after indexing.
  • Candidate generation. Narrowing the field so something more careful can look at fewer items. This one matters most, and it's the next section.

The blocking pattern: resemble first, rule second

The best stacks don't choose. They use embeddings to find candidates and a decision model to rule on them. Entity resolution is the canonical case: comparing every record to every other record is impossible at scale, so a cheap step (shared keys, phonetic codes, embedding neighbors) proposes plausible pairs, and only those pairs get the question "are these the same company?" That page owns the full pipeline; the principle generalizes.

Search is the same shape. Retrieval (keyword or embedding) builds a shortlist; a ruling reorders it by what the user actually meant. One cataloged build, jevsearch, skips embeddings entirely for site search: keyword hits stream first, then Jev re-ranks by intent, with no vector database. The builder reports $0.26 per 1,000 uncached searches and a 278 ms median, uncached over the network (as reported). Another, semantic code search without a vector index, passes a shortlist to Jev in a single batched call.

If you want to see the comparison measured rather than argued, jev-search-rerank-eval evaluates Jev reranking against lexical, embedding, and fusion baselines on Chinese and English retrieval, including an analysis of judge circularity. That's one builder's evaluation on their data; there are no official Jev benchmarks, and your corpus is the only one that settles it for you.

Embeddings vs LLM classification: cost and speed

Embeddings are cheaper per item once indexed, full stop. A lookup is close to free; a ruling is a model call. Jev's reported per-verdict costs are fractions of a cent (see the cost breakdown), which is cheap for a shortlist of 50 and wasteful for a corpus of 10 million. That asymmetry is exactly why the blocking pattern exists.

Speed follows the same logic. Vector search at scale is fast. Rulings on a short list are fast enough for interactive use, per the jevsearch receipt above. Rulings on everything, every query, are not the design.

Frequently asked questions

Are embeddings or LLM classification better?

Embeddings are better for similarity, retrieval, and topic grouping; a decision model is better for categories defined by conditions and exceptions. Most production systems use embeddings to shortlist and a ruling to decide.

What's the difference between semantic similarity and judgment?

Similarity measures how alike two texts are; judgment answers a specific question about one text. "Resembles refund tickets" and "asks for a refund" diverge exactly on the edge cases that matter.

Can Jev replace my vector database?

For small shortlists and site search, some builders skip embeddings, like jevsearch. For retrieval over large corpora, keep the index and use Jev on what it returns.

What is blocking?

A cheap first pass that shrinks all possible pairs or candidates down to plausible ones, so the expensive ruling only runs where it matters. The entity resolution guide covers it in depth.

Should I replace my embedding classifier?

Only where it fails on conditional categories. Keep it where topics are stable, and add rulings for the cases with "unless" in them; can Jev replace my classifier? walks through the call.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.