shipwithjev

Blog / Recipes / FIG. 137

Five Cascade Architectures, Compared

Five cascade architecture examples from real Jev builds: escalation, routing, shortlist, pick-then-write and gatekeeper patterns, compared.

"Put the cheap model first" is the whole idea of a cascade, and the LLM routing guide already makes the case for it, thresholds and all. What that guide doesn't do is show how differently builders have actually wired it. This page collects five cascade architecture examples from the directory, each built around Jev, the decision model from TypeSafe AI, and compares where the cheap step sits and what it hands off.

All numbers are as reported by each builder and linked to the entry. They measure different tasks, so read them as evidence each shape works, not as a leaderboard.

The five variants at a glance

VariantWhat Jev decidesWho does the restExample
EscalationThe answer, with confidenceA frontier model takes the unsure sliceFraud detection
RouterWhich model should handle thisThe chosen modeljev-router
Shortlist, then judgeWhich candidates matterCheap retrieval found them firstjevsearch, Tocsin
Pick, then writeWhich action or toolA generative model fills in textWebMCP benchmark
GatekeeperWhether the expensive step runs at allNothing, most of the timeSmart-home fast path

1. Escalation: answer first, escalate the doubt

The canonical shape. Jev answers every item; low-confidence items go up a tier. The fraud detection build is the reference: 100 emails in 1.42 seconds, unsure cases sent to Kimi K3, 96 of 100 correct for about $0.07, as reported. Use this when the task is itself a decision and you mostly need a safety net.

The design lever is the threshold, and confidence thresholds covers how to set one.

2. Router: decide who answers

The most familiar of the model routing patterns. Here Jev never answers the task; it picks the model that will. jev-router sends each Claude Code task to the cheapest model that can do it. pi-jev-router asks what a task needs, then chooses the cheapest OpenRouter model on the quality-and-price frontier that meets it. The Qwen and Sonnet router splits between a local and a hosted model and shows every routing decision in a UI, which is the part most routers skip and shouldn't.

Use this when your tasks are generative but vary wildly in difficulty.

3. Shortlist, then judge

Cheap, dumb retrieval narrows the field; Jev judges what's left. jevsearch streams keyword hits first, then re-ranks by intent, reported at $0.26 per 1,000 uncached searches with no vector database. Nader Dabit's Gmail intent search suggests letting embeddings pull candidates first for a very large inbox. Oko reports 0.48 against BM25's 0.21 for needed code within an 8k-token budget.

The extreme version compresses instead of retrieving. Tocsin grouped 22.8 million log lines into 11,812 patterns and asked Jev about each pattern once: six minutes, $0.64, 123 patterns worth a look, as reported, against an estimated $1,120 to run an LLM over every line.

4. Pick, then write

Jev chooses, a generative model composes. This is the cleanest answer to "but Jev can't write." In the WebMCP benchmark, Jev picked each tool and Mercury 2.5 wrote the arguments: 49 of 49 tasks at about 112 times lower model cost than the frontier comparison, per the team. Jev driving the page alone solved 25 of 49, which is exactly why the split exists. pi-jev does the same for code: Jev picks the file excerpts, a local model writes.

Use this whenever the output is prose or code but the hard part is choosing.

5. Gatekeeper: decide whether to spend at all

The least visible and possibly the biggest saver. The smart-home fast path calls Jev before its LLM; device requests switch hardware immediately, and anything uncertain falls through to the LLM. wakegate decides whether an event deserves waking a sleeping agent, skipping only when confident. doc-router asks which PDF pages actually need OCR.

Notice the asymmetry all three share: skipping is only allowed on a confident verdict, and the default is to do the expensive thing. Anything irreversible should never hang on a lone verdict.

Choosing between them

Most production systems end up combining two: a gatekeeper or shortlist in front, escalation behind. Pick by asking where your cost actually goes. If it's volume, shortlist or gate. If it's hard cases, escalate. If it's model choice, route. If it's text, pick then write. Reducing LLM costs walks through that audit.

Frequently asked questions

What is a cascade architecture in AI?

A pipeline where a fast, cheap model handles or triages every request and only a subset reaches a slower, pricier model. The core theory is in the LLM routing guide.

Which cascade variant saves the most money?

It depends on where your spend goes; the largest reported gap here is Tocsin's $0.64 against an estimated $1,120, from compressing before asking. Task differences make direct comparisons unreliable.

Can Jev be the only model in a cascade?

For purely decision-shaped tasks, several builds run Jev alone. Once any step needs generated text or open-ended reasoning, a chat or frontier model has to own that step.

Is model routing the same as a cascade?

Routing is one variant: the decision model picks the destination instead of answering. Escalation, shortlisting and gatekeeping are the others.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.