Blog / Triage & routing / FIG. 111
Escalation Design Patterns
Eight LLM escalation patterns for decision pipelines, from confidence gates to never-approve gates, with builder receipts and the anti-patterns.
The quality of an AI pipeline is decided less by how often the model is right than by what happens when it isn't sure. LLM escalation is that "what happens": the rules that move a case from a cheap automatic verdict to a bigger model, a human, or a safe default. Most teams implement exactly one escalation rule (low confidence goes to a person) and discover the rest the hard way. This page is the catalog. Examples use Jev, TypeSafe AI's decision model, which per ecosystem documentation returns a probability per choice, the raw material most of these patterns run on.
Custody: the core cheap-first, expensive-on-doubt cascade is LLM routing, and where to draw the confidence line is the confidence thresholds guide. This page covers the patterns built on and around them.
The eight escalation patterns
1. The confidence gate. The baseline, stated in one line because routing owns it: verdicts above your threshold proceed, verdicts below it go up a tier. Hassan's fraud cascade is the reference receipt: Jev judged 100 emails in 1.42 seconds, the unsure cases went to Kimi K3, and the pipeline got 96 of 100 right for about $0.07, as reported. Every other pattern here is a refinement or a guard around this one.
2. The stakes bypass. Some categories escalate regardless of confidence. A ticket classified as "security report" or "legal threat" with 0.99 probability still goes to a person, because the cost of an automated mistake there is out of proportion to the savings. Implement it as a policy table keyed on the verdict's category, not on its probability. In support triage, this is the rule that stops a P0 riding on one unaudited call.
3. Wake the expensive model. Instead of escalating uncertain cases, escalate interesting ones. The decision model watches everything and only hands off when something needs a generative response. The receipt: a builder has Jev monitor more than 25 customer WhatsApp groups in real time, deciding whether something needs his attention, such as an urgent problem or an open order; when it does, an LLM writes him the message (build, as reported). The expensive model's cost scales with events, not with traffic.
4. The graduated ladder. Escalation doesn't have to be binary. jev-logtriage has Jev score collapsed log batches for noise, severity, and whether an operator should act; plain code then maps the answers onto five rungs: suppress, watch, review, notify, or page. Per its entry, nothing is executed. The model supplies judgment, code supplies policy, and each rung has a different cost of attention.
5. Escalate but never approve. Give the model authority in one direction only. jev-gates is a set of seven gates for Claude Code (rules, scope, intent, done, claims, proof, and commit honesty) that escalate but never approve, per the entry. The model can stop things and summon a human; it cannot wave anything through on its own. This asymmetry is the right default wherever an agent can take actions.
6. The unattended allowlist. The mirror image: a narrow class of cases may proceed automatically, and everything else waits. MetalBear's jev-auto-approve reads a pull request diff and asks Jev whether it's the kind of change that can go in unattended; everything else waits for a person, per the entry. The question is about eligibility for autonomy, not about quality, which is a much easier thing to rule on well.
7. The disagreement trigger. Ask two differently-framed questions about the same judgment (or ask two models), and escalate whenever they disagree. Disagreement catches confident-but-wrong verdicts that a confidence gate misses, because two framings rarely fail identically. It costs a second verdict, which at builder-reported decision-model prices is usually cheap insurance on consequential routes.
8. The "unclear" lane and the outage lane. Two fallbacks every pipeline needs. First, if your choice set includes UNCLEAR (the default set does, per ecosystem documentation), route UNCLEAR to review directly instead of forcing it into a real category. Second, decide in advance what happens when the model is unreachable: queue and wait, degrade to rules, or fail closed to a human. The wrong default chosen during an outage becomes a week of cleanup.
LLM escalation anti-patterns in human in the loop routing
Re-rolling until confident. Re-asking a low-confidence case until it clears the bar converts uncertainty into fake certainty. Escalate; don't reshuffle.
The escalation black hole. A review queue without an owner, a service level, or a daily count is just a slower way of dropping cases.
No feedback loop. Every escalated case is a labeled example of your cheap tier's weakness. Review them weekly and decide whether the fix is a sharper question or a category that should stay escalated forever.
Irreversible actions on a lone verdict. Payments, deletions, sends, and account changes need a verification step or a human, at any confidence. The ergonomics of the human side, queue design and reviewer load, are covered in human-in-the-loop design.
Frequently asked questions
What is LLM escalation?
The set of rules that moves a case from an automatic model verdict to a stronger model, a human, or a safe default, usually triggered by low confidence, high stakes, or model disagreement.
Should every low-confidence verdict go to a human?
Not necessarily: a bigger model can take the uncertain slice first, with humans reserved for what it can't resolve or for high-stakes categories. The fraud cascade receipt used a bigger model as the second tier.
How do I escalate high-stakes cases that the model is confident about?
Use a stakes bypass: a policy table that sends certain categories to a person regardless of probability. Confidence measures the model's certainty, not the cost of being wrong.
How much human review should escalation produce?
As little as calibration allows and as much as stakes require; watch the escalation rate as an operational signal. Setting the line on data is covered in the confidence thresholds guide.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.