shipwithjev

Blog / Guardrails / FIG. 76

Add a Guardrail to Your Chatbot This Afternoon

Add chatbot guardrails in an afternoon: check each reply with Jev before it ships, route unsure answers to a fallback, log every verdict.

Your chatbot is one creative answer away from a screenshot. Chatbot guardrails are the fix: before a generated reply reaches a user, a second model asks a few closed questions about it and blocks, rewrites, or escalates on the answers. Jev, TypeSafe AI's decision model, is cheap and fast enough to check every single reply rather than a sample.

The full architecture (which checks belong in the battery, how failure handling should work across a product) is the job of LLM guardrails. This page gets one working guardrail into production by dinner.

The pattern in one sentence

Your chat model writes, Jev judges, your code decides. Jev never writes the reply and never rewrites it; it doesn't generate text. It answers questions like "Does this reply promise a refund?" and returns a probability per answer.

The best public evidence for a judge-then-escalate setup is Hassan's fraud pipeline: Jev sorted 100 emails in 1.42 seconds, sent the uncertain ones to Kimi K3, and the full pipeline got 96 of 100 right for about $0.07, as reported (build). Different task, same shape: a fast verdict on everything, a heavier fallback for the unsure slice. A Mastra moderation build reports about 0.4 seconds median latency for its input check (build), which is the kind of overhead a chat reply can absorb. Both are builder-reported; no official Jev benchmarks exist.

The afternoon recipe

  1. Pick one failure you actually fear (15 minutes). Not "bad answers". Something specific: promising refunds, giving medical dosing, quoting prices that aren't on the pricing page, mentioning a competitor. One guardrail done well beats six done vaguely.

  2. Write the output check as a closed question (20 minutes). For example: "Does this assistant reply commit the company to a refund, credit, or discount? YES / NO / UNCLEAR." Include the user's message as context so the judge can see what was asked. Question craft matters more than anything else here.

  3. Wrap your reply path (45 minutes). The check sits between generation and delivery.

# pseudocode, not real API syntax
reply   = chat_model.generate(conversation)
verdict = jev.choose(REFUND_PROMISE_Q, pack(user_msg, reply), ["YES", "NO", "UNCLEAR"])

if verdict.top_choice == "NO" and verdict.top_probability >= 0.9:
    send(reply)
else:
    send(SAFE_FALLBACK)          # e.g. "Let me connect you with the team."
    queue_for_human(conversation, reply, verdict)

Per ecosystem documentation, the model is typesafe-ai/jev via the Vercel AI Gateway with per-choice probabilities. The real syntax is at docs.typesafe.ai.

  1. Set the threshold conservatively (10 minutes). Only ship replies the judge is confident are safe. Everything else gets the fallback. You'll over-block at first, which is the correct way to be wrong on day one.

  2. Log every verdict (20 minutes). Store the reply, the verdict, the confidence, and what was sent. Your first week of logs is the dataset you'll tune the threshold on.

  3. Test with ugly inputs (30 minutes). Write 20 conversations designed to trigger the failure and 20 that shouldn't. Check the guardrail catches the first set and mostly passes the second.

Input side vs output side

The recipe above checks what your bot says. Checking what users send (jailbreaks, instruction smuggling, "ignore your previous instructions") is a separate layer with its own failure modes, covered in prompt injection detection. Most teams need both eventually. Start with whichever failure would hurt more this month.

Safe AI replies without killing the product

The obvious objection: "Won't this make my bot annoying?" If you block too much, yes. That's why the fallback should be graceful ("let me get someone who can help") rather than a refusal, and why you tune from logs rather than vibes. Most replies to most questions are boring and pass cleanly.

The other objection: "Can't the chat model just follow the system prompt?" Sometimes. A guardrail exists for the times it doesn't. A separate judge with one narrow question is harder to talk out of its job than the model that's already mid-conversation with a persuasive user.

Non-negotiables

Irreversible actions need more than a verdict. If your bot can issue refunds, cancel orders, or change accounts, a single guardrail pass is not authorization. Require explicit user confirmation plus a server-side policy check, and a human for anything above a set amount.

High-stakes domains need humans in the loop. Health, legal, money, hiring: the guardrail reduces risk, it doesn't make the bot qualified. Keep an audit log and a human review path. None of this is legal advice.

Frequently asked questions

How much latency do chatbot guardrails add?

One builder reports around 0.4 seconds median for a Jev moderation check, as reported. Run the check in parallel with any streaming UI work to hide most of it.

Should the guardrail rewrite bad replies?

Jev can't rewrite; it only judges. Send a safe fallback or regenerate with the chat model, then judge again, and cap retries so you don't loop.

How many guardrail checks should a chatbot have?

Start with one for your worst failure, then add checks as logs reveal new ones. LLM guardrails covers how to build out the full battery.

Does an output guardrail stop prompt injection?

Partly, since it catches bad replies whatever caused them, but it doesn't screen hostile input. Pair it with an input check from prompt injection detection.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.