shipwithjev

Blog / Monitoring / FIG. 117

Monitoring Judge Drift

Model drift monitoring for verdict pipelines: track verdict distribution shift, escalation rate and a frozen canary set to tell which drift you have.

A judge pipeline rarely breaks loudly. Nothing throws, latency looks fine, the dashboard stays green, and three weeks later someone notices the spam rate halved without anyone fixing spam. That's drift: the relationship between your inputs, your questions, and your verdicts shifting while every health check passes.

Classic model drift monitoring was built for trained classifiers watching feature distributions. Verdict pipelines built on a decision model like Jev need a slightly different kit, because there are more moving parts: the world, your questions, the hosted model, and whoever is probing your filter. The production guide lists the four signals worth watching in one paragraph. This page owns the depth: what each signal means, and how to tell which kind of drift you're actually looking at. The broader LLM monitoring picture, beyond judge drift, has its own guide.

Five kinds of drift, one symptom

When verdicts shift, one of these moved:

Input drift. The world changed. A product launch floods support with a new ticket type; a holiday shifts the email mix. Verdicts change because inputs changed, and the judge may be doing exactly the right thing.

Question drift. Someone reworded a question, tweaked a boundary clause, or added a label. Every rewording is a new instrument. If verdicts aren't tagged with question version, this one masquerades as everything else.

Model drift. Hosted models get updated. An independent behavior study in the directory names the specific version it tested, Jev 1.13.0, which is the useful reminder: the model behind your pipeline has versions, and they can change under you. Check docs.typesafe.ai for how versions are pinned or announced.

Adversarial drift. Someone is learning your judge. Spam is the canonical case: senders iterate until messages pass, so the caught rate falls while the true spam rate doesn't.

Upstream drift. The boring one, and the most common. A parser starts returning empty bodies, an integration truncates text, a language detector breaks. The judge dutifully rules on garbage.

The signals, ranked by what they tell you

Verdict distribution shift. The label mix per day against a trailing baseline. It's the first alarm and the least specific: it tells you something moved, not what. Compare windows (this week vs the last four), not single days, and alert on large relative changes per label.

Probability shape. Watch the share of verdicts in the uncertain band and the median top-choice probability, not just the winners. A pipeline whose labels hold steady while confidence quietly sags is drifting toward the edge of what its questions can handle.

Escalation rate, in both directions. The single most useful signal on this page. A rising escalation rate means questions and reality are diverging: new input types the judge can't place confidently. A rate falling toward zero is worse, because it usually means the gate broke, a threshold got edited, or inputs went empty and started producing confident nonsense. Alert on both edges.

Input statistics. Length, language, empty-field rate, source mix. Cheap to compute, and the fastest way to spot upstream drift before you blame the model.

Human-audit agreement. The only signal that measures correctness rather than change. A weekly sample of verdicts checked by a person, tracked as a percentage over time. Everything else tells you where to look; this tells you whether it matters.

The frozen canary set

Distribution signals can't separate input drift from model drift, because both change the numbers. The fix is a canary: a fixed set of a few hundred inputs, never updated, re-run on a schedule with pinned questions.

If canary verdicts move while nothing in your config changed, the model moved. If the canary holds steady while production shifts, your inputs moved. That one comparison settles most drift arguments in minutes. Since verdicts can vary slightly run to run, compare the canary's distribution and agreement rate against its own history, not individual verdicts.

The cost argument for skipping it doesn't survive the receipts: 500 emails were classified for 3.5 cents, as reported, so a daily canary of that size is pennies a month. Seed it from your gold set so it carries human labels, and it doubles as a daily accuracy check.

A diagnosis order that saves hours

When an alert fires, check in this order, cheapest first:

  1. Did any question, label, or threshold change? Diff the config. Question drift is self-inflicted and the fastest to confirm.
  2. Did input statistics change? Empty rates, lengths, sources. Upstream breakage shows here.
  3. Did the canary move? If yes with no config change, suspect a model update and check the provider's changelog.
  4. Is it concentrated in one source or segment? Adversarial drift clusters: one sender domain, one signup path.
  5. Is it just the world? Check with the humans who know the business before calling it a bug.

Then resist the urge to auto-correct. Silently re-tuning thresholds to restore last month's verdict mix erases real change along with the drift. Open an incident, decide deliberately, and add the cases that surfaced it to your eval set. Cost belongs on the same dashboard, and the cost monitoring guide covers it; a spend spike and a verdict shift on the same day are often the same bug.

Frequently asked questions

What is judge drift?

It's a shift in a verdict pipeline's outputs caused by changes in inputs, question wording, the hosted model, adversaries, or upstream data, often while every health check still passes.

What's the single best drift signal for a verdict pipeline?

Escalation rate, watched in both directions: rising means questions no longer fit reality, falling toward zero usually means the gate or the inputs broke. The production guide lists it among the four core signals.

How do I tell model drift from input drift?

Re-run a frozen canary set with pinned questions on a schedule. If the canary moves with no config change, suspect the model; if it holds while production shifts, your inputs changed.

How often should drift checks run?

Distribution and escalation metrics daily, the canary daily or weekly, and human-audit agreement weekly. Cheap signals often, expensive signals regularly.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.