Blog / Recipes / FIG. 110
Question Versioning in Practice
Prompt versioning for judge questions: what counts as a breaking change, how to tag verdicts, and how to roll out a reworded question safely.
Prompt versioning sounds like bureaucracy until the first time a dashboard lies to you. For judge pipelines it's worse than for chat prompts, because a judge question is a measuring instrument: change the wording and every number it produces afterward is on a different scale. This page is the practice in depth: what counts as a version, what counts as a breaking change, and how to roll one out without forking your metrics. Examples use Jev, TypeSafe AI's decision model, whose closed-answer questions make versioning unusually clean.
Where this sits: Jev in production states the rule (questions are production config, version everything) in one section. Writing the questions well is how to write judge questions. This page is the version control in between, and checking whether a new version actually helped is prompt testing.
The metrics-fork story
Every team that skips versioning tells some version of the same story. What follows is a composite, not a specific build.
A support team runs an escalation question: "Does the ticket state that a production system is down?" Escalations sit around a steady rate for weeks. Someone notices tickets saying "we can't log in" aren't escalating, and helpfully adds a parenthetical: "(Login failures for all users count.)" Correct fix. Nobody tags it. Two weeks later the escalation rate is noticeably higher, a manager asks what broke, and the team spends a sprint checking upstream systems, the helpdesk integration, and the model, before someone finds the one-line diff. The world didn't change. The instrument did, and every chart silently spliced two instruments together.
The cure is boring: every verdict carries the version of the question that produced it, and no chart ever mixes versions without saying so.
What prompt versioning actually covers
Version the whole decision surface, not just the sentence:
- Question text, including boundary clauses and examples
- Choice set: the labels, their order, their wording, and whether an "unclear" option exists
- Context shape: which fields you send and how they're formatted
- Threshold and routing policy attached to the question
- Model identifier, as reported by the API
That last one gets forgotten. An independent behavior study examined framing sensitivity and failures on a specific Jev version, 1.13.0, per its entry. Whatever its findings, the fact that researchers pin to a version number is the lesson: model updates are changes to your instrument too, and your verdict log should record the model identifier the API reports.
Judge question version control: breaking vs non-breaking
Borrow semantic versioning, adapted for instruments.
Breaking (major): anything that can move verdicts on unchanged inputs. New or removed labels, reworded criteria, new boundary clauses, changed context fields, a new model. Treat it like an API contract change: new version, new baseline, no comparisons across the line without a bridge.
Non-breaking (minor): changes that shouldn't move verdicts but might, such as rewording a label's display name or reformatting context. Treat them as breaking until a re-run of your calibration set shows the distribution didn't move.
Metadata-only (patch): comments, owner, documentation. No re-run needed.
The honest default: if you're not sure, it's breaking. Wording effects in judge questions are routinely larger than people expect.
Rolling out a new version safely
- Diff it. Store questions as files in version control, with the version in the filename or header, and review changes like code.
- Calibrate it. Run the new version on the same labeled set as the old one, using the 100-case method, and record both agreement numbers.
- Shadow it. Run old and new side by side on live traffic for a while, acting only on the old. The disagreement rate between versions is your bridge: it tells you how much of any future metric shift is instrument change.
- Cut over, and annotate. Switch the acting version, and mark the cutover date on every dashboard that uses it.
- Update everything keyed on the question. Cache keys must include the question version, or a reworded question keeps serving old rulings. jevcache keys on model, schema, and state for the same reason, per its entry. Thresholds get re-fit. Eval baselines reset.
Automated question optimization makes this discipline more important, not less. Tools like jev-align optimize a question against human labels; every optimized output is a new version with a new error rate, and it deserves the same diff, calibration, and shadow run as a hand edit.
Where versions live in the log
Every verdict record carries its question version, choice-set version, threshold version, and model identifier. That's what turns "why did this metric move?" from archaeology into a query. The full record schema is in verdict logging and audit trails.
Frequently asked questions
What is prompt versioning for LLM judges?
Tracking every change to a judge's question text, choice set, context shape, thresholds, and model as a numbered version, and tagging each verdict with the version that produced it. It keeps metrics comparable across changes.
Is rewording a question a breaking change?
Assume yes until a calibration re-run shows verdicts didn't move. Small wording edits can shift decisions on unchanged inputs, which is exactly what forks metrics.
Should I version the model as well as the question?
Yes: record the model identifier the API reports with every verdict. Model updates can change verdicts just like wording edits do.
How do I compare metrics across question versions?
Run old and new versions in shadow on the same traffic and measure their disagreement rate. Use that bridge when interpreting trends, and never splice versions silently; Jev in production covers the surrounding ops.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.