Blog / Evals & judging / FIG. 165
Prompt Testing: Versioned, Measured, Boring
Prompt testing measures whether a prompt change helps or hurts before it ships: gold sets, versioning, side-by-side runs, and gates for regressions.
Prompt testing is running a prompt change against a fixed set of real cases before it ships, so you know whether it helped, hurt, or just moved the errors around. The goal is to make prompt changes boring: versioned, measured, and reversible.
For judge pipelines, the "prompt" is the question wording, and the stakes are higher. A reworded question silently changes every verdict downstream, including the ones that decide what a human ever sees.
Why prompt changes need tests
Small wording changes shift outputs in ways nobody predicts. The fix for one bad case quietly breaks three good ones. And prompts that worked last month can behave differently after a model update, with no edit on your side.
Eyeballing five outputs after a change catches none of this. A fixed test set catches most of it.
The prompt testing loop
- Version every prompt: an ID, the exact text, and the date. Question versioning has the scheme.
- Keep a gold set of 100 to 200 real cases with expected outcomes (gold set method).
- Run old and new side by side on the same set.
- Compare: agreement with expected outcomes, per-case differences, and shifts in the answer distribution.
- Gate: ship only if nothing critical regressed past your threshold.
- Log the version with every production output, so later changes in behavior can be traced to a specific edit.
What to measure
For closed outputs, measure agreement with expected labels, per-class precision and recall, and shifts in confidence. For generated outputs, score properties with closed-question judges. For judge questions themselves, measure agreement with human labels before and after the change, per the calibration guide.
Read every case that flipped. Aggregate scores hide the story; flipped cases tell it.
Common prompt testing mistakes
- Testing on the examples you wrote the prompt from. They'll pass. That's overfitting, not evidence.
- Changing the prompt and the model at once. One variable per change, or you can't tell which one helped.
- Skipping re-tests after model updates. The prompt didn't change, but its behavior did.
- Deleting failing cases instead of fixing the prompt that fails them.
Tools for prompt testing
Test runners like promptfoo and DeepEval automate the loop and plug into CI; the evaluation framework guide compares them by job. A script, a spreadsheet, and a set of judge questions works fine for small suites. LLM testing covers the wider test stack prompt tests sit inside.
Frequently asked questions
What is prompt testing?
Running prompt changes against a fixed set of real cases before shipping, to catch regressions and confirm improvements.
How many cases do I need to test a prompt change?
100 to 200 real cases catch most regressions. Keep a separate must-pass list for critical edge cases.
Should I A/B test prompts in production?
Only after offline testing passes. A small live split can confirm the result; it shouldn't be the first test.
How do I test prompts for an LLM judge?
Measure agreement with human labels before and after the wording change, and version both texts. Tips on wording in what Jev is.
What's the most important prompt testing habit?
Change one thing at a time, and log the version with every output.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.