Blog / Evals & judging / FIG. 164
LLM Testing: A Practical Guide
LLM testing checks that model-backed features behave before and after every change. Four test layers, writing cases that matter, and CI wiring.
LLM testing is how you check that a model-backed feature does what it should, before release and after every change to the prompt, the model, or the data. It borrows from software testing but can't copy it: outputs vary between runs, correctness is often a judgment call, and a test that passes today can fail after a model update nobody announced.
Why LLM testing is different
Three things break ordinary test habits:
- Nondeterminism. The same input can produce different outputs, and even closed-answer models wobble on borderline cases (is Jev deterministic?).
- No single correct string. Two good answers can share almost no words.
- Change without code change. Provider-side updates shift behavior under you.
The response is to test properties instead of exact strings, run borderline cases more than once, and score judgment calls with closed questions.
The four test layers
Assertions. Deterministic checks: valid JSON, required fields present, length limits, no banned content. Cheap, exact, run on everything.
Behavioral regression tests. A gold set of real cases, each paired with the property that must hold, scored by a judge. Run on every prompt or model change. The eval-suite recipe builds one in a day.
Adversarial tests. Injection attempts, instructions hidden in user content, jailbreak patterns. Instruction-bleed defenses lists the cases worth keeping.
Production tests. Canary sets re-judged on a schedule and sampled judging of live traffic, because the pre-release suite can't see tomorrow's inputs.
Writing test cases that matter
Source cases from real traffic and past failures, not from the examples you used to write the prompt; those will always pass. Each case needs three parts: the input, the property that must hold, and the closed question that checks it. The gold set method covers sampling and labeling.
Keep the judge questions versioned. A reworded judge changes every score, which is its own kind of test failure (prompt testing covers the loop).
What LLM testing costs
The judging is the cheap part. Cataloged builds run thousands of verdicts for cents, as reported in the X-post analysis build, so a 200-case suite with a few questions per case costs pocket change per run. The expensive part is a person labeling the gold set once, and keeping it current.
Wiring LLM tests into CI
Run assertions and the regression suite on every prompt or model change. Fail the build when agreement drops past a set threshold on critical cases. Quarantine flaky cases instead of deleting them, and review them weekly, the way QA teams already handle flaky tests.
Frequently asked questions
How do you test an LLM?
Use deterministic assertions where possible and closed-question judges elsewhere, against a gold set of real cases, on every change.
Can LLM tests be deterministic?
Assertions can. Judged tests vary slightly, so use thresholds and repeat borderline cases rather than trusting one run.
How many test cases do I need?
Start with 100 to 200 real ones and add every production failure you find.
What's the difference between LLM testing and LLM evaluation?
Testing gates releases with pass or fail thresholds; evaluation measures quality more broadly. The suites overlap heavily.
Should the same model judge its own outputs?
Avoid it. A different model family reduces self-preference bias. See what Jev is for one cross-family option.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.