shipwithjev

Blog / Recipes / FIG. 85

An Eval Suite in One Day, Honestly

Build an LLM eval suite in one day: 50 eval cases, a judge grading setup that's been checked against you, and a run cheap enough for every commit.

Most teams don't skip evals because they disagree with them. They skip them because "build an LLM eval suite" sounds like a quarter-long platform project, and there's a feature due Friday. It isn't. A useful suite fits in one working day if you accept that version one will be small, a little ugly, and far better than vibes.

The theory of what evals are and why they layer lives on the LLM evals page. This page is a schedule. The judge in it is Jev, the decision model from TypeSafe AI (what that means), because grading is a closed-answer job and closed answers are what it does.

Morning: 50 real cases, written as questions

Not synthetic. Pull 50 real inputs from logs, support threads, or whatever your feature actually receives. Aim for a spread:

  • about 30 ordinary cases,
  • about 15 that previously went wrong or drew a complaint,
  • about 5 adversarial ones: weird formatting, prompt-injection attempts, empty input.

Run your current system on all 50 and save the outputs. That snapshot is your baseline, and the whole day depends on it.

Fifty feels small. It is. It's also enough to catch regressions that matter.

Then, in hour three, every eval case becomes a few closed questions about the output. Not "rate this 1-10," which drifts, but checks a stranger could answer:

  • Does the answer address the question actually asked? (yes/no)
  • Does it state anything not supported by the provided source? (yes/no)
  • Does it follow the required format? (yes/no)
  • Is the tone appropriate for a customer: yes, borderline, no? (choice)

Keep each question about one thing. The judge-questions guide is the reference for wording, and it will save you from the classic mistake of asking two things in one question.

Midday, hour 4-5: grade by hand first

This is the step people skip, and it's the step that makes the suite honest. Before any model grades anything, you answer every question for all 50 outputs yourself. It's tedious. Do it anyway.

Your hand grades are the answer key. Without them, the judge grading setup has nothing to be checked against, and you're trusting a model to grade a model with no referee.

Afternoon, hour 6: point the judge at it

Now run Jev on the same questions over the same 50 outputs. The block below is pseudocode; per ecosystem documentation, Jev is reached through the Vercel AI Gateway as typesafe-ai/jev, and docs.typesafe.ai has the real client syntax.

# pseudocode, not real API syntax
for case in cases:
    verdict = judge(
        text = case.input + case.output,
        questions = EVAL_QUESTIONS
    )
    compare(verdict.answers, case.hand_grades)
report agreement per question

Then look at agreement question by question, not as one average. A question where the judge agrees with you 48 times out of 50 is ready. A question at 35 of 50 is badly worded or genuinely subjective; rewrite it or keep it human-graded.

Cost won't be the constraint. Ian Nuttall ran 3,282 posts with eight questions each, roughly 26,000 verdicts, for $0.1282 in 8 minutes 34 seconds, as reported (build). Your 50 cases times four questions is a rounding error by comparison.

Afternoon, hour 7: plant defects

A suite that always passes proves nothing. Take ten good outputs and break them on purpose: add an unsupported claim, drop a required field, make the tone rude. Does the judge catch them?

This is the best cheap trick in evals, and there's a public receipt. Mike Taylor ran Jev over 37 documents with 21 questions each, 1,709 judgments for under a cent, and reports it found six of the seven defects he had planted; only a larger model found the seventh (build). Note what that result says: good, not perfect, and the miss is information. Whatever your planted defects slip past, route those checks to a stronger model or a human.

Last hour: wire the LLM eval suite to every change

Put the run in CI or a script you run before every prompt or model change. Output: per-question pass rate compared to the baseline, plus the list of cases that flipped. Flipped cases are what you read; averages are what you report. The wider set of test layers this suite slots into is covered in LLM testing.

Write down three things in the repo: which questions are judge-graded, which are human-graded, and the date you last checked the judge against your hand grades. When you change a question, re-check it. That's the entire maintenance plan for version one.

By 6 p.m. you have fifty real cases, a question set with measured agreement, a list of checks the judge can't do, and a run that costs pennies. No official Jev benchmarks exist, so this suite is also your evidence about how well it judges your domain. It isn't finished. It's started, which is the part that usually never happens.

Frequently asked questions

How many eval cases do I need to start?

Fifty real cases is enough to catch meaningful regressions on day one. Grow the set by adding every new failure you see in production.

Can I trust an LLM judge for grading?

Only per question, and only after checking it against your own hand grades. Questions with low agreement stay human-graded; the LLM evals page covers the escalation pattern.

What makes good eval cases?

Real inputs with a mix of ordinary, previously failed, and adversarial examples. Each case is graded by several closed questions that each check one thing.

Why plant defects?

Planted defects show what the judge misses before production does. Any check that lets them slip should go to a stronger model or a human.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.