Blog / Recipes / FIG. 118
Building Gold Sets That Stay Gold
How to build a gold standard dataset for LLM judges: sourcing, blind labeling, adjudication, and the rules that keep eval ground truth from rotting.
Every eval number you report is a comparison against something. That something is your gold set: the inputs where you know the right answer, because a person decided it and wrote down why. Get it wrong and every downstream metric is precise, reproducible, and meaningless.
A gold standard dataset is the least glamorous asset in an AI pipeline and the one that decides whether the others can be trusted. The evals guide owns the suite around it: assertions, judge verdicts, spot checks, cadence. This page owns the ground truth itself: how to build it, and how to stop it rotting.
What makes a label gold
A gold label isn't just a correct answer. It's a correct answer with provenance. At minimum, each case needs:
- The input, frozen exactly as the judge will see it.
- The answer, from your closed choice set.
- Who decided, and how: two independent labelers, or one expert, plus an adjudicator for disagreements.
- A one-line reason, especially for boundary cases. Future you will not remember why a sarcastic refund request counted as a complaint.
- The question version the label answers. Reword the question and the label may no longer apply.
Pseudocode for the record shape, not any real schema:
// pseudocode: one gold case
case_id, input_text, question_id, question_version,
gold_answer, labeler_a, labeler_b, adjudicated_by,
reason, source_segment, added_on, status (active | superseded)
Sourcing: stratify, don't just sample
A random sample of production traffic gives you mostly easy cases, which makes every judge look brilliant. Build the set in layers:
- The boring middle, sampled at random, so you don't trade median quality for edge cases.
- Boundary cases, pulled deliberately from low-confidence verdicts, where the question's edges actually get tested.
- Adversarial and bleed probes: text that mentions a category without being it, and instances that never name their category.
- Every bug you ever shipped. This is the regression-case rule, and it's non-negotiable: when a verdict goes wrong in production and someone notices, that input becomes a permanent gold case with the corrected label. Never delete it.
- Segment coverage: every language, source, and customer tier you serve, or your accuracy number is really a number about your biggest segment.
The evals guide puts the working floor at a few hundred cases. Size per question, not per suite: 300 cases split across ten questions is 30 per question, which is vibes with decimal points.
Labeling: blind, doubled, adjudicated
Label blind. The data labeling guide describes model-first labeling, where a model drafts and humans audit, and it's the right workflow for bulk datasets. Gold is the exception. Humans who see the model's answer first tend to agree with it, and then your gold set measures the model against itself. Labelers decide first; model verdicts get compared afterwards.
Label twice. Two people, independently, on every case you can afford. Their disagreement rate is the most honest number in your pipeline: if two careful humans agree only 80 percent of the time on a question, no judge will reliably beat 80 percent, and the fix is rewording the question, not swapping the model.
Adjudicate and record. A third person resolves disagreements and writes the reason. Disputed cases are your best boundary-clause material; fold the resolutions back into the question wording.
Keep models out of the answer key. A model-written key rewards whichever judge agrees with that model. The self-preference deep dive covers why that corrupts comparisons in ways that look like real results.
How gold sets rot, and how to stop it
Policy changes. The business redefines "urgent", and last quarter's labels are now wrong. When a definition changes, bump the question version, re-adjudicate affected cases, and mark old labels superseded rather than overwriting them.
Overfitting. Tune questions against the same cases long enough and they'll ace the set while missing new inputs. Keep a sealed holdout slice, 20 percent or so, that nobody looks at while iterating. Open it only to confirm a change.
Staleness. Inputs from a year ago may not look like today's traffic. Add fresh production samples every month, and track accuracy on new cases separately from the long-standing ones.
Silent deletion. Someone removes "confusing" cases to tidy up, and the hardest tests disappear. Cases get superseded with a reason, never deleted.
A useful model for the discipline comes from the jev-evaluation repository: the builder wrote down 28 predictions before collecting any data, then ran nine adversarial experiments against them, 123,805 requests for $12.69, as reported, publishing what held and what didn't. Fix the answer before you look at the results. That's the whole gold-set ethic in one habit.
What it costs, honestly
The judge runs are close to free: 3,282 posts at eight questions each cost $0.1282, as reported, so re-running a few hundred gold cases on every change costs effectively nothing. The expensive part is human time, and it should be. A few hundred carefully adjudicated cases per question is a few days of focused work, and it's the only part of the pipeline a model can't do for you. For Jev pipelines, gold is where the humans earn their place.
Frequently asked questions
What is a gold standard dataset?
It's a set of inputs with human-decided correct answers, recorded with who decided, why, and against which question version. Every accuracy number in an eval suite is measured against it.
How big should an eval gold set be?
A few hundred cases per question as a working floor, stratified across boring, boundary, adversarial, and regression cases. The evals guide covers how the set fits into the wider suite.
Can I use model-generated labels as ground truth?
Not for gold. Model drafts are fine for bulk training data with human audit, but gold labels should be decided blind by humans, or the set ends up measuring agreement with the drafting model.
What is the regression-case rule?
Every production mistake someone catches becomes a permanent gold case with the corrected label, and is never deleted. Over time it turns your worst days into your best tests.
How do I keep a gold set from going stale?
Version questions and supersede outdated labels instead of overwriting them, add fresh production samples monthly, and keep a sealed holdout slice to catch overfitting.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.