shipwithjev

Blog / Calibration / FIG. 115

Self-Preference Bias: The Deep Dive

LLM self-preference bias, in depth: why judges favor their own family's outputs, where it hides in evals, and how cross-family judging fixes it.

Ask a model to grade its own homework and it tends to find the homework excellent. That's LLM self-preference bias: a judge rating outputs from itself, or from its model family, more favorably than equivalent outputs from elsewhere. The LLM-as-a-judge guide lists it alongside fluency, position, and verbosity bias in a single line. This page is the long version, because self-preference is the one bias that quietly corrupts the numbers you use to decide which model to ship.

Two mechanisms, one symptom

"The model likes itself" is a cute summary and an incomplete one. Two different things produce the same inflated score.

Familiarity. A model's own phrasing, structure, and habits look normal to it. Outputs that match its style read as fluent and correct; outputs from a different family read as slightly off. The judge isn't flattering itself on purpose. It's mistaking "sounds like me" for "sounds right".

Shared blindness. The more dangerous one. A model can't reliably catch errors it would make itself. If the generator hallucinated a plausible fact the judge also believes, the judge passes it. Same training lineage, same gaps, same confident wrongness. This is correlated error, and it's why self-judging suites can report steady improvement while real quality stalls: the judge and the generator are failing together, and the grade stays green.

Familiarity inflates scores. Shared blindness hides failures. Only the second one can put a bug in production with a passing eval next to it.

Where it hides

Your eval suite. The obvious place. If your product runs on one chat family and your eval suite grades with the same family, every regression the two share is invisible.

Model selection bake-offs. Comparing your current model against a competitor, graded by a judge from your current model's family, is a rigged race. The incumbent wins ties it shouldn't.

Rewrite-until-approved loops. A generator iterating until a same-family judge approves is optimizing for that judge's taste. The loop converges on "what this family likes", which is not the same as "what users need".

Answer keys and gold sets. The subtle one. If a model writes the reference answers, every judge that agrees with that model looks accurate. A dev.to write-up on a reported one-point win for Jev over GPT-5.6 Luna (67.8 percent to 66.8 percent, as reported in TypeSafe's own evals per the article) makes exactly this point in its headline: other frontier models wrote the answer key. Whoever writes the key has a vote in every score. The fix lives in building gold sets with human adjudication.

Training data. Label with one family, train on the labels, evaluate with the same family, and you've laundered one model's preferences through three stages. The data labeling guide covers the hygiene rule.

Self-preference also has cousins: judges can prefer outputs by origin in general, not just their own. One builder used Jev to audit model recommendations and reported that US models were placed first 91.5 percent of the time, even where benchmarks favored others (as reported; we haven't reproduced it). Judges have opinions about who made things. Assume yours does too until you've measured.

The cross-family rule

The fix is simple to state: the judge should come from a different model family than anything it grades. Different vendor, different training lineage, different failure modes. Errors then stop being correlated, so a mistake the generator makes has a real chance of being caught.

Decision models add an interesting property here. Jev, TypeSafe AI's decision model, doesn't generate prose at all, so when it judges chat-model outputs it's never grading its own writing style; there's no house style for it to prefer. That removes the familiarity mechanism by construction. It doesn't prove the absence of shared blindness, though: TypeSafe hasn't published training lineage in detail, and "architecturally different" is a hypothesis to test, not a certificate. Consistency is a separate question from bias, but it helps: LangChain reported Jev's quality-score variance at 92 to 913 times lower than GPT-5.6 Luna, Terra, and Claude Sonnet 4.6 on continuous scoring, as reported. A steady instrument makes bias easier to measure, because the noise isn't drowning it.

Measuring it on your own stack

You don't need a paper to check this. Three cheap tests:

  1. The swap test. Take outputs from two model families on the same inputs. Grade both with each candidate judge. If a judge's preference flips depending on which family it belongs to, you've found self-preference.
  2. The human anchor. On a human-labeled sample, compute each judge's agreement with humans split by which family generated the output. A judge that agrees with humans 90 percent on its own family and 78 percent on others is telling you something.
  3. The disagreement matrix. Run two or three cross-family judges on the same batch and look at where they disagree. Disagreements cluster around the cases where one judge's blind spots live.

At builder-reported verdict prices, running two judges on everything costs less than one meeting about whether to. Keep the human anchor permanent; it's the only test that doesn't depend on any model's opinion.

Frequently asked questions

What is self-preference bias in LLMs?

It's the tendency of a model acting as judge to rate outputs from itself or its model family more favorably than comparable outputs from other families. It inflates scores and, worse, hides errors the judge and generator share.

How do I avoid self-preference bias in evals?

Judge with a different model family than the one being graded, anchor on a human-labeled sample, and never let a same-family model write your answer keys. The evals guide covers the suite around those rules.

Is a decision model immune to self-preference?

Not proven immune. A model that doesn't generate prose can't prefer its own writing style, but shared blind spots are still possible, so measure with the swap test before trusting any judge.

Does self-preference matter for classification, not just grading?

Yes, whenever the thing being classified was produced by a model. Labeling model outputs with the same family that produced them carries the same correlated-error risk.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.