In brief

Since 2023, the LLM-as-Judge paradigm has become widespread: an LLM scores the outputs of another LLM, or even its own. GPT-4 used as a judge achieves an agreement rate of over 80% with human preferences — a figure comparable to inter-human agreement, and high enough to justify replacing part of manual annotation. But the same article that establishes this result documents three systemic biases. The subsequent literature added a fourth, then showed that correction attempts remained fragile. This is not an alignment problem in the broad sense: it is a metrology problem. When the measurement instrument is made of the same material as what it measures, independence is structurally limited.

In short: think of a thermometer made from a sample of the air it has to measure. It produces a reading, but that reading reflects the properties of the material more than those of the measured object. The LLM-judge has exactly that defect: it evaluates a generated text with biases that are also its own — preference for length, for its own style, overconfidence. Metrology stops being objective.


Four documented biases

1. Position bias

The order in which responses are presented to the LLM judge influences the verdict. Wang et al. (2023) tested this phenomenon on 80 queries: by simply changing the presentation order, Vicuna-13B beats ChatGPT in 66 out of 80 cases when ChatGPT is the judge. The precise mechanism — primacy bias, recency bias, or attention artifact — remains debated. What is established: the ranking is not stable in the face of input permutation.

2. Verbosity bias

LLM judges systematically prefer longer responses, regardless of their actual quality. Dubois et al. (2024) quantified this bias on AlpacaEval: after controlling for length by regression, the correlation between LLM-judge scores and human preferences rises from 0.94 to 0.98. This bias has a direct practical implication: models trained on LLM-judge signals learn to produce long responses rather than useful ones.

3. Self-preference bias

A model used as a judge favors its own outputs. Panickssery et al. (2024) establish causality, not merely correlation: GPT-4 and Llama 2 can distinguish their own generations from those of other models with significant accuracy. And there is a linear correlation between self-recognition capacity and the strength of the preference bias. In other words, the models most capable of recognizing their own style are also the most biased in their judgment.

4. Systematic overconfidence

Prasad and Nguyen (2025) organized adversarial debates between LLMs and measured self-declared win probabilities. The observed average is 72.9% — against a rational baseline of 50% for a two-participant debate. In 61.7% of debates, both models simultaneously claim a win probability greater than or equal to 75%. Only one can be right.

In short: these four biases — position, verbosity, self-preference, overconfidence — are not marginal quirks but structural properties. They add up. An LLM judge without correction returns a score that looks like an objective grade and is in reality a cocktail of signals unrelated to actual quality.


Calibration: what models know about themselves

Kadavath et al. (2022, Anthropic) showed that large models are generally well calibrated on multiple-choice questions in adapted formats. When asked to estimate their own probability of knowing an answer — the P(IK) signal — they generalize partially across tasks. This calibration deteriorates out of distribution, but the baseline remains solid.

The problem arises when calibration meets iteration. Huang et al. (2025) show that iterative self-improvement — a model that evaluates itself and then retrains on its own judgments — amplifies overconfidence if no regulatory mechanism is introduced. The Expected Calibration Error increases with each cycle. This result has direct consequences for pipelines that rely on self-evaluation loops, notably RLHF. [NON VERIFIED — generalization to RLHF architectures other than those tested in the paper is not documented in the collected sources.]


Self-correction: partial independence

One approach to work around biased self-evaluation is to force the model to verify its own claims via separate questions, without access to the initial response. This is the principle of Chain-of-Verification (Dhuliawala et al., 2023, Meta): generate verification questions, answer them independently, then revise.

The limitation is structural. The model that verifies and the model that generated share the same parameters, and therefore the same encoded biases. The independence obtained is partial. [NON VERIFIED — this limitation is discussed in the literature but without direct quantification in the collected sources.]

Gulli (2025, Ch.4) formalizes this problem under the name cognitive self-revision bias and proposes the Producer-Critic pattern: two distinct agents, one generates, the other critiques with a fresh perspective dedicated to error detection. Separation of roles reduces the bias, but does not eliminate it if both agents derive from the same base model.


Debiasing: real gains, fragile robustness

Several approaches reduce measured biases. Wang et al. (2023) propose a three-step framework — multi-evidence calibration, balanced position calibration, human in the loop. Other work (FairJudge, RBD, AutoRubric) improves accuracy by 10 to 18% on average on static benchmarks.

But Ding et al. (2026) show that evaluation rubrics — often presented as the solution to bias — themselves constitute an attack surface. Subtle rubric modifications induce directional preference drift and reduce accuracy by up to 27.9%, while passing standard benchmark validation. Debiasing methods are robust to random perturbations, not to targeted adversarial perturbations.


The state of measurement in 2026

Zhou et al. (2026) introduced JudgeBiasBench, a benchmark covering 12 representative bias types. Their results show that current judges exhibit significant and diverse bias patterns. BiasScope (Lai et al., 2026) automatically detects biases at scale and reveals that even powerful LLMs used as evaluators display error rates above 50%. Fang et al. (2026) add an emergent bias: LLM judges increasingly prefer LLM-generated text over human text, regardless of actual quality.

Lau (2026) documents a reproducibility problem distinct from position biases: at temperature=0, different models produce substantially varied scores for the same inputs. This is no longer an intra-model bias problem — it is an inter-model variance problem.

The reference result of the field — over 80% agreement between GPT-4 and human preferences — is itself contested. Feng et al. (2025) show via the Sage benchmark that LLMs exhibit significant reliability problems in structured settings. The central question remains open: does high agreement with humans reflect high quality, or shared biases between humans and LLMs?


Key takeaways

  • LLM judges exhibit four documented biases: position, verbosity, self-preference, overconfidence. They accumulate and interact.
  • Model calibration on their own knowledge is generally correct on structured tasks, but degrades out of distribution and under iteration.
  • Self-correction shares the same parameters as the initial generation. The Producer-Critic pattern with two distinct agents reduces this problem without eliminating it.
  • Debiasing techniques improve scores on static benchmarks but remain vulnerable to targeted adversarial perturbations.
  • In 2026, error rates of the best LLM judges on specialized benchmarks still exceed 50%.