In short

When a LLM evaluates the quality of an output, it accepts more than three invalid outputs out of four as valid: the true negative rate (TNR) measured across 14 models is below 25% (Zhao et al., 2025). An automated review system built on GPT-4o reaches 69% balanced accuracy on scientific papers, dropping to 66% out-of-distribution (Lu et al., 2026). Entirely fabricated papers — with no real experiments — pass LLM filters in up to 82% of cases (Dang et al., 2025). These figures raise a precise methodological question: can the reliability of an LLM judge be quantified, and if so, how?

In short: think of an overly lenient grader. They accept 96% of correct papers (TPR=96%) — fine — but they also accept 75% of bad ones (TNR=25%). On a stack where most papers are correct, their pass rate looks high. But if you use this grader to filter fraudulent papers, three out of four slip through. That is exactly the LLM judges’ signature: structurally permissive, masked by an overall agreement that flatters.


The FPR of the automated judge

The LLM-as-judge paradigm — using a large language model to score or rank outputs in place of a human evaluator — has established itself since 2023 as a cost-effective alternative to human annotation. Zheng et al. (2023), introducing MT-Bench and Chatbot Arena, show that GPT-4 achieves over 80% agreement with human preferences. This figure is frequently cited as validation of the paradigm.

The problem is that an 80% global agreement says nothing about the distribution of errors. In binary classification, the relevant metric is the breakdown into true positive rate (TPR — valid outputs correctly accepted) and true negative rate (TNR — invalid outputs correctly rejected). This breakdown reveals a severe asymmetry.

Zhao et al. (2025) measure the full TPR/TNR profile across 14 LLMs evaluated on 366 Python programs. Result: TPR exceeds 96%, TNR stays below 25%. The direct translation is an implicit false positive rate (FPR) above 75% on invalid outputs. LLMs accept defective outputs as correct in more than three cases out of four.

This imbalance has a name in the literature: agreeableness bias. LLMs are structurally trained to produce cooperative, validating outputs. This behavior carries over to the judge role: rating an output negatively requires resisting an agreement pressure that the model has internalized during training. Both synthesis surveys (Gu et al., 2024; Li et al., 2024) converge on this point: the high TPR / low TNR imbalance is the dominant empirical pattern in the LLM-as-judge literature.

The practical consequence is immediate. An automated evaluation pipeline that uses a LLM as a quality filter does not reject bad outputs — it lets them through. The judge no longer acts as a filter; it acts as a validation buffer.


69% accuracy: the AI Scientist baseline (Nature 2026)

Lu et al. (2026) publish in Nature the results of Sakana AI’s AI Scientist system. This system generates research ideas, runs simulated experiments, writes papers, and performs automatic review via an Automated Reviewer. This reviewer is a large-scale LLM-as-judge case, with a rigorous protocol modeled on NeurIPS guidelines.

The reviewer ensembles five independent evaluations and acts as Area Chair. The model used is GPT-4o. The authors explicitly note that all other tested models exhibit a “positivity bias” or refuse to conform to the required output format — GPT-4o is the only usable model in this role.

Performance measured on OpenReview papers (2017–2024): 69% balanced accuracy (defined as the average of TPR and TNR). On a clean 2025 dataset — papers posterior to the model’s likely training window — balanced accuracy drops to 66%. The 3-point absolute loss out-of-distribution indicates partial dependence on superficial features of papers seen during training.

Balanced accuracy is the correct metric here, and the choice is not incidental. Dror et al. (2024) show that balanced accuracy is equivalent to the Youden J statistic (BA = (J+1)/2), and that alternative metrics are biased by class imbalance. A judge with TPR=96% and TNR=25% achieves a very high apparent accuracy when valid outputs dominate the dataset — but its balanced accuracy is only 60.5%. Balanced accuracy makes visible the FPR masked by standard metrics.

The AI Scientist’s 69% should be read in this context: it is a system optimized for this task, with a rigorous protocol, that still sits 31 accuracy points from perfect in balanced accuracy on ML conference papers.


Forced calibration multiplies discriminative power

An LLM judge produces noisy scores. Chen et al.’s (2026) “Noisy but Valid” framework formalizes this intuition: an imperfect judge can remain statistically useful if its error parameters (true TPR and FPR) are estimated empirically, then integrated into a corrected hypothesis test.

The operating principle is straightforward. A small human calibration set is assembled — between 50 and 200 manually annotated examples. This set allows estimation of the judge’s true TPR and FPR on the target domain. A variance-corrected decision threshold is then applied to a large dataset labeled by the judge alone. Chen et al. show that this procedure can increase statistical power beyond direct evaluation on the same human budget: noisy LLM judgments, once calibrated, carry more information than human labels alone on an identical budget.

Zhao et al.’s (2025) regression approach operates on the same principle, with a different implementation: post-hoc correction of judge scores calibrated on a small human set. RMSE (root mean square error) reduction reaches 72% in some configurations. The correction does not fix the model’s structural bias — it estimates it and compensates for it numerically.

Minority-veto is a third technique, simpler to implement. In an ensemble of LLM judges, a few “invalid” votes suffice to force the ensemble label to “invalid” — even if the majority votes “valid.” This asymmetric rule directly counterbalances agreeableness bias by penalizing false positives more heavily than false negatives. Zhao et al. (2025) and both synthesis surveys (Gu et al., 2024; Li et al., 2024) identify minority-veto and calibrated regression as the two most effective approaches for raising TNR.

Another lever is verbosity bias correction. Dubois et al. (2024) show that the Spearman correlation between raw AlpacaEval and Chatbot Arena is 0.93–0.94. After statistically controlling for response length, this correlation rises to 0.98. A single correction factor — length — adds 4 to 5 rank-correlation points. Length-controlled win rate became the default in AlpacaEval 2.0 precisely because verbosity bias was exploitable: a model could inflate its win rate without any quality improvement, simply by producing longer responses.

In short: correcting an LLM judge’s bias has an upfront cost (manually annotating 50-200 examples) that looks heavy, but it is what makes the judge usable at scale. Without this calibration, your metrics are decorative. With it, you turn a biased instrument into an approximately reliable measurement — not perfect, but discriminative.


The cost of permissiveness

Permissiveness in an LLM judge is not merely a statistical anomaly. It has concrete consequences when the judge operates inside an automated pipeline.

Dang et al. (2025) test this scenario directly. Their BadScientist system generates entirely fabricated research papers — with no real experiments — using five manipulation strategies named TooGoodGains, BaselineSelect, StatTheater, CoherencePolish, and ProofGap. These papers are submitted to multi-model LLM review systems (o3, o4-mini, GPT-4.1, calibrated against ICLR 2025 standards). Result: fabricated papers are accepted in up to 82% of cases.

The most striking detail is not the acceptance rate itself. It is that LLM reviewers frequently flag integrity problems in their textual evaluations — and then assign acceptance scores anyway. The authors call this pattern the “concern-acceptance conflict”: the model identifies the problem but does not draw the obvious conclusion. The authors’ conclusion is unambiguous: “LLMs are not ready for autonomous review.”

Li et al. (2025) document a complementary form of this problem in LLM-REVal. In a multi-turn simulation protocol, LLM reviewers systematically inflate scores for LLM-generated papers compared to human papers. The identified mechanism is a linguistic feature bias: the characteristic writing style of LLM-generated text is favored independently of scientific content. Conversely, human papers containing explicit language about risks or fairness are systematically underscored, even after revision.

These two studies converge on the same diagnosis: in a fully-LLM automated evaluation pipeline, LLM-generated outputs benefit from a structural advantage that the content does not justify. Permissiveness is not uniformly distributed — it is directional, and it favors precisely those outputs that stylistically resemble LLM text.

For practical applications — model selection, training data filtering, RAG system evaluation — this orientation implies systematic contamination: pipelines that use LLMs to both generate and evaluate their supervision data tend to over-represent LLM-style outputs, regardless of their actual quality. Both synthesis surveys (Gu et al., 2024; Li et al., 2024) call this phenomenon “preference leakage.”


How to measure LLM judge reliability

The literature converges on a set of concrete methodological recommendations.

Use balanced accuracy, not accuracy or F1. Dror et al. (2024) show that a judge with TPR=96% and TNR=25% achieves very high apparent accuracy when valid outputs dominate the dataset, but a balanced accuracy of only 60.5%. F1 is also biased by class imbalance and can select judges with a catastrophic FPR while appearing performant. Balanced accuracy — the average of TPR and TNR — is the only metric that penalizes both types of error equally.

Build a human calibration set. Chen et al. (2026) demonstrate that 50 to 200 manually annotated examples suffice to estimate a judge’s true TPR/FPR parameters on the target domain. Without this calibration set, the judge’s parameters are unknown and statistical corrections cannot be applied. The performance gap between perfectly known parameters and empirically estimated parameters — Chen et al.’s “Oracle Gap” — quantifies the potential gain from even minimal annotation effort.

Measure TNR independently of TPR. The common practice of reporting global agreement hides the TPR/TNR asymmetry. A judge that accepts everything has a TPR of 100% and majority-class accuracy — but a TNR of 0%. Any LLM judge evaluation protocol must include an explicit TNR measurement on a subset of confirmed invalid outputs.

Apply minority-veto or calibrated regression. When an ensemble of judges is used, minority-veto — a few negative votes suffice to force rejection — directly compensates for agreeableness bias. Post-hoc calibrated regression (Zhao et al., 2025) reduces RMSE by 72% in some configurations and requires only a small human calibration set.

Test out-of-distribution. The AI Scientist’s 69% balanced accuracy on 2017–2024 papers drops to 66% on 2025 papers. Lu et al. (2026) flag this training data contamination risk. Any LLM judge benchmark must include an evaluation on data posterior to the model’s likely training window.

Monitor linguistic feature bias. Li et al. (2025) document that LLM-style texts benefit from a systematic advantage independent of content. In a pipeline where evaluated data is itself generated by LLMs, this bias amplifies permissiveness rather than correcting it.


Which correction fits which context?

ContextRecommended correctionWhy
Human calibration set available (50-200 ex.)Calibrated regression (Zhao et al.)-72% RMSE, immediate gain against judge noise
Ensemble of several LLM judgesMinority-vetoAsymmetric in favour of rejection — offsets agreeableness directly
Free-length evaluationLength-controlled win rate (Dubois)+4-5 Spearman points, now the AlpacaEval 2.0 default
No human calibration availableRefuse LLM-as-a-Judge on that dimensionWithout a calibration set, TPR/FPR are unknown — no correction is possible

In plain terms: correcting an LLM judge’s bias costs an upfront effort — hand-annotating 50 to 200 examples — that looks heavy, but it is what makes the judge usable at scale. Without that calibration, your metrics are decorative. With it, you turn a biased instrument into an approximately reliable measurement — not perfect, but discriminating.


Key takeaways

  • Across 14 tested LLMs, TNR is below 25%: more than three invalid outputs out of four are accepted as valid (Zhao et al., 2025). Agreeableness bias is not a marginal phenomenon — it is the empirical norm.

  • Balanced accuracy is the only metric that makes this problem visible. An 80% global agreement with humans (Zheng, 2023) and a balanced accuracy of 60.5% (Dror et al., 2024, on the TPR=96%/TNR=25% profile) are compatible — the first metric hides what the second reveals.

  • An optimized automated review system — five ensembled LLM judges, NeurIPS protocol, GPT-4o model — caps out at 69% balanced accuracy (Lu et al., 2026). This is the current practical ceiling under the best documented conditions.

  • Entirely fabricated scientific papers pass LLM filters in up to 82% of cases (Dang et al., 2025). The “concern-acceptance conflict” — the model flags the problem but does not reject it — is the central mechanism.

  • Statistical correction is possible and effective: calibrated regression on a small human set (72% RMSE reduction, Zhao et al., 2025), minority-veto, verbosity bias correction (+4–5 Spearman correlation points, Dubois et al., 2024). These techniques do not fix the structural bias — they compensate for it numerically.