In short

When a lab announces that its model “beats the previous one on MMLU,” what should we make of it? Benchmarks are the measuring instruments of AI, but like any instrument, they have blind spots. Data contamination, score saturation, biases in automated evaluators: LLM evaluation research is traversed by fundamental debates. Understanding these limitations means understanding what the rankings actually measure — and what they conceal.

In short: an LLM benchmark is the equivalent of a standardized exam in school — a useful device, but one that only measures what it can phrase as multiple choice. When all students score 90%, the exam no longer ranks. When the topic leaked before the test, the grade no longer measures understanding. LLM benchmarks in 2026 accumulate both pathologies — hence the proliferation of new instruments (GPQA, Humanity’s Last Exam) and the importance of reading scores cautiously.


The foundational benchmarks

The basic idea is straightforward: submit a model to a set of questions with known answers, count the correct responses, compare. The first major benchmarks all appeared between 2018 and 2022.

MMLU (Massive Multitask Language Understanding, Hendrycks et al., ICLR 2021) is the most celebrated. It covers 57 disciplines — mathematics, law, medicine, history, computer science — in multiple-choice format, from middle-school level to professional expertise. The intuition: a model that excels across 57 domains possesses broad understanding. In 2021, GPT-3 barely scored 20 points above chance. Three years later, frontier models all exceeded 88 to 90%.

GSM8K (Cobbe et al., OpenAI, 2021) contains 8,500 math problems at the elementary-to-middle-school level, each requiring 2 to 8 reasoning steps. Initial goal: test whether models can chain simple inferences. In 2021, the best models failed systematically. By 2024, frontier models exceeded 90 to 95%.

HumanEval (Chen et al., OpenAI, 2021) evaluates code generation: 164 programming problems with unit tests. The criterion is binary — the code runs and passes the tests, or it doesn’t. Codex solved 28.8% of problems on the first attempt. GPT-4 now exceeds 85%.

TruthfulQA (Lin, Hilton, Evans, ACL 2022) takes a different approach: 817 questions designed to elicit errors typical of humans — false beliefs, cultural myths, popular but inaccurate claims. A counterintuitive result at its publication: larger models were initially less truthful, with GPT-3 reaching only 58% versus 94% for human experts.

BIG-Bench (Srivastava et al., 2022) is a collective construction: 450 authors, 132 institutions, 204 tasks spanning linguistics, mathematics, biology, physics, and social biases. It includes a “Hard” subset (BBH) identifying the 23 tasks most resistant to models.

In short: MMLU tests general culture in MCQ, GSM8K school arithmetic, HumanEval code with unit tests. Three different instruments, three angles. A model that succeeds on one can fail on another — the MMLU score is not a universal IQ, it is a measure of aptitude for a specific format.


From automated scoring to human judgment

Automated benchmarks have a structural limitation: they compare outputs against predefined reference answers. For open-ended tasks — writing, explaining, conversing — this approach reaches its limits. Two alternative methods have emerged.

MT-Bench (Zheng et al., NeurIPS 2023) proposes 80 multi-turn questions covering 8 categories: writing, role-play, extraction, reasoning, mathematics, coding, science. The main innovation is the concept of LLM-as-a-Judge: using a powerful LLM — GPT-4 — as an automated evaluator in place of human judges. Observed agreement with expert human annotations exceeds 80%, which made the approach popular. The same foundational paper identifies several structural biases that we will return to.

Chatbot Arena (Chiang et al., ICML 2024) opts for large-scale human evaluation. Real users pose questions to two anonymous models and vote for the best one. The ranking uses the Bradley-Terry statistical model — the same used for chess and tennis rankings. By early 2024, the platform had collected more than 240,000 votes from 90,000 users across more than 100 languages. It has become the most-cited leaderboard by labs when justifying their announcements.


Three problems that invalidate scores

Contamination

The problem is structural: test questions end up in training data, intentionally or not. On the APPS programming benchmark, StarCoder-7B achieves a score 4.9 times higher on leaked examples than on others. For GSM8K, studies have measured accuracy drops of up to 13% after removing contaminated examples. What the score measures in such cases is no longer reasoning — it is memorization. Two surveys document the scale of the phenomenon (arXiv:2406.04244, 2024; arXiv:2502.17521, 2025).

Saturation

When all frontier models exceed 90% on a benchmark, it no longer distinguishes anything. MMLU is saturated above 88%. GSM8K is saturated beyond 95%. The benchmark stops being informative precisely when it is needed most — to compare the best models against each other.

The research community’s response has been to raise the bar: GPQA (doctoral-level questions in chemistry, biology, physics), LiveCodeBench (problems published after the training cutoff date), and Humanity’s Last Exam (2025), explicitly designed to resist current models with 2,500 expert questions. SWE-Bench, for its part, evaluates agents capable of fixing real bugs in real code repositories.

Construct validity

The deepest problem is the least visible: does a benchmark measure what it claims to measure?

Apple provided a striking demonstration with GSM-Symbolic (Mirzadeh et al., ICLR 2025). The researchers took GSM8K problems and simply replaced the numerical values with different ones, or added an irrelevant clause to the problem statement. Models with excellent GSM8K scores collapse on these variants. Conclusion: these models are not solving math problems — they are recognizing patterns seen during training.

More broadly, a 2025 study (Frontiers in Artificial Intelligence) analyzed 445 NLP benchmarks: only 16% use rigorous scientific methods. About half claim to measure abstract concepts like “reasoning” or “usefulness” without precisely defining those terms.

In short: out of 445 published benchmarks, about 70 (16%) are solidly built. The others measure something, but exactly what remains unclear. When a paper announces “our model improves reasoning by 12 points,” the first question is not “how did it do it?” but “does the benchmark used really measure reasoning?”.


The biases of LLM-as-a-Judge

Automated evaluation by LLM is convenient, but the founding MT-Bench paper itself identifies several biases:

  • Position bias: the judge model favors the response placed first in the comparison.
  • Verbosity bias: longer responses are preferred, even if they are redundant.
  • Self-enhancement bias: an LLM favors its own outputs when evaluating them.

Subsequent work (Koo et al., arXiv:2410.02736, 2024) catalogued up to 12 distinct types of bias. Swapping the presentation order of two responses can reverse the verdict in more than 10% of cases for coding tasks.

Chatbot Arena is also not immune to criticism: voters are not representative of professional users, long and confident responses are preferred regardless of their accuracy, and factual verification plays no role in voting criteria.


Decision matrix — which benchmark for which need

ContextRecommendationWhy
Model selection for specific use casePrivate benchmark built on real use casesEliminates contamination risk + measures what truly matters.
Quick comparison of frontier modelsHumanity’s Last Exam, GPQA, LiveCodeBenchUnsaturated, designed to resist 2025-2026 models.
Open-task evaluation (writing, reasoning)LLM-as-a-Judge with order rotation80%+ agreement with expert humans if rotation neutralizes position bias.
Real user preference evaluationChatbot Arena (critical reading)Massive sample (240K+ votes) but documented verbosity bias.
Regression detection after fine-tuningMMLU-Pro, GSM-SymbolicUnsaturated variants that still detect measurable gaps.
Research paper, claim “our model reasons better”GSM-Symbolic + adversarial variantsDistinguishes memorization from true reasoning (Apple ICLR 2025).

Key takeaways

  • Historical benchmarks (MMLU, GSM8K) are now saturated for frontier models: all exceed 90%, they no longer serve to distinguish the best.
  • Data contamination turns a reasoning evaluation into a memorization evaluation. Scores observed on contaminated examples can be several times higher than real scores.
  • GSM-Symbolic (Apple, ICLR 2025) showed that excellent math scores can collapse as soon as the problem statement is slightly altered — a sign that models follow patterns, not general reasoning.
  • LLM-based evaluation (LLM-as-a-Judge) is practical but biased: position, verbosity, and self-enhancement can reverse verdicts.
  • Chatbot Arena brings a large-scale human dimension, but its voters are not representative and do not evaluate factual accuracy.