In short
An AI system is not deterministic. The same input, at two different moments, can produce two distinct outputs. On top of this, there are model updates, pipeline evolutions, infrastructure changes. Without a fixed measurement instrument, it is impossible to know whether the system is improving or degrading. Regression testing for AI systems addresses this need: a frozen reference corpus, executed at regular intervals, compared against a baseline. Our tests show that a corpus of around fifty cases, broken down by difficulty, is sufficient to detect most drifts — and to reveal behavioral biases that global scores conceal.
Why measure systematically
Classical software is deterministic: same inputs, same output. A test that passes today passes tomorrow, unless the code changes. A bug, when it appears, is an obvious regression.
An AI system does not work that way. Three sources of variation coexist:
- Model stochasticity. Even at low temperature, an LLM can produce slightly different outputs on the same prompt. The effect becomes significant on edge cases. Lau (2026, arXiv:2603.04417) documents substantial variance between repeated calls at
temperature=0depending on the model. - Silent evolution. Model providers update their endpoints without always announcing the changes. The same model identifier may cover several internal versions over time.
- Pipeline drift. An AI agent is rarely a single call: it is a chain of instructions, tools, and prompts. A fix in one component can degrade behavior elsewhere, without the author of the fix noticing.
Without repeated measurement on a frozen set of inputs, you are flying blind. Regression testing is not a guarantee — it is the instrument that signals when an investigation is needed.
In short: a classical test tells you “it works or not”. An AI regression test tells you “where we stand compared to the last measurement”. The delta — gain or degradation — is more informative than the absolute score, because a raw score of 0.90 can be very good or catastrophic depending on the previous score (0.75 or 0.98).
Building a reference corpus
A useful reference corpus follows four principles.
Real cases, not synthetic ones. The corpus must reflect situations encountered in production. A general benchmark like MMLU is useful for ranking models against each other, but it says nothing about the reliability of a specialized agent facing its own inputs. Cases are extracted from logs, anonymized if necessary, and frozen.
Decomposition by difficulty. Mixing all cases into a single average hides what matters. A useful decomposition: easy (must always pass, clear regression), medium (normal working zone, where improvement happens), hard (known system limits, not always solvable).
Trap cases. A specific category deserves separate treatment: cases where the correct answer is to do nothing or refuse. An over-acting agent costs as much as an under-acting one. Without trap cases, the corpus only measures production — not restraint.
Minimum size. Below ~50 cases, scores are too noisy to be comparable. Our tests use a corpus of 50 cases spread across three distinct operations (18, 18, and 14 cases respectively). The distribution by level is at least 10 cases per level, with 11 explicit trap cases.
Once the corpus is built, it is frozen. Modifying a case after the fact invalidates the historical comparison. Additions go into a separately numbered version.
In short: decomposition by difficulty is what makes the test informative. A global score of 90% can mean “100% easy + 80% medium + 80% hard” (expected hard regression, not alarming) OR “70% easy + 100% medium + 100% hard” (critical regression on fundamentals). Without decomposition, both situations look like the same number.
Running the test
The execution itself requires a few precautions.
Mask expected answers before execution. If the ground truth is accessible to the system at prediction time — in the same file, in a cache, in a prompt — the observed scores no longer measure the ability to solve, but the ability to copy. Masking happens outside the system’s perimeter: expected answers are removed from the sent package, kept separately for final scoring.
Run cold. A system that retains a history of previous runs is not in the same conditions as a fresh system. The protocol: restart the system without preloaded context, load only the corpus inputs, capture predictions, then score.
Score automatically. Scoring must be deterministic and independent of the measured system. For classification, this is trivial. For generation, define measurable criteria (fidelity, presence of fields, format constraints) and write a scorer that evaluates them mechanically. LLM-based evaluation is possible but introduces its own biases — this pattern is described in the article on self-evaluation.
Three-phase protocol. In order: (1) preparation of the masked package, (2) cold prediction by the system, (3) scoring. Separation ensures that each phase is independently auditable.
Our tests on this protocol yield a global score of 0.938 on 50 cases, compared to an earlier run that peaked at 0.909 — a delta of +0.029 attributable to correcting a ground truth contamination in the previous version.
In short: ground truth masking is not an over-cautious precaution. Without it, the system can “see” the right answer in its context (cache, neighboring file, poorly isolated prompt) and copy it — you then measure copying, not solving. Any protocol without explicit masking is suspect by default.
What the test reveals
The global score is the least informative of the results. The real insights come from decomposition.
By operation. A global score of 0.94 can mask one operation at 1.00 and another at 0.88. Our tests observe exactly this pattern: routing at 1.000, insertion at 0.889, propagation at 0.911. The fragile operation — insertion — concentrates attention for subsequent revisions.
By difficulty. A system can reach 100% on easy cases and drop on hard ones. This is expected. What is alarming is the reverse: an average score on the easy cases (regression on what used to pass) while the hard ones hold. This pattern signals a recent and targeted degradation, more serious than a historical limit.
Trap cases. Our tests reach 100% on the 11 trap cases. A ratio below 90% on this category is a strong signal: the system is over-acting, producing when it should not, filling in when it should stay silent.
Frequent surprise: discovering behavioral biases. More than the score, the list of failed cases describes the behavior. Our tests revealed a specific tendency on borderline insertion and propagation cases: the system adds elements that were not requested — a classifier that proposes two labels where only one is expected, a propagation agent that touches a second uninvited scope. This over-inclusion bias is not visible in the global score and does not appear on easy cases. Without decomposition, it would remain invisible until a production failure.
Ribeiro et al. (2020, ACL) theorize this approach under the name behavioral testing: tests do not measure a score, they describe expected behavior across families of cases. Amershi et al. (2019, ICSE-SEIP) identify this practice as one of nine distinctive features of software engineering applied to machine learning.
Which test format for which need
| Context | Recommendation | Why |
|---|---|---|
| Production AI pipeline, frequent updates | Frozen 50-200 case corpus, auto scoring, weekly run | Quick detection of silent regressions, acceptable cost. Dominant pattern. |
| Frontier model in development (R&D) | Large corpus (1k+ cases) + public benchmarks (MMLU, GSM8K, etc.) | Exhaustive coverage before publication, even at extra cost. |
| Tasks without clear ground truth (creation, advisory) | LLM-as-judge evaluation + rubric | No single right answer, the score measures compliance with criteria. Judge bias risk. |
| Trap cases / refusals / safety | Dedicated 20-50 case corpus where the right action = doing nothing | Detects overconfidence, sycophancy, safety hallucinations. Often forgotten. |
| Fast regression test (smoke test) | 5-10 critical case corpus, run on each commit | Minimal safety net, < 30s per run, catches gross failures. |
Key takeaways
- An AI system without a reproducible reference corpus is unverifiable over time. Regression testing is not a guarantee — it is the instrument that signals when an investigation is needed.
- A useful corpus includes real cases, decomposed by difficulty, with an explicit category of cases where doing nothing is the correct answer.
- The protocol — ground truth masking, cold execution, independent scoring — conditions the validity of the result. A run without masking measures copying, not solving.
- Decomposing the score by operation and by difficulty reveals what the average conceals: fragile zones, recent regressions, behavioral tendencies.
- Discovering latent biases — over-inclusion, overconfidence, format dependency — is part of the protocol, not a failure. The test exists to expose them.