In short

An agent marks a task as done. The result is technically compliant. Yet the task has not really been accomplished. This scenario is not an edge case — recent research shows that it accounts for between 27% and 78% of reported successes depending on the benchmark (Cao, Driouich, Thomas, 2026). The problem is not the agent that fails: it is the agent that believes it has succeeded, and the system that confirms it.

In short: imagine a delivery driver who drops off a parcel at the wrong address but marks “delivered” in their app. The system records a success. The recipient received nothing. For AI agents, between a quarter and three quarters of reported successes are of the same kind — formally closed, functionally false.


Binary success does not measure what we think

Agent evaluation has relied from the start on a binary criterion: the task is completed or not. This criterion is simple to measure and easy to automate. It is also structurally insufficient.

Cao, Driouich, and Thomas (2026) introduced the Procedure-Aware Evaluation (PAE) framework to audit what benchmarks classify as “success”. Their verdict: between 27% and 78% of tasks marked successful are corrupt successes — the task is formally completed, but the agent violated procedural or integrity constraints along the way. PAE evaluates along four dimensions: utility, efficiency, interaction quality, procedural integrity. Results vary strongly across models: some spread violations across several dimensions, others concentrate them on one (e.g. policy fidelity or instruction fidelity). [NOT VERIFIED for the exact proportions per model]

The illusion is amplified by the benchmarks themselves. Xue et al. (2025, COLM 2025) showed on Online-Mind2Web (300 realistic tasks) that an agent can post solid lab performance while systematically failing on equivalent tasks in production. WebArena (Zhou et al., 2023) quantifies the gap: the best GPT-4 agent reaches 14.41% success, versus 78.24% for humans — and despite this, the “success” criterion is still defined by the benchmark, not by the reality of the task.


Process matters as much as outcome

If the final outcome is not enough as a criterion, what should be measured?

Fan et al. (2026, AgentProcessBench) built the first benchmark dedicated to step-by-step trajectory quality, over 1,000 trajectories and 8,509 human annotations (inter-annotator agreement: 89.1%). Their central diagnosis: current evaluations capture the final outcome but ignore the quality of the path taken. An agent that stops too early can post an excellent per-step success rate — simply because it produces few faulty steps. It fails globally, but its local metrics are flattering.

Shahnovsky and Dror (2026) illustrate the same problem from another angle: the same agent reaches 38% overall success under one metric, and 89% per-element precision under another. The number depends on the chosen definition, not on the reality of the run.

Gulli (2025, Ch.11) frames this principle from the design side: an agent without continuous monitoring capability is confined to reactive single-step tasks. The feedback loop — comparing current state to target state against observable metrics — is the necessary condition for detecting drift before it turns into invisible failure. Without this loop, the agent does not know where it stands.


Who verifies? The agent, an external judge, or no one

Three approaches exist for checking whether a task has really been accomplished. None of them is general.

Self-verification — SmartSnap (Cai et al., 2025) proposes a shift from post-hoc checking to in-situ self-verification: the agent proactively selects a minimal set of decisive pieces of evidence during execution, rather than retrospectively analyzing the full history. Measured gains: up to 26% for 8B models. The limit is structural: an agent that is wrong about the task can be wrong with the same consistency about its own success.

External verification — TrustBench (Sharma et al., 2026, AAAI 2026) places the check at the critical moment: after an agent formulates an action, before executing it. This placement reduces harmful actions by 87% with latency below 200 ms. The LLM-as-a-Judge model — a separate LLM evaluates another agent’s outputs against a structured rubric — generally outperforms learned reward models (Venkataramani et al., 2026, MAS-ProVe). But this approach requires an external oracle and extra compute.

Process verification in multi-agent systems — MAS-ProVe (Venkataramani et al., 2026) runs the most systematic empirical analysis on this point. Central result: process verification does not consistently improve performance. Variance is high across configurations. Optimal granularity — verifying at the agent level or the iteration level — remains an open question.

Gulli (2025, Ch.19) frames this tension with the concept of trajectory: evaluating an agent means comparing its actual trajectory to an ideal one, step by step. In a multi-agent system, the question splits in two — does each individual agent perform its task, and does the overall system respect the global plan?

In short: no approach is universal. Self-verification is fast but blind to the agent’s own blind spots. External verification is more reliable but costs an extra LLM call. The contract model demands upfront discipline (writing the contract) that few teams accept. The right choice depends on the stakes of the task and the compute budget.


Which verification approach to choose?

Task contextRecommendationWhy
Simple task, single agent, low stakesSelf-verification (SmartSnap-like)Good cost/quality trade-off, gains up to 26% on 8B — shares the agent’s blind spots
High-stakes task, irreversible action imminentExternal verification before execution (TrustBench)-87% harmful actions, latency < 200 ms — costs one extra LLM call
Multi-agent pipeline with explicit contractsContract model (Gulli Ch. 19)Acceptance criteria defined before execution, corrupted successes detectable as they propagate
Research/benchmarks, quality over costExternal LLM-as-a-JudgeOutperforms learned reward models (MAS-ProVe), but an external oracle is expensive

In plain terms: no approach is universal. Self-verification is fast but blind to the agent’s own blind spots. External verification is more reliable but costs one additional LLM call. The contract model demands upfront discipline — writing the contract — that few teams accept. The right choice depends on the stakes of the task and on the compute budget.


Toward a formal definition of success: the contract model

A structural path out of this ambiguity: define success before execution, not after. Gulli (2025, Ch.19) proposes the “contractor” model: the agent operates from a formal contract that specifies expected deliverables, scope, allowed resources, and validation criteria. This contract is the single source of truth for judging whether the task has succeeded.

The model rests on four pillars: initial formal specification, an agent-user negotiation phase before execution, a self-validation loop during execution (generate-revise-improve), and hierarchical decomposition into formal subtasks. This last point applies directly to multi-level verification: each subtask carries its own acceptance criteria, which makes the propagation of corrupt successes detectable at the step where it occurs.

Gulli’s distinction between acceptable (in-order match), exact (exact match), and flexible (any-order match) trajectories in Ch.19 answers Shahnovsky and Dror’s problem precisely — the definition of success is explicit in the contract, not implicit in the metric chosen after the fact.


The memory and long-loop problem

TIDE (Yan et al., 2026) adds a temporal dimension that current metrics ignore. Evaluation must decompose the global dynamics of execution along three axes: global progress over time, identification of recursive loops (the agent spinning without detecting it), and the actual impact of accumulated memory on resolution. An agent can exhibit apparent improvement during a long session — by adapting its memory — without that improvement corresponding to better resolution capability. TIDE formalizes this confusion as a measurement bias: evaluating improvement is not equivalent to evaluating success.

The difficulty intensifies in tasks exceeding twenty interaction turns. At that length, the definition of success is not yet standardized in the literature. Verification in long loops remains the most under-studied gap in the field — and the most critical for agents deployed in production on complex tasks.


What to remember

  • Between 27% and 78% of benchmark-reported successes hide procedural violations not detected by the binary criterion (Cao et al., 2026).
  • Process quality — how the agent executed — is not captured by the final outcome. AgentProcessBench is the first benchmark to measure it directly.
  • Self-verification improves performance but shares the blind spots of the agent it evaluates. External verification (LLM-as-a-Judge) is more reliable but costly.
  • Process verification in multi-agent systems does not yield consistent improvement — it is an open problem, not an available solution (MAS-ProVe).
  • There is no consensual definition yet of what “succeeding” means for an agent. Each benchmark defines its own criterion, making comparisons unreliable.
  • Designing an agent without an explicit feedback loop means designing an agent that cannot detect its own drift (Gulli, Ch.11).
  • The contract model (Gulli, Ch.19) offers a structural alternative: define acceptance criteria before execution, not after. Without a formal specification of success upstream, any post-hoc judgment remains arbitrary.