In short

An LLM agent does not return a single response: it executes a sequence of steps, calls tools, and makes intermediate decisions. Evaluating this behavior goes far beyond classical unit testing. The available approaches — trajectory metrics, LLM-as-a-Judge, continuous monitoring, formal specifications — answer different questions and each carries its own blind spots.

In short: testing a Python function means giving it an input and verifying the output. Testing a cook is different: you check the final dish, but also the cleanliness of the work surface, the order of preparations, total time, coherence between intent and execution. Agent evaluation does the same. It observes the trajectory (which tools, in what order), the quality of the deliverable, and drift over time. No single method covers all three axes.


Why agentic evaluation is a distinct problem

A classical software test verifies a deterministic output from a fixed input. An agent operates in a dynamic, probabilistic environment: the same request can produce different execution paths depending on tool state, available memory, or how the context is framed. Performance can degrade gradually without any explicit error being raised.

Three properties make agentic evaluation structurally different:

  • Non-determinism: the agent chooses its tools and the order of its actions at each execution.
  • Intermediate opacity: an error can occur at any step, not just in the final output.
  • Temporal drift: the input data distribution shifts, external APIs evolve, and performance degrades without a clear alert signal.

“Evaluating intelligent agents goes beyond traditional tests to continuously measure their effectiveness, efficiency, and adherence to requirements in real-world environments.” — Gulli et al. (2025), Ch.19


Trajectory metrics: comparing the actual path to the ideal path

An agent’s trajectory is the sequence of steps it executes to reach a goal. Trajectory evaluation compares this actual path against a reference path (expected tool trajectory) defined when test cases are created.

Three levels of rigor are available depending on context:

LevelDefinitionTypical use
Exact matchThe tool sequence must be identical to the referenceCritical processes, high reliability required
In-order matchExpected tools appear in the right order, but extra steps are toleratedPartially flexible workflows
Any-order matchAll expected tools are called, in any orderTasks where order is unconstrained

A typical test case format includes: the user request, the expected trajectory, reference intermediate responses, and the expected final response.

Main limitation: defining the ideal trajectory assumes you know the correct behavior in advance. For open-ended or creative tasks, this reference is hard to establish and may over-constrain valid alternative behaviors.

In short: the trajectory metric is the equivalent of a medical check-up sheet. “Did the doctor measure blood pressure? Temperature? In the usual order?” If yes, the exam is compliant. This grid works well when the right protocol is known. It blocks when there are several valid ways to solve the problem.


LLM-as-a-Judge: qualitative evaluation by an evaluator model

Some dimensions — relevance, clarity, usefulness, absence of bias — cannot be reduced to a comparison against a fixed reference. LLM-as-a-Judge uses a language model as an evaluator: it receives the agent’s output and a structured rubric, then produces scores and a justification.

A typical rubric defines several criteria (e.g., completeness, factual accuracy, tone, absence of hallucination), each scored on a scale (1–5), with a justification request and structured JSON output to facilitate aggregation.

Strengths:

  • Capable of evaluating complex text outputs without manual references.
  • Scalable: applies to thousands of outputs without systematic human annotation.
  • Flexible: the rubric adapts to the domain (legal, medical, technical).

Limitations:

  • The evaluator model inherits the biases of the evaluated model if both share the same training base.
  • The rubric must be empirically calibrated: a vague rubric produces scores with little actionable value.
  • Inference cost adds to that of the evaluated agent.
  • Results are not reproducible if the evaluator’s temperature is not fixed.

Continuous monitoring and drift

An agent in production does not behave like an agent in development. Drift refers to the gradual degradation of performance due to changes in input distribution, API updates, or shifts in user context.

Continuous monitoring records, in a structured way, at each interaction: latency, token count, tools called, final result, errors raised. This data feeds persistent systems (JSON logs, time-series databases, observability platforms) that detect anomalies by comparison against a baseline.

Common metrics to watch:

  • Latency per step: detect bottlenecks and abnormally slow tool calls.
  • Error rate by type: distinguish tool errors, reasoning errors, and timeouts.
  • Trajectory distribution: a trajectory that grows longer without justification may signal degradation.
  • Token consumption: an agent consuming more and more context may signal a formulation problem.

Main limitation: monitoring detects symptoms, not causes. Distinguishing model drift from a legitimate change in expected behavior requires human interpretation.


The Contractor model: toward formal specifications

The previous approaches evaluate after execution. The Contractor model proposes a different approach: embedding evaluation criteria into the task definition itself, as a formal contract.

The contract specifies: explicit deliverables, scope of action, authorized resources, and validation criteria. It rests on four pillars:

  1. Formal specification: the contract is the single source of truth. No ambiguity about what constitutes success.
  2. Pre-execution negotiation: the agent dialogues with the user to clarify the contract before starting.
  3. Iterative execution with self-validation: the agent generates, revises, and improves in a loop until the contract criteria are satisfied.
  4. Hierarchical decomposition: complex tasks are broken into assignable sub-contracts, each independently verifiable.

“The contractor framework reimagines AI interaction by embedding principles of formal specification, negotiation, and verifiable execution directly into the agent’s core logic.” — Gulli et al. (2025), Ch.19

Strengths: reduces ambiguity, enables auditability, facilitates decomposition of complex orchestrations.

Limitations: formal specification assumes the user knows precisely what they want — which is not always the case. The negotiation phase adds latency. The model is less suited to exploratory tasks where success criteria emerge during execution.


Evaluating multi-agent systems

When multiple agents collaborate, evaluation extends to coordination. Four questions structure this evaluation:

  1. Cooperation: does information pass correctly from one agent to the next?
  2. Plan adherence: does the system follow the planned decomposition without blocking?
  3. Routing: is the right agent selected for the right task?
  4. Scalability: does adding an agent improve or degrade overall performance?

“You can examine how well each individual agent performs its specific job, but you must also evaluate how the entire system is performing as a whole.” — Gulli et al. (2025), Ch.19

A multi-agent system can pass every individual evaluation while failing collectively. End-to-end evaluation remains essential.

Decision matrix: which evaluation method to choose

ContextRecommendationWhy
Deterministic workflow with known steps (ETL, audit)Trajectory metrics (exact match)Stable reference, granular control, low cost.
Creative or open-ended output (writing, advice)LLM-as-a-Judge with calibrated rubricEvaluates subjective qualities, scalable across thousands of examples.
Continuous service production (24/7 monitoring)Continuous monitoring + historical baselineDetects drift without blocking execution, targeted alerts.
Critical task with clear specification (legal, medical)Contractor modelUpstream criteria, distributed validation, complete audit trail.
Multi-agent system in pre-prodEnd-to-end evaluation + 4 Gulli questionsVerify cooperation, plan, routing, scalability together.
R&D phase, fast comparison of optionsLight LLM-as-a-Judge + 50-case sampleFast, sufficient to arbitrate if one approach is notably better.

Key takeaways

  • Trajectory metrics (exact, in-order, any-order match) provide objective evaluation but require a pre-defined reference.
  • LLM-as-a-Judge covers qualitative dimensions but inherits the evaluator model’s biases and requires a calibrated rubric.
  • Continuous monitoring detects drift in production; it does not replace structured evaluation during development.
  • The Contractor model embeds success criteria into the task specification — effective for critical contexts, less suited to open-ended tasks.
  • In multi-agent settings, individual evaluation is not enough: system coordination and scalability must be tested separately.