In short
An agent system cannot reliably evaluate itself: the agent that produces a result applies the same biases as the one verifying it. Multi-agent auditing solves this problem by separating production and verification into distinct agents. Scanner agents inspect deliverables, trajectories, or behaviors; a consolidator agent aggregates and flags deviations. This architecture enables continuous monitoring without blocking the main execution pipeline.
In short: an accountant who checks their own books will never find their own systematic errors — they redo the same additions in the same way. An external auditor finds them, because they look with another eye and apply other rules. Multi-agent auditing reproduces this principle: verification is delegated to agents that did not write the report, with their own inspection grids.
The self-evaluation problem
A monolithic agent tasked with verifying its own outputs suffers from a structural flaw: it uses the same model, the same representational biases, and sometimes the same context as the one that produced the output being evaluated. If the model systematically misinterpreted an instruction, it will fail to detect the error in exactly the same way it made it.
This problem is documented in the multi-agent systems literature. Gulli (2025) describes the Producer-Critic pattern as a direct response: separating generation and evaluation into distinct agents “to eliminate self-revision biases.” An independent critic agent applies predefined criteria and can identify deviations the producer would never have flagged.
Recent research confirms this diagnosis. The Agent-as-a-Judge framework (arXiv:2410.10934) shows that a dedicated evaluator agent significantly outperforms classical LLM-as-a-Judge and approaches the reliability of human evaluation on complex tasks. The key advantage: the evaluator agent can inspect intermediate steps of the process, not just the final result.
Architecture: scanners and consolidator
The dominant pattern for automated auditing follows a two-level structure.
Level 1 — Scanner agents
Multiple specialized agents operate in parallel, each focused on a specific audit criterion:
- Trajectory scanner: compares the actual execution path to the expected trajectory (tools called, step order, decisions made).
- Compliance scanner: verifies that outputs conform to a defined schema (format, mandatory fields, business constraints).
- Drift scanner: detects progressive performance degradation by comparing current metrics against reference baselines.
- Quality scanner: evaluates subjective dimensions (clarity, relevance, consistency) via a structured rubric.
This specialization is deliberate. A scanner focused on a single criterion applies explicit rules, reducing the risk that it rationalizes an error instead of flagging it.
Level 2 — Consolidator agent
The consolidator agent receives reports from all scanners and produces a structured summary. Its mission includes three operations:
- Aggregate scanner signals into an overall diagnosis.
- Flag contradictions between scanners (two scanners giving opposing evaluations of the same deliverable).
- Prioritize alerts by criticality for the operator.
Gulli (2025, Ch.19) formalizes this logic in his multi-agent evaluation grid, which poses four questions: is cooperation between agents effective? Is the overall plan being followed? Is the right agent assigned to the right task? Does adding agents improve or slow down the system?
In short: the “scanners + consolidator” architecture is the organization of a financial audit. Several specialized inspectors (one on VAT, one on compliance, one on cash flow) work in parallel without talking to each other. The chief auditor synthesizes their reports, manages contradictions, ranks anomalies. Each one sees their angle; none has the power to mask a problem by convergence of view with the others.
The Contractor model: auditing integrated into execution
A more advanced evolution consists of integrating auditing directly into the executing agent’s logic, rather than externalizing it as post-processing. Gulli (2025, Ch.19) names this paradigm the Contractor model.
Its four pillars:
| Pillar | Role in auditing |
|---|---|
| Formal contract | Explicit specification of deliverables and validation criteria before execution |
| Dynamic negotiation | The agent dialogues with the user to clarify criteria before starting |
| Execution with self-validation | Generate → revise → validate loop at each step |
| Decomposition into sub-contracts | Subtasks inherit the validation criteria of the parent contract |
The benefit for auditing: each subtask carries its own success criteria, making verification distributed and traceable. A consolidator agent can inspect completed sub-contracts to reconstruct the audit of the entire pipeline.
Concrete applications
Code auditing
In a code generation pipeline, multiple scanners operate in parallel: a syntactic scanner (compilation, typing), a security scanner (known vulnerable patterns), a style scanner (convention compliance). The consolidator aggregates and blocks the merge if a critical scanner fails. The producer agent is never involved in its own verification.
Documentary compliance auditing
For generating regulatory documents (contracts, financial reports), auditor agents independently verify the presence of mandatory clauses, the consistency of cited figures, and the absence of prohibited formulations. AgentAuditor (arXiv:2602.09341) applies this principle by building a reasoning tree over critical divergence points, outperforming majority voting and LLM-as-a-Judge across six tested architectures.
Drift monitoring in production
A monitor agent compares the metrics of each session (latency, error rate, trajectory compliance) against a historical baseline. Gulli (2025, Ch.19) describes this pattern as “continuous monitoring”: data is recorded in persistent systems (JSON logs, InfluxDB, Datadog), and an agent periodically analyzes trends to alert before a degradation becomes critical.
Strengths and limitations
Strengths
- Reduced self-evaluation bias: an independent auditor agent does not share the blind spots of the producer agent.
- Parallelization: multiple scanners operate simultaneously without blocking the main execution pipeline.
- Traceability: each audit decision is associated with an agent, a criterion, and evidence — which facilitates debugging.
- Scalability: adding a new audit criterion amounts to adding a scanner, without modifying the rest of the pipeline.
Limitations
Who audits the auditor? This is the central question this pattern does not resolve on its own. A poorly specified or biased scanner agent will systematically produce false positives or false negatives. The quality of the audit depends on the quality of the rubrics and criteria defined by humans.
Computational cost: each scanner consumes tokens. A pipeline with ten scanners multiplies evaluation cost tenfold. Trade-offs are necessary in production: exhaustive auditing on critical cases, lightweight auditing on routine cases.
Coordination: if scanners share a context or depend on the same data, parallelizing them introduces consistency risks. The consolidator must explicitly handle contradictions between scanners, at the risk of producing an incoherent synthesis.
Subjective evaluation: for qualitative criteria (clarity of a response, appropriate tone), auditor agents remain LLMs with their own evaluation biases. A structured rubric reduces but does not eliminate this variability. [UNVERIFIED for inter-rater concordance rates under real production conditions.]
Decision matrix: which audit pattern to choose
| Context | Recommendation | Why |
|---|---|---|
| Simple pipeline, 1-3 deterministic validation criteria | Minimal Producer-Critic | 1 external critic, low overhead, sufficient to break self-validation. |
| Complex pipeline, ≥ 4 heterogeneous criteria (security + quality + compliance) | Scanners + consolidator | Parallel specialization, traceability per criterion, scalable. |
| Critical production requiring upstream auditing | Contractor model | Upstream criteria, distributed validation, sub-task inheritance. |
| Continuous post-deployment monitoring (slow drift) | Monitor agent + historical baseline | Early detection, alert before visible degradation. |
| Very high-stakes decisions (medical, legal) | Hybrid agent + human audit | Agents filter, human decides as last resort. |
In short: everything depends on the cost of an undetected error. On an internal pipeline with fast feedback, a single external critic is enough. On a system that produces legal contracts, you stack scanner upon scanner, and keep a human in the terminal loop. The rule: audit depth must be proportional to the cost of the most expensive error.
Key takeaways
- Separating production and verification into distinct agents structurally reduces self-evaluation bias.
- The scanners + consolidator pattern enables parallel multi-criteria auditing without blocking the main execution pipeline.
- The Contractor model integrates audit criteria into execution contracts, making validation distributed and traceable.
- The fundamental limitation remains human: audit quality depends on the rubrics and criteria defined upstream.
- Who audits the auditor? The operational answer is a combination of periodic human supervision and second-order metrics on the auditors themselves.