In short

An agentic AI system — where a model can plan, delegate, and execute tasks — raises a governance question: who validates when the AI both designs and executes? A practical answer is to distribute three distinct roles: AI copilot (design), technical executor (execution), human operator (decision and validation). This decision triangle rests on a fundamental asymmetry: only the human is independent of the underlying model.


The problem the triangle solves

In a minimal agentic system — a single model that receives an objective and executes it — validation is absent. The model produces a plan and applies it. If the plan is flawed, nothing intercepts it.

The temptation is to add a second agent to check the first. But if both rely on the same base model, their agreement does not constitute independent validation: they are two instances of the same reasoning. This is the convergence bias: two LLMs derived from the same weights converge toward similar conclusions, even in the face of an error. In practice, the error detection rate via intra-model cross-validation drops below 30% on structural reasoning tasks [order of magnitude estimate], compared to 70-90% when validation involves a structurally different actor [order of magnitude estimate].

The solution is not to add more agents. It is to introduce a structurally different actor: a human whose judgment does not depend on the model.

In plain terms: adding a second LLM to check the first is like asking someone to reread their own draft with a second pair of glasses. The glasses change, the head does not. Only a third actor — independent of the model — can break the loop.


The three roles of the triangle

In plain terms: three actors, three non-interchangeable responsibilities. The copilot thinks, the executor acts, the human validates. As soon as one encroaches on another’s role, the guarantee of independence collapses.

Distribution matrix by context

ContextDominant RoleWhy
Decomposition of a vague objective, plan design, post-hoc analysisAI copilot (design)Fluent linguistic reasoning, low marginal cost (~0.01-0.10 € per decomposition [order of magnitude estimate]), no side effects.
Mechanical application of a validated plan, filesystem execution, API calls, versioningTechnical executor (execution)Auditable scope, traced actions, minimal latency (~1-10 s per atomic action [order of magnitude estimate]).
Pre-execution validation, ambiguity arbitration, irreversible merge, external publicationHuman operator (decision)Only actor independent of the model — breaks the convergence bias. High cost (~2-15 min of attention per breakpoint [order of magnitude estimate]), therefore reserved for structural decisions.

Typical relative weight over an orchestrated session: approximately 60-70% of elapsed time in execution (executor), 15-25% in design (copilot), 10-20% in validation (human) [order of magnitude estimate, measured on the Matteo / Claude.ai / CC triptych in internal R&D]. As soon as the human share drops below 5%, the system drifts toward intra-model self-validation — the convergence loop closes back up.

The AI copilot — WHAT and WHY

The AI copilot handles design. It decomposes an objective into subtasks, formalizes hypotheses, designs prompts, analyzes results. It operates in natural language, without access to file systems or execution tools.

This constraint is deliberate: the copilot does not touch the artifacts produced. It thinks. It does not act.

Strength: the separation between design and execution forces an explicit statement of the plan before it is applied. The human operator can read, critique, and validate before anything is executed.

Limit: the AI copilot remains a language model. It can formalize an incorrect decomposition with the same fluency as a correct one. Its internal consistency is not a guarantee of correctness.

The technical executor — HOW

The technical executor takes a validated decomposition and applies it. It has access to the filesystem, tools, APIs, and version control. Its scope of initiative is minimal: it executes, flags anomalies, archives results.

The structuring rule is the following: the executor never independently creates a complex decomposition. If a task exceeds the scope of an atomic action, it escalates to the human operator rather than decomposing on its own initiative.

Strength: confining the executor to execution allows auditing the actions performed independently of the reasoning that produced them.

Limit: the executor and the copilot often rely on the same base model. Their mutual agreement — the copilot validating what the executor produced — does not constitute independent verification.

The human operator — decision and validation

The human operator defines objectives, validates decompositions before execution, and merges results. They are the only point in the system whose judgment does not depend on the underlying model.

This role does not consist of supervising every action at the microsecond level. It consists of positioning validation checkpoints at the moments when the system makes structural decisions: before executing a plan, before publishing a deliverable, before applying an irreversible change.

Strength: the human breaks the convergence loop. Even if the copilot and the executor converge on the same error, the human can intercept it.

Limit: this role assumes that the operator understands sufficiently what the system produces to validate it non-formally. A surface-level validation — “looks fine” — does not fulfill the function.


The decision flow in practice

The triangle operates as a sequential flow with mandatory checkpoints:

  1. The human operator defines the objective.
  2. The AI copilot decomposes the objective into subtasks and designs the execution prompts.
  3. The human operator validates the decomposition.
  4. The technical executor applies the plan, task by task.
  5. The human operator validates the deliverables before any final action (merge, publication, deployment).

This flow does not assume continuous supervision. It assumes structured breakpoints where the human is in a position to block an action before it becomes irreversible. In practice, 2 to 4 validation checkpoints are sufficient per orchestrated session of medium size (1-3 hours of work) [order of magnitude estimate] — beyond that, the human becomes the bottleneck; below that, the flow drifts toward unsupervised execution.

In plain terms: no need to watch every keystroke. The human intervenes only at the moments when the system is about to do something that cannot be undone — publish, merge, deploy, delete. Between two validation checkpoints, the copilot and the executor work autonomously.


Why the role separation holds

The separation of roles is not an organizational convention. It is a reliability property of the system.

A language model cannot validate its own outputs independently. It can critique, reformulate, improve — but its critique is produced by the same process as its production. This is a fundamental limit, not a bug fixable by a larger model. Intra-model self-critique tests typically plateau at 40-60% error detection on complex tasks [order of magnitude estimate], whereas an independent actor — human or third-party system — reaches 80-95% on the same tasks [order of magnitude estimate].

The decision triangle works around this limit by externalizing validation to a structurally different actor. It is not that the human is “better” than the model on every task. It is that the human is independent — and independence is what makes validation useful.

The practical rule that follows: the AI copilot never validates results produced by the technical executor. This validation belongs to the human operator. If the human delegates this validation to the copilot, the triangle is broken. Across approximately 100 orchestrated sessions observed in internal R&D [order of magnitude estimate], cases where human validation was short-circuited exhibit an undetected structural error rate 3 to 5 times higher than sessions with effective validation [order of magnitude estimate].

In plain terms: the triangle holds because one of the three vertices is different in nature — not smarter, just independent. It is the same logic as an external auditor: they are not a better accountant than the CFO, but they do not depend on the CFO.


Key takeaways

  • An agentic system without independent validation cannot detect its own structural errors.
  • Convergence bias makes validation by a second agent derived from the same model ineffective.
  • The triangle distributes three roles: design (AI copilot), execution (technical executor), validation (human operator).
  • The human operator is not a continuous supervisor — they are a structured breakpoint at irreversible decisions.
  • The separation holds if and only if the human effectively validates the deliverables, not formally.