In brief

The Reflection pattern allows an AI agent to evaluate its own output and correct it iteratively, without human intervention at each cycle. In its most robust form — the Producer-Critic model — two distinct roles are assigned: one agent generates, another critiques. This separation eliminates a structural bias in self-revision. The loop is not infinite: explicit stopping criteria regulate the number of iterations.


The problem: why an agent cannot effectively proofread itself

An LLM that generates a text, then is immediately asked to correct it, suffers from a known bias: it tends to validate what it just produced. This is not a malfunction — it is a property of the generation process. The model that produced the response has already activated certain inference paths; re-reading with the same model in the same context state often reproduces the same blind spots.

This bias is analogous to the proofreading phenomenon in humans: an author re-reads their text and sees what they meant to write, not what they actually wrote.

Reflection, as formalized in the agents literature, addresses this problem. It introduces a structured feedback loop: generate → evaluate → correct → iterate if necessary.


The Producer-Critic model

The Producer-Critic is the most refined form of the Reflection pattern. It rests on an explicit separation of roles:

RoleFunctionBias eliminated
ProducerGenerates the initial response according to the task—
CriticEvaluates the producer’s output with a fresh perspectiveSelf-complacency bias

Gulli (2025) formulates this principle: “The Critic agent approaches the output with a fresh perspective, dedicated entirely to finding errors and areas for improvement.”

The critic is not a spell-checker. It evaluates against structured criteria defined in advance: factual accuracy, internal consistency, completeness, adherence to instructions. These criteria must be explicit — a generic critic (“is this correct?”) is less effective than one with a formalized control checklist.

In practice, both roles can be fulfilled by the same base model with different prompts, or by two distinct models. The separation is logical, not necessarily physical.

In short: the critic is not a spell-checker — it is an evaluator equipped with an explicit grid (accuracy, consistency, completeness, instructions followed). Asking “is this correct?” produces less value than asking “check these 5 specific criteria on the output”.


The feedback loop: structure and mechanics

The reflective loop follows a four-step cycle:

  1. Execution — the producer generates an initial version
  2. Evaluation — the critic analyzes the output against the defined criteria
  3. Refinement — the producer integrates the critique and produces a revised version
  4. Iteration decision — continue or stop based on whether the quality criterion has been met

At each turn, the conversation history grows richer: task, intermediate output, critique. Gulli notes that this accumulation of context improves the quality of successive iterations — the agent evaluates its output not in isolation, but taking into account previous cycles.


Stopping criteria: the infinite loop problem

Without an explicit stopping condition, a reflective loop can run indefinitely. Each iteration consumes tokens, latency, and inference budget. Two complementary mechanisms regulate this:

Satisfaction criterion: an explicit signal indicates that the target quality has been reached. In code, this can take the form of a special token (CODE_IS_PERFECT) that the critic emits when no correction is needed. For text, a threshold score on the defined criteria.

Maximum iteration count: a safety limit independent of quality. Even if the critic continues to identify imperfections, the process stops after N cycles. This ceiling prevents over-refinement loops where each iteration produces diminishing marginal gains at increasing cost.

Both mechanisms must coexist. A satisfaction criterion alone can be compromised if the critic is poorly calibrated; a ceiling alone ignores actual quality.

In short: without an explicit ceiling (typically 3-5 iterations), a Producer-Critic loop can “polish” an output indefinitely, each iteration bringing 0.1 quality points for 100% additional cost. The pattern is profitable only if marginal improvement exceeds marginal cost — the decay is fast after 3 cycles.


Self-reflection vs. Producer-Critic: when to use which

There is a lighter variant: self-reflection, where a single agent evaluates its own output via a second LLM invocation with a critique prompt. This is less infrastructure-intensive than a two-agent system.

ApproachAdvantageLimitation
Self-reflectionSimple, fewer tokens, single modelSelf-revision bias partially present
Producer-CriticGenuinely fresh perspective, reduced biasDouble LLM call cost, more complex orchestration

Self-reflection suits cases where initial quality is already high and surface-level refinement is sought (phrasing, formatting). Producer-Critic is preferable when the task is high-stakes, when errors are difficult to detect in close context, or when the domain requires evaluation against specialized criteria (legal, medical, technical).


Coupling with memory

The effectiveness of reflection depends on the agent’s ability to account for previous cycles. An agent without session memory restarts each iteration from scratch: it may reproduce the same errors it just corrected.

Gulli notes that “preserving the conversational history provides crucial context for the evaluation phase, allowing the agent to assess its output not in isolation, but within the framework of previous interactions.”

In practice, this means the context window grows with each iteration. On long loops or complex tasks, this expansion can approach the model’s limits or degrade evaluation quality (the critic agent gets “lost” in the history). This is a direct operational constraint.


Concrete applications

Code generation: the producer generates a function, the critic checks the logic, edge cases, and specification compliance. Several iterations can correct bugs not caught in the first pass.

Structured writing: the producer generates a report, the critic checks argumentative consistency, completeness, and absence of contradictions.

Information extraction: the producer extracts data from a document, the critic verifies coverage and accuracy by re-reading the source document.

Exception handling: coupled with error detection, reflection allows an agent to diagnose a failure and generate a recovery strategy.


Strengths and limitations

What the pattern provides:

  • Measurable quality improvement on complex tasks
  • Detection of errors that direct evaluation misses
  • Adaptable: works with a single model or specialized models
  • Composable with other patterns (planning, exception handling)

What the pattern does not solve:

  • The critic itself can have blind spots or be poorly calibrated
  • Each iteration adds latency and cost — marginal gains diminish
  • The growing context window can degrade coherence on long loops
  • If evaluation criteria are poorly defined, the critic optimizes for the wrong metrics
  • Over-correction exists: an overly strict critic can degrade an initially correct response

Key takeaways

  • Reflection introduces an iterative self-evaluation loop: generate, critique, refine.
  • The Producer-Critic model separates generation and evaluation to reduce self-revision bias; the critic’s fresh perspective is its core value.
  • Explicit, structured evaluation criteria are necessary — a critic without a control checklist is unreliable.
  • Stopping criteria (quality signal + iteration ceiling) are mandatory to avoid infinite loops and over-optimization at increasing cost.
  • The cost of the pattern is real: increased latency, additional tokens, growing context window. This must be weighed against the expected quality gain.