In short

Fire-and-consolidate is an orchestration pattern that breaks a task into independent units, processes them in parallel across N agents, then aggregates all results in a single consolidation pass. The consolidation phase only starts after all parallel agents have completed. Field reports on this type of architecture show very high completion rates, often above 95%.

In plain terms : five to ten agents are launched in parallel on separate sub-tasks, we wait for them to submit their work, then a single “editor” agent assembles everything into a synthesis. It is teamwork without a meeting: each in their own corner, a reporter at the end.


Origin: when a spontaneous pattern becomes an architecture

The pattern does not arise from a design decision. It emerges from a practical observation: when multiple agents work in parallel on independent tasks, their outputs naturally converge toward a consolidable format. The final synthesis step imposes itself.

This movement — from emergence to formalization — is characteristic of robust engineering patterns. They are not invented; they are observed, named, then codified. Once named, the pattern becomes prescriptive: it can be replicated intentionally, adapted, and tested.

The distinction between an “emergent pattern” and a “chosen architecture” has concrete consequences. An emergent pattern has been validated by practice before being documented. Its scope of application is better understood than that of a pattern designed a priori.


Mechanics: two phases, one rule

In plain terms : phase 1, everyone works at the same time; phase 2, a single agent re-reads everything and writes the final report. Phase 2 is not allowed to begin before phase 1 ends.

The pattern is structured into two strictly sequential phases.

Phase 1 — Parallel (the “fire”)

The orchestrator launches N agents simultaneously. Each agent receives a complete, autonomous prompt: it does not need the others to do its work. Each writes its output to a dedicated file. There is no communication between agents during this phase.

The independence rule is strict: if the tasks are interdependent, the pattern does not apply. Parallelism is only possible when work units are separable without loss of information.

Phase 2 — Sequential (the “consolidate”)

Once all agents from phase 1 have finished, a consolidating agent reads all the outputs and produces the final synthesis. This phase does not begin until the last parallel agent has delivered its work.

If a phase 1 agent fails, consolidation proceeds on the available partial results, with explicit flagging of missing elements. The system degrades gracefully rather than failing completely.

The central rule: phase 2 waits for phase 1. Without this synchronization constraint, the consolidator would work on an incomplete corpus without knowing it.

Observed order of magnitude: for 8 agents whose unit task takes 2 to 3 minutes, a sequential run would last ~20 minutes; in parallel, phase 1 completes in ~3 minutes (time of the longest task), and consolidation typically adds 1 to 2 minutes. The wall-clock latency gain is roughly a factor of N, where N is the number of agents [order of magnitude estimate].


When to apply it

In plain terms : we pull out the pattern starting at three independent work units. Below that, it is simpler to process serially; above it, we gain time without complicating the logic.

The trigger criterion is twofold: the work units must be independent of each other, and their number must justify the coordination overhead. In practice, the relevant threshold is three units or more.

Below that, direct sequential processing is less costly and just as fast. Above it, the latency gain becomes significant: N tasks processed in parallel take the time of the longest, not the sum of all. In typical reported deployments, between 5 and 15 parallel agents per wave are observed, with a practical ceiling around 20 before consolidation itself becomes a bottleneck [order of magnitude estimate].

The pattern adapts to several types of work: reading and extracting from multiple sources, batch document generation, systematic corpus auditing, large-scale classification. In all these cases, the common structure is the same — N independent units, then a synthesis.


Limits and risks

In plain terms : the pattern is solid on parallelism, fragile on the final synthesis. Three major risks: uniformization of outputs, absence of self-correction between agents, and total dependence on a good initial prompt.

The herding effect in consolidation

The main risk does not come from the parallel phase but from the consolidation. A consolidating agent faced with N similar outputs tends to smooth out differences rather than preserve them. Minority nuances — sometimes the most informative — disappear in the synthesis.

This phenomenon intensifies when parallel agents have received the same base prompt: their outputs are structurally similar, which reinforces the consolidator’s temptation to merge them without discrimination.

Partial countermeasure: explicitly instruct the consolidator to flag divergences rather than resolve them. The synthesis then becomes a convergence/divergence report rather than a blind merge.

The absence of mutual correction

Each phase 1 agent works in a silo. A factual error in one agent’s output will not be corrected by its peers — they do not read it. If the same error appears in multiple outputs, the consolidator may treat it as established fact.

This limit is structural: it is the direct trade-off of parallelism. Allowing agents to correct each other would introduce dependencies that would break the independence required in phase 1.

Dependence on the initial prompt

The quality of parallel outputs depends entirely on the quality of the initial prompt. An ambiguous prompt produces N divergent interpretations that consolidation cannot always reconcile. An incomplete prompt produces N incomplete outputs that consolidation rarely fills in.

Fire-and-consolidate amplifies prompt quality — in both directions. A good prompt produces N good outputs that consolidate cleanly. A poor prompt produces N disparate outputs whose consolidation is difficult or misleading.


Variant: multi-wave orchestration

In plain terms : if a single wave is not enough, we chain several waves. The consolidation of wave 1 becomes the brief of wave 2. Beyond 2 waves, the coordination cost exceeds the gain.

When the volume of work exceeds what a single wave can handle, or when the results of the first wave must inform the second, the pattern extends into several sequential waves.

The structure remains identical at each wave: N parallel agents, then consolidation. What changes is that the consolidation of wave N becomes the starting point for wave N+1. The orchestrator analyzes gaps after each wave and adjusts the prompt for the next one.

Experience shows that two waves suffice in most cases, typically covering 80 to 90% of observed use cases [order of magnitude estimate]. Beyond that, coordination costs increase without a proportional gain in final synthesis quality.


Key takeaways

  • Fire-and-consolidate launches N agents in parallel (phase 1) then aggregates their results in a single pass (phase 2), with strict synchronization between the two phases.
  • Applicable only when work units are independent; the practical trigger threshold is three units or more.
  • The consolidation phase is the main point of fragility: risk of smoothing out divergences (herding effect) and propagation of uncorrected errors.
  • The quality of the initial prompt is decisive — the pattern amplifies the strengths and weaknesses of the framing.
  • In multi-wave mode, consolidation of each wave informs the next; two waves cover the majority of use cases.