In short
We orchestrated 35 tooled LLM agents in parallel to migrate 80 articles in 15 minutes, with a 100% success rate. The academic literature tests at most 9 tooled agents under controlled conditions (Kim et al., 2025). This guide documents the concrete architecture, positions this deployment against the state of the art, and analyzes an unexpected emergent behavior of the orchestrator — along with the control mechanisms that turn autonomy into reliability.
In short: imagine 35 translators sitting in a large room, each with a different text to translate. None of them talk to each other. A team leader handed them the shared glossary and the validation grid. When each one hands in their copy, the leader rereads, sets the margins, builds the final PDF. That is exactly what a “fire-and-consolidate” deployment does: you avoid upstream coordination because the tasks do not cross. If they did, it would no longer be the same mechanic.
The three-phase architecture
A deployment at this scale is structured around three distinct phases.
Preparation (orchestrator alone)
Before any dispatch: reading the full inventory (80 articles with source, destination, and category), reading the pre-computed cross-link mapping, cleaning up 5 obsolete components, detecting and resolving 2 assignment conflicts (articles assigned to two different agents). This phase takes a few minutes but prevents collisions during the parallel phase.
Parallel execution (35 agents)
The 35 agents were launched in 3 batches of ~12 agents — not in a single block. Beyond 12–15 simultaneous agents, the risk of API throttling increases significantly.
The terminal execution trace shows three successive waves: 12 agents (Actors + Fundamentals), then 11 agents (Infrastructure + Tools + early Techniques), then 12 agents (remaining Techniques + Issues). Each line displays the agent name and its “Running in the background” status — 35 lines in total, all running simultaneously at the end of the dispatch.
Each agent received: its list of assigned articles (2 to 4), a link to a shared instructions file on disk (71 lines), and the relevant cross-link entries.
Consolidation (orchestrator alone)
After all 35 agents returned: counting produced articles (76 out of 80 — the 4 missing being planned merges), correcting 9 broken cross-links, deleting 27 obsolete files, building the site (85 pages, 0 errors). This phase took ~7 minutes — nearly half the total duration. Consolidation does not scale as well as parallel execution.
Three conditions for scaling beyond 20 agents
Experience across about twenty multi-agent deployments (2 to 35 agents) shows that three conditions must be met simultaneously.
Sub-task homogeneity. Each agent performs the same operation: read a source article, enrich its metadata, review the content, write the result. The prompt is nearly identical from one agent to another — only the assigned articles differ.
Complete input provided. No agent needed to search for information. The exhaustive inventory, the cross-link mapping, the detailed instructions — everything was provided upfront. Each agent’s uncertainty was zero.
Total write decoupling. Each agent writes to a different file path. No shared write targets. Assignment conflicts were detected and resolved before dispatch — a single file shared in write mode across 35 agents would create race conditions that are impossible to debug.
In short: these three conditions are not optional “best practices”. If one breaks, the deployment derails. Without homogeneity, the prompt grows until cache is lost. Without complete input, each agent fetches its context with its own decisions and coordination must restart. Without write decoupling, two agents silently overwrite the same file and you only see it at the broken build two hours later.
What the literature says
Two worlds of agents
Multi-agent LLM research divides into two barely comparable categories.
On one side, large-scale simulations deploy thousands of lightweight agents — a single LLM call per agent, with no tools or complex reasoning. The Generative Agents study (Park et al., 2023) simulates 25 agents in a social environment for “thousands of dollars.” AgentSociety (2025) scales to over 10,000 agents generating 5 million interactions. MacNet (Qian et al., 2024, ICLR 2025) supports more than 1,000 agents in a directed acyclic graph, with performance following a logistic growth law.
On the other, tooled agents — with filesystem access, code execution, and multi-turn reasoning loops. The most systematic study (Kim et al., 2025, Google/UW) tests 180 configurations on 1 to 9 agents, 5 architectures, and 3 LLM families. Key results:
| Metric | Value |
|---|---|
| Error amplification (independent agents) | 17.2× |
| Error amplification (centralized coordination) | 4.4× (containment) |
| Token overhead (independent → hybrid) | 1.6× to 6.2× |
| Centralized coordination gain (parallelizable tasks) | +80.8% |
| Degradation (sequential reasoning) | -39% to -70% |
The MAST dataset (Cemri et al., 2025) analyzes 1,600 failure traces across 7 frameworks and identifies 14 failure modes. Coordination failures account for 36.9% of cases. Conclusion: “the performance gains of multi-agent systems on popular benchmarks are often minimal.”
The documentation gap
No published study documents a controlled deployment between 10 and 50 tooled agents. Kim et al. cap at 9. Simulations (MacNet, AgentSociety) use lightweight agents that are not comparable. ChatDev (Qian et al., 2023, ACL 2024) deploys 7 specialized roles with a 60% success rate on non-trivial tasks — but in a sequential workflow, not in parallel.
Why fire-and-consolidate bypasses the ceiling
Kim et al. identify a saturation threshold: beyond ~45% of single-agent performance, multi-agent coordination produces negative returns. This ceiling arises from inter-agent communication overhead.
The fire-and-consolidate pattern (orchestrator dispatches → agents execute independently → orchestrator consolidates) eliminates this overhead: agents never exchange with each other. Coordination overhead is zero by design — but only because the tasks are intrinsically independent.
What puts this in perspective
The 100% success rate is tied to the conditions (independent tasks, complete input, zero write coupling), not to any intrinsic superiority. As soon as tasks require inter-agent coordination, the Kim et al. ceiling applies in full. This is a single execution (N=1), not replicated.
When the orchestrator improvises
What the execution trace shows
The terminal screenshot captures a specific moment: the orchestrator announces “I am creating a shared instructions file for the agents, then dispatching the 35 agents.” It writes a 71-line file to disk containing the target path, the frontmatter format, and the review instructions. It then launches the first batch of 12 agents in the background.
This behavior was not prescribed.
What the original prompt prescribed
The operator’s initial prompt was over 400 lines, structured in three phases (preparation, parallel, consolidation). It contained a section “Instructions common to all agents” — a block of ~50 lines that the orchestrator was supposed to inject into each sub-agent’s prompt. With 35 agents, this represented ~1,750 lines of duplicated content.
The orchestrator identified this inefficiency and resolved it autonomously: extract the instructions into a shared file on disk, then give each agent a link to that file instead of the full text.
Observed benefits:
- Lighter prompts: each agent prompt contains only its assigned articles
- Single source of truth: a correction in the file applies to all agents in subsequent batches
- Inter-batch adjustability: if batch 1 reveals a problem, the instructions can be corrected before batch 2
The dual reading
This behavior constitutes an emergent optimization: the agent develops an efficiency strategy not foreseen by the operator. The gain is real — the deployment is more robust and more token-efficient.
But it is also a partial loss of visibility: the operator does not know, at dispatch time, that the deployment structure has changed. If the instructions file contained an error, all 35 agents would propagate it — whereas in the original design, the error could have been detected when reading the prompt.
Bounding emergent autonomy
The preceding observation raises a structural question: how do we benefit from the agent’s optimization intelligence without losing visibility into its decisions?
The pre-execution challenge
One effective mechanism is to impose an explicit reflection phase before any action. The agent must identify one practical obstacle and one potentially false implicit assumption in the plan it is about to execute. These two elements are recorded at the top of the deliverable, visible to the operator.
In the 35-agent deployment case, this challenge had identified the risk of assignment conflicts — two articles assigned to the same agent in the original prompt. This allowed 2 conflicts to be corrected before dispatch. Without this mechanism, two agents would have written to the same file, with the second silently overwriting the first.
Architectural self-evaluation
A second mechanism intervenes after the planning phase: the agent evaluates its own deployment architecture. Is the number of agents right? Should certain groups be split? Are there potential file conflicts?
This self-evaluation does not replace human validation — it prepares it. The operator receives a structured diagnosis instead of an opaque plan, and can intervene on specific points rather than reviewing everything.
From unconstrained to bounded
| Without mechanism | With mechanism |
|---|---|
| Agent optimizes silently | Agent states its doubts, then optimizes |
| Operator discovers changes after the fact | Operator validates before execution |
| Errors propagate to N agents | Errors are detected before dispatch |
The principle: the agent may propose improvements to the plan, provided it signals them explicitly. Emergent optimization (shared instructions file) remains desirable — but it must be visible.
Decision matrix: when to use massive fire-and-consolidate
| Context | Recommendation | Why |
|---|---|---|
| Independent tasks, complete input, ≥ 20 work units | Fire-and-consolidate 20-35 agents | Kim et al. ceiling bypassed, linear throughput gain. |
| Inter-agent coordination needed (shared planning, mutable state) | Centralized coordination 5-9 agents | ~9-agent ceiling Kim et al.; beyond, negative returns. |
| Long sequential tasks (each step depends on the previous result) | Pipeline 1-3 agents | Sequential multi-agent degrades performance -39% to -70%. |
| Very large volume (≥ 100 units), simple agents (1 LLM call each) | MacNet/AgentSociety simulation 100-1000 agents | Lightweight agents tolerate scale; tooled agents do not. |
| Doubt about task independence | 5-agent pilot before scaling | Verify the absence of race conditions on a small batch. |
Metrics
| Metric | Value |
|---|---|
| Agents deployed | 35 (all same model) |
| Success rate | 100% (35/35) |
| Total duration | 15 min 26 sec |
| Min / max agent duration | ~110 sec / ~403 sec (ratio 3.7×) |
| Articles produced | 76 (80 − 4 planned merges) |
| Cross-links corrected in consolidation | 9 |
| Estimated tokens | ~1.5 million (~43K per agent) |
| Final build | 85 pages, 0 errors |
The duration gap (ratio 3.7×) comes from unequal load: the slowest agent handled 4 articles with complex merges, the fastest only 2 simple ones.
Key takeaways
- 35 tooled agents in parallel occupy a documentation gap: the literature tests at most 9 tooled agents (Kim et al., 2025) or 1,000+ lightweight agents (MacNet) — the two are not comparable
- The fire-and-consolidate pattern bypasses the coordination ceiling by eliminating inter-agent communication — but only for intrinsically independent tasks
- At large scale, the orchestrator develops unprescribed optimization strategies (shared instructions file) — an efficiency gain that implies a loss of visibility for the operator
- Control mechanisms (pre-execution challenge, architectural self-evaluation) turn emergent autonomy into bounded autonomy — the agent can optimize, provided it signals its choices
- Three scaling conditions: task homogeneity, complete input, zero write coupling. If any one is missing, the success rate degrades