In short
A single agent working on a long software task hits two limits: execution duration and the ability to maintain a coherent view of a complex repository. Having multiple agents work in parallel seems natural, but the interference between their modifications frequently fails at integration. Geng and Neubig (CMU, 2026) propose CAID, a multi-agent coordination system inspired by git primitives, and demonstrate gains of 6 to 26 percentage points on long-horizon software engineering benchmarks.
The problem: why naive parallelization fails
When two agents modify the same repository simultaneously without an isolation mechanism, conflicts are inevitable. One agent renames a function; the other writes code that calls the old name. Each produces a correct result in isolation — but the integrated whole is broken, and fixing it often requires a complete revision of at least one agent, not a simple patch.
Empirical studies identify this problem as the primary failure mode of multi-agent systems on shared artifacts [Cemri et al., 2025; Khatua et al., 2026]. Existing approaches — role-based pipelines, hierarchical orchestrators — address task decomposition, but not the physical coordination of concurrent modifications.
In short: telling agents “try not to step on each other” does not work — performance falls BELOW the single agent. Giving each agent its own physical copy of the repository (worktree) and explicitly merging at the end via
git mergeis the difference between “naive parallel that degrades” and “CAID that improves by 6 points”.
CAID: three git primitives as foundations
CAID (Centralized Asynchronous Isolated Delegation) structures coordination around three mechanisms directly modeled on human software development practices.
1. Dependency graph and centralized delegation
Before any dispatch, a manager agent builds a dependency graph of the repository. Each node is a unit of work; each edge indicates that a task depends on another. A task is only assigned to an engineer once all its dependencies are already integrated into the main branch.
The manager decomposes the work into groups of tasks corresponding to the maximum number of parallel engineers. Strongly interdependent files are grouped and assigned to the same engineer to reduce conflict risk. Manager-to-engineer communication passes through a structured JSON protocol, not free text — making responsibility boundaries programmatically verifiable.
2. Isolation via git worktrees
Each engineer works in their own git worktree, a physically isolated copy of the repository. Parallel modifications cannot overwrite each other. Certain shared files (e.g., __init__.py) are marked read-only for engineers.
Experience shows that this physical isolation is not substitutable by verbal instructions. In a “soft isolation” configuration where engineers share a workspace and receive instructions to avoid conflicts, performance drops below the single agent on PaperBench (55.5% vs. 57.2%). With worktrees, the same system reaches 63.3%.
3. Integration via git merge and self-verification
When an engineer finishes, they commit their changes. The manager merges to the main branch via git merge. In case of conflict, it is the engineer who produced the conflicting commit who is responsible for resolving it — they pull the main branch into their worktree, resolve locally, and resubmit.
Before committing, each engineer runs the tests covering their modified files. Any failed test or runtime exception must be resolved before submission. This distributed self-verification avoids concentrating the revision load on the manager.
Measured results
CAID is evaluated on two long-horizon benchmarks:
- Commit0: implement Python libraries (tinydb, minitorch, jinja…) from a skeleton, with unit tests. Only passing all tests counts.
- PaperBench: reproduce the main contributions of a conference paper (implementation + running experiments).
On Claude Sonnet 4.5, CAID (4 engineers) improves the pass rate from 53.1% to 59.1% on Commit0-Lite, and from 57.2% to 63.3% on PaperBench. For MiniMax 2.5, the gain on PaperBench is from 10.4% to 36.7%.
Doubling iterations is not enough
A natural question: couldn’t a single agent close the gap by simply having more time? The authors test max_iterations = 100 vs. 200. The result is clear: doubling the iteration budget produces marginal gains and, in some cases, degrades performance (Claude Sonnet 4.5’s score on PaperBench decreases with 200 iterations). CAID’s gains on the same reference budget are systematically far superior.
The “fallback” strategy is ineffective
A practical strategy would be to launch a single agent first, then switch to multi-agent if it fails. The data shows that this approach accumulates the costs of both configurations without proportional gain. On PaperBench with Claude Sonnet 4.5: the sequential approach reaches 66.8% but multiplies execution time (3884s vs. 2080s) and cost ($12.6 vs. $9.3), for only 3.5 additional points over direct CAID. On Commit0-Lite with MiniMax 2.5, the final score is identical (57.0%) in both cases, but the total cost doubles.
When to use CAID versus the alternatives
| Context | Recommendation | Why |
|---|---|---|
| Isolated software task, estimated < 1 h | Single agent | Coordination is pure overhead, with no parallelism benefit |
| Long-horizon task (Commit0-style), 4 independent modules | CAID, 4 engineers + worktrees | +6 points over a single agent, net gain once coordination is amortised |
| Task with strong coupling between files | Single agent with a high iteration budget | CAID degrades — integration conflicts outweigh parallelism gains (the N=8 simpy case) |
| Paper reproduction (PaperBench) | CAID, 4 engineers with manager review | +6 points, maximum robustness, justifies the coordination overhead |
| Soft isolation (verbal instructions) | To be avoided | Performance below a single agent — physical coordination is necessary |
Factors that make or break coordination
The number of engineers is not monotone
More agents does not imply better performance. On Commit0-Lite, 4 engineers improve the score over 2, but 8 engineers degrade it. On PaperBench, going from 2 to 4 or 8 engineers brings no measurable gain while increasing time and cost.
The problem with 8 engineers, illustrated on the simpy repository, is explicit: the manager assigns different functions from the same events.py file to multiple engineers. The modifications are logically separable, but create file-level integration conflicts. The pass rate drops from 92.1% (N=4) to 44.3% (N=8). The lesson: the number of agents must match the intrinsic modularity of the task, not simply be maximized.
Delegation determines the trajectory
Two CAID runs on the same repository (minitorch), with different prompts, produce 8.7% and 34.3% pass rates. The difference comes down to a single file: autodiff.py, critical for dependent tests. Run 2 assigns it explicitly; run 1 never assigns it. The manager’s ability to identify high-impact dependencies is decisive.
Verification vs. efficiency: a measurable trade-off
Three prompt configurations are compared on a subset of Commit0-Lite:
- Manager review each round: 60.2% pass rate, 3689s runtime
- Self-verification by the engineer: 55.1%, 2244s
- Efficiency-first mode: 54.0%, 1909s
There is a clear trade-off between integration robustness and execution speed. Self-verification is a reasonable balance; systematic manager review improves quality but significantly burdens coordination.
Limits
The API cost of CAID is systematically higher than the single agent, and total execution time does not decrease proportionally to parallelization — integration remains sequential and gated on tests. Delegation relies on prompt heuristics, not learned policies: an incorrect decomposition produces locally correct but globally difficult-to-integrate outputs. The framework remains specific to software engineering; extending it to other domains (document synthesis, planning) would require redefining the isolation and verification mechanisms.
Key takeaways
- The branch-and-merge pattern, implemented via
git worktreeandgit merge, is the central primitive that makes multi-agent collaboration on shared artifacts reliable. - Physical isolation (separate worktrees) systematically outperforms isolation by verbal instruction, especially when dependencies are implicit.
- Doubling a single agent’s iteration budget does not compensate for the structural limit of sequential processing on long-horizon tasks.
- The optimal number of agents is constrained by the modularity of the task and the manager’s delegation capacity — beyond that, integration conflicts cancel out parallelism gains.
- The “single agent first, multi-agent as fallback” strategy is dominated by direct multi-agent execution on both cost and time.