In short

When building a system with multiple LLM agents, the instinct is to want the best model possible. Choose GPT-4o over GPT-4o-mini, Opus over Sonnet. More powerful = more reliable, right?

The data says otherwise. What determines the reliability of a multi-agent system is not the model used — it is how the agents are organized, coordinated, and constrained. Architecture first. Model, second.


Explanation

What we mean by “architecture”

A multi-agent system consists of several LLMs working together: one decomposes a task, others execute it in parallel, a final one consolidates the results. Each agent has a role, tools, and a limited scope.

“Architecture” refers here to the set of structural decisions: how agents share the work, how they communicate, in what order they execute, how errors are handled and reported. It is not the model’s code — it is how the pieces fit together.

Why architecture dominates

A more powerful LLM generates more precise text. But a multi-agent system rarely fails because of a badly predicted token. It fails because an agent receives a truncated context, because a dependency between tasks is not respected, because a badly formatted intermediate result breaks the next step in the chain, because two agents write to the same file without coordination.

These failures are structural. They do not disappear by upgrading to the next model.

The MAESTRO research (arXiv 2601.00481) measured this empirically across 12 multi-agent systems: system architecture is the dominant factor for reproducibility, cost, and accuracy — ahead of model choice. The experiments show that performance varies significantly from run to run on the same system, and that this variance is explained by structural choices, not by the raw capabilities of the model.

The capability gain paradox

Princeton University published a study in February 2026 evaluating 14 models on two benchmarks with 12 reliability metrics (consistency, robustness, predictability, safety). Result: recent capability gains produced only modest reliability improvements (“Towards a Science of AI Agent Reliability”, arXiv 2602.16666).

In other words: models are becoming more capable, but not proportionally more reliable. Reliability depends on something else — and that “something else” is architecture.

The root cause of failures

A third study, “Characterizing Faults in Agentic AI” (arXiv 2603.06847), analyzed 13,602 issues and pull requests on real agentic systems. Its conclusion: the majority of failures stem from a mismatch between probabilistically generated outputs and deterministic interface constraints.

Plainly stated: the model generates probabilistic text (it never produces exactly the same output twice), but the rest of the system expects deterministic inputs (valid JSON, an exact filename, a response in a precise format). When the interface between the two is poorly designed, the system breaks. Regularly. And a more powerful model does not solve this problem — an architecture with clear validations and contracts does.

Dispatch by cognitive type, not by model

A practical consequence of this principle: it is better to choose a model based on what an agent must do cognitively, not based on the number of files it has to read or the perceived “complexity” of the task.

Our tests show that the discriminating factor between models is not quantitative. It is not “lots of data → big model”. It is: “mechanical extraction → lightweight model is sufficient”, “critical judgment, nuance, contradiction detection → more analytical model required”. An agent tasked with listing facts from a document does not need the same model as an agent tasked with designing an experiment protocol.


Concrete examples

Sonnet ≈ Opus on analytical tasks

Our tests (30 comparative runs, 10 tasks of increasing difficulty) detected no observable quality difference between Sonnet and Opus on analytical tasks: extraction, summarization, source comparison, contradiction detection, protocol design. On the most complex task, Opus used 3 times more tool calls than Sonnet (15 vs 5), took 42% longer (195s vs 138s), and produced a deliverable of not significantly different quality.

This is not a criticism of Opus — it is confirmation that for this range of analytical tasks, the limiting factor is not model capability. It is prompt quality, scope clarity, and the structure of the input data.

Haiku holds further than expected

On extraction and summarization tasks (T1 to T5 in our tests), Haiku produces usable results 100% of the time. It continues to produce usable — but minimal — results up to the most difficult task (T10, protocol design across 5 files). The threshold is not where we expected it.

The difference between Haiku and Sonnet is not “works / does not work”. It is a gradient of depth: Haiku is correct and minimal, Sonnet is correct and nuanced. For a data collection agent, Haiku is sufficient. For an agent tasked with identifying methodological tensions, it is no longer enough.

Mistral 7B: the limit was not capability

In a comparative documentary research test, Mistral 7B local (via Ollama) achieved 64% coverage versus 100% for Sonnet and Opus. That is not the most revealing gap. The most revealing: Mistral ignored format constraints in 2 out of 3 cases (expected sections, Markdown structure). The outputs were unusable not because the model was “less intelligent”, but because it did not follow structural instructions.

The main problem was not factual quality — it was instruction-following. An architectural constraint (validate the output format before passing to the next agent) would have detected the failure immediately.

12 parallel agents without degradation

Our tests show that 12 Sonnet agents can run in parallel without rate limits or quality degradation (16 seconds total). This is not trivial: it means that a well-architected system can massively parallelize without needing more powerful models. The performance comes from the dispatch architecture, not from the individual model.


Key takeaways

  • Architecture determines reliability, not the model. Three independent research teams (MAESTRO, Princeton, Berkeley/CMU) arrive at the same conclusion through different methods. Our own observations point in the same direction.
  • Choosing a more powerful model does not fix an architectural problem. If agents receive malformed contexts, if task dependencies are not respected, if errors are not detected and reported — a better model changes none of that.
  • Dispatch by cognitive type is more effective than dispatch by “complexity”. The question is not “is this file large?” but “does this agent need to exercise critical judgment?”. Mechanical extraction → lightweight model. Nuanced analysis → analytical model. The rest is over-engineering.
  • Cost savings are real and immediate. On common analytical tasks, Sonnet produces the same quality as Opus at lower cost. Haiku is sufficient for extraction. The compute savings do not come from the model chosen — they come from the precision of the dispatch.