In short
Across 15 multi-agent orchestration deployments and 174 agents launched in parallel, the tests recorded a single incident — classified as “partial”, not “failed”. The reported success rate is 99.4%. Before drawing conclusions about the robustness of these systems, three questions deserve an answer: who defines success? what is actually measured? and what is left out?
In plain terms : 99.4% success across 174 agents sounds like a chain of 174 links that almost perfectly holds. But the chain’s solidity depends as much on what is measured as on what holds — and here, “holds” means “delivered something”, not “delivered something correct”.
What the 99.4% figure really says
A multi-agent orchestration system decomposes a complex task into subtasks assigned to independent agents running in parallel. Under this protocol, each agent is given a bounded scope, produces a textual deliverable, and the orchestrator consolidates the results.
Out of 174 individual agent runs, 173 produced a deliverable deemed acceptable. One was marked “partial” — the agent had returned an incomplete but non-empty output — and none was declared a failure. The 99.4% rate is therefore a declared completion rate, not a quality rate.
The distinction matters: an agent that generates plausible but factually wrong text is counted as a success. An agent whose network access was blocked mid-execution — a case observed during a collection test — is counted as a success too, as long as its output remains usable in the operator’s eyes.
In plain terms : the 99.4% is a delivery rate, not a quality rate. Like a courier who drops off the package — without guaranteeing that the contents match what was ordered.
The declaration bias: the operator judges their own system
In these tests, the operator is the one evaluating each agent: success, partial, or failure. There is no third-party judge, no uniformly applied automated criterion, no independent audit of deliverables.
This type of bias is well documented in the literature on automated-system evaluation [NOT VERIFIED in this specific context]. The operator who designed the system is naturally inclined to interpret ambiguous cases in favor of success. A deliverable “sufficient for the intended task” and a deliverable “of reproducible quality” are two different things.
This dynamic is not unique to these tests. In academic LLM benchmarks, authors who define and measure the success criterion themselves consistently report better results than those submitted to external evaluation — the documented average gap reaches several percentage points [NOT VERIFIED]. Rigor here would consist of defining what counts as a failure in writing before the run, then sticking to it even when the output is “almost good enough”.
In the studied case: the single PARTIAL agent was observed during a test processing 123 sources. The line between PARTIAL and success could have been drawn differently — and the overall rate would have shifted. If two additional agents had been classified as PARTIAL across all tests, the headline “174 agents, 0 failures” would become “174 agents, 3 partials”.
In plain terms : when the referee is also a player, the line between “good” and “almost good” shifts in the convenient direction. It is not cheating — it is a structural bias no operator spontaneously avoids.
What the log does not measure
The test registry captures process metrics: number of agents launched, completion status, approximate duration. It does not capture:
- The intrinsic quality of outputs. An agent can produce text that is structured, coherent, and wrong. No systematic factual verification is included in the tracking.
- Hallucinations. If an agent invents a source, misquotes a fact, or extrapolates beyond its data, none of that appears in the reliability metrics.
- Silent degradations. An output that answers the prompt but drops 30% of the relevant information is counted as a success.
- Context incidents. During one test, several agents encountered network-access blocks. The incident is mentioned in notes, not in the failure counter.
These blind spots are structural: they follow from having defined “reliability” as “the agent produced something” rather than as “the agent produced something correct”.
In plain terms : the counter records the presence of a deliverable, not its accuracy. It is counting the letters dropped in the mailbox without ever verifying they contain the right message.
What a rigorous measurement should include
Existing benchmarks for agentic systems — AgentBench, WebArena, SWE-bench — share one trait: they define success through an external criterion that can be verified automatically. Did the agent accomplish the task as an independent oracle can confirm? Not: does the operator find the output satisfactory?
Transposed to multi-agent orchestration, this would require at a minimum:
- A written success criterion before the run, not adjusted after the fact.
- Automated or third-party verification on a sample of deliverables (factual accuracy, coverage, internal consistency).
- Tracking of quality degradations, even when the agent formally “completes” its task.
- Traceability of non-blocking incidents — interrupted network access, silent truncation, unavailable sources — which do not show up in the failure counter but affect final quality.
None of these elements are exotic. They represent standard practice for any serious regression test. Their absence from the available metrics does not invalidate the observed results, but it circumscribes their scope.
In plain terms : a serious measurement requires an external referee, a criterion written before the match, and a record of the plays that do not score but wear the team down. Here, all three are missing.
The conditions that favor success
The tests were conducted under controlled conditions: well-bounded tasks, agents with narrow scope, textual deliverables requiring no external verification. This setup mechanically favors high rates.
The observed scale-ups — up to 22 agents in parallel — produced no declared failures. But they were also not tested on high-error-risk tasks: no agent was responsible for a critical decision, an irreversible action, or an output subject to independent external validation.
The robustness of an orchestration system in less controlled environments — unstable APIs, ambiguous data, objective success criteria — remains an open question that these figures do not settle.
Another factor to consider: the selection of tasks tested. The documented deployments targeted cases where multi-agent orchestration yields an obvious gain. Hard cases — tasks with strong interdependence between agents, deliverables requiring global consistency, real-time constraints — were not included in the tests documented here. The 99.4% rate is therefore also a reflection of the chosen perimeter, not just of system robustness.
In plain terms : testing on flat terrain does not prove a vehicle can handle mountains. The 99.4% describes what was attempted, not what could be.
What to remember
- The 99.4% rate measures completions declared by the operator, not the quality or correctness of the outputs produced.
- The only incident classified as “partial” could have been classified as “failure” under a stricter definition — the overall rate depends on the line that was drawn.
- Hallucinations, silent degradations, and network-access incidents are not integrated into the reliability metrics.
- Scaling to 22 parallel agents without a declared failure is a positive signal, but obtained on tasks with low risk of verifiable error.
- Before generalizing these results, what would be needed: a third-party judge, objective success criteria, and tests on tasks with external verification.