In short
In 2025, a team from UC Berkeley analyzed 1,642 execution traces of multi-agent LLM systems and extracted a structured taxonomy: MAST (Multi-Agent System Failure Taxonomy). It identifies 14 failure modes grouped into three categories — system design, inter-agent misalignment, and task verification. These modes are not rare edge cases: they occur regularly in production, sometimes causing unanticipated operational costs (ZenML documents one case at $47,000 over eleven days). Understanding this taxonomy means having a shared vocabulary for diagnosing, correcting, and preventing agent failures.
In short: think of MAST as the medical nomenclature of agent failures. Before it, you said “it doesn’t work” — like a doctor saying “they’re sick”. With it, you can name precisely FM-1.1 (the agent ignores constraints), FM-2.6 (correct reasoning but divergent action), FM-3.3 (incorrect verification). Diagnosis becomes possible, and so does treatment.
The MAST taxonomy — a systematic classification
A multi-agent LLM system (MAS) is a set of autonomous language models that collaborate to accomplish a task: one agent plans, another executes code, a third verifies the results. This type of architecture is increasingly common in automation applications, but the failure mechanisms specific to it remained poorly documented.
Cemri, Pan, Yang et al. (2025) fill this gap with MAST. The method is grounded in Grounded Theory: researchers annotated 150 execution traces without predefined categories, allowing patterns to emerge from the data. Inter-annotator agreement reached kappa = 0.88 (Cohen’s kappa), considered strong agreement in information science. The final corpus — MAST-Data — contains 1,642 annotated traces from seven open-source frameworks (MetaGPT, ChatDev, HyperAgent, OpenManus, AppWorld, Magentic, AG2) on coding, mathematics, and general agent tasks.
The taxonomy identifies 14 failure modes, distributed across three categories:
| Category | Observed frequency |
|---|---|
| FC1 — System design issues | 41.8% |
| FC2 — Inter-agent misalignment | 36.9% |
| FC3 — Verification and termination | 21.3% |
To make this taxonomy operational at scale, the authors developed an automated annotation pipeline (LLM-as-a-Judge) that achieves 94% precision and kappa = 0.77 compared to expert human annotations. MAST, MAST-Data, and the annotator are published as open source on GitHub.
An important finding: comparative analysis across four models (GPT-4, Claude 3, Qwen2.5, CodeLlama) shows that performance variations exist but do not follow a simple size-based logic. On MetaGPT with programming tasks, GPT-4o generates 39% fewer FC1-type errors than Claude 3.7 Sonnet in the analyzed traces. More broadly, the authors conclude that these failures are indicative of fundamental design flaws in MAS architectures, not merely artifacts specific to the tested frameworks.
Category 1: system design issues (41.8%)
The first category groups failures whose root cause lies in how the system and its agents are designed: task specifications, roles, context management. It represents more than 40% of observed failures, making it the most frequent source of problems.
FM-1.1 — Non-adherence to task specifications (15.7%)
This is the most frequent failure mode across all categories. An agent receives a task with explicit constraints — output format, functional scope, limits — and partially or completely ignores them. The result is produced, but it does not match what was requested.
Concrete example: an agent tasked with generating a CSV report with precisely named columns produces a JSON file. The task is “completed” from the agent’s perspective, but the result is unusable.
This mode points to a specification or constraint-transmission problem. The diagnostic question: is the task prompt precise enough that a human reader could not interpret it differently?
FM-1.2 — Non-adherence to role specifications (11.8%)
Each agent in a MAS is assigned a role: planner, executor, verifier, and so on. This failure mode occurs when an agent overflows its assigned role — for example, a “planner” agent that begins directly executing actions instead of delegating, or a “verifier” agent that modifies the result rather than simply evaluating it.
The problem is often subtle: the agent produces something useful, but in a register that is not its own. This creates responsibility conflicts between agents and makes debugging difficult, since it is no longer clear which agent is responsible for which decision.
The diagnostic question: do the role instructions for each agent explicitly define what the agent should not do?
FM-1.3 — Step repetition (13.2%)
An agent re-executes an already-completed step without valid reason. This mode is particularly costly in tokens and time, and can mask the absence of a properly implemented working memory. It appears frequently in combination with other modes: HyperAgent notably suffers from FM-1.3 combined with FM-3.3 (incorrect verification), creating a loop where the agent verifies incorrectly, fails, and restarts indefinitely.
The diagnostic question: does your system maintain an explicit registry of completed steps, accessible to the agent before it plans the next action?
FM-1.4 — Conversation history loss (1.5%)
This mode occurs when the context window is unexpectedly truncated: the agent reverts to a prior conversational state, as if recent exchanges had not taken place. Its raw frequency in the MAST corpus is low, but its impact can be severe: decisions made before the truncation are invisible to the agent, which may re-engage an already-resolved subtask or contradict previously validated decisions.
The ZenML analysis (2025) specifies that “context rot” — the progressive degradation of contextual coherence — begins between 50,000 and 150,000 tokens in context windows, regardless of the theoretical limits declared by providers.
The diagnostic question: does your system have a compression or summarization mechanism for history before reaching context limits?
FM-1.5 — Unawareness of termination conditions (1.9%)
An agent unable to recognize when a task is truly complete. It continues producing outputs, calling tools, soliciting other agents — even when the objective has been reached. This mode can generate infinite loops with the associated costs. The case documented by ZenML (two agents in a recursive loop for eleven days) illustrates what the absence of an explicit termination criterion can cost.
The diagnostic question: does each task in your system have a success criterion verifiable by the agent itself, without human intervention?
Category 2: inter-agent misalignment (36.9%)
The second category concerns failures that emerge from interactions between agents. A single agent may function correctly — it is in collaboration that problems appear. This category covers six distinct modes.
FM-2.1 — Conversation reset
A conversation between agents is unintentionally reset: the established collaboration context is lost. Agents resume their exchange from a prior state, potentially with contradictions relative to decisions already made. This mode is particularly insidious because it can go unnoticed if the final result is superficially correct.
FM-2.2 — Failure to request clarification
An agent receives an ambiguous instruction and, instead of requesting clarification, chooses an interpretation — often without signaling that an ambiguity exists. The silent decision propagates through the entire pipeline. This mode is documented in the qualitative analysis by Abdin et al. (2024) under the term “premature action without grounding”: the agent acts before sufficiently anchoring its understanding of the task.
The diagnostic question: are your agents explicitly instructed to signal and block when facing ambiguity, rather than choosing a default interpretation?
FM-2.3 — Task derailment (7.4%)
An agent deviates from the initial objective. It produces irrelevant or even counterproductive actions that no longer have any connection to the assigned task. This mode differs from FM-1.1 (non-adherence to specifications) in that it emerges dynamically during execution — often triggered by poorly managed context or by the outputs of other agents. The agent does not ignore initial constraints; it gradually loses sight of them.
FM-2.4 — Information retention (0.85%)
An agent possesses information relevant to other agents but does not share it. This can occur due to the absence of an explicit sharing mechanism, or because the agent incorrectly assesses the relevance of the information to its neighbors. Information retention is difficult to detect: what is not transmitted does not appear in traces.
The diagnostic question: does your architecture explicitly define which information each agent must report to the orchestrator or other agents?
FM-2.5 — Ignoring other agents’ inputs (1.9%)
An agent receives contributions or recommendations from another agent and does not account for them in its decision. This mode may produce results that appear coherent but ignore constraints or knowledge brought by the rest of the system. It is particularly problematic in architectures where a specialized agent (domain expert) is supposed to enrich the decisions of a generalist agent.
FM-2.6 — Reasoning-action divergence (13.2%)
This mode is the second most frequent in FC2. An agent expresses correct reasoning — it identifies the right plan, the right decision — but its actual action does not correspond to that reasoning. The model “says” one thing and “does” another. This discrepancy is particularly difficult to diagnose because the reasoning traces (chain-of-thought) appear correct, masking the actual failure.
Abdin et al. (2024) document a related phenomenon: “over-helpfulness,” where an agent substitutes missing entities (default values, invented parameters) without signaling the approximation, producing a superficially valid but substantively incorrect action.
The diagnostic question: does your system verify that the actions actually executed correspond to the intentions expressed in the reasoning trace?
Category 3: verification and termination (21.3%)
The third category groups failures related to quality assessment and task closure. These modes occur at the end of the pipeline — or precisely because the end of the pipeline is poorly managed.
FM-3.1 — Premature termination
An agent stops execution before the task is truly complete. It declares success — or simply ceases to act — while necessary steps remain to be accomplished. The AppWorld framework is particularly affected by this mode in the MAST traces, likely due to its star topology without a predefined workflow: without an explicit structure indicating the remaining steps, the agent has no reference for assessing its progress.
This mode is the opposite of FM-1.5 (inability to stop): the problem here is stopping too early rather than too late.
FM-3.2 — Absent or incomplete verification
No mechanism guarantees the accuracy, completeness, and reliability of the produced result. The agent completes the task and transmits the result without validation. MAST shows that frameworks with explicit verifiers (MetaGPT, ChatDev) generate fewer total failures than those without. But the presence of a verifier is not sufficient: the authors document a case where ChatDev generates a program that passes superficial verification but contains runtime bugs, because the verifier does not validate against the actual rules of the task.
The diagnostic question: are your verification criteria aligned with the actual constraints of the task, or only with surface heuristics (correct syntax, valid format)?
FM-3.3 — Incorrect verification
The verification mechanism exists, but it incorrectly validates an incorrect result. This is the most difficult mode to detect: the system reports success, logs indicate that verifications passed, and yet the result is wrong. HyperAgent particularly suffers from this mode in combination with FM-1.3: the agent verifies incorrectly, concludes erroneously that the result is wrong, and repeats the step — creating a loop.
IBM Research and UC Berkeley applied MAST to IT-Bench (NeurIPS 2025), a benchmark of real IT automation tasks. Result: agents solve only 11.4% of SRE (Site Reliability Engineering) scenarios, 25.2% of CISO scenarios, and 25.8% of FinOps scenarios. These figures concretely illustrate the cumulative impact of FC3 modes under near-real conditions.
The practitioner’s checklist
The 14 MAST modes translate into operational questions. Here is a diagnostic checklist to go through before deploying a multi-agent system, or when diagnosing a failure.
Task design (FC1)
- Are the constraints for each task formulated unambiguously, with positive and negative examples when necessary? (FM-1.1)
- Do the role instructions for each agent explicitly define responsibilities and limits — what the agent must not do? (FM-1.2)
- Does the system maintain an explicit registry of completed steps, visible to each agent before planning the next action? (FM-1.3)
- Is there a summarization or compression mechanism for context before reaching window limits? (FM-1.4)
- Does each task have a success criterion verifiable by the agent itself, without human intervention? (FM-1.5)
Inter-agent collaboration (FC2)
- Are agents explicitly instructed to block and signal when facing ambiguity, rather than choosing a default interpretation? (FM-2.2)
- Does the architecture explicitly define which information each agent must transmit to others? (FM-2.4)
- Does a mechanism verify that the actions actually executed correspond to the reasoning expressed in the traces? (FM-2.6)
- Do specialized agents have a structured communication channel so that their contributions are accounted for? (FM-2.5)
Verification and closure (FC3)
- Does each task have an explicit definition of completeness, accessible to the agent during execution? (FM-3.1)
- Are verification criteria aligned with the actual constraints of the task, not just with surface heuristics? (FM-3.2)
- Does the system have regression tests for cases where verification might incorrectly conclude success? (FM-3.3)
Limitations to keep in mind
MAST is built on seven open-source frameworks. Its representativeness for proprietary enterprise systems remains to be established. The documented interventions — prompt engineering, addition of verifiers — improve performance by 14 to 15.6% according to available studies, but remain insufficient for autonomous production deployment without supervision. The taxonomy itself could evolve with new architectures (extended reasoning, persistent memory, multimodal tools). Finally, MAST identifies categories of failure but not always the responsible agent in a distributed system — a causal attribution problem that research addresses separately (arXiv:2603.25001).
In short: the MAST checklist is a prevention tool, not a guarantee. Out of 14 modes, two dominate: FM-1.1 (15.7%) and FM-2.6 (13.2%). If your prompt and architecture handle those two, you cover nearly 30% of structural failures. The rest is incremental work — gradually covering the modes that surface in your specific domain. No team ships a system at 14/14, but 8-10/14 is already significantly more stable than the industry average.
Key takeaways
- 14 modes, 3 categories. MAST is the first systematic taxonomy of failures in multi-agent LLM systems, built on 1,642 annotated traces with kappa = 0.88. It covers system design (41.8%), inter-agent misalignment (36.9%), and task verification (21.3%).
- Design problems dominate. FM-1.1 (non-adherence to task specifications, 15.7%) and FM-2.6 (reasoning-action divergence, 13.2%) are the most frequent modes. These are architecture and specification problems, not model bugs.
- Verification is not a guarantee. Frameworks with explicit verifiers perform better on average, but FM-3.3 shows that a verifier can validate an incorrect result. The quality of verification criteria matters as much as their presence.
- Current interventions are insufficient. Prompt engineering and the addition of verifiers improve performance by 14 to 15.6% in available studies, but success rates under real conditions remain low (11.4% on SRE tasks in IT-Bench).
- MAST is a diagnostic tool, not a guarantee. The taxonomy makes it possible to name and anticipate failures. It does not resolve the broader question: at what performance threshold is a multi-agent system acceptable for autonomous production deployment?