In short

In a multi-agent pipeline, sub-agents fail. Sometimes loudly, sometimes silently. Our tests across several distinct architectures — isolated agents, coordinated teams, peer-to-peer communication — surface five distinct failure modes. For each one: how to recognize it, and what to do concretely.

In short: think of a team leader who delegates to interns without setting up a reporting procedure. When an intern gets stuck, forgets an instruction, or misses the point, the leader does not know — they discover the problem at final delivery, or never. The failure modes listed here describe the five most common ways for a sub-agent to fail silently.


Why sub-agents fail differently

A sub-agent is not a script. It is a language model that receives a task, calls tools, and returns a result as free text. This chain introduces specific failure points that are absent from a classic API call.

Another particularity: the parent agent sees only the sub-agent’s final message. Not the intermediate steps, not the tools called, not the hesitations. If something goes wrong mid-execution, the parent must infer it from the received text — not from a structured error code.

Our observations cover three distinct configurations: isolated agents launched in parallel, teams of agents with a shared task list, and direct agent-to-agent communication via mailbox. Failure modes vary by configuration, but some are cross-cutting.


Mode 1 — The agent loops indefinitely

What happens. The agent receives a contradictory or underspecified task. It attempts an action, notices the result does not satisfy the expected condition, retries, and repeats. Without a stop mechanism, it can spin indefinitely.

Concrete example. Our tests submitted a structurally impossible task to an agent: simultaneously satisfying two mutually exclusive conditions. The agent ran three iterations (A → B → A → detection) before stopping. It identified the logical contradiction itself and declared the impossibility. This behavior is not guaranteed — it depends on the model and the instructions provided.

How to detect it. Abnormally long execution time is the first signal. In architectures with communication, the absence of a completion message after a reasonable delay indicates a problem. Without communication, the parent receives nothing — total silence.

Mitigation. The only configurable safeguard is in the instructions given to the agent: explicitly specify “if you attempt the same action more than twice without a different result, stop and report the deadlock.” The API offers no maxTurns parameter — our tests confirm this on the tested version. Without an explicit stop instruction in the prompt, looping behavior depends on an internal system timeout that is undocumented and not configurable.


Mode 2 — The silent error (fire-and-forget)

What happens. The agent sends a message to a recipient that does not exist, or is no longer active. The send returns success: true. The message disappears. No error is raised.

Concrete example. Testing inter-agent communication, we sent a message to a nonexistent agent (SendMessage to an unregistered name). Result: success: true, "Message sent to inbox". No error. The message was never processed.

How to detect it. The absence of a response or acknowledgment is the only signal. In an architecture without explicit confirmation, this error is not automatically detectable. You need either a parent-side timeout (“if no response within X seconds, consider the send lost”), or an application-level acknowledgment mechanism.

Mitigation. Design exchanges as fire-and-forget with confirmation: the destination sub-agent systematically responds “received” at the start of its execution. The parent waits for this confirmation before continuing. Without it, after a delay, the parent retries or declares failure.


Mode 3 — The error reported as free text (non-parseable)

What happens. The sub-agent encounters an error — missing file, permission denied, impossible task. It reports this to the parent. But this signal is natural-language prose written by the model, not a structured error code. The parent must interpret narrative text to decide what to do.

Concrete example. During an error-handling test, an agent tasked with reading a nonexistent file received the tool’s error message ("File does not exist") and relayed it to the parent in a narrative summary: “I could not read the file because it does not exist. Here is the exact error message: […]”. The parent receives prose, not an object {status: "error", type: "file_not_found", path: "..."}.

In architectures using mailbox communication, the same behavior appears: the failing agent sends a text message describing the situation. The protocol provides no error message type — the sub-agent chooses how to report.

How to detect it. Look for text patterns in received messages: “failed”, “error”, “impossible”, “does not exist”, “permission denied”. Fragile but functional if agents are instructed to use conventional keywords.

Mitigation. Explicitly request a structured return format in the prompt: “If you encounter an error, begin your message with FAILURE: followed by the error type, then describe the situation.” This does not solve the root problem — the model can always deviate — but significantly reduces ambiguity.

In short: sub-agents never return a structured error object (like {status: "error", code: 404}). They return natural text describing what happened. The parent must read and interpret that text. To reduce ambiguity, imposing an agreed prefix (FAILURE:) turns a message into a detectable signal.


Mode 4 — Context loss on restart

What happens. A sub-agent is stopped then restarted to continue a task. The context (conversation history) is partially preserved, but the model used at restart may differ from the one used at initial launch.

Concrete example. Our tests on agent resumption via identifier revealed unexpected behavior: an agent initially launched with a lightweight model is resumed by the parent using the parent’s model (more powerful). The conversation context is correctly restored, but the model identity is not. This mid-execution model switch is silent — nothing indicates it in the interface.

Why this is a problem. A pipeline that assumes behavioral consistency between phase 1 and phase 2 of an agent may receive stylistically or qualitatively different outputs. Regression tests on this type of pipeline are complex to interpret.

How to detect it. Explicitly verify the active model after each resumption, if the API exposes it. Document the model used in intermediate outputs.

Mitigation. If model continuity is critical, avoid agent resumption and prefer relaunching a fresh agent with the history encoded in the prompt. Higher cost, more predictable behavior.


Mode 5 — Unblocking without wakeup

What happens. In an architecture with task dependencies, a blocked task “unblocks” automatically when its prerequisites are satisfied — but the agent in charge of that task does not automatically resume. It remains waiting until an explicit signal is sent.

Concrete example. In a test with three coordinated agents, the agent responsible for synthesis was waiting for completion of the two upstream agents. When those finished, the dependencies resolved in the task list. But the synthesis agent remained idle. An explicit message had to be sent (“your prerequisites are satisfied, you may proceed”) for it to resume.

The documentation suggests that unblocking is automatic. Our observations show that task unblocking and agent wakeup are two distinct events: the first is automatic, the second is not.

How to detect it. An agent in a waiting state that produces nothing after its dependencies are lifted. In an architecture with notifications, it will send repeated “available” signals without progressing.

Mitigation. After each upstream task completion, the parent (or coordinator) explicitly checks which agents are unblocked and sends each a start signal. Do not assume that dependency resolution triggers automatic wakeup.


Summary table

Failure modeSignal at parentAuto-detectablePrimary mitigation
Infinite loopSilence or excessive delayDifficult (timeout)Stop instructions in the prompt
Silent error (fire-and-forget)No responseNo without acknowledgmentExplicit application-level confirmation
Non-parseable returnAmbiguous narrative textPartial (keywords)Return format enforced in the prompt
Context loss on restartSilentNo without inspectionAvoid resumption, relaunch fresh
Unblocking without wakeupNo progressWith active monitoringExplicit parent signal after unblocking

Key takeaways

  • Sub-agents never return structured error codes. Every failure arrives as free text. The parent interprets — it does not parse.
  • The absence of a return (silence) is the hardest failure signal to detect. A looping agent and a lost message produce the same symptom at the parent level: nothing.
  • Coordination mechanisms (dependencies, blocked tasks) operate at the task-list level, not at the agent level. An agent does not wake itself up when its task unblocks.
  • The primary mitigation is in the instructions given to the agent: return format, stop rule, receipt confirmation. What the prompt does not specify, the model improvises.
  • A robust pipeline treats the absence of a response as an error, not as a silent success.