In short

Red teaming is the practice of deliberately attacking a system to uncover its vulnerabilities before an adversary does. When the target is an isolated large language model, red teaming focuses on the prompt — jailbreaks, injections, alignment bypasses. When the target becomes an agentic system (an LLM that calls tools, talks to other agents, reads a persistent memory and rests on a chain of software dependencies), the attack surface changes in nature: the LLM is no longer the main target, merely one of the components to protect. A paper published in May 2026 by Dreadnode (Dheekonda et al., arXiv:2605.04019v1) illustrates that shift by measuring an attack success rate (ASR) of 85 % on Llama Scout 17B across 7,727 attempts, with no human-written code. This article maps the new discipline: why it differs from classic red teaming, which attacks it classifies, which normative frameworks are emerging, and how to audit a system in practice.

In plain terms: red teaming a lone LLM is testing a lock; red teaming an agentic system is testing an entire house — lock, windows, cellar, postal worker, delivery drivers and the subcontractors’ key rings.


Understanding the problem: the LLM is no longer alone

Picture a consultant answering questions from an office, in isolation. To fool them, you talk to them: you rephrase the question, you lie about who you are, you wear them down with dilemmas. That is red teaming a solitary LLM — a verbal duel, and a well-studied one.

Now give that consultant a phone, a calendar shared with other consultants, access to a continuously updated client file, and a contract with a dozen external subcontractors. To fool them, you no longer need to talk to them at all: you can booby-trap a file they will read, corrupt a subcontractor who feeds them information, or leave a message in the shared calendar that their colleagues will take as authentic. An agentic system resembles that second consultant.

Vocabulary

  • Agentic system: an LLM embedded in a perceive → reason → act loop, able to call external tools, read from and write to a memory, and possibly coordinate other agents.
  • Attack surface: the set of entry points through which an adversary can influence a system’s behaviour. For an isolated LLM it is essentially the prompt; for an agent it is also every tool, every memory source and every peer.
  • Red teaming: a structured offensive practice that simulates attacks to expose a system’s vulnerabilities. Distinct from fuzzing (random inputs) and from classic penetration testing (which targets infrastructure).
  • Attack success rate (ASR): the percentage of attempts that achieve a defined adversarial goal (making the model produce a forbidden answer, making it perform an unauthorised action).
  • Tool use: an LLM’s ability to call external functions (web search, code execution, database) by emitting structured calls that the environment executes on its behalf.

Three observations make this discipline distinct from classic red teaming. First, vulnerabilities no longer show up only in generated text: an agent can be silently diverted without a single suspicious word appearing in its output. Second, attacks can cross temporal boundaries: a message poisoned today can influence a decision several days later, after a memory rewrite. Third, third-party components are part of the perimeter: a typo-squatted Python package, a tampered JSON schema or a poorly vetted document in a knowledge base can compromise the agent without touching the model at all.

In plain terms: an isolated LLM is compromised through words; an agentic system is also compromised through what it reads, what it calls, who answers it, and what its dependencies install in its environment.


Why red teaming changes when the agent joins the room

Four structural reasons explain the widened perimeter.

Third-party tools multiply attackable handlers. Every tool wired to an agent (through MCP, the Model Context Protocol — a standard interface between LLMs and external functions — or through proprietary tool calling) is an entry point. The JSON schema describing the tool, the code that runs it, and the channel carrying the result are all three manipulable. A loosely constrained tool accepts parameters it should have rejected. A tool replaced by a malicious namesake hands control back silently. The LLM, for its part, simply did its job: it called what it was given.

Persistent memory creates long attack windows. A solitary LLM forgets everything between two conversations. An agent can write to a memory (session summaries, knowledge base, shared state) and read it back later. An attacker who manages to insert a falsehood into that memory plants a time bomb: the next task that consults the area silently inherits the bias. The literature calls this knowledge poisoning — distinct from classic prompt injection because it does not require being present at the moment of the attack.

Inter-agent communication introduces an unauthenticated trust bus. When several agents collaborate they exchange messages: subtasks, partial results, completion signals. If that bus is not cryptographically signed, any process with write access can impersonate a legitimate agent. That is peer spoofing. And even if every agent is healthy, a single adversarial output in an N-agent consolidation can be enough to skew the synthesis — a phenomenon documented as consensus poisoning.

The supply chain extends from software dependencies to models. An agent in production pulls Python packages, Node modules, a base Docker image, and often one or more models downloaded from a public hub (Hugging Face, for instance). Every link is a potential target: slopsquatting (a package hallucinated by an LLM during code generation, which an attacker then publishes on PyPI under the predicted name), dependency confusion (an internal package whose name is taken publicly with a higher version), model merge backdoor (a public model carrying malicious behaviour triggered by a specific token).

An empirical study by Mulla et al. (2025, arXiv:2504.19855) covering 214,271 attack attempts by 1,674 participants across 30 LLM challenges shows that automated approaches reach 69.5 % success against 47.6 % for manual efforts — automation does not replace offensive intelligence, but it widens coverage.

In plain terms: the target diffuses across every component. The LLM is no longer the boss to defeat — it is the central link in a chain whose every other link is independently attackable.


The Dreadnode taxonomy — 5 attack fronts

The Dreadnode paper (Dheekonda, Pearce, Landers, 2026) introduces one of the first operational taxonomies dedicated to agentic red teaming. Out of 38 catalogued adversarial transformation modules, five categories are specific to agentic systems — meaning they would make no sense against an isolated LLM.

Front 1 — MCP and tool attacks

Tool poisoning: a malicious tool description is injected into the catalogue presented to the LLM. The model, which picks its tools by reading their descriptions, is manipulated upstream of the tool call itself. Tool shadowing: a new tool is registered whose name collides with a legitimate one, which is silently replaced. Rug pull: a benign tool is shown at audit time, then the code is swapped at actual execution (a practice known in the crypto ecosystem, transposed to agents). Schema poisoning: the validating JSON schema is left able to accept parameters it should have rejected (unconstrained additionalProperties, over-broad enumerations, any types).

In plain terms: the tool layer is a new attack skin. No need to jailbreak the LLM — it is enough to change what it believes its tools to be.

Front 2 — Multi-agent exploits

Cross-worker prompt infection: a compromised agent slips a malicious sub-prompt into its output; when the consolidator reads that output, it executes the hidden instruction. Peer spoofing: an attacker emits a message that mimics a legitimate agent’s format without having produced any real work — consolidation takes it in as an authentic output. Consensus poisoning: in an orchestration where several agents vote or contribute to a synthesis, a single adversarial node can be enough to bend the conclusion, especially if the consolidation weights contributions naively.

Front 3 — Attacks on reasoning

CoT backdoor: a particular trigger in the prompt bends the model’s chain of thought toward hidden behaviour, without changing outputs on any other input (chain-of-thought being the step-by-step reasoning the model exposes). Reasoning hijack: a false but plausible-looking reasoning step is injected into the context, and the model carries on from it as if it were its own. Reasoning DoS (denial of service): a prompt is built to force the model into generating excessively long reasoning before answering, saturating the context window or the token budget.

Front 4 — Browser agents

When an agent drives a web browser (clicks, typing, DOM reading), it inherits the classic web vulnerabilities plus a few new ones. Visual prompt injection: malicious text hidden inside an image. AI ClickFix: a page designed to make the agent click a trap a human would have avoided. Navigation hijack: a domain impersonating a legitimate site. This front falls outside the scope of most internal orchestration-oriented audits, but it grows in importance as agents gain autonomy over navigation.

Front 5 — Supply chain

Package hallucination: an LLM generates an import for a package that does not exist; an attacker publishes that package under the predicted name, with a malicious payload. Model merge backdoor: a public model merged from several open-weights models can inherit hidden behaviour injected by one of its parents. Dependency confusion: an internal package whose name is also available publicly, at a higher version, can be pulled preferentially by some package managers.

Summary matrix

FrontConcrete exampleDefence difficulty
MCP & tool attacksJSON schema without additionalProperties: falseMedium — runtime validation + manifest hash
Multi-agent exploitsUnsigned completion signal between workersMedium — HMAC + content-addressable hash
Reasoning attacksCoT backdoor triggered by a rare tokenHigh — statistically hard to detect
Browser agentText hidden in an image, AI ClickFixHigh — a restrictive browser sandbox is required
Supply chainLockfile without SHA-256 hashesLow (effort) but high (impact) — unified pinning

Worth noting: the same study reports that multi-turn attacks (Crescendo) and graph-based ones (Graph of Attacks with Pruning, GAP) reach 100 % ASR on Llama Scout, against 96 % for the tree-based Tree of Attacks with Pruning (TAP) method — and that the none transformation (no prompt mutation at all) already reaches 80 % ASR, revealing alignment gaps that require no adversarial sophistication whatsoever to exploit.

In plain terms: the sophistication of the transformations is not the main variable. Many attacks succeed because elementary structural countermeasures are missing — not because the attacker deployed an exotic technique.


OWASP ASI01-10 — the emerging normative framework

OWASP (Open Worldwide Application Security Project) — the non-profit foundation that has maintained the famous OWASP Top 10 of web risks since 2003 — launched its Agentic Security Initiative (ASI) in 2024, a framework dedicated to risks specific to agentic systems. Where the OWASP LLM Top 10 (2025) covers model-centric vulnerabilities (prompt injection, sensitive information disclosure, excessive agency), ASI addresses risks that only appear once the LLM is embedded in an agent architecture.

Of the framework’s ten categories, eight are regularly cited in the accessible literature:

  • ASI01 — Tool Manipulation: compromise of the tool layer (poisoning, shadowing, rug pull, schema poisoning).
  • ASI02 — Knowledge Poisoning: insertion of falsehoods into a knowledge base or persistent memory.
  • ASI03 — Excessive Agency: an agent holding more rights than its task requires (file deletion, financial transactions, system writes).
  • ASI04 — Trust Boundary Violation: unauthorised crossing of boundaries between agents, environments or accounts.
  • ASI05 — Reasoning Hijack: diversion of the reasoning chain by injecting plausible false steps.
  • ASI06 — Memory Poisoning: durable pollution of the context window or of session summaries.
  • ASI09 — Supply Chain: compromise through software dependencies, models or third-party skills.
  • ASI10 — Model Integrity: attacks on the integrity of the model itself (model merge, malicious fine-tuning, weight swapping).

Categories ASI07 and ASI08 complete the framework on aspects that were not detailed in the literature analysed for this article — a normal trait of a young framework still stabilising.

In plain terms: ASI01-10 is the OWASP Top 10 equivalent for agents. It is a shared vocabulary for describing agentic risks without reinventing the words each time.

Two adjacent frameworks complete the picture. MITRE ATLAS (Adversarial Threat Landscape for AI Systems) provides a nomenclature of numbered techniques (AML.T0049 for supply chain compromise, AML.T0051 for prompt injection). NIST AI RMF 1.0 structures governance around three functions — MAP (risk mapping), MEASURE, MANAGE. The three families are not competitors: OWASP ASI names the attack classes, MITRE ATLAS catalogues the precise techniques, NIST AI RMF frames the governance.


An internal audit methodology — the adversarial Octopus

Auditing a real agentic system means mapping its surface, confronting it with an attack taxonomy, and prioritising countermeasures. A reproducible methodology deployed inside the lab’s R&D setting follows this outline.

Decomposition. An orchestrating agent splits the perimeter into disjoint zones (tool layer, multi-agent layer, reasoning layer, supply chain layer). For each zone, a dedicated worker runs a targeted audit in parallel: inventory of entry points, mapping to the Dreadnode taxonomy, mapping to OWASP ASI, mitigation proposals. A final consolidator aggregates the reports, deduplicates findings, prioritises mitigations and flags open tensions.

Anti-convergence through scope diversification. Each worker receives a distinct module to reduce correlation between reports. Without diversification, several instances of the same model converge on the same biases — a documented effect that makes “double validation” by two identical LLMs of little use in security.

Honest acknowledgement of model convergence. Even with disjoint scopes, several workers derived from the same base model share structural biases (sycophancy, similar severity ranking, common vocabulary for describing vectors). Independent triangulation remains indispensable: a human operator, an external reference (paper, CVE, advisory) or a model from another family must confirm or refute the critical findings.

A recent audit conducted this way inventoried 41 surfaces spread across the tool, multi-agent and reasoning-plus-supply-chain layers. Out of 42 logical findings, 39 were consolidated after deduplication. Two structurally critical findings stood out: the absence of cryptographic authentication on inter-agent signals (peer spoofing exploitable as soon as any process holds local write access), and the absence of an append-only mode on long-term knowledge bases (opening the door to persistent knowledge poisoning). Five mitigation families were prioritised and nine architectural tensions were opened for a later session.

In plain terms: an agentic audit is not a classic pentest. It combines surface inventory, taxonomy-based attack classification, and the explicit acknowledgement that the base model’s convergence limits what one can assert alone.


Patterns observed in practice

Three patterns recur often enough to deserve specific attention.

Pattern 1 — Overly permissive schemas. JSON schemas describing tools rarely declare additionalProperties: false at the root and inside nested objects. As a result, a tool call can carry unforeseen parameters that server-side code may read (logs, forwarding to sub-tools, persistence). That is an open door for schema poisoning — the attacker does not need to alter the tool’s code, only to exploit the fields the schema lets through.

Pattern 2 — Unauthenticated inter-agent communication. Many orchestrations exchange plain JSON files (completion signals, intermediate outputs) with no cryptographic signature. A local attacker — or a compromised subprocess — can write a fake completion signal and convince the orchestrator that a worker finished its work with a given result, when no worker ran at all. The canonical countermeasure is an HMAC (Hash-based Message Authentication Code) signature with a shared key stored outside the repository, plus a content-addressable hash of the output file quoted in the signal.

Pattern 3 — Unlocked Python dependencies. requirements.txt files that pin only ranges (>=3.0.0,<4.0.0) rather than exact versions with hashes open the door to several supply chain attacks: slopsquatting (a package hallucinated by an LLM during code generation, then published by an attacker), dependency confusion (accidental priority given to a public package sharing an internal package’s name), silent substitution at the next install.

Order of magnitude for locking: pinning a base Docker image by digest fits on one line (FROM python:3.11-slim@sha256:<digest>). Transitively locking a typical Python project with pip-compile --generate-hashes produces a file of several hundred SHA-256 hashes — one representative R&D project observed internally holds around 1,500. The cost is marginal in effort; the impact is high in security.

In plain terms: three doors are usually left open by default — over-permissive schemas, unsigned signals, unhashed dependencies. Closing them requires no sophisticated cryptography, only disciplined application.


Actionable mitigations

Five families of countermeasures cover the majority of observed findings. None is sufficient alone; the protective effect comes from their combination.

Mitigation 1 — Supply chain locking by hashes. Generate a lockfile with SHA-256 hashes (pip-compile --generate-hashes on the Python side, package-lock.json with integrity on the Node side). Pin the base Docker image by immutable digest. Automatically verify in continuous integration that no package is pulled without a referenced hash.

Mitigation 2 — HMAC signature on inter-agent communication. Every completion signal emitted by a worker carries a signature computed from a master key stored outside the repository. The orchestrator rejects, at read time, any signal whose signature is invalid. The SHA-256 hash of the associated output file is included in the signal so that nobody can swap the output between writing and consolidation.

Mitigation 3 — Extended integrity manifest. A reference file lists every critical artefact (tool schemas, prompt templates, skills, internal conventions) with its SHA-256 hash. An audit script, run at session start or pre-commit, detects any silent drift (untracked modification of a protected file). The pattern has long existed for sensitive configuration files on a public site; it generalises to the long-term reference material of an agentic platform.

Mitigation 4 — Append-only mode and SHA-256 audit log on knowledge bases. Every mutation of a knowledge base generates a signed entry in an append-only journal (before/after hash, author, session). A pre-commit git hook refuses the mutation or deletion of an existing entry unless explicitly instructed. This measure makes silent modification of a past fact impossible — without preventing the addition of new facts or documented requalification.

Mitigation 5 — Targeted defensive hardening. Strict schema validation (additionalProperties: false everywhere), a ban on non-whitelisted shell calls (never shell=True with a command read from external configuration), systematic XML escaping on every variable substitution in a template, non-collision assertions when registering dynamic tools.

Effort × impact matrix

MitigationEffortImpactPractical priority
5 — Targeted defensive hardeningS (low)HighImmediate quick win
1 — Lockfile + Docker digestS (low)HighImmediate quick win
2 — Inter-agent HMACM (medium)CriticalLocalised architecture
3 — Extended manifestM (medium)HighLocalised architecture
4 — Append-only knowledge baseL (large)CriticalWorkflow overhaul

Operating heuristics.

  • Start with the quick wins. Mitigations 1 and 5 close most of the doors left open by default without any architectural change.
  • Do not confuse PARTIAL with complete. A procedural convention (“review rule before commit”) is not a cryptographic countermeasure. The two coexist; neither replaces the other.
  • Document your triangulations. A critical finding sourced only by internal convergence (several LLMs from the same model) is not triangulated — it needs an external reference (CVE, paper, advisory) or a human operator.

Limits and outlook

This article describes a discipline still taking shape, and several limits deserve to be named explicitly.

Static audit vs real offensive scan. The methodology presented rests on code reading, surface inventory and taxonomic classification. It does not reproduce an active exploit: no simulated peer spoofing, no tool poisoning injected into a test environment, no supply chain compromise reproduced. A finding marked critical on the basis of a structural absence of countermeasure is not a demonstrated exploit — it is a design risk. Validation by empirical exploit is a later stage, more costly and riskier to orchestrate internally.

Model convergence. Automated agentic red teaming tends to converge when the workers come from the same base model. Same analytical biases, same severity ranking, same vocabulary — the apparent “double validation” is in reality a disguised correlation. Mitigating that convergence requires either a mix of models from distinct families (Opus + Sonnet + an open-weights model), or systematic external human triangulation on critical findings.

Young frameworks. OWASP ASI01-10, the Dreadnode taxonomy and MITRE ATLAS are frameworks still stabilising. A different breakdown may emerge over the next two or three years, particularly as browser agents and multi-agent architectures distributed across several organisations gain ground.

On the outlook side, several directions are already visible. Continuous red teaming (integrating audits into the CI/CD pipeline rather than running one-off campaigns) is beginning to spread. Automation of supply chain audits through dedicated tools (pip-audit, npm audit, manifest hash scanners) is progressing. And the crystallisation of cross-team conventions (HMAC on signals, append-only knowledge bases, global manifest hash) should make it possible to compare audits between projects without reinventing everything each time.


Conclusion

Agentic red teaming is not a cosmetic extension of classic red teaming. It is a distinct discipline that takes note of a change in nature: the LLM becomes one of the components to protect, and the attack surface extends to everything that speaks to the model, everything the model calls, and everything in the software chain that keeps them running. The Dreadnode taxonomy (5 fronts), the OWASP ASI framework (10 categories) and Octopus-style audit methodologies provide a first reproducible foundation — still young, still convergent, still dependent on human triangulation for its critical verdicts. For a team running agents in production, the return on a handful of quick wins (hash-based lockfile, signed signals, integrity manifest, schema hardening) far exceeds their implementation cost. The rest — model convergence, real offensive scanning, stabilised frameworks — belongs to a research programme that remains open.