In short

The analysis of 1,200 deployments in production (the real environment where software used by end customers runs, as opposed to a laboratory prototype) — source ZenML, 2025 — delivers an unambiguous finding: the experimentation phase is over, the engineering phase begins. Adoption figures vary widely depending on the source — 11% at Deloitte, 57% at LangChain — but this dispersion reveals a definition problem more than a contradiction. The cases that work share one common trait: narrow agents (one agent = one precise task in a bounded domain, not a generalist assistant), single-domain, under human supervision.

In plain terms: putting an LLM agent into production is like opening a restaurant. A prototype is cooking for your friends on a Sunday — bugs are forgiven. Production is serving 500 covers a night, every night, with health inspections, accounts to keep and customers who complain. The demo promise and operational reliability are two different trades.


Why the figures diverge

In plain terms: the figures say the same thing with different rules. A survey that counts “a team has one agent running” will count everyone. A survey that demands “formal governance + broad adoption” will count almost no one. Both are honest, they are measuring two different stages.

57% of organizations have agents in production according to LangChain (n=1,300+). Deloitte (2026) measures only 11%. Both figures are correct — they are not measuring the same thing.

LangChain counts any functional deployment (an agent running for at least one real use case, even within a single team), including limited-scope or single-team ones. Deloitte measures deployment at the organizational scale (widespread adoption with central governance, security policies, audit), with governance and widespread adoption. The gap reveals that many organizations have agents running somewhere, but very few have reached systematic deployment (integrated into the standard workflow, with maintenance, monitoring and operational resilience).

McKinsey (Nov. 2025, n=1,993) refines the picture: 88% of organizations use AI in at least one function, but only 23% are actively deploying an agentic system (a set of LLM agents that chains decisions and tool calls to accomplish a complete task semi-autonomously) at scale. Gartner (Aug. 2025) estimates that fewer than 5% of enterprise applications currently embed agents, with a projection of 40% by end of 2026.

The consensus is not on the numbers. It is on the trend: the transition from prototype to maintained production system is the real challenge.


Six patterns that distinguish teams that ship

In plain terms: the teams that succeed are not the ones with the “best” model, but those that treat the agent as industrial software — with guardrails, tests, monitoring and engineer discipline.

ZenML analyzed 1,200 cases in its LLMOps Database (an open database documenting production LLM deployments; the equivalent of a shared industry feedback loop). Six patterns distinguish teams that actually deploy from those stuck in demo mode.

Technical vocabulary for this section

  • Context engineering: the discipline of choosing which information to inject into the model’s window, when, and in what form. Close to what a librarian does for a researcher: sorting the right sources before reading.
  • Prompt engineering: the fine-grained writing of textual instructions given to the model. Useful, but insufficient at scale if the injected information is wrong to begin with.
  • Token: the elementary unit of text for an LLM, roughly 3–4 characters in English. The “context window” is measured in tokens.
  • Guardrails: automated safeguards that prevent the agent from taking certain actions (wrong tool, excessive cost, missing permission).
  • Circuit breaker: the electrical circuit breaker applied to software. If a counter explodes (too many requests, too much cost), the circuit cuts off automatically before wider damage.
  • Shadow testing: “agent that watches without acting” mode. The agent produces decisions in parallel with the human process, they are compared, adjustments are made, and activation only happens once precision reaches the target.
  • Observability: the capacity to see what the system is doing in real time (which calls, which costs, which errors). Without observability, a bug in production stays invisible.

1. Context engineering over prompt engineering. The degradation of the context window (temporary memory space in which the model reads and writes — it does not grow indefinitely without loss of quality) begins between 50,000 and 150,000 tokens regardless of the model’s theoretical capacity. High-performing teams architect information injection (just-in-time: the model is only given the info it needs now, not the whole file; tool masking: tools not relevant to the current step are hidden) rather than refining prompts.

2. Guardrails in infrastructure, not in instructions. Architectural safety constraints — circuit breakers (software circuit breakers), double-layer permissions (the agent has a requested right + a right actually verified by a second system), session isolation (what one user does does not affect another) — replace prompt-based approaches for critical systems. Saying “be careful” in the prompt is not enough: the guardrail must be in the code that wraps the agent, not in text the model can ignore.

3. Narrow agents under human supervision. Production agents operate as single-domain specialists, not as autonomous entities. Ramp (expense management) handles 65% of approvals autonomously [UNVERIFIED — single source ZenML].

4. Shadow testing before production. Agents are first run in shadow mode (the agent computes a decision in parallel but the human decides; comparison happens after the fact) on real transactions, compared against human decisions, and only activated once the target precision threshold is reached.

5. Evaluation as a standard engineering practice. 89% of teams with production agents have implemented observability (instrumentation that logs every call, every cost, every error — source LangChain, 2025). Only 52% have formal evaluations (reproducible test sets, with quantified metrics and regression comparison). The most mature teams build reference datasets and automated evaluation systems.

6. Software engineering before model selection. ZenML states explicitly: “Software engineering fundamentals — not frontier models (the most powerful models on the market at a given time, typically GPT-4, Claude, Gemini latest generation) — remain the primary predictor of success.” Optimizing the network and infrastructure generates more value than switching to a newer model.


Concrete examples: architectures that work

In plain terms: three real cases, three different approaches, one constant — the scope is bounded and responsibilities are clear. Nobody is trying to make an agent that does everything.

Stripe — one-shot agents with a narrow scope

Stripe’s “Minions” architecture is built on stateless agents (each call is isolated, the agent keeps no memory between two requests — a bit like a form to be filled each time, not an ongoing discussion): one agent, one task, one LLM call. For complex compliance workflows (logical chains of tasks), a directed acyclic graph or DAG (a diagram where tasks chain in a precise order without ever looping back, like a family tree that never loops) chains these specialized agents. Result on fraud detection: precision went from 59% to 97% for large merchants.

Uber — network of agents specialized by stage

Uber migrated code at scale using LangGraph (an open-source framework for orchestrating multiple LLM agents in a graph, with shared state and conditional logic), assigning a distinct agent to each step of the test pipeline (automatic chaining of code verification steps: lint = style analysis, build = compilation, test = running automated tests). The lint agent couples deterministic static analysis (a tool that reads the source code and applies fixed rules, always the same result for the same code) with LLM for ambiguous cases — a so-called “hybrid” pattern (the LLM is used only when the deterministic tool cannot decide) that limits LLM calls to situations where they genuinely add value.

JPMorgan — institutional deployment

JPMorgan’s LLM Suite (internal JPMorgan platform giving access to several LLM models through a unified interface for employees, with security controls and audit) onboarded (integrated into the platform, with training and access) 200,000 users in 8 months, with 450 use cases in production. AI budget 2024: approximately $1.3 billion (out of a total $17 billion tech budget). Productivity gain on code estimated at +10–20%.


Documented failures

In plain terms: failures do not come from a model that “hallucinates”. They come from agents that loop endlessly without a guardrail, or from agents that talk to each other without verifying what they say.

Uncontrolled cost escalation

A case from the ZenML Database illustrates the risk of an infinite loop bug (the agent calls a tool again without an exit condition; each call has a cost, costs add up until a human intervenes): weekly costs went from $127 to $47,000 in four weeks. Six weeks of correction work. Agents with access to billing systems require explicit circuit breakers (hard-coded numerical limits: if cost > X, we stop without asking).

Generic multi-agent systems do not reach production

Cemri, Pan, Yang et al. (UC Berkeley, arXiv March 2025) analyzed 1,600+ execution traces (complete recordings of what the agent did at each step: which tool called, which input, which output, which error) across 7 different frameworks (software libraries providing the structure and tools to build an agentic system), including GPT-4, Claude 3, and CodeLlama. Result: 14 distinct failure modes (categories of dysfunction, each with its own root cause), classified into three categories — system design, inter-agent misalignment (two agents working on the same task drift in contradictory directions without realizing it), task verification.

On ChatDev (open-source multi-agent framework — several agents collaborating via message exchange to accomplish a complex task), the correct correction rate is 25%. Even with optimization interventions (prompt engineering + improved orchestration), the gain is only +14 percentage points — insufficient for real production.

The critical distinction: generalist multi-agent systems (designed to handle all kinds of tasks) exhibit prohibitive failure rates. Bounded-scope multi-agent systems (the exact domain covered is defined in advance and the rest is excluded) and specialized (Uber, Stripe) work.

The Gartner prediction on cancellations

Gartner (2025) forecasts that at least 40% of agentic AI (autonomous systems where LLM agents make decisions and act without human validation at each step) projects will be canceled by end of 2027. The identified cause is not technical: escalating costs, insufficient business value, and absence of risk controls (organizational and technical devices that limit financial and reputational exposure: quotas, audits, cross-validation). Deloitte (2026) corroborates: only 21% of organizations have a mature governance model (written rules, assigned roles, formalized approval and revocation processes) for autonomous agents.


ROI: self-reported vs independent

In plain terms: when numbers come from a survey where the company declares itself, they are flattering. When numbers come from a third party auditing the books, they drop. The wide gap (171% vs < 5%) does not mean someone is lying — it means we are not measuring the same thing, and no study imposes a standard yet.

Self-reported surveys (surveys where the companies themselves answer without external verification of their figures) show high figures: average return on investment or ROI (ratio between generated gain and invested amount, expressed as a percentage; a ROI of 171% means that for 100 € invested, 271 € is recovered) of 171% according to some aggregations, 74% of executives declaring ROI in the first year [UNVERIFIED — sources without published methodology]. These figures should be treated with caution: they come primarily from vendor-commissioned studies or surveys without independent audit (verification by a third party with no interest in the announced result).

McKinsey (independent source, n=1,993) is notably more restrained: 39% of organizations attribute an impact on their operating profit (what the company actually earns after paying its current costs, also called EBIT — earnings before interest and taxes) to AI. Among them, the majority report less than 5% of operating profit.

Both positions coexist in the literature without reconciliation. The absence of a standardized definition of “agent in production” and the absence of independent audits largely explain the dispersion.


Key takeaways

  • Adoption figures (11% Deloitte vs 57% LangChain) measure different realities: systematic organizational deployment versus any functional instance. Both are correct.
  • Successful production agents are narrow, single-domain, and supervised — not autonomous and generalist.
  • UC Berkeley (MAST, 2025) documented 14 distinct failure modes in multi-agent systems, with correction rates as low as 25% on generalist frameworks.
  • Gartner predicts 40% cancellations by 2027 for governance and insufficient business value reasons — not for technical reasons.
  • Self-reported ROI (171% average) contrasts sharply with independent McKinsey measurements (EBIT impact < 5% for the majority). No academic source corroborates the high levels.