In brief
A classic LLM receives an input and produces an output — then stops. An agentic system chains perception, reasoning, and action in a loop: it decomposes an objective, calls tools, observes results, and readjusts its plan until completion. This paradigm shift, which moved from experimentation to industrial production between 2023 and 2025, transforms language models into systems capable of acting in the world — with real gains and new risks.
In short: a classic LLM is a consultant who answers your question and then hangs up. An agentic system is an assistant who opens the calendar, makes the calls, takes notes, checks the result, starts over if necessary — until the mission is done. The model remains the same brain; what changes is that we give it hands to act and eyes to verify.
The fundamental loop: perception, reasoning, action
An AI agent rests on a repeating cycle that the literature summarizes in three steps.
Perception: the agent receives a state of the world — a user request, a tool result, a message from another agent, or the contents of a database. This state is encoded in the model’s context.
Reasoning: the model analyzes the state, decides on the next action, and, in advanced architectures, produces an explicit chain of thought before acting.
Action: the agent executes a concrete action — API call, web request, database write, code generation, or delegation to another agent. The result modifies the state of the world and becomes the next perception.
This loop structurally distinguishes an agent from an LLM in question-answer mode. A conversational LLM produces text; an agent produces effects. The same difference separates an assistant that describes a course of action from one that executes it directly.
In short: the perception-reasoning-action loop is exactly what a cook does when tasting their sauce. They perceive (the taste), they reason (is salt missing?), they act (add, stir), then they taste again. An agentic LLM does the same thing with its context instead of taste buds: each action changes the state, each new state feeds the next decision.
Four building blocks
Tool use
The ability to call external tools — search engines, APIs, code interpreters, databases — is the prerequisite for any agentic architecture. Without tools, the agent is limited to what its context contains. Toolformer (Schick et al., Meta AI, 2023) demonstrated that a model could learn on its own when and how to call tools. Since then, function calling has been natively integrated into the main commercial APIs.
Memory
Agents operate with several complementary types of memory. Short-term memory corresponds to the active context of the model’s window. Long-term memory relies on vector databases or databases that the agent can query. Some architectures also distinguish semantic memory (general facts), episodic memory (traces of past interactions), and procedural memory (reusable procedures). The management of these memories remains an active area: a 2025 review of memory mechanisms in multi-agent systems documents the absence of consolidated standards.
Planning
For long tasks, the agent must decompose an objective into ordered sub-tasks. The ReAct pattern (Yao et al., ICLR 2023) interleaves reasoning traces and tool calls in the same sequence, which reduces hallucinations compared to pure reasoning. The plan-and-execute pattern separates global planning (produced by a powerful LLM) from local execution (delegated to a lighter, less expensive component). Benchmarks show that plan-and-execute architectures can achieve 92% completion rates with a 3.6× speedup over sequential ReAct execution.
Orchestration
When multiple agents cooperate, their exchanges must be coordinated. The orchestrator decides who does what, in what order, and how results are aggregated. In centralized architectures, a pilot agent distributes tasks to specialized agents. In decentralized architectures, agents negotiate directly with each other.
In short: these four building blocks work like the crew of a racing sailboat. Tool use is the hands that hoist the sail. Memory is the logbook — we note what we have already tried. Planning is the navigator who calculates the route. Orchestration is the skipper who distributes the roles. Remove a block, and the boat drifts.
Architectures: single agent vs. multi-agent
Single agent
A single agent has a set of tools and manages its own perception-reasoning-action loop. This architecture is simple to deploy, easy to debug, and sufficient for the majority of tasks. It reaches its limits on very long tasks (context saturation), tasks requiring heterogeneous expertise, and parallelizable tasks.
Multi-agent
A multi-agent system distributes work among specialized agents that operate in parallel or in sequence. Efficiency depends less on the number of agents than on the quality of coordination protocols: unstructured exchanges between agents produce noise rather than value. Recent frameworks (LangGraph, CrewAI, AutoGen) offer different primitives — state graph, organizational roles, asynchronous conversation — but none has emerged as a universal standard.
Design patterns
Three patterns structure the majority of agentic architectures:
- ReAct: Thought / Action / Observation loop, suited for moderately long tasks requiring frequent adjustments.
- Plan-and-execute: global planning followed by step-by-step execution, suited for long tasks with a fixed objective.
- Reflection: the agent critiques its own outputs and improves them without modifying its weights. Reflexion (Shinn et al., NeurIPS 2023) achieves 91% success on HumanEval without retraining.
Decision matrix: which pattern to choose
| Context | Recommendation | Why |
|---|---|---|
| Task < 10 steps, unpredictable environment (web scraping, exploration) | ReAct | Step-by-step responsiveness, adjusts to observations in real time. |
| Long task (≥ 20 steps) with clear objective (reporting, ETL) | Plan-and-execute | Pre-computed global plan, deterministic execution, 3.6× speedup. |
| Creative or technically demanding output (code, writing) | Reflection | Iterative self-critique, +20-30 points on HumanEval without retraining. |
| Heterogeneous specializations (research + analysis + writing) | Centralized multi-agent | An orchestrator dispatches to specialized agents. |
| Parallelizable volume (audit 100 files, corpus classif) | Multi-agent fan-out | Independent workers, final consolidation, linear throughput. |
Positioning: neither RAG nor fine-tuning
Agentic AI does not replace existing approaches — it complements them.
RAG and agentic AI are complementary. RAG connects an LLM to a knowledge base to ground its responses in verifiable facts. Agentic AI takes RAG as a tool: an agent can decide to query a vector database, then call an API, then synthesize both results. RAG answers “what does this base know?”; the agent answers “how do I reach this objective with the available resources?”.
Fine-tuning and agentic AI are orthogonal. Fine-tuning modifies the model’s weights to specialize its intrinsic capabilities. Agentic AI leaves the weights intact and adds capabilities through tooling and orchestration. You can run a fine-tuned model in an agentic architecture — the two stack.
Prompt engineering and agentic AI operate at different levels. Prompting techniques (chain of thought, few-shot) operate at the scale of a single call. Agentic architecture operates at the scale of a complete pipeline of several tens or hundreds of calls.
Emerging standards: MCP and A2A
Interoperability is the structural problem of agentic systems: until recently, each framework defined its own interfaces.
MCP (Model Context Protocol), published by Anthropic in November 2024 and now under Linux Foundation governance, standardizes how an agent accesses external tools and resources. MCP operates vertically: it defines the contract between an agent and its tools.
A2A (Agent-to-Agent Protocol), launched by Google in April 2025 and transferred to the Linux Foundation in June 2025, standardizes communication between agents — discovery, delegation, conflict resolution. A2A operates horizontally: it defines the contract between cooperating agents.
The two protocols are complementary. In December 2025, the Agentic AI Foundation (AAIF) brings OpenAI, Anthropic, Google, Microsoft, AWS, and Block together around these standards.
Limitations
Amplified hallucination: in a multi-step loop, a reasoning error at step N propagates and amplifies at subsequent steps. A 5% error rate per step becomes greater than 60% over 20 composed steps. Studies on deployed systems document hallucination rates ranging from 0.7% to 29.9% depending on models and tasks.
Computational cost: an agentic task can generate tens to hundreds of LLM calls. Documented incidents in 2025 report agents entering recursive loops, generating six-figure cloud bills. The token budget is a concrete operational constraint, not a theoretical one.
Evaluation complexity: classic benchmarks (perplexity, multiple-choice accuracy) do not capture agentic performance. Specialized benchmarks (SWE-bench, WebArena, GAIA) are more relevant but contested: 2025 work (Kapoor et al.) shows that agents that do nothing succeed at 38% of tasks on certain benchmarks, due to a lack of robust criteria.
Security: indirect prompt injection — malicious instructions hidden in data processed by the agent (emails, web pages) — is the most documented vulnerability. It allows redirecting an agent toward unauthorized actions. The EchoLeak exploit (CVE-2025-32711) against Microsoft Copilot in 2025 is a real-world demonstration.
In short: an autonomous agent amplifies BOTH the strengths AND weaknesses of the model that drives it. Where an erroneous chatbot response costs one message to correct, an erroneous agent decision can trigger 50 API calls behind the scenes, drain a cloud budget, or execute an irreversible action. Human supervision remains the only truly robust safety net.
Key takeaways
- An AI agent is an LLM placed in a perception-reasoning-action loop with access to tools, memory, and a planning mechanism.
- The ReAct, plan-and-execute, and reflection patterns cover the majority of current agentic architectures.
- Agentic AI is complementary to RAG and orthogonal to fine-tuning — the three approaches stack in hybrid architectures.
- MCP and A2A are the emerging interoperability standards, both under Linux Foundation governance since late 2025.
- The main limitations are error propagation, computational cost, evaluation difficulty, and security vulnerabilities specific to agents.