In short

A large language model, as used in a chatbot, produces a response from an input — then stops. An LLM agent goes further: it breaks down an objective into steps, calls external tools (search engines, APIs, code interpreters), observes the results, and adjusts its plan accordingly. This shift from passive model to active agent rests on three technical building blocks — tool calling, planning, and agent coordination — that took shape between 2022 and 2025. Benchmarks reveal real capabilities but also a substantial gap from human performance on realistic tasks.

In short: an LLM in chatbot form is a tour guide who describes the route without moving. An LLM agent is a guide who walks with you, pushes the doors, checks the map, asks for directions, and starts over if you took the wrong street. The vocabulary is identical; the difference is that there are real footsteps on the ground.


From response to action: tool calling

Until 2022, LLMs produced text without being able to interact with the outside world. The Toolformer paper (Schick et al., Meta AI, 2023) changed that. The idea is straightforward: a model can learn, from examples, when to call a tool, how to formulate the call, and how to integrate the response into its reasoning. The tools tested were basic — calculator, search engine, translation — but the demonstration was decisive: a 6.7 billion parameter model with tool access sometimes outperformed much larger models without tools.

OpenAI, Google, and Anthropic integrated this paradigm directly into their APIs from 2023–2024. The mechanics are as follows: the model generates a structured function call (a JSON object describing the tool to call and its parameters), the system intercepts it, executes the actual call, and returns the result to the model. This is today’s industry standard, referred to as function calling.

In 2025, effort shifted toward interoperability: the Model Context Protocol (MCP) developed by Anthropic standardizes how tools are described and invoked, independently of the model using them.


Thinking to act: the ReAct loop

Knowing how to call a tool is not enough to accomplish a complex task. Planning is also required — deciding the sequence of actions. This is the purpose of ReAct (Yao et al., 2022, published at ICLR 2023), one of the most cited architectures in the field.

ReAct’s principle is to interleave reasoning and action within the same token sequence. At each step, the agent produces a thought trace (what it understands about the situation), decides on an action (calling a tool, making a query), then observes the environment’s response. This cycle — Thought / Action / Observation — allows the agent to remain coherent over several dozen steps while adapting its plan to each new piece of feedback.

On text navigation and e-commerce navigation tasks, ReAct outperforms imitation and reinforcement learning approaches by +34% and +10% in success rate respectively.

Reflexion (Shinn et al., NeurIPS 2023) adds a layer of self-critique: after an episode, the agent generates a verbal reflection on its errors, which it stores in memory to guide future attempts. Without modifying its weights, a model equipped with Reflexion achieves 91% success on HumanEval (a code generation benchmark), compared to 80% for GPT-4 without this mechanism at the time.

LATS (Language Agent Tree Search, Zhou et al., 2023) goes further by exploring multiple trajectories in parallel via tree search inspired by Monte Carlo. The result: 92.7% on HumanEval. The cost: a compute budget multiplied 3–10× compared to ReAct. This performance/cost trade-off is one of the central debates in the field.

In short: ReAct is the guide who speaks aloud while walking — each thought is a comment, each action a door pushed open. Reflexion is the same guide who takes notes after each mistake to avoid repeating it. LATS is ten copies of the guide trying different paths in parallel and picking the best one. Performance climbs, the bill climbs too.

Decision matrix: which planning pattern

ContextRecommendationWhy
Medium task (5-20 steps), fast iterations neededReActReference standard, base cost, traceable, sufficient in most cases.
Code generation or repeatable tasks with possible critiqueReflexion+11 pts HumanEval (80→91%), no retraining, error memory.
Hard search, multiple valid paths to exploreLATS+1-2 pts vs Reflexion, but 3-10× the compute cost. Reserve for critical cases.
Composed pipeline, heterogeneous expertiseMulti-agents (AutoGen / MetaGPT)Role specialization, structured artifacts, parallelism.

Coordinating multiple agents

Some tasks benefit from being distributed across specialized agents. Two frameworks have structured the field.

AutoGen (Wu et al., Microsoft Research, 2023) conceptualizes agents as asynchronous conversational entities. Multiple agents — assistant, code executor, orchestrator — can converse and co-resolve a task. AutoGen introduced terminology that remains in common use today.

MetaGPT (Hong et al., ICLR 2024) adopts an organizational metaphor: each agent embodies a role (product manager, architect, engineer, QA) and follows standard operating procedures encoded in its instructions. Agents exchange structured artifacts (specifications, diagrams) rather than free text, reducing ambiguities at the agent interface.

Recent comparative analyses (2025) show that frameworks diverge in their priorities: LangGraph favors traceability and persistent memory; CrewAI favors role-based coordination; AutoGen favors conversational flexibility. There is no universal solution — the choice depends on the constraints of the use case.


What benchmarks reveal

To measure progress, the field has developed agent-specific evaluations — different from classical benchmarks that test static capabilities.

SWE-bench (Jimenez et al., ICLR 2024) submits agents to 2,294 real GitHub issues. The agent must produce a patch that passes regression tests. In 2024, the best systems resolved around 12% of instances; in 2025, the best exceed 50% on the verified version of the benchmark.

WebArena (Zhou et al., ICLR 2024) offers 812 navigation tasks on functional websites. The best GPT-4 agent achieves 14.4% success; a human achieves 78.2%.

GAIA (Mialon et al., Meta/HuggingFace, ICLR 2024) poses real-world questions requiring combined reasoning, web navigation, and tool calls. Humans achieve 92%; GPT-4 equipped with plugins reaches only 15%. This 77-point gap summarizes the chasm between agents’ perceived capabilities and their actual performance on composite tasks.

These figures must be read carefully. Research from 2025 (Kapoor et al., arXiv 2507.02825) identifies serious validity problems in agent benchmarks: agents that do nothing succeed on 38% of tasks in certain benchmarks, due to insufficiently robust evaluation criteria. Automated judges make arithmetic errors that skew rankings. These biases can distort performance measurements by 100% in relative terms.

In short: agent benchmarks are incomplete road maps. On SWE-bench (GitHub bug repair), the best systems went from 12% to 50% in one year — real progress. But on GAIA (everyday questions), humans are at 92%, GPT-4 at 15%. And even those 15% are sometimes suspect: an agent that does nothing wins 38% of certain tests. The direction is right, the gap to humans remains massive.


Concrete risks

The agentic paradigm introduces risks absent from classical conversational models.

Cascading fragility is the most documented problem. An error at step N propagates to subsequent steps. On tasks spanning 20 to 50 steps, the compounding error rate can render the agent unusable. Real-world incidents in 2025 illustrate the concrete cost: a development agent deleting a production database, a commercial agent making an unauthorized purchase.

Prompt injection is a threat specific to agents. When an agent processes external data — emails, web pages, files — that data may contain malicious instructions that hijack its behavior. In a multi-agent system, an injected instruction can propagate from one agent to another (Cohen et al., arXiv 2410.07283, 2024). A documented exploit against Microsoft Copilot in 2025 demonstrated this attack under real-world conditions.

An analysis from the AI Agent Index 2025 (arXiv 2602.17753) reveals that delegating security to end users is the dominant practice among deployed systems — an approach researchers consider insufficient.


Key takeaways

  • An LLM agent combines tool calling, multi-step planning, and — in advanced architectures — coordination between specialized agents.
  • The ReAct loop (Thought / Action / Observation) remains the reference architecture for agentic planning; variants such as Reflexion and LATS improve performance at the cost of higher computational overhead.
  • Benchmarks show rapid progress (SWE-bench: from 12% to 50% in one year) but also a persistent gap with humans on realistic tasks — GAIA: 15% vs. 92%.
  • The benchmarks themselves are contested: validity problems documented in 2025 suggest that some scores overestimate real capabilities.
  • Cascading fragility and prompt injection are the two most concrete risks for agents deployed in production.