In short

Not all LLM systems are agents. The distinction is not cosmetic: it rests on the presence or absence of a perception-planning-action loop. Gulli (2025) proposes four levels of sophistication, from L0 (reasoning only) to L3 (teams of specialized agents), with each level adding a capacity for interaction with the environment. An arXiv research paper (Huang, 2024) independently identifies five levels along a similar axis. Both frameworks converge on one point: what makes an agent is not the size of the model, but the architecture surrounding it.


The loop as the distinguishing criterion

A standard LLM transforms an input into an output — once, with no feedback. An agent, by contrast, operates in an iterative loop: it perceives its environment (user input, tool results, memory state), plans the next steps toward an objective, and then acts (generates text, calls a tool, delegates to another agent). It observes the result, then starts again.

This perception-planning-action loop (sometimes called the PSAA loop: Perception-Scan-Action-Learning) is the central criterion. A system that implements it is an agent. A system that does without is an LLM pipeline — useful, but fundamentally different.

“An AI agent is a system designed to perceive its environment and take actions to achieve a specific goal.” — Gulli (2025), p. 14

The loop is not binary: it can be partial, infrequent, or heavily constrained by a human. Hence the value of a level-based classification.

In short: a Q&A system on static documentation = LLM pipeline, not an agent. A system that queries an API in real time and adapts its answer = agent. The difference: one waits for the world in its initial context; the other goes to fetch it along the way. It is an architectural threshold, not a continuum.


The four levels (L0 → L3)

L0 — Pure reasoning engine

The LLM generates responses from its pre-trained knowledge. No tools, no persistent memory, no interaction with the outside world.

Capability: produce coherent text, reason through a problem within the limits of the context window. Structural limit: knowledge is frozen at the training cutoff. The model can neither verify information nor execute actions.

Typical example: a Q&A chatbot on static documentation, a text reformulation assistant.

L1 — Connected solver (tool use)

The agent has access to external tools: search engine, APIs, code interpreter, database. It can call these tools, observe their results, and integrate them into its reasoning.

Decisive transition: the model can go beyond its internal knowledge. An L1 agent can look up a real-time price, execute code, or read a file. Limit: the agent does not strategically manage what it places in its context. On long tasks, the context becomes saturated or incoherent.

Typical example: a research assistant that runs web queries and summarizes results. Most consumer AI assistants with search access operate at this level.

L2 — Strategic solver (context engineering)

The agent applies what Gulli calls context engineering: the strategic selection, packaging, and management of relevant information at each step. Rather than accumulating all results in the context window, it filters, summarizes, and structures information to maintain coherence over long tasks.

“Context engineering is the discipline that accomplishes this by strategically selecting, packaging, and managing the most critical information from all available sources.” — Gulli (2025), p. 89

L2 also introduces proactivity: the agent can solicit feedback on its own processes, generate critiques of its outputs, and refine its plan without explicit instruction (the Reflexion pattern).

Decisive transition: the agent manages its own cognitive state, not just tool calls. Limit: an L2 agent remains a single system. Very broad or heterogeneous tasks will saturate it.

Typical example: Google DeepResearch — it plans a research task, identifies gaps along the way, and replans before producing the final report.

L3 — Multi-agent systems

Multiple specialized agents collaborate, coordinated by an orchestrator agent. Each agent has a defined role, its own context, and communicates with others through defined interfaces.

“Complex challenges are often best solved not by a single generalist, but by a team of specialists working in concert.” — Gulli (2025), p. 119

This level rejects the “super-generalist agent” model in favor of a division of labor analogous to that of a human organization. Gulli identifies six collaboration topologies: single agent, peer-to-peer network, supervisor, supervisor-tool, hierarchical, and custom topology.

Decisive transition: complexity is no longer absorbed by a single model but distributed across multiple specialized units. Limit: inter-agent coordination introduces new failure points. Communication and synchronization costs scale quickly. A shared ontology between agents is necessary to avoid misunderstandings.

In short: between L0 and L3, typical operational cost is multiplied by 10 to 100×. L0 = one LLM call (1-3 seconds, ~10 input/output tokens). L1 = call + 1-3 tools (5-15 seconds). L2 = full strategy (30 seconds to several minutes). L3 = multi-agent orchestration (minutes to hours for complex tasks). Choosing the right level is also choosing a budget.


Summary table

LevelDesignationKey capabilityMain limit
L0Reasoning engineGeneration from the model aloneFrozen knowledge, no action
L1Connected solverTool calls, external accessNon-strategic context management
L2Strategic solverContext engineering, self-improvementSingle agent, saturation on broad tasks
L3Agent teamSpecialization, division of laborComplex coordination, synchronization costs

When does an LLM become an agent?

The L0 → L1 transition is the sharpest: as soon as a system can call an external tool and integrate its result into reasoning, the perception-planning-action loop is present in its minimal form.

The L1 → L2 transition is more gradual. It occurs when the agent makes decisions about what it retains in its context, not just what it calls. A practical proxy: if the agent can identify that a piece of information has become useless and set it aside, it is operating at L2.

The L2 → L3 transition is architectural. It requires establishing an inter-agent communication protocol, shared memory, or an explicit delegation mechanism. An agent that calls itself recursively is not an L3 system — it is an L2 agent with recursion.

These boundaries remain [UNVERIFIED] in the sense of formal standardization: the level-based classification is a pedagogical framework, not an established industry standard.

Which level for which need

ContextRecommendationWhy
FAQ, Q&A on fixed documentationL0 (LLM alone)No tools, latency < 2s, minimal cost. Sufficient if documentation fits in the context window.
Real-time assistant (web search, document reading, calculations)L1 (tool use)Predefined tools, acceptable latency. Dominant production pattern 2025-2026.
Deep research, multi-step project (DeepResearch, Manus AI)L2 (context engineering)LLM manages its own context, capable of replanning. Higher cost and latency but robust.
Complex heterogeneous automation (analysis + design + review + correction)L3 (multi-agents)Specialization, parallelization, resilience. Reserved for cases where L2 saturates.

Limits of the framework

The L0-L3 taxonomy is intentionally simplified. Other dimensions exist:

  • Degree of autonomy within the loop: a system can be L2 in architecture yet require human validation at every step (human-in-the-loop). Effective autonomy is distinct from architectural sophistication.
  • Memory: the levels do not distinguish short-term memory (context window) from long-term memory (vector database). An L1 agent with persistent memory may outperform an L2 agent without memory on certain tasks.
  • L3 system topologies: six collaboration models exist according to Gulli (2025). The “L3” classification groups systems whose complexity varies by a factor of 10.

Other taxonomies exist. Vellum AI (2025) proposes six levels by separating “context awareness” and “goal orientation.” The arXiv paper (Huang, 2024) distinguishes rule-based agents, reinforcement learning agents, and LLM-based agents across five levels. Gulli’s L0-L3 framework remains the most widely used in practice-oriented implementation literature.


Key takeaways

  • The perception-planning-action loop is the criterion that distinguishes an agent from an LLM pipeline. Without it: L0.
  • L1 adds tools. L2 adds strategic context management. L3 adds collaboration between specialized agents.
  • Each level adds capabilities and new failure points.
  • The classification is pedagogical, not normative. The same system may operate at different levels depending on the task.
  • Effective autonomy (degree of human oversight) is a dimension orthogonal to architectural sophistication.