In brief
An AI agent is not a binary state — it is a graduation of capabilities. The L0–L3 taxonomy proposed by Gulli (2025) organizes this progression according to the nature of the interaction between the LLM and its environment: from raw text generation (L0) to collaborative multi-agent systems (L3). Understanding these levels allows choosing the right architecture for a given problem, and lucidly evaluating what an “agentic” system actually means.
L0 — The bare reasoning engine
At level 0, the LLM operates alone, without tools, without external memory, without access to the environment. It responds solely from its pre-trained knowledge. Generation is deterministic relative to the input: same prompt, same context, same likely response.
This level corresponds to the most common use case — a conversational chatbot with no branching toward APIs or databases. The limitation is structural: the model’s knowledge is frozen at its training cutoff date; it can neither verify nor act.
Yu Huang (arXiv:2405.06643) categorizes L0 as “No AI” in his own scale, reserving the term “agent” from the point where a system is capable of autonomous decision-making. Lilian Weng’s formulation confirms this dividing line: a minimal agent requires LLM + memory + planning + tool use — L0 meets none of these additional criteria.
In short: at L0, the LLM is a sophisticated tutor. It knows its lesson but has no phone, no calculator, no notebook. The training cutoff date (often 6 to 24 months ago) makes it incapable of any recent information.
L1 — The connected solver
The transition to L1 occurs when the LLM gains access to external tools: search engines, APIs, vector databases (RAG), code execution. It can now go beyond its internal knowledge and act on real-time data.
Gulli describes L1 as a “connected solver”: the agent receives a mission, queries its environment, and produces an enriched response. The model remains guided by largely predefined paths, however — it does not freely choose its overall strategy.
Anthropic labels this class “workflows” in its guide “Building Effective Agents”: “systems where LLMs and tools are orchestrated through predefined code paths.” The 5 identified patterns (prompt chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer) all fall within this register. HuggingFace takes a more fluid reading: agency is a continuous spectrum, and L1 represents a segment of that spectrum where the LLM partially influences the control flow.
L2 — The strategic solver
L2 introduces three additional capabilities absent at L1: active context engineering, continuous proactivity, and self-improvement via feedback.
Context engineering — defined by Gulli as “the discipline of strategically selecting, packaging, and managing the most critical information from all available sources” — enables the agent to build its own informational environment at each step, rather than receiving a fixed context. The agent no longer merely responds: it filters, aggregates, and reformats information before reasoning.
Continuous proactivity means the agent anticipates future needs (weather data, calendar availability) without waiting for an explicit instruction. Self-improvement via feedback closes the loop: the agent solicits evaluation of its own outputs and adjusts its behavior accordingly.
This level corresponds to what Anthropic calls an “augmented LLM”: LLM + retrieval + tools + memory, where the LLM dynamically controls the use of each of these components. The L1/L2 boundary lies at the degree of control the model exercises over its own process.
In short: what distinguishes L1 from L2 is not access to tools — it is who decides. At L1, code decides which tools to call in what order, the LLM executes. At L2, the LLM decides when and with what content to call which tool — it pilots its own process.
L3 — Multi-agent systems
L3 crosses a qualitative threshold: instead of a single versatile agent, several specialized agents collaborate under the coordination of an orchestrator agent. Gulli rephrases Conway’s logic — “complex challenges are often best solved not by a single generalist, but by a team of specialists working in concert.”
The typical architecture: a manager agent decomposes the task, delegates to specialized agents (research, writing, validation, design), then integrates the results. Each sub-agent optimizes a sub-problem; the global system solves what a single agent would not reliably solve.
Two emergent properties distinguish L3 from L2. First, resilience: if a sub-agent fails, the orchestrator can reassign or retry without rebuilding the entire pipeline. Second, effective parallelization: independent sub-tasks run simultaneously, reducing overall latency.
Gulli also mentions “metamorphic” systems — multi-agent architectures capable of creating, duplicating, or removing agents based on performance analysis during execution. [NOT VERIFIED — this capability is not documented in the academic sources consulted; it comes from the Gulli chapter alone.]
In short: L3 is not “L2 replicated N times”. It is an architectural leap — the ability to break a problem into sub-tasks delegated to specialized agents (research, writing, validation), then to stitch their results back together. The added orchestration cost is considerable: the majority of use cases do not justify it.
Which level to choose by need
| Context | Recommendation | Why |
|---|---|---|
| Q&A chatbot on stable knowledge | L0 (LLM alone) | No tools, minimal latency, lowest cost. Good option if documentation is fixed. |
| Real-time information retrieval, simple RAG | L1 (tool use) | Predefined tools (web search, RAG), deterministic orchestrator code. 80% of production cases. |
| Long tasks like deep research, exploratory project | L2 (context engineering) | The LLM decides what to keep, what to compress, when to replan. Higher cost but robustness over time. |
| Heterogeneous workflows (analysis + writing + design + review) | L3 (multi-agents) | Specialization by role, parallelization, resilience. But coordination is costly — reserved for cases where L2 saturates. |
Alternative taxonomies: why L0–L3 is not the only framework
The Gulli taxonomy is not isolated, but it is not universal either.
Yu Huang (arXiv:2405.06643) proposes an L0–L5 scale inspired by SAE autonomous driving levels. His critical dividing line falls between L2 (imitation or reinforcement learning) and L3 (LLM as cognitive controller) — corresponding precisely to the pivot between pre-LLM systems and LLM-driven systems. Levels L4 (autonomous learning) and L5 (personality and multi-agents) extend the framework beyond Gulli’s scope.
Mitchell et al. (arXiv:2502.02649, HuggingFace) define 5 functional levels, from simple instruction executor to autonomous code generator, with an explicit normative position: levels 4–5 present increasing risks that justify, in their view, not developing them.
Feng, McDonald, and Zhang (arXiv:2506.12469) take an orthogonal angle: their 5 levels measure the user’s role in the interaction (operator, collaborator, consultant, approver, observer) rather than the agent’s technical capabilities. Their central thesis — “autonomy is a property that can be designed independently of capability” — implies that an L3-capable agent can deliberately operate with L1-level autonomy if constrained to consult the user before each action.
Google DeepMind (arXiv:2311.02462) adds a further distinction: autonomy is not “determined” by capabilities, merely “unlocked.” This point is relevant for deployments: an L3 system is not condemned to maximum autonomy.
HuggingFace takes an even more fluid position in its official course (smolagents + agents-course): “Agency is not a discrete, 0 or 1 definition: instead, agency evolves on a continuous spectrum, as you give more or less power to the LLM on your workflow.” The operational definition adopted by HuggingFace — “AI Agents are programs where LLM outputs control the workflow” — measures agency by the share of the control flow that the LLM effectively influences, as opposed to deterministic code. This reading makes L0–L3 boundaries permeable: a pipeline can be L1 for some branches and L2 for others depending on configuration.
The arXiv:2601.12560 survey formalizes agents as POMDPs (Partially Observable Markov Decision Processes) and organizes the field into 6 orthogonal dimensions (components, cognitive architecture, learning, multi-agents, environments, evaluation/safety). This framework does not hierarchize levels but decomposes them — useful for diagnosing an existing system, less so for communicating a pedagogical progression.
Key takeaways
- L0 to L3 describes a progression from isolation (pure generation) to collaboration (coordinated multi-agents), with two main pivots: access to tools (L0→L1) and dynamic control of the process by the LLM itself (L1→L2).
- Context engineering — strategic selection and packaging of information at each step — is the central competency that differentiates L2 from L1; it is not automatically acquired with tools.
- L3 does not mean “total autonomy”: the question of the autonomy level granted to a multi-agent system is a design decision distinct from its technical capabilities (Feng et al., DeepMind).
- Multiple taxonomies coexist in the literature with different divisions (L0–L5 in Yu Huang, 5 functional levels in Mitchell et al.); Gulli’s L0–L3 is one operational framework among others, practitioner-oriented.
- An important negative result: no taxonomy consulted provides empirical validation on benchmarks — the levels remain conceptual categories at this stage, not verified performance measures.