In brief

An AI agent is not a binary state — it is a graduation of capabilities. The L0–L3 taxonomy proposed by Gulli (2025) organizes this progression according to the nature of the interaction between the LLM and its environment: from raw text generation (L0) to collaborative multi-agent systems (L3). Understanding these levels allows choosing the right architecture for a given problem, and lucidly evaluating what an “agentic” system actually means.


L0 — The bare reasoning engine

At level 0, the LLM operates alone, without tools, without external memory, without access to the environment. It responds solely from its pre-trained knowledge. Generation is deterministic relative to the input: same prompt, same context, same likely response.

This level corresponds to the most common use case — a conversational chatbot with no branching toward APIs or databases. The limitation is structural: the model’s knowledge is frozen at its training cutoff date; it can neither verify nor act.

Yu Huang (arXiv:2405.06643) categorizes L0 as “No AI” in his own scale, reserving the term “agent” from the point where a system is capable of autonomous decision-making. Lilian Weng’s formulation confirms this dividing line: a minimal agent requires LLM + memory + planning + tool use — L0 meets none of these additional criteria.

In short: at L0, the LLM is a sophisticated tutor. It knows its lesson but has no phone, no calculator, no notebook. The training cutoff date (often 6 to 24 months ago) makes it incapable of any recent information.


L1 — The connected solver

The transition to L1 occurs when the LLM gains access to external tools: search engines, APIs, vector databases (RAG), code execution. It can now go beyond its internal knowledge and act on real-time data.

Gulli describes L1 as a “connected solver”: the agent receives a mission, queries its environment, and produces an enriched response. The model remains guided by largely predefined paths, however — it does not freely choose its overall strategy.

Anthropic labels this class “workflows” in its guide “Building Effective Agents”: “systems where LLMs and tools are orchestrated through predefined code paths.” The 5 identified patterns (prompt chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer) all fall within this register. HuggingFace takes a more fluid reading: agency is a continuous spectrum, and L1 represents a segment of that spectrum where the LLM partially influences the control flow.


L2 — The strategic solver

L2 introduces three additional capabilities absent at L1: active context engineering, continuous proactivity, and self-improvement via feedback.

Context engineering — defined by Gulli as “the discipline of strategically selecting, packaging, and managing the most critical information from all available sources” — enables the agent to build its own informational environment at each step, rather than receiving a fixed context. The agent no longer merely responds: it filters, aggregates, and reformats information before reasoning.

Continuous proactivity means the agent anticipates future needs (weather data, calendar availability) without waiting for an explicit instruction. Self-improvement via feedback closes the loop: the agent solicits evaluation of its own outputs and adjusts its behavior accordingly.

This level corresponds to what Anthropic calls an “augmented LLM”: LLM + retrieval + tools + memory, where the LLM dynamically controls the use of each of these components. The L1/L2 boundary lies at the degree of control the model exercises over its own process.

In short: what distinguishes L1 from L2 is not access to tools — it is who decides. At L1, code decides which tools to call in what order, the LLM executes. At L2, the LLM decides when and with what content to call which tool — it pilots its own process.


L3 — Multi-agent systems

L3 crosses a qualitative threshold: instead of a single versatile agent, several specialized agents collaborate under the coordination of an orchestrator agent. Gulli rephrases Conway’s logic — “complex challenges are often best solved not by a single generalist, but by a team of specialists working in concert.”

The typical architecture: a manager agent decomposes the task, delegates to specialized agents (research, writing, validation, design), then integrates the results. Each sub-agent optimizes a sub-problem; the global system solves what a single agent would not reliably solve.

Two emergent properties distinguish L3 from L2. First, resilience: if a sub-agent fails, the orchestrator can reassign or retry without rebuilding the entire pipeline. Second, effective parallelization: independent sub-tasks run simultaneously, reducing overall latency.

Gulli also mentions “metamorphic” systems — multi-agent architectures capable of creating, duplicating, or removing agents based on performance analysis during execution. [NOT VERIFIED — this capability is not documented in the academic sources consulted; it comes from the Gulli chapter alone.]

In short: L3 is not “L2 replicated N times”. It is an architectural leap — the ability to break a problem into sub-tasks delegated to specialized agents (research, writing, validation), then to stitch their results back together. The added orchestration cost is considerable: the majority of use cases do not justify it.

Which level to choose by need

ContextRecommendationWhy
Q&A chatbot on stable knowledgeL0 (LLM alone)No tools, minimal latency, lowest cost. Good option if documentation is fixed.
Real-time information retrieval, simple RAGL1 (tool use)Predefined tools (web search, RAG), deterministic orchestrator code. 80% of production cases.
Long tasks like deep research, exploratory projectL2 (context engineering)The LLM decides what to keep, what to compress, when to replan. Higher cost but robustness over time.
Heterogeneous workflows (analysis + writing + design + review)L3 (multi-agents)Specialization by role, parallelization, resilience. But coordination is costly — reserved for cases where L2 saturates.

Alternative taxonomies: why L0–L3 is not the only framework

The Gulli taxonomy is not isolated, but it is not universal either.

Yu Huang (arXiv:2405.06643) proposes an L0–L5 scale inspired by SAE autonomous driving levels. His critical dividing line falls between L2 (imitation or reinforcement learning) and L3 (LLM as cognitive controller) — corresponding precisely to the pivot between pre-LLM systems and LLM-driven systems. Levels L4 (autonomous learning) and L5 (personality and multi-agents) extend the framework beyond Gulli’s scope.

Mitchell et al. (arXiv:2502.02649, HuggingFace) define 5 functional levels, from simple instruction executor to autonomous code generator, with an explicit normative position: levels 4–5 present increasing risks that justify, in their view, not developing them.

Feng, McDonald, and Zhang (arXiv:2506.12469) take an orthogonal angle: their 5 levels measure the user’s role in the interaction (operator, collaborator, consultant, approver, observer) rather than the agent’s technical capabilities. Their central thesis — “autonomy is a property that can be designed independently of capability” — implies that an L3-capable agent can deliberately operate with L1-level autonomy if constrained to consult the user before each action.

Google DeepMind (arXiv:2311.02462) adds a further distinction: autonomy is not “determined” by capabilities, merely “unlocked.” This point is relevant for deployments: an L3 system is not condemned to maximum autonomy.

HuggingFace takes an even more fluid position in its official course (smolagents + agents-course): “Agency is not a discrete, 0 or 1 definition: instead, agency evolves on a continuous spectrum, as you give more or less power to the LLM on your workflow.” The operational definition adopted by HuggingFace — “AI Agents are programs where LLM outputs control the workflow” — measures agency by the share of the control flow that the LLM effectively influences, as opposed to deterministic code. This reading makes L0–L3 boundaries permeable: a pipeline can be L1 for some branches and L2 for others depending on configuration.

The arXiv:2601.12560 survey formalizes agents as POMDPs (Partially Observable Markov Decision Processes) and organizes the field into 6 orthogonal dimensions (components, cognitive architecture, learning, multi-agents, environments, evaluation/safety). This framework does not hierarchize levels but decomposes them — useful for diagnosing an existing system, less so for communicating a pedagogical progression.


Key takeaways

  • L0 to L3 describes a progression from isolation (pure generation) to collaboration (coordinated multi-agents), with two main pivots: access to tools (L0→L1) and dynamic control of the process by the LLM itself (L1→L2).
  • Context engineering — strategic selection and packaging of information at each step — is the central competency that differentiates L2 from L1; it is not automatically acquired with tools.
  • L3 does not mean “total autonomy”: the question of the autonomy level granted to a multi-agent system is a design decision distinct from its technical capabilities (Feng et al., DeepMind).
  • Multiple taxonomies coexist in the literature with different divisions (L0–L5 in Yu Huang, 5 functional levels in Mitchell et al.); Gulli’s L0–L3 is one operational framework among others, practitioner-oriented.
  • An important negative result: no taxonomy consulted provides empirical validation on benchmarks — the levels remain conceptual categories at this stage, not verified performance measures.