In short
An AI agent manages three distinct memory types, imported from cognitive psychology: semantic memory (knowledge about the world), episodic memory (past experiences), and procedural memory (know-how and instructions). The CoALA framework (Sumers et al., 2023) formalizes this taxonomy for LLM-based agents by adding a fourth component, working memory, which links the three others to the ongoing decision cycle.
In short: a human seeing a “café” sign mobilizes three distinct memories — they know what a café is (semantic), they remember walking past this café yesterday (episodic), they know how to order an espresso (procedural). And all of this is filtered by their working memory, the immediate scratchpad. A persistent AI agent, to function coherently, must reconstruct these four layers — that is what CoALA formalizes.
Why distinguish memory types
An LLM without a memory architecture is a stateless system: every request starts from scratch. As soon as an agent must maintain coherence across multiple exchanges, learn from past mistakes, or generalize acquired behaviors, it becomes necessary to explicitly separate what belongs to the immediate context from what must persist.
The distinction is not arbitrary. It reflects concrete technical constraints: the context window is bounded, costly, and volatile; external storage is persistent but accessing it requires an explicit retrieval policy. Choosing which memory type hosts which information directly determines the agent’s capabilities and limitations.
Working memory: the hub of the decision cycle
Working memory contains the active information during the current decision cycle. It aggregates inputs from the other memory types and directs outputs toward storage or action execution.
Concretely, it corresponds to the LLM’s context window: what is there is immediately available, but capacity is limited and nothing survives the session without an explicit persistence mechanism. ReAct (Yao et al., 2022) illustrates this role: the reasoning traces interleaved between actions constitute a form of externalized, temporary working memory. This approach produced a 34% improvement on ALFWorld and 10% on WebShop compared with classical imitation and reinforcement methods.
In short: working memory is the active desk. Anything not on it is immediately inaccessible. The LLM context window plays this role, but with a physical limit: 200K tokens is about 150,000 words, the equivalent of a short novel. Beyond that, you must either forget (what falls off) or store elsewhere (consolidation toward persistent memory).
Semantic memory: knowledge about the world
Semantic memory stores the agent’s factual and conceptual knowledge — about the world and about itself. It can be initialized from external databases (encyclopedias, technical manuals) or built dynamically through autonomous agent learning.
Unlike classical read-only RAG systems, CoALA provides for both reads and writes: the agent actively enriches its knowledge base. The Generative Agents of Park et al. (2023) illustrate the mechanism: agents generate higher-level reflections from their observations and store them to guide future behavior — a spontaneous consolidation process.
LD-Agent (Li et al., 2024) combines episodic and semantic memory in a dialogue system: a long-term memory bank stores interlocutor profiles and persistent information, while a short-term bank captures the ongoing session. A topic-based retrieval approach improves access precision compared with raw vector similarity. [UNVERIFIED: direct comparison of topic-based retrieval vs. vector on common benchmarks]
Episodic memory: past experiences
Episodic memory records the events experienced by the agent: input-output pairs from past interactions, decision trajectories, session history. It serves two main purposes: providing few-shot examples during planning, and enabling learning without weight modification.
Reflexion (Shinn et al., 2023) demonstrates this second use directly: an episodic buffer stores textual reflections after each attempt; the agent uses this verbal feedback to improve its strategy on the next attempt, without fine-tuning. Measured result: 91% accuracy on HumanEval, compared with 80% for GPT-4 at the time of publication. [UNVERIFIED: direct comparison of episodic buffer vs. fine-tuning on general tasks beyond code and ALFWorld]
MemGPT (Packer et al., 2023) approaches episodic memory from an operating systems perspective: a recall storage preserves the interaction history (equivalent to a system log), distinct from archival storage (persistent and queryable) and the active context (equivalent to RAM). Interrupts manage transfers between levels, enabling processing of documents that exceed the context window size.
Procedural memory: know-how and instructions
Procedural memory encodes the agent’s behavior. CoALA distinguishes two forms:
- Implicit: knowledge encoded in the LLM’s weights, inherited from training.
- Explicit: procedures written in the agent’s code (actions, decision policy).
This second form is modifiable at runtime — which makes it powerful and risky. CoALA explicitly notes that writing to procedural memory is “significantly more risky” than modifying episodic or semantic memory, because an error can introduce bugs into the agent’s behavior.
Voyager (Wang et al., 2023) mitigates this risk by implementing procedural memory as a library of executable, versionable, and composable code skills. Acquired skills accumulate without overwriting previous ones, avoiding the catastrophic forgetting observed during weight fine-tuning. [UNVERIFIED: generalization of this approach beyond environments where code is directly executable and verifiable like Minecraft]
The Reflection pattern (Gulli, 2025, Ch.8) offers a third path: the agent consults its current instructions and the session history, generates improved new instructions, then persists them — a controlled iteration on explicit procedural memory.
Tensions and open questions
Two taxonomies coexist without consensus. CoALA advocates 4 types (working + episodic + semantic + procedural). MemGPT adopts 3 OS-inspired levels (in-context / archival / recall). No paper proposes a direct comparative benchmark between the two frameworks on the same tasks. [UNVERIFIED]
Retrieval remains an open problem. Three paradigms are active in the literature: dense vector similarity (dominant in RAG), symbolic rules (mentioned in CoALA), topic-based retrieval (LD-Agent). No comparative benchmark on common corpora and tasks. [UNVERIFIED]
Interaction between memory types is poorly formalized. The coordination mechanisms — when to escalate from working memory to episodic, when to consolidate episodic into semantic — remain unspecified outside CoALA, which remains a conceptual framework without a shared reference implementation.
In short: CoALA gives a useful vocabulary to think about an agent’s memory, but no standard framework yet orchestrates these four layers in production. Each team (LangChain, Letta/MemGPT, Mem0, Zep) proposes its own orchestration. Choosing a framework is also choosing an opinion on when to consolidate, when to forget, when to retrieve.
Decision matrix — which memory for which need
| Agent need | Memory type | Typical implementation |
|---|---|---|
| Keeping the thread of a long conversation (≤ session) | Working (context window) | Rolling summary, prompt caching |
| Storing user preferences across multiple sessions | Semantic | Vector store + user key |
| Learning from past mistakes without fine-tuning | Episodic | Reflexion-style textual buffer + retrieval |
| Capitalizing reusable skills | Explicit procedural | Versioned function library (Voyager pattern) |
| Audit/traceability (who said what when) | Episodic with timestamp | Structured append-only journal |
| Stable domain knowledge (encyclopedia, manuals) | Semantic | Classic RAG over reference corpus |
What to remember
- An AI agent structures its memory into four components: working (active context), semantic (knowledge), episodic (experiences), procedural (behaviors). The taxonomy comes from cognitive psychology, formalized for LLMs by CoALA (Sumers et al., 2023).
- Episodic memory can replace fine-tuning for iterative improvement: Reflexion achieves 91% on HumanEval without modifying the weights.
- Explicit procedural memory (code) is modifiable but risky: an error can corrupt the agent’s behavior.
- Two competing taxonomies (CoALA vs. MemGPT) coexist without a common benchmark.
- The coordination mechanisms between memory types remain the primary unresolved gap.