In short
A survey published in February 2026 (Yang et al., arXiv:2602.05665) catalogues and classifies all graph-based memory approaches for AI agents. The authors propose a taxonomy of six cognitive types, analyze four main storage structures, and identify six distinct retrieval operators. Their central conclusion: graph representations outperform vector and key-value approaches on four measurable dimensions — a finding that informs architecture choices for any persistent agent system.
In short: think cartographer vs photographer. The photographer (vector memory) takes pictures of every street — useful to recognize a scene. The cartographer (graph memory) plots streets, intersections, flows. To answer “which path goes from A to B,” the cartographer wins; to recognize a sign, the photographer wins. Persistent AI agents need both, but for a long time we have built only with photos.
The 6 types of agent memory
The survey’s taxonomy distinguishes two broad families. Knowledge Memory is a passive, static repository of objective, global, and verifiable information. Experience Memory is a proactive, dynamic personal log of interactions and their outcomes.
These two families break down into six cognitive types:
- Semantic: general world knowledge, stable facts, definitions. Corresponds to what the model “knows” independently of any usage context.
- Procedural: skills and action rules — how to accomplish a task, which tools to use, in what order.
- Associative: latent links between concepts, indirect relationships that are not made explicit in the raw data but can be inferred.
- Working: immediate scratchpad, the current state of the ongoing task. Corresponds to the active context window.
- Episodic: chronological sequences of past sessions — what was done, in what order, with what results.
- Sentimental: emotional tone of interactions, user preferences, approval or rejection signals accumulated over time.
This classification is not purely theoretical: each type corresponds to different storage and retrieval requirements, which justifies the use of distinct structures to represent them.
The survey also organizes the memory lifecycle into four phases: Extraction (building representations from interactions), Storage (persistence in the chosen structure), Retrieval (access at runtime), and Evolution (updating and reorganization over time).
Why graphs beat vectors
Vector approaches — embeddings, top-k semantic similarity, vector databases — dominate current RAG architectures. They answer the question “which content is semantically close to this query?” well. They answer poorly questions of a different nature: why are two facts related? What is the hierarchy between concepts? Did event A precede or follow event B?
Yang et al. identify four dimensions on which graphs outperform linear, vector, and key-value approaches:
1. Explicit relationship modeling. A graph stores relationships as named edges between nodes, making causal reasoning direct. A vector encodes the relationship implicitly in the embedding space — it can be lost or blurred during retrieval.
2. Hierarchical organization. Graphs allow structuring information at multiple levels of abstraction, from granular facts up to high-level themes. A flat embedding does not naturally support this hierarchy.
3. Temporal and dynamic structuring. Graph edges can carry temporal metadata. Bi-temporal graphs (see next section) distinguish when a fact was true in the real world from when it was recorded — a distinction impossible with a vector.
4. Efficient structured retrieval. Graph traversal follows explicit relationships between entities, producing verifiable reasoning paths. Pure vector similarity can surface thematically close content that is structurally unrelated.
In short: if the agent must answer “why is A linked to B?”, a graph gives it the direct path (A → relation X → C → relation Y → B). A vector at best gives it “A and B are thematically close.” For a legal or medical assistant, the difference is between a traceable explanation and a silent correlation.
Storage structures
The survey distinguishes several storage architectures, each suited to a different memory profile.
Classic Knowledge Graph
The standard KG stores triplets (head entity, relation, tail entity), built by LLM extraction with conflict detection. It is suited to static long-term memory: stable facts, world rules, definitions. It can be enriched with temporal metadata for dated episodic facts.
Hypergraphs — HyperGraphRAG
A hypergraph generalizes the binary graph by allowing hyperedges that connect an arbitrary number of nodes simultaneously. This preserves n-ary relationships without fragmenting them into binary pairs — a relationship involving three entities simultaneously (for example “X collaborated with Y on project Z”) is not faithfully represented by two separate edges.
HyperGraphRAG exploits this property with “dual retrieval”: the agent can query both isolated entities and complete hyperedges, depending on the required granularity.
Bi-temporal model — Graphiti
Most memory systems overwrite obsolete facts during an update. The bi-temporal model distinguishes two time axes:
- Valid time: when the event occurred in the real world.
- Transaction time: when the information was recorded in the database.
Graphiti implements this model: when a fact becomes false, it is not deleted but temporally invalidated. The full history remains accessible, enabling questions such as “what did the agent know at a given moment?” or “what was the situation before this update?”.
TReMu pushes this reasoning further by distinguishing the session timestamp from the absolute time inferred from the event, with support for date arithmetic for neuro-symbolic reasoning.
MemoTime organizes temporal knowledge graphs (TKGs) in a “Tree of Time” that enforces chronological ordering and prevents logical hallucinations of the effect-before-cause type.
Hybrid architecture — Optimus-1
Optimus-1 explicitly separates the two broad memory families into two distinct structures:
- An HDKG (Hierarchical Directed Knowledge Graph) for world rules and stable facts.
- An AMEP (Abstracted Multimodal Experience Pool) for trajectories and interaction history.
This separation reflects the survey’s taxonomic distinction: knowledge memory and experience memory have different properties (stability vs. dynamism, global vs. personalized) that justify distinct storage structures.
Retrieving the right information
The survey identifies six retrieval operators, which can be combined according to need:
- Semantic: top-k vector similarity on node or edge embeddings.
- Rule-based: symbolic filters (entity type, time range, explicit attribute).
- Temporal: time windows, decay functions (more recent facts receive higher weight).
- Graph-based: intra-layer traversal (traversal within a single structure) or inter-layer (crossing different structures).
- RL-based: adaptive policies (PPO, GRPO) that learn when and how to query memory.
- Agent-based: planning-feedback loop with API calls — the agent decides itself which part of its memory to consult.
Two traversal strategies compete in the ecosystem:
Entity-centric (Mem0): retrieval starts from a central entity and explores its immediate relationships. Precise, but limited to the neighborhood of the target entity.
Breadth-first (Zep): progressive expansion from the entry point, surfacing broader context but potentially less focused.
H-MEM takes a different approach: index-based routing in intra-layer traversal, which directs retrieval without scanning the entire graph.
A complementary technique — the post-retrieval strategy — generates an intermediate representation (theme, intention, draft structure) before retrieval, reducing sensitivity to superficial query formulations. The generation of latent memory tokens pushes this idea into the model’s latent space.
Open challenges
Despite the described advantages, the survey identifies six challenges the community has not yet resolved:
Scaling. Graph structures are costly to maintain in production as the volume of nodes and edges grows. Existing indexes and traversal algorithms show limits on large enterprise-scale graphs.
Real-time consistency. In distributed systems, keeping a memory graph consistent when multiple agents write simultaneously is an unsolved problem. Edge conflicts and concurrent updates produce inconsistencies that are difficult to detect and correct.
Completeness/latency trade-off. Exhaustive retrieval (broad graph exploration) guarantees better contextual coverage, but increases latency. Targeted retrieval is fast but may miss relevant relationships. Optimal thresholds depend on the domain and task.
Memory quality measurement. Current metrics evaluate memory indirectly, via performance on downstream tasks. No direct measure of a memory representation’s quality yet exists — relationship accuracy, completeness, freshness.
Cross-domain generalization. Memory architectures optimized for one domain (medical agents, game agents, conversational agents) generalize poorly outside their training context. A unified framework remains to be built.
Interpretability of learned retrieval policies. RL-based policies are effective but opaque. Understanding why an agent chose to consult a particular part of its memory — and verifying that this choice was relevant — is difficult with current explanation methods.
In short: moving from vectors to graphs is not a trivial upgrade. Building an agent KG requires tooling (triplet extraction, deduplication, temporal management) and increases maintenance cost. For a simple chat, vectors suffice; for a persistent agent that must reason and explain, the graph investment pays off.
Decision matrix — which memory structure for which agent
| Agent profile | Recommendation | Why |
|---|---|---|
| Short-term conversational assistant (chat ≤ 100 turns) | Vector store + rolling summary | Sufficient, simple to deploy, no need for explicit relations. |
| Legal/medical assistant (mandatory traceable reasoning) | Knowledge Graph + Graphiti (bi-temporal) | Verifiable reasoning paths; history preserved for audit. |
| Multi-entity agent with n-ary relations (collaborations, projects) | HyperGraphRAG | Hyperedges avoid fragmentation of multi-entity relations. |
| Game/simulation agent (learns from trajectories) | Optimus-1 (HDKG + AMEP) | Separates stable rules (HDKG) from accumulated experience (AMEP). |
| Large-scale industrial agent (> 10M nodes) | Vector store + hybrid retrieval | Graphs still costly to scale in massive production (open challenge). |
| Agent that must date its knowledge (audit, compliance) | Graphiti (bi-temporal) | Distinction between valid time and transaction time, impossible with vectors. |
What to remember
- The Yang et al. (2026) survey proposes the first structured taxonomy of graph-based memory for agents: 6 cognitive types, 4 main storage structures, 6 retrieval operators.
- Graphs outperform vectors on four concrete dimensions: explicit relationships, hierarchy, temporality, structured traversal. This is not a vague qualitative advantage — it is a difference in functional capability.
- The bi-temporal model (Graphiti) solves a problem that vector databases cannot address: preserving the history of successive states of a fact without losing traceability.
- Hybrid architectures (Optimus-1) explicitly separate knowledge memory and experience memory, reflecting the survey’s fundamental taxonomic distinction.
- Six open challenges remain, including the absence of direct metrics for evaluating memory quality — which makes comparison between systems difficult and optimization indirect.