In short
Five LLM-maintained knowledge base systems, designed independently, converge on the same architecture: Markdown files as substrate, a centralized index for navigation, automated maintenance as a survival condition. This is not a literature review — it is an empirical observation. The convergence is structural, not about implementation details.
In short: think of a librarian who has to organize their own library without help. With fifty books, they manage. With five thousand, they give up. That is exactly the fate of personal wikis before LLMs. When the assistant can classify, index, and correct on its own at near-zero cost, the librarian — researcher, financier, or engineer — can finally keep their library alive.
The observation
In April 2026, five independent LLM-driven knowledge management systems are publicly documented. Their authors did not coordinate. Their use contexts are different: personal research, an agent framework, an open-source tool, corporate finance, LLM research.
And yet, when you place the architectures side by side, the same pattern appears.
The five systems:
- LLM Wiki — Andrej Karpathy (2026). Architecture formalized in a public gist. Three layers: immutable sources, LLM-maintained wiki, schema co-evolved by human and LLM.
- Deep Agents Memory — LangChain (2026). Memory module of the Deep Agents framework. Markdown files via
edit_file, three scopes (agent, user, org), hot-path and background consolidation. - agentmemory — rohitg00 (2026). Open-source persistent memory system for coding agents. Four consolidation tiers, hybrid BM25/vector/graph retrieval, auto-cleanup.
- CFO Brain — Jon Dorsey (2026). Personal stack: Claude + Granola + Obsidian for a CFO. Five Obsidian folders, two Claude Code skills (/call, /today).
- Our lab — labo-llm.fr (2025-2026). Knowledge State in Markdown files, centralized index, structural lint, chronological lab notebook, human operator in the loop.
The table
This is the core piece. Six comparative axes, five systems.
| Axis | LLM Wiki (Karpathy) | Deep Agents (LangChain) | agentmemory (rohitg00) | CFO Brain (Dorsey) | Our lab |
|---|---|---|---|---|---|
| Memory substrate | .md files in a directory | Files via edit_file, configurable backends | Embedded DB + git snapshots | .md files in Obsidian | .md files in git repo |
| Navigation | index.md → drill-down per page | Workspace index, loaded into system prompt | Hybrid retrieval BM25 + vectors + knowledge graph (RRF) | Obsidian wiki-links + index files | KNOWLEDGE_STATE.md → KS_{AXIS}.md |
| Maintenance / lint | Periodic lint: contradictions, orphan pages, stale claims, gaps | Not documented | TTL, contradiction detection (Jaccard > 0.9), low-value eviction, cascade | Not documented | lint_ks.py: 9 detection categories |
| Logging | append-only log.md, parsable format (grep) | LangSmith traces (tool calls) | Not documented | Not documented | Chronological lab_notebook/ |
| Consolidation | Ingest = fanout 10-15 pages, good answers re-injected | Hot path (synchronous) + background (cron/scheduled, separate agent) | Automatic LLM compression across 4 tiers (Working → Episodic → Semantic → Procedural) | Manually triggered CC skills (/call processes transcripts, /today generates brief) | Semi-automated: CC produces claims and fiches, operator validates |
| Human / LLM division | Human = sourcing, exploration, framing. LLM = summarization, cross-referencing, ranking, bookkeeping | Human = task definition. Agent = autonomous execution | Automatic (hooks, cron). Human = initial configuration | Human = sourcing (Granola captures). CC = processing and propagation | Operator = decides and validates. CC = executes bookkeeping |
Three constants run across all five systems:
-
Markdown as substrate. Four out of five use .md files directly. The fifth (agentmemory) uses a database but exposes results as structured text. The format is readable by both humans and LLMs without conversion.
-
A centralized index. Each system maintains an entry point that lets the LLM know what exists before searching. The approach varies — index file, wiki-links, hybrid search — but the function is identical: prevent the LLM from working blind.
-
The LLM does the bookkeeping. The division of labor converges: humans provide the material and ask the questions, the LLM does the classification, cross-referencing, and update work that nobody wants to do manually.
In short: convergence is not a fashion effect but a response to a structural constraint. Markdown is readable by both populations — humans and models — the index acts as a table of contents, and bookkeeping is precisely the work the LLM does at zero marginal cost where the human is overwhelmed.
What it implies
Karpathy states the core point: humans abandon wikis when the maintenance cost grows faster than the value. The LLM solves this problem because the marginal cost of maintenance tends toward zero.
This is the thesis of Vannevar Bush’s Memex (1945) — a personal knowledge system with “associative trails” between documents — whose unsolved problem was precisely sustained maintenance. Eighty years later, the LLM provides the missing mechanism.
For our research, the convergence across five sources reinforces a signal: the boundary between deterministic and probabilistic, at the heart of our work, materializes in these architectures. The substrate (files, index, lint) is deterministic — verifiable, versioned, auditable. The processing (ingestion, consolidation, cross-referencing) is probabilistic — delegated to the LLM, not guaranteed. The five systems manage this boundary differently, but none eliminates it.
The human/LLM division of labor is not a design choice — it is a consequence of this boundary. The human intervenes where determinism is necessary (validation, sourcing, decisions). The LLM intervenes where probabilism is sufficient (summarization, classification, maintenance).
The gaps — what we don’t know
Convergence is a signal, not proof. Several limitations prevent drawing conclusions.
No cross-system benchmark. None of the five systems is formally compared to the others. agentmemory publishes internal metrics (Recall@10 = 64.1%, 92% token reduction compared to context-dumping), but on its own test set (240 observations, 30 sessions). The other four systems publish no quantitative measurements.
Incomparable scales. LLM Wiki works “surprisingly well” at ~100 sources and a few hundred pages [UNVERIFIED]. Our lab operates on ~550 claims and ~23 experiment fiches. agentmemory is benchmarked on 240 observations. CFO Brain is a personal system with two skills. Deep Agents is a framework with no published usage data. We are comparing architectures, not performance.
Everything is qualitative except agentmemory. Karpathy publishes no figures. Dorsey describes his system as “VERY early” and admits not knowing whether his index files are necessary. Our lab measures structure (lint) but not retention. Only agentmemory offers formal metrics — and they cover retrieval, not consolidation.
Consolidation remains an open problem. Three divergent approaches coexist: fully automatic (agentmemory, 4-tier compression), semi-automatic (Deep Agents, hot path + cron), manually triggered (CFO Brain, skills). No data allows us to say which produces the best long-term retention.
Automatic cleanup is not validated. agentmemory automatically deletes old memories with low importance. The false-positive rate of this eviction is not documented. The other four systems do not perform automatic cleanup — which does not mean that is better, just that we don’t know.
Key takeaways
- Five independent systems converge on the same architecture: Markdown files, centralized index, LLM as maintainer.
- The convergence is structural, not about implementation — consolidation mechanisms diverge significantly.
- The pattern solves the historical Memex problem (Bush, 1945): sustained maintenance of a personal knowledge base.
- No cross-system benchmark exists. Comparison remains qualitative and architectural.
- The signal is strong but data is insufficient to move from observation to conclusion.