In brief

MemEvolve (Zhang et al., arXiv:2512.18746, December 2025) identifies a structural limitation of current memory systems for AI agents: their architecture remains fixed even as the agent’s knowledge base progresses. To overcome this asymmetry, the paper proposes a simultaneous dual evolution — the experiential knowledge base and the memory architecture co-evolve via a bi-level optimization process. Across four complex agent benchmarks, the gain reaches 17.06% in performance, with confirmed generalization cross-task and cross-LLM model.


The problem: static memory architectures

Modern AI agents rely on memory systems to accumulate experience, retain facts, and retrieve relevant context. These systems are operationally sophisticated — they handle encoding, storage, retrieval, and updates. But their architecture remains static: one chooses a configuration (graph type, retrieval strategy, consolidation policy), and it applies uniformly regardless of the task or the agent’s maturity level.

Zhang et al. designate this phenomenon as the staticity of the memory system itself. The agent evolves — it accumulates experience, refines its responses, specializes — but the architectural framework organizing this evolution never adapts itself. This is an asymmetry: the architecture serves as a lever to evolve the agent without ever questioning its own calibration.

The practical consequence is concrete: a memory system well-parameterized for a web navigation task may be suboptimal for multi-step reasoning, or for an agent operating on a new LLM. Memory architecture is not neutral — it constrains what the agent can learn and how it can retrieve that knowledge.

In short: today, you choose the memory architecture at startup and never come back to it. MemEvolve turns that architecture into a living parameter — the library rearranges its shelves based on the books arriving and the questions being asked.


The dual evolution

MemEvolve addresses this diagnosis with a dual-evolution principle: simultaneously co-evolving two levels.

Lower level — the experiential knowledge base. This is the classical level: the agent interacts with its environment, accumulates trajectories, and enriches its memory. This level is present in most existing memory systems.

Upper level — the memory architecture itself. This is MemEvolve’s novelty: the architecture (the encode, store, retrieve, and manage operations — their configuration, combination, and hyperparameters) is treated as an optimizable parameter rather than a design constant.

The two levels interact through a bilevel optimization scheme: the upper loop explores and evaluates candidate architectural configurations, while the lower loop accumulates experience under each of these configurations. Information flows upward — the performance obtained under a given architecture informs the upper loop about what to retain, modify, or eliminate.

The tournament process

The concrete implementation follows a competitive multi-phase protocol:

  1. Generation: N candidate memory systems are independently produced, each representing a distinct architectural configuration.
  2. Initial tournament: N+1 systems (the N candidates plus the existing baseline system) are evaluated on a common task set. This is a head-to-head competition on neutral ground.
  3. Final: The best systems from the initial tournament are tested on additional tasks, to confirm the robustness of the selection and avoid overfitting to the initial evaluation tasks.

This elimination structure ensures that the retained system is not only superior on a particular task set but also generalizes. This is a methodologically important point: many meta-learning approaches suffer from overfitting on their calibration benchmark.

In short: MemEvolve has N “architecture candidates” play a tournament against a baseline system. The winners go through a final on unseen tasks to confirm they are not just good at the qualifier. The absolute winner gets +17% performance where a fixed architecture plateaus.


EvolveLab: a unified testbed

To conduct these experiments rigorously and reproducibly, Zhang et al. developed EvolveLab, a unified codebase that distills 12 representative memory systems into a common modular design space.

The design space is structured around four fundamental operations:

OperationRole
EncodeTransform raw data (observations, interactions) into storable representations
StoreOrganize and persist representations in memory (structure, indexing)
RetrieveRecover relevant elements based on a query or context
ManageMaintenance policy: consolidation, forgetting, reorganization, updates

By mapping 12 heterogeneous systems to these four axes, EvolveLab makes configurations comparable and combinable. This is a contribution in its own right: the absence of a common design space is one of the reasons it is difficult to compare existing memory systems and identify what explains their performance differences.

EvolveLab also serves as a playground for MemEvolve’s optimization loop: exploring the architectural space amounts to exploring combinations of configurations across these four dimensions.


Results

Across four complex agent benchmarks, MemEvolve produces a performance improvement of up to 17.06% compared to baseline systems. Evaluations cover SmolAgent and Flash-Searcher agents — two distinct agent architectures, which is important for validating the scope of the results.

Two generalization properties are confirmed:

  • Cross-task: improvements are not limited to the tasks used in the selection tournament. The optimized architecture transfers to tasks not seen during optimization.
  • Cross-LLM: the selected architectural configurations remain advantageous when the underlying LLM is changed. This suggests that MemEvolve identifies robust structural properties of memory architecture, rather than parameters over-fitted to a particular model.

Cross-model generalization is the most theoretically significant result: it indicates that the optimal memory architecture is not entirely dependent on the LLM exploiting it. There is a structural component of architectural quality that transcends the model.

In short: if you change the LLM behind the agent, the good memory architectures selected by MemEvolve remain good. This suggests there are intrinsically superior structures, independent of the model exploiting them — much like a Dewey classification remains useful regardless of which librarian uses it.

When to use MemEvolve

ContextRecommendationWhy
Stable production agent, known tasksFixed memory architectureOptimization cost unjustified, marginal gain
Migration to new LLM (e.g., provider change)Run MemEvolve on representative tasksPre-MemEvolve architecture likely suboptimal under new model
Rapidly evolving domain (web, open research)Periodic re-optimizationTask space drifts, architecture must follow
Single benchmark with limited API budgetManual approach + ablationMemEvolve costly: N candidates × tournament × final

Limitations and perspectives

Process complexity. Bi-level optimization with a multi-stage tournament is computationally costly. Generating N candidate systems, evaluating them on agent tasks, then running a final — each iteration represents a significant inference budget. The paper does not precisely document this cost, making it difficult to assess feasibility for deployments under strict budget constraints.

Discrete design space. EvolveLab reduces the diversity of memory systems to 12 configurations in a 4-dimensional operational space. This is a necessary simplification to make optimization tractable, but it structurally excludes configurations outside this space. Future architectural innovations will not be automatically covered.

The bilevel optimization claim warrants verification. In the sources collected for this article, the term comes from secondary sources (alphaXiv summaries, WebAgentlab comments). Reading the full paper is necessary to confirm whether the bilevel formalism is explicitly adopted or whether this is an external characterization of the process.

Link with Meta-Harness. MemEvolve fits within a broader pattern observable in the 2025 literature: systems that use an outer loop to optimize configuration parameters of an inner loop. Meta-Harness (system prompt optimization via meta-agent) follows similar logic. The difference is that MemEvolve applies this pattern to the memory architecture itself rather than to agent instructions. These approaches converge toward a common question: which parameters of an agent system can be automatically optimized, and at what cost?


Key takeaways

  • Current agent memory architectures are static: they allow the agent to learn but do not learn to adapt themselves to task contexts.
  • MemEvolve proposes a simultaneous dual evolution: experiential knowledge base (lower level) and memory architecture (upper level), coupled via bilevel optimization.
  • The tournament process generates N candidate configurations, pits them against the baseline on common tasks, then tests the best against additional tasks.
  • EvolveLab unifies 12 representative memory systems in a modular space with 4 operations (encode, store, retrieve, manage), making configurations comparable and combinable.
  • The 17.06% gain generalizes cross-task and cross-LLM — suggesting MemEvolve identifies robust structural properties rather than over-fitted parameters.