MindMemOS: A Portable and Self-Evolving Memory Operating Layer for AI Agents

arXiv:2608.12428 · cs.AI, cs.IR, cs.IT, math.IT · Submitted 2026-08-12 · Read on arXiv

Kaichao Liang, Yuqi Cui, Hao Kong, Xinyuan Huang, Guohaotian Hou, Qingcan Kang, Liang Chen, Yiyang Yin, Ke Ye, Jiaquan Guo, Da Chen, Lingan Zeng, Yixing Peng, Rong Yao, Shixiong Kai, Mingxuan Yuan

Noah's Ark Lab, Huawei Technologies

cs.AI, cs.IR, cs.IT, math.IT

Submitted: 2026-08-12

Updated: 2026-08-14

Comments: 35 pages,14 figures

Code: https://github.com/mindscale-noah/MindMemOS

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: MindMemOS is a portable and self-evolving memory operating layer for AI agents, designed to address the rigidity of existing memory systems that remain fixed after development.

Terminology

Summary

MindMemOS is a portable and self-evolving memory operating layer for AI agents, designed to address the rigidity of existing memory systems that remain fixed after development. It organizes open-world information using a unified entity–property–time structure, supporting scenario-adaptive memory modeling, higher-order pattern discovery, autonomous memory refinement, and continuous skill evolution.

The system's architecture decouples agent integration, memory algorithms, and memory structure. Its memory model is a three-dimensional graph with entity, property, and time dimensions, enabling traceable navigation across interconnected entities and evolving information. Two memory generation algorithms are provided: MindVanilla, which processes dialogue without a predefined schema using turn-aware input processing and recall-aware extraction, and MindSchema, which uses episode segmentation, memory generation, entity fusion, and graph merge guided by a scenario-adaptive schema.

A compact search module combines sparse (BM25) and dense (embedding) matching with reciprocal rank fusion, supporting bidirectional traversal over entity, property, relational, and temporal associations, orchestrated by an agentic LLM-based controller.

Four evolution mechanisms are introduced:

  1. MindMemEvolve: A validation-driven evolutionary search algorithm that optimizes memory schemas using task-specific evaluation signals. It employs error-informed mutation, exploratory mutation, crossover, and selection to adapt entity and property definitions and discover higher-order patterns. The optimization objective is to maximize a Judge score over a training set, using an LLM-guided evolutionary algorithm with population size K, elite fraction α, and cumulative training steps H.

  2. Dreaming: An offline consolidation mechanism that selects unconsolidated add records, constructs entity-centered scopes, detects issues (conflicts, duplication, complementary fragments, low-value content, ambiguous relationships), and applies conservative mutation plans to merge redundant records, resolve conflicts, and preserve provenance.

  3. Feedback: Converts corrective user signals into memory-maintenance actions. Explicit feedback revises stored memories through natural-language corrections, while implicit feedback detects correction signals from ordinary dialogue, classifies their persistence scope (task-temporary, scenario-specific, long-term), and selectively converts them into memory updates.

  4. MindSkillEvolve: Transforms agent execution trajectories into reusable and progressively refined skills. It collects trajectories bound to specific skill versions, analyzes them for effective strategies and recurring failures, and generates versioned skill updates through unsupervised or score-guided refinement.

Experimental results show:

  • On LOCOMO, MindSchema achieves 94.03% overall accuracy, ahead of EverOS (93.05), Zep (85.22), and MemOS (80.76), with particular strength in Single-hop (96.79) and Multi-hop (93.97) reasoning, and the highest Open-domain score (82.29).

  • On PersonaMem, MindSchema achieves 70.63% overall accuracy, a gain of 3.06 percentage points over EverOS (67.57), driven by improvements in Recall Sha. (81.40 vs. 74.42) and Suggest. (47.31 vs. 35.48).

  • On MemoryAgentBench FactConsolidation, dreaming improves accuracy while compressing roughly one fifth of active memories. For gpt-4o-mini, single-hop accuracy improves from 0.635 to 0.738, multi-hop from 0.118 to 0.180, and overall from 0.377 to 0.459, with AMCR values of 21.4%, 19.4%, and 20.4%. For gpt-5-mini, single-hop improves from 0.900 to 0.920, multi-hop from 0.190 to 0.250, and overall from 0.545 to 0.585, with AMCR values of 23.5%, 21.5%, and 22.5%. After dreaming, MindMemOS outperforms Infini Memory on all metrics.

  • On SpreadsheetBench, MindSkillEvolve-Unsup. improves success rate from 51.3% to 55.3%, and MindSkillEvolve-Sup. further improves it to 57.2%, a 9.2 percentage point improvement over the initial-skill baseline (48.0%).

Case studies illustrate dreaming's conflict resolution (archiving superseded memories and recording supersedes edges), feedback's retrieval enrichment (shifting answers from incorrect to correct options), MindMemEvolve's schema evolution (expanding from 2 to 49 dynamic properties, with first-order and higher-order properties emerging autonomously), and skill evolution's shift from tool API instructions to executable error-avoidance rules, with supervised evolution adding boundaries to prevent overgeneralization.

The paper concludes that MindMemOS achieves state-of-the-art performance under reported evaluation protocols, particularly on multi-hop questions requiring synthesis of evidence across multiple memories, and outlines future work on experience-to-skill evolution and file-system-level memory management.

Improvements for AI systems

Improvements to AI Systems:

  1. Dynamic Memory Schema Evolution – Replace fixed memory schemas with an evolutionary algorithm (MindMemEvolve) that mutates, crosses over, and selects entity/property definitions based on task-specific judge scores. The improved AI system can autonomously discover higher-order properties (e.g., from 2 to 49 properties) and adapt its memory structure to new domains without manual redesign.

  2. Offline Memory Consolidation via Dreaming – Implement a periodic consolidation phase that detects conflicts, duplicates, and low-value memories, then merges or archives them while preserving provenance via supersedes edges. The improved AI system can compress 20% of active memory, boost multi-hop reasoning accuracy (e.g., from 0.118 to 0.180 for gpt-4o-mini), and reduce retrieval noise.

  3. Bidirectional Temporal-Entity Graph Traversal – Use a three-dimensional memory graph (entity–property–time) with sparse-dense hybrid search (BM25 + embeddings + reciprocal rank fusion) and an LLM controller for agentic navigation. The improved AI system can trace evolving facts across time, answer multi-hop questions requiring synthesis (e.g., 93.97% on LOCOMO multi-hop), and handle open-domain queries with 82.29% accuracy.

  4. Explicit and Implicit Feedback-Driven Memory Repair – Convert user corrections (both direct natural-language edits and implicit signals from dialogue) into memory maintenance actions, classified by persistence scope (task-temporary, scenario-specific, long-term). The improved AI system can self-correct stored facts in real time, shifting answers from incorrect to correct options in retrieval tasks, and avoid repeating past errors.

  5. Skill Evolution from Execution Trajectories – Transform agent action logs into versioned, reusable skills via unsupervised or score-guided refinement (MindSkillEvolve). The improved AI system can learn from its own successes and failures, generating executable error-avoidance rules (e.g., spreadsheet manipulation success rate from 48.0% to 57.2%), and prevent overgeneralization by adding boundaries during supervised evolution.

  6. Schema-Guided Memory Generation – Use scenario-adaptive schemas (MindSchema) with episode segmentation, entity fusion, and graph merge to structure memories during dialogue. The improved AI system can achieve 94.03% overall accuracy on LOCOMO and 70.63% on PersonaMem, with significant gains in recall (81.40 vs. 74.42) and suggestion (47.31 vs. 35.48) tasks.

  7. Provenance-Preserving Conflict Resolution – When memories conflict, archive superseded records and add explicit supersedes edges rather than deleting data. The improved AI system can maintain a traceable history of fact changes, enabling auditability and better handling of contradictory information over time.

  8. Unified Memory Operating Layer – Decouple agent integration, memory algorithms, and memory structure into a portable layer. The improved AI system can be plugged into any agent architecture, reusing the same memory optimization mechanisms across different LLMs (e.g., gpt-4o-mini, gpt-5-mini) and tasks without retraining.

Abstract

Memory is a core component of AI agents, enabling them to accumulate experience, maintain personalization, and adapt over long-term interactions. However, existing memory systems often remain fixed after development, limiting their ability to adapt their memory models, organization strategies, and procedural knowledge through continued use. We present MindMemOS, a portable and self-evolving memory operating layer that organizes open-world information using a unified entity property timestructure. MindMemOS supports scenario-adaptive memory modeling, higher-order pattern discovery, autonomous memory refinement, and continuous skill evolution. Its MindMemEvolve algorithm employs validation-driven evolutionary search to optimize memory schemas for target scenarios, whiledreaming consolidates accumulated memories by merging redundant records and resolving conflicts. In addition, implicit corrective feedback serves as a human-in-the-loop signal for identifying and revising potentially inaccurate or misaligned memories. Its MindSkillEvolve algorithm further transforms agent execution trajectories into reusable and progressively refined skills. MindMemOS achieves 94.03% accuracy on LOCOMO and 70.63% on PersonaMem. MindSkillEvolve improves SpreadsheetBench success by 9.2 percentage points over the initial-skill baseline.

Sources

Related papers