LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation
Dongfang Li, Zixuan Liu, Junmai Wang, Jiahe Huang, Fuhao Li, Bonian Jia, Baotian Hu, Min Zhang
Harbin Institute of Technology, Shenzhen
cs.CL
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: 34 pages, 5 figures
Code: https://github.com/memodb-io/memobase
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: LYCHEE MEMORY V2 is an efficient long-term memory framework for LLM agents that replaces turn-level consolidation with semantic segment-level consolidation.
Terminology
Summary
LYCHEE MEMORY V2 is an efficient long-term memory framework for LLM agents that replaces turn-level consolidation with semantic segment-level consolidation. Instead of consolidating every interaction, LYCHEE MEMORY batches multiple exchanges into segments and encodes each finalized segment into context-independent typed memory records. Segment-level batching lowers LLM encoding frequency, while semantic boundary detection helps preserve coherent event-level and temporal evidence compared with fixed-window batching. The resulting records are organized with lightweight structured indexes for query-planned evidence retrieval.
The core idea is that past experience should re-enter future reasoning with the proper granularity, temporal relation, and evidence strength. Instead of consolidating every turn, LYCHEE MEMORY batches multiple exchanges and invokes the LLM once for each finalized segment, thereby reducing construction frequency. Within this segment-level batching scheme, semantic surprise and cohesion signals determine where segments end, helping preserve coherent event boundaries rather than relying on mechanical fixed windows. Each segment is then encoded into context-independent typed memory records that contain natural-language statements, memory types, entities, topics, temporal scopes, and links to the original dialogue evidence. To maintain continuity across segments without re-reading the full history, the system carries lightweight disambiguation feedback, including entity aliases and reference relations, into later consolidation steps. LYCHEE MEMORY further organizes records in an append-only structured memory store with entity, topic, temporal, event-frame, and entity-topic indexes. At query time, a planner decomposes the user question into multiple recall routes and retrieves evidence from semantic memory, structured indexes, temporal indexes, and raw turns, followed by route-level reranking, fusion, and diversity-aware selection. This design improves evidence coverage without relying on blind context expansion.
The method consists of four main components. First, online semantic segmentation groups the conversation stream into coherent segments and triggers segment-level encoding only when a segment is finalized. The semantic surprise score is defined as s t = 1 - max(sim(e t, c k), sim(e t, h k)), where sim is cosine similarity. The cohesion drop d t induced by x t is d t = max(0, Coh(S k) - Coh(S k ∪ x t)). The final boundary score combines semantic surprise, cohesion drop, token pressure, and turn-count pressure: p t = σ(b + w s φ(s t) + w c d t + w l L t + w n N t). A segment is finalized when p t exceeds a fixed threshold δ or a hard token cap is reached. This design keeps the boundary decision fully embedding-based, so mid-segment exchanges are buffered without LLM inference. The reduction in encoding calls comes from consolidating at segment rather than exchange granularity: the number of calls is proportional to the number of segments S rather than the number of exchanges T.
Second, segment-level memory encoding converts each finalized segment into memory records that can be retrieved and interpreted without the original dialogue context. For a finalized segment S k, the encoder receives the segment text together with a compact reference context ρ k from recent segments and produces a set of memory records R k and an updated disambiguation state d k: (R k, d k) = Encode(S k, ρ k). The encoding prompt asks the LLM to perform three operations in a single pass: extract atomic information units, resolve coreference and elliptical mentions, and normalize relative temporal expressions using the session timestamp. Each record r i is represented as r i = (id i, τ i, text i, E i, K i, T i, src i), where id i is an internal record identifier, τ i is the memory type, text i is a self-contained natural-language statement, E i is the entity set, K i is the topic-tag set, T i records normalized event or validity times when available, and src i stores provenance links to the source turns. The memory type τ i is selected from a finite schema covering facts, preferences, events, constraints, procedures, failure patterns, and tool affordances. The reference context for the next segment is constructed as ρ k+1 = (d k; Recent(R k-m:k)), where Recent selects a bounded set of recent record summaries from the same session, truncated to a fixed budget.
Third, structured evidence organization organizes encoded records so that retrieval can satisfy both semantic and structured constraints. For each record r i, the organizer inserts the record into a vector store and updates a structured store. The structured store maintains five classes of evidence nodes: entity nodes, topic nodes, entity-topic nodes, temporal nodes, and event-frame nodes. Each evidence node stores a searchable text representation, an optional embedding, and pointers to linked memory records. Because all nodes are derived from record metadata, this phase requires only embedding and bookkeeping operations, and the write-side LLM cost remains dominated by segment-level encoding.
Fourth, plan-guided multi-route retrieval retrieves evidence for a query using one planning step followed by deterministic recall. Given a query q and recent dialogue context H q, the planner outputs a structured plan Π(q, H q) = (y, R 1,..., R m), where y is the question type and each route R j = (g j, Q j, C j, T j) contains a route goal, one or more search queries, structured constraints, and optional temporal constraints. Each route executes four recall channels in parallel: direct record recall, evidence-node recall, temporal recall, and raw-turn recall. Candidates from all routes are merged with reciprocal-rank fusion: RRF(d) = Σ j 1/(κ + rank j(d)). Candidates are optionally reranked by a lightweight cross-encoder within each route before route-level lists are fused. The fused candidates are then passed through diversity-aware selection so that evidence from different routes is preserved. The only generative query-time operation is the planning call; all subsequent recall, expansion, reranking, fusion, and selection steps are embedding lookups, structured filtering, or arithmetic scoring.
Experiments using GPT-4.1-Mini show that LYCHEE MEMORY achieves state-of-the-art performance, reaching 89.22% on LoCoMo and 92.20% on LongMemEval-S. Compared with A-Mem, it reduces construction tokens by 86.0% on LoCoMo and 75.9% on LongMemEval-S without increasing query-time token usage. On LoCoMo, LYCHEE MEMORY's largest gains over A-Mem with GPT-4.1-Mini appear in multi-hop reasoning (87.23% vs. 59.93%, +27.3 pp) and open-domain questions (67.71% vs. 42.71%, +25.0 pp). It also reaches 93.34% on single-hop questions, +20.1 pp above A-Mem (73.25%). On temporal reasoning, LYCHEE MEMORY reaches 86.60%, +13.7 pp above A-Mem (72.90%) and +38.9 pp above MemoryOS (47.66%). On LongMemEval-S, LYCHEE MEMORY's three largest gains over A-Mem with GPT-4.1-Mini occur in temporal reasoning (87.22% vs. 52.63%, +34.59 pp), preference tracking (90.00% vs. 63.33%, +26.67 pp), and multi-session reasoning (87.97% vs. 61.65%, +26.32 pp). It also reaches 97.44% on knowledge-update questions, compared with 82.05% for A-Mem.
In terms of cost, on LoCoMo, LYCHEE MEMORY uses only 204.1K construction tokens, 86.0% and 86.6% lower than A-Mem's 1459.9K and Mem0's 1520.8K respectively, and 58.3% lower than TiMem's 489.5K. On LongMemEval-S, LYCHEE MEMORY uses 304.7K, 75.9% lower than A-Mem's 1264.3K and 50.9% lower than TiMem's 620.9K. Despite substantial accuracy gains, LYCHEE MEMORY's query tokens do not increase. On LoCoMo, LYCHEE MEMORY uses 4.01K query tokens, lower than A-Mem's 5.56K (-27.9%) and TiMem's 10.71K (-62.6%). On LongMemEval-S, it uses 8.88K, lower than A-Mem's 15.46K (-42.6%) and TiMem's 11.36K (-21.8%).
Ablation studies show that replacing segment-level batching with eager construction causes accuracy to drop from 89.22% to 81.88% (-7.3 pp) while construction tokens increase from 204.1K to 849.9K (+316%). Fixed-window consolidation retains batching and uses slightly fewer construction tokens than the full system (174.7K vs. 204.1K), but its accuracy drops to 82.40% (-6.8 pp), with the largest losses on multi-hop (87.23% → 74.82%) and open-domain questions (67.71% → 56.25%). Replacing typed, self-contained records with summary-level records reduces construction tokens to 99.7K, while accuracy drops to 80.78% (-8.4 pp) and temporal reasoning falls from 86.60% to 73.83% (-12.8 pp). Removing cross-segment reference context leaves construction tokens nearly unchanged (189.6K vs. 204.1K) but reduces accuracy to 81.56% (-7.7 pp). Using record-vector retrieval only, query tokens remain comparable to the full system (4.09K vs. 4.01K) but accuracy drops to 81.75% (-7.5 pp). Removing the query planner reduces query tokens to 2.26K but accuracy drops to 83.38% (-5.8 pp). Jointly disabling fusion, reranking, and diversity-aware selection reduces accuracy from 89.22% to 66.62% (-22.6 pp). Boundary-threshold sensitivity analysis shows that overall accuracy ranges from 88.18% to 89.22% when varying the threshold δ from 0.30 to 0.70, a total variation of 1.04 percentage points.
The paper's contributions are threefold: identifying memory-construction granularity as an important efficiency lever for long-term agent memory; introducing LYCHEE MEMORY, a semantic segment-level memory construction framework that encodes each finalized segment into typed, self-contained records with bounded cross-segment context; and demonstrating a stronger accuracy–cost trade-off on LoCoMo and LongMemEval-S, improving long-term memory QA accuracy while reducing construction-token cost and query-time token usage relative to A-Mem.
Improvements for AI systems
Improvements to AI systems:
-
Implement semantic segment-level memory consolidation instead of turn-level or fixed-window consolidation. The AI system batches multiple exchanges into coherent segments based on semantic surprise (s t = 1 - max(sim(e t, c k), sim(e t, h k))) and cohesion drop signals, reducing LLM encoding calls from T (number of turns) to S (number of segments). This cuts construction tokens by 86% while improving accuracy by 7.3 percentage points over eager turn-level construction.
-
Encode memories as typed, self-contained records with fields for memory type (facts, preferences, events, constraints, procedures, failure patterns, tool affordances), entities, topics, normalized temporal scopes, and provenance links to source turns. This enables context-independent retrieval and interpretation, improving temporal reasoning by 12.8 percentage points over summary-level records.
-
Carry lightweight disambiguation feedback across segments (entity aliases and reference relations) rather than re-reading full history. This maintains continuity across segments and improves accuracy by 7.7 percentage points compared to omitting this cross-segment context.
-
Organize memories in an append-only structured store with five evidence-node classes: entity, topic, entity-topic, temporal, and event-frame nodes. This enables structured filtering alongside semantic retrieval, supporting multi-hop reasoning and open-domain questions with gains of +27.3 and +25.0 percentage points respectively over A-Mem.
-
Use plan-guided multi-route retrieval where a single LLM planning call decomposes the query into multiple recall routes (each with goal, search queries, structured constraints, temporal constraints). Execute four parallel recall channels (direct records, evidence nodes, temporal, raw turns), then fuse with reciprocal-rank fusion and apply diversity-aware selection. This improves evidence coverage without blind context expansion, boosting accuracy by 5.8 percentage points over planner-free retrieval.
-
Apply diversity-aware selection after fusion to preserve evidence from different routes, preventing over-concentration on a single retrieval path. Jointly disabling fusion, reranking, and diversity selection causes accuracy to drop by 22.6 percentage points.
-
Use bounded cross-segment reference context (recent record summaries truncated to a fixed budget) during encoding to resolve coreference and elliptical mentions without full history re-reading. This keeps construction costs low while maintaining coherence.
What the improved AI system can do:
-
Achieve state-of-the-art long-term memory QA accuracy: 89.22% on LoCoMo and 92.20% on LongMemEval-S using GPT-4.1-Mini.
-
Reduce memory construction token usage by 86.0% on LoCoMo and 75.9% on LongMemEval-S compared to A-Mem, while also reducing query-time tokens by 27.9% and 42.6% respectively.
-
Excel at multi-hop reasoning (87.23%), open-domain questions (67.71%), single-hop questions (93.34%), and temporal reasoning (86.60%) on LoCoMo.
-
Track user preferences over long sessions (90.00% accuracy), handle knowledge updates (97.44%), and reason across multiple sessions (87.97%) on LongMemEval-S.
-
Maintain robust performance across boundary-threshold variations (accuracy variation of only 1.04 percentage points when threshold changes from 0.30 to 0.70).
-
Operate with only one generative LLM call per segment during construction and one planning call per query, with all other operations being embedding lookups, structured filtering, or arithmetic scoring.
Sources
- Memanto: Typed Semantic Memory with Information-Theoretic Retrieval for Long-Horizon Agents
- ScrapMem: A Bio-inspired Framework for On-device Personalized Agent Memory via Optical Forgetting
- MemFlow: Intent-Driven Memory Orchestration for Small Language Model Agents
- Beyond Semantic Organization: Memory as Execution State Management for Long-Horizon Agents
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- G-Long: Graph-Enhanced Memory Management for Efficient Long-Term Dialogue Agents
- LightMem: Lightweight and Efficient Memory-Augmented Generation
- Organize then Retrieve: Hierarchical Memory Navigation for Efficient Agents
- MemRouter: Memory-as-Embedding Routing for Long-Term Conversational Agents
- Infini Memory: Maintainable Topic Documents for Long-Term LLM Agent Memory
- MemRefine: LLM-Guided Compression for Long-Term Agent Memory
- MemReranker: Reasoning-Aware Reranking for Agent Memory Retrieval
- TiMem: Temporal-Hierarchical Memory Consolidation for Long-Horizon Conversational Agents
- EMBER: Efficient Memory via Budgeted Evidence Retention for Long-Horizon Agents
- EviMem: Evidence-Gap-Driven Iterative Retrieval for Long-Term Conversational Memory
- MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models
- Evaluating Very Long-Term Conversational Memory of LLM Agents
- What Deserves Memory: Adaptive Memory Distillation for LLM Agents
- MemGPT: Towards LLMs as Operating Systems
- DMF: A Deterministic Memory Framework for Conversational AI Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering