Learning What Not to Forget: Long-Horizon Agent Memory from a Few Kilobytes of Learning

arXiv:2606.20954 · cs.CL, cs.AI · Submitted 2026-06-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Learning What Not to Forget: Long-Horizon Agent Memory from a Few Kilobytes of Learning".

Jane: The paper was written by Nusrat Jahan Lia and Aritra Mazumder from Institute of Information Technology, University of Dhaka and University of Utah.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: Now that we understand why memory management was such a problem, let’s look at what LRE actually does according to the paper's summary.

Jane: The researchers are proposing a system called LRE, which is designed to be much more intelligent than standard methods like FIFO or simple decay models.

Lu: It learns which historical units are critical—the load-bearing ones—and keeps them by verbatim extraction, meaning it doesn't paraphrase or lose the exact content.

Meng: This is a massive difference from using an LLM to compress the memory, because we avoid lossy compression where important details might be summarized away.

Lalam: It seems like LRE allows the AI to maintain its state in a way that preserves its own past, which is vital for complex, multi-step tasks.

Tom: But how does it decide what's relevant without a human telling it?

Jane: That’s where the learned relevance score comes in—it calculates a probability of relevance for every single unit of history.

Lu: The genius of that is that this scoring function relies only on the prefix history, meaning the AI makes its decision proactively based on past state alone.

Meng: It doesn't need to wait for a specific downstream question; it just needs to look at the history and decide if this unit is critical, which makes it highly proactive and deployable.

Lalam: I think that ability to be query-blind suggests that the AI has developed an internal sense of its own operational requirements.

Improvements: Tom: We've seen how LRE works, but let’s talk about the tangible results—how much better is it actually performing compared to other methods?

Jane: The data shows that LRE maintains accuracy comparable to keeping the entire history, which is a huge achievement given its efficiency.

Lu: And we are seeing real performance gains; specifically, on easier tasks, it exceeds the no-eviction baseline by twenty-seven percent relative gain.

Meng: Even better, this reduction in context size is achieved with zero compressor calls, meaning there’s no extra latency or computational load from running large language models just to summarize the memory.

Lalam: That’s a relief for system reliability; avoiding those costly LLM calls makes LRE incredibly practical and avoids potential failure points in the complex workflow.

Tom: It also seems to solve problems that other methods completely miss, like in AppWorld tasks where LRE completed fourteen tasks that no other run policy could achieve.

Jane: And it’s doing this with a much smaller footprint; the paper shows LRE requires only a few kilobytes of learning—less than the size of a single photo—to function effectively.

Lu: That’s a huge win for efficiency; we are looking at up to fifty-two percent reduction in peak context size, which is massive for scaling up AI systems.

Meng: From an engineering view, this reduction is substantial because it's not just a theoretical benefit; it's a practical way to handle the growth of long-running agent memory.

Lalam: I think this allows the AI to maintain its operational state with much greater stability and less resource strain.

Conclusion: Tom: We have seen how LRE is both efficient and effective, so let’s bring all this together and wrap up the discussion on "Learning What Not to Forget: Long-Horizon Agent Memory from a Few Kilobytes of Learning."

Jane: It seems like we've established that this method is not just a clever heuristic but a genuine advance in how AI handles its own memory, offering real predictive power.

Lu: The fact that it can be trained on the system’s own behavior, without needing expensive human annotation, is an incredible breakthrough for future development in autonomous systems.

Meng: We are looking at a scalable architecture here; LRE provides a deployment-ready solution that minimizes inference overhead while maximizing reliability.

Lalam: I feel this research has the potential to change how we design agents entirely, allowing them to possess true long-term coherence and intelligence in their memory.

Tom: That’s certainly an ambitious goal for the AI community right now.

Jane: It's clear that this paper presents a solution that is both powerful and incredibly efficient for long-horizon AI agents who operate over many steps.

Lu: It’s especially impressive how well it works across different domains, from conversational memory to complex agentic tool use.

Meng: The practical implications of achieving this level of performance with minimal compute cost are huge, and it's very encouraging for the industry.

Lalam: This is truly a wonderful example of learning what not to forget, ensuring that our AI systems have the stable, reliable memory they need to improve our culture and interactions.

Wrap-up: Tom: We’ve explored the title, the mechanism of LRE, its impressive performance gains, and what it does for the future.

Jane: It’s clear that this paper presents a solution that is both powerful and incredibly efficient for long-horizon AI agents.

Lu: I am particularly excited about how we can use self-supervised labels to unlock new possibilities in agent training, moving beyond traditional human annotation methods.

Meng: The practical implications of achieving this level of performance with minimal compute cost are huge, and it's very encouraging for the industry's AI infrastructure.

Lalam: It is a wonderful example of learning what not to forget, ensuring that our AI systems have the stable, reliable memory they need to improve our culture and interactions.

Institute of Information Technology, University of Dhaka · University of Utah

cs.CL, cs.AI

Submitted: 2026-06-18

Updated: 2026-09-03

Code: https://github.com/microsoft/acon

Importance score: 84/100

The gist: The paper introduces a novel approach, Learning Relevance Encoding (LRE), designed to solve the critical problem of managing context in long-horizon agent interactions where full context retention

Key concepts

LRE (Long-Horizon Agent Memory)
LRE is a proposed system for AI memory management that learns which historical units are critical to keep. Unlike standard methods or LLM compression, it preserves the exact content of important history without paraphrasing. It allows the AI to maintain its operational state effectively over long tasks.
Learned Relevance Score
This is a function used by LRE to determine which parts of the past are important. The AI calculates a probability of relevance for every unit of history based only on the current prefix history, making it proactive and deciding what to keep without needing a human to define importance.

Terminology

Summary

The paper introduces a novel approach, Learning Relevance Encoding (LRE), designed to solve the critical problem of managing context in long-horizon agent interactions where full context retention becomes computationally prohibitive. LRE proposes a mechanism that proactively determines which past experiences are most relevant for future task success, allowing agents to operate within strict memory budgets while maintaining high performance. This work is crucial because it provides a deployable policy that shifts the focus from simply retaining all data to intelligently curating the most impactful information, significantly improving efficiency in complex conversational and tool-use scenarios.

Learning Relevance Encoding (LRE) Mechanism

LRE operates fundamentally outside of the core transformer model, which distinguishes it from token-level methods like H2O or SnapKV. The system's unit of eviction is defined as the agent's action, observation pair, while the conversational unit is the dialogue turn or session. The method relies on a learned scorer that operates on a compact set of causal features designed to be inexpensive, model-agnostic, and available at decision time. Crucially, LRE achieves its relevance assessment by requiring only logged behavior, eliminating the need for complex model-internal signals. The system's core claim is that future utility can be learned from these causal signals available at write time.

Performance in Constrained Environments

The empirical evaluation demonstrates LRE's superior performance, particularly when operating under resource constraints. On the LoCoMo benchmark (token-F1), LRE variants significantly outperform baseline policies:

  • LRE-self-supervised achieved an accuracy of 0.276 with a context token count of 26.1k.

  • LRE-supervised achieved an accuracy of 0.296 with a context token count of 38.8k, showing that LRE leads the budget-constrained policies.

Similarly, on LongMemEvalS (LLM-judge), where the full context was found to be suboptimal (the full context is worst), LRE variants clustered with other deployable policies, demonstrating robust performance even when the ideal retention strategy is unknown.

Computational Efficiency and Budget Savings

LRE offers substantial operational cost savings by minimizing reliance on expensive neural forward passes. The method requires zero neural forward passes and scores resident units on a CPU, contrasting sharply with dense salience or LLMLingua-2 encoders, which require one forward pass per unit and are 102 –103 × slower. At an iso-recall threshold of 80% gold recall, LRE reaches the target using significantly fewer resources:

  • LRE-self-supervised achieved a budget saving of 50.8% compared to recency.

  • LRE-supervised achieved a budget saving of 41.5% compared to dense salience (BGE).

Deployment Considerations and Limitations

The authors provide several caveats regarding deployment, emphasizing that LRE is a reference implementation rather than a fixed specification. Key limitations include:

  • Feature-set Dependence: While the learned scorer uses an inexpensive feature set, the authors caution that Different deployments may expose different signals... [and] should measure and validate its own candidate features from logged behavior before deployment.

  • Budget-Bounded Performance: LRE is described as complementing, rather than replacing, unbounded retention. The claim of Pareto-efficiency holds only when the budget is binding.

  • Unit Granularity: The current unit of eviction is text; the system has not been extended to handle multimodal agents whose state includes images or audio.

Improvements for AI systems

Based on this material, the primary area for improvement is moving LLM deployment from naive full-context retention to Intelligent, Cost-Aware Context Management. The goal is to maintain high downstream performance while drastically reducing inference cost and computational overhead.

Here are the specific improvements and what the resulting AI system can achieve:


This system integrates a learned, external memory mechanism into the standard LLM inference pipeline, making context retention proactive rather than reactive.

  • Implementation of the LRE Scoring Module: Integrate a dedicated, lightweight scoring module outside the main transformer block. This module takes logged historical units (the (action, observation) pairs) and outputs a Utility Score for each unit based on learned causal features (as described by LRE).

  • Specificity: This scorer must be trained to predict future utility (e.g., answer recall or task success) given the current state, rather than just predicting similarity or recency.

  • Dynamic Context Budgeting Unit: Implement a real-time cost monitor that tracks the token cost of the context window against a predefined operational budget (Budget Op). This unit dictates how much context can be retained.

  • Specificity: When Context Tokens > Budget Op, the Memory Orchestrator triggers an eviction cycle using the LRE scores to prune units until the budget is met, ensuring Pareto-efficiency within operational constraints.

  • Hybrid Retention Layer: Design a multi-tiered memory structure:

  1. Short-Term (KV Cache): Standard transformer mechanism for immediate context.

  2. Mid-Term (Active Buffer): A small, high-priority buffer holding the most critical, LRE-scored units that fit within the current token window limit (Budget Op).

  3. Long-Term (External Memory Bank): A persistent database storing all historical units, indexed by their LRE Utility Score and associated metadata (e.g., task ID, session turn).

  • Targeted Self-Supervised Fine-Tuning: Instead of using a fixed answer-recall threshold (tau=0.4), the system must be trained in an iterative loop:
  1. Run a set of logged interactions on the LLM (Host Model).

  2. Calculate performance metrics (e.g., Recall Answer).

  3. Adjust the internal LRE scoring weights (theta LRE) to maximize Recall Answer given a fixed budget, treating the score itself as a trainable loss component during fine-tuning (i.e., optimizing for utility, not just feature correlation).

  • Hardware-Aware Deployment Profiling: For deployment, integrate profiling tools that estimate the real serving cost of different strategies (CPU vs. T4/GPU). The system must dynamically select the memory strategy (e.g., LRE on CPU vs. Dense Salience on GPU) based on the current cost tolerance set by the business unit.

The resulting ADUMO system transforms an LLM from a memory-limited black box into a Sustained, Cost-Optimized Reasoning Agent.

  1. Achieve Contextual Persistence Beyond Window Limits: It can maintain coherence and deep understanding across extremely long interactions (thousands of turns) where standard LLMs would forget initial premises due to context window overflow.

  2. Guarantee Performance Under Budget Constraints: Critically, it guarantees that even when forced to operate under severe token budgets (e.g., mobile deployment or high-volume API calls), the retained context is the most valuable information necessary for the next action, maximizing downstream task success (LoCoMo/LongMemEvalS performance) while minimizing cost.

  3. Provide Transparent Cost-Benefit Analysis: The system doesn't just run; it reports its efficiency: To achieve a predicted Recall Answer of X, we retained Y units, costing only Z% of the operational budget. This allows engineering teams to make financially informed deployment decisions.

  4. Support Mixed-Modality Evolution: The architecture is explicitly designed for extension. By abstracting the unit as a structured, scored memory block, it can readily incorporate multimodal data (images, audio embeddings) into its scoring mechanism simply by expanding the unit representation and cost function, without requiring retraining of the core LLM.

Sources

Related papers