TransMem: Transforming Hidden States into Memory for Large Language Models
Haodong Lei, Junming Liu, Yirong Chen, Pinlong Cai, Botian Shi, Ding Wang, Hongsong Wang
cs.MA, cs.CL
Submitted: 2026-07-31
Comments: 12 pages, 4 figures
Code: https://github.com/Haodong-Lei-Ray/TransMem
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language model (LLM) agents increasingly operate over long interaction histories, where effective reasoning requires identifying and exploiting task-relevant evidence distributed across past
Terminology
Abstract
Large language model (LLM) agents increasingly operate over long interaction histories, where effective reasoning requires identifying and exploiting task-relevant evidence distributed across past observations and actions. However, useful information encoded in previously computed representations is often underutilized during subsequent generation. We propose TransMem, a lightweight inference-time parametric memory module that transforms sparse historical hidden states from a frozen LLM backbone into reusable memory representations. TransMem uses a lightweight gating network to dynamically apply the latent intervention to the current hidden states, without repeatedly encoding the preceding context. To learn transferable memory utilization rather than task-specific knowledge, we introduce evidence-conditioned self-distillation. A memory-augmented student processes the full context and matches the predictive distribution of an evidence-only teacher that shares the same frozen backbone. Experiments on LoCoMo, HotpotQA, and MemoryAgentBench demonstrate consistent improvements across different model architectures and scales. TransMem yields gains of 11.58--29.25 F 1 on LoCoMo and 10.20--13.03 F 1 on HotpotQA, while improving the average MemoryAgentBench accuracy from 29.54% to 40.00%. These results establish sparse historical hidden states as an effective and efficient memory substrate for long-context LLM agents. Our code is available at https://github.com/Haodong-Lei-Ray/TransMem.
Sources
- Understanding LoRA as Knowledge Memory: An Empirical Analysis
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- DeepSeek-V3 Technical Report
- Rethinking Memory in LLM based Agents: Representations, Operations, and Emerging Topics
- The Llama 3 Herd of Models
- Parametric Memory Decoding for Zero-Shot Routing in LoRA-Based External Parametric Memory
- Hindsight is 20/20: Building Agent Memory that Retains, Recalls, and Reflects
- MemCoT: Test-Time Scaling through Memory-Driven Chain-of-Thought
- delta-mem: Efficient Online Memory for Large Language Models
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- SimpleMem: Efficient Lifelong Memory for LLM Agents
- MemVerse: Multimodal Memory for Lifelong Learning Agents
- MemGPT: Towards LLMs as Operating Systems
- MeMo: Memory as a Model
- Qwen2.5 Technical Report
- Qwen3 Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- EffGen: Enabling Small Language Models as Capable Autonomous Agents
- Rethinking How to Remember: Beyond Atomic Facts in Lifelong LLM Agent Memory
- MemSifter: Offloading LLM Memory Retrieval via Outcome-Driven Proxy Reasoning
Related papers
- Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control
- You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents
- Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems
- MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization
- PeroMAS: A Multi-agent System of Perovskite Material Discovery
- StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning