PSD: Pseudo Self-Distillation of Memory Representation Capabilities for LLM Agents
cs.IR, cs.CL, cs.LG
Submitted: 2026-09-20
Updated: 2026-09-20
License: http://creativecommons.org/licenses/by/4.0/
The gist: Memory systems are becoming a core component of LLM agents, but constructing and maintaining memory remains expensive because it relies on repeated calls to large proprietary language models.
Terminology
Abstract
Memory systems are becoming a core component of LLM agents, but constructing and maintaining memory remains expensive because it relies on repeated calls to large proprietary language models. This cost creates a major barrier to deploying memory-enhanced agents at scale. In this paper, we present Pseudo Self-Distillation (PSD), a framework that enables small language models (SLMs) to construct hierarchical memory representations by distilling behavior from a strong black-box oracle through a multi-stage training pipeline. Standard distillation methods require access to teacher logits or hidden states, which closed models do not expose. Unlike conventional self-distillation settings, where supervision is derived from a model's own predictions, sampled rollouts, or aggregated outputs, PSD enables a single-model distillation setup while channeling external oracle knowledge through the prompt. PSD uses a single small model in two roles: a teacher that sees a privileged prompt containing the oracle's answer as reference context, and a student that sees only the task prompt. The student learns to reproduce the teacher's output distribution, absorbing oracle-guided behavior into its own weights without accessing the oracle's internals. On LoCoMo, PSD-trained Qwen3-0.6B, 1.7B, and 4B match or exceed GPT-4.1-mini on downstream retrieval at a fraction of the deployment cost, with off-policy PSD achieving the strongest results across most conditions. We further show that this memory-construction capability transfers out-of-distribution to LongMemEval, despite the students being trained exclusively on LoCoMo with no exposure to LongMemEval data.
Sources
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Retrieval-Augmented Generation with Graphs (GraphRAG)
- Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision
- Distilling the Knowledge in a Neural Network
- Reinforcement Learning via Self-Distillation
- Stable On-Policy Distillation through Adaptive Target Reformulation
- TinyBERT: Distilling BERT for Natural Language Understanding
- Memory OS of AI Agent
- Rethinking On-Policy Self-Distillation for Thinking Models
- Sequence-Level Knowledge Distillation
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
- What Deserves Memory: Adaptive Memory Distillation for LLM Agents
- Evaluating Very Long-Term Conversational Memory of LLM Agents
- MemGPT: Towards LLMs as Operating Systems
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- Purified OPSD: On-Policy Self-Distillation Without Losing How to Think
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
- MIRIX: Multi-Agent Memory System for LLM-Based Agents
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG