Making Prospective Memory SLM-Shaped: Typed Intention Stores for Small-Model Agents
cs.AI
Submitted: 2026-09-01
Updated: 2026-09-01
License: http://creativecommons.org/licenses/by/4.0/
The gist: Prospective memory means carrying out a deferred intention at the right future cue while other work continues.
Terminology
Abstract
Prospective memory means carrying out a deferred intention at the right future cue while other work continues. Benchmarks now isolate it as an agent skill, yet frontier LLMs still struggle: the best published PM-Bench scaffold reaches only 65.1% Set-F1. We argue that this loop is schema-constrained state tracking rather than open-ended reasoning, and that small models can execute it when the action space is typed. We propose the Prospective Intention Store (PIS) that puts lifecycle logic in code and scoped language work on the model. The scaffold is agentic and training-free: no selector fine-tuning and no trajectory distillation. On PM-Bench, DeepSeek-Chat with PIS reaches 82.9% Set-F1. On Gemma-E2B, Set-F1 is only 4.2% without a store and at most 6.6% under seven retrospective memories, while PIS reaches 66.2%. PIS further reaches 70.1% Set-F1, where retrospective memory methods stay at most 54.4%. PIS sets a new state of the art on this benchmark and enables small models to surpass the published large-model scaffold.
Sources
- Small Language Models are the Future of Agentic AI
- Octopus v2: On-device language model for super agent
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- DeepSeek-V3 Technical Report
- TinyAgent: Function Calling at the Edge
- LightMem: Lightweight and Efficient Memory-Augmented Generation
- Gemma 2: Improving Open Language Models at a Practical Size
- DuoMem: Towards Capable On-Device Memory Agents via Dual-Space Distillation
- Memory OS of AI Agent
- PM-Bench: Evaluating Prospective Memory in LLM Agents
- Evaluating Very Long-Term Conversational Memory of LLM Agents
- MemGPT: Towards LLMs as Operating Systems
- Gorilla: Large Language Model Connected with Massive APIs
- Qwen3 Technical Report
- Small Language Models for Agentic Systems: A Survey of Architectures, Capabilities, and Deployment Trade offs
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
- A-MEM: Agentic Memory for LLM Agents
- Lightweight LLM Agent Memory with Small Language Models
- TriggerBench: Investigating Prospective Memory for Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection