Latent Preference Modeling for Multi-Session Personalized Tool Calling
cs.CL, cs.AI
Submitted: 2026-04-20
Updated: 2026-09-08
Comments: Under review. 25 pages, 13 figures, 14 tables. v2: expanded benchmark and analysis
Code: https://github.com/HYU-NLP/PRefine
Project page: https://langchain-ai.github.io/langmem
License: http://creativecommons.org/licenses/by/4.0/
The gist: Users often omit essential details in their requests to LLM-based agents, resulting in under-specified inputs for tool use.
Terminology
Abstract
Users often omit essential details in their requests to LLM-based agents, resulting in under-specified inputs for tool use. This poses a fundamental challenge for tool-augmented agents, as API execution typically requires complete arguments, highlighting the need for personalized tool calling. To study this problem in a more realistic setup, we present Multi-Session Personalized Tool Calling (MPT), a benchmark comprising 4,695 instances over 459 multi-session interaction histories that cover three challenges: Preference Recall, Induction, and Transfer. We further propose PRefine, a test-time memory method that maintains the user's latent preference as a textual hypothesis revised through a generate-verify-refine loop. Across five LLMs, existing memory systems underperform full-history prompting; PRefine outperforms all baselines and alone surpasses it on Preference Transfer. These results indicate that memory for personalized agents must abstract behavior into preferences, rather than simply archive it.
Sources
- T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic Planning
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks
- Advancing and Benchmarking Personalized Tool Invocation for LLMs
- CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
- FunctionChat-Bench: Comprehensive Evaluation of Language Models' Generative Capabilities in Korean Tool-use Dialogs
- LatentCRS: A Variational EM Framework for Bridging Semantics and Behavior in LLM-based Conversational Recommendation
- SimpleMem: Efficient Lifelong Memory for LLM Agents
- MultiSessionCollab: Learning User Preferences with Memory to Improve Long-Term Collaboration
- Generative Agents: Interactive Simulacra of Human Behavior
- Towards Scalable Multi-domain Conversational Agents: The Schema-Guided Dialogue Dataset
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Memex(RL): Scaling Long-Horizon LLM Agents via Indexed Experience Memory
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
- A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications
- MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering