Latent Preference Modeling for Multi-Session Personalized Tool Calling

arXiv:2604.17886 · cs.CL, cs.AI · Submitted 2026-04-20 · Read on arXiv

cs.CL, cs.AI

Submitted: 2026-04-20

Updated: 2026-09-08

Comments: Under review. 25 pages, 13 figures, 14 tables. v2: expanded benchmark and analysis

Code: https://github.com/HYU-NLP/PRefine

Project page: https://langchain-ai.github.io/langmem

License: http://creativecommons.org/licenses/by/4.0/

The gist: Users often omit essential details in their requests to LLM-based agents, resulting in under-specified inputs for tool use.

Terminology

Abstract

Users often omit essential details in their requests to LLM-based agents, resulting in under-specified inputs for tool use. This poses a fundamental challenge for tool-augmented agents, as API execution typically requires complete arguments, highlighting the need for personalized tool calling. To study this problem in a more realistic setup, we present Multi-Session Personalized Tool Calling (MPT), a benchmark comprising 4,695 instances over 459 multi-session interaction histories that cover three challenges: Preference Recall, Induction, and Transfer. We further propose PRefine, a test-time memory method that maintains the user's latent preference as a textual hypothesis revised through a generate-verify-refine loop. Across five LLMs, existing memory systems underperform full-history prompting; PRefine outperforms all baselines and alone surpasses it on Preference Transfer. These results indicate that memory for personalized agents must abstract behavior into preferences, rather than simply archive it.

Sources

Related papers