APeB: Benchmarking Personalization Ability of Large Language Model Agents
cs.AI, cs.HC
Submitted: 2026-07-03
Updated: 2026-08-27
Comments: NA
Code: https://github.com/bytedance/deer-flow
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLM-powered agents struggle with personalization when users issue raw, underspecified queries.
Terminology
Abstract
LLM-powered agents struggle with personalization when users issue raw, underspecified queries. In this setting, agents must infer latent intent, extract preferences from noisy interaction histories, and select among competing alternatives. Existing benchmarks rarely test this capability, as they often rely on user-refined queries or simplified histories. We introduce personalized product search (PPS), a testbed for agentic personalization under raw queries and diverse histories. We construct Agent Personalized Benchmark (APeB) from action logs, pairing underspecified intents with rich histories and user-viewed candidate items. Evaluating state-of-the-art LLMs with multi-step agent workflows, we find that models handle explicit queries well but struggle with early-stage queries requiring intent and preference discovery. Rubric analysis attributes this gap mainly to ineffective history use. A simple history-aware query-refinement pipeline, VQRA, yields consistent gains, highlighting the need for dedicated history-utilization modules in personalized agents.
Sources
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
- Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- OneRec: Unifying Retrieve and Rank with Generative Recommender and Iterative Preference Alignment
- Rethinking Memory in LLM based Agents: Representations, Operations, and Emerging Topics
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- A Deep Look into Neural Ranking Models for Information Retrieval
- OneRec-Think: In-Text Reasoning for Generative Recommendation
- Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders
- Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
- DeepShop: A Benchmark for Deep Research Shopping Agents
- Mem-PAL: Towards Memory-based Personalized Dialogue Assistants for Long-term User-Agent Interaction
- On the Way to LLM Personalization: Learning to Remember User Conversations
- Evaluating Very Long-Term Conversational Memory of LLM Agents
- Recommender Systems with Generative Retrieval
- Shopping Queries Dataset: A Large-Scale ESCI Benchmark for Improving Product Search
- A Survey of WebAgents: Towards Next-Generation AI Agents for Web Automation with Large Foundation Models
- BPR: Bayesian Personalized Ranking from Implicit Feedback
- Integrating Summarization and Retrieval for Enhanced Personalization via Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection