A Behavioral Trait Leaks into Preferences: Diagnosing Trait Interference in LLM User Simulators

arXiv:2609.25572 · cs.AI · Submitted 2026-09-22 · Read on arXiv

cs.AI

Submitted: 2026-09-22

Updated: 2026-09-22

Comments: CIKM 2026 short

DOI: 10.1145/3799682.3839914

Code: https://github.com/chaehyun1/PQA

License: http://creativecommons.org/licenses/by/4.0/

The gist: LLM-based user simulators aim to bridge the offline-online gap in recommender evaluation by emulating users through injected traits, where preference attributes determine what a user engages with and

Terminology

Abstract

LLM-based user simulators aim to bridge the offline-online gap in recommender evaluation by emulating users through injected traits, where preference attributes determine what a user engages with and a behavioral activity trait governs how long they browse. However, we show this intended trait independence collapses during simulation, causing two failures: (i) Trait Interference, where amplified activity distorts preference boundaries and forces interactions with mismatched items to sustain browsing, and (ii) Evaluation Invalidity, where satisfaction scores inflate with activity-driven page counts despite taste mismatches, biasing evaluation toward trait distributions rather than recommender performance. To resolve this, we propose PQA, a page-level quality anchoring method that guides simulators using a personalized anchor reflecting each user's intrinsic preference standard. By assessing whether a page meets this standard before further browsing, PQA enables proactive exits from low-quality pages, letting the activity trait retain its intended role of modulating browsing depth within preference-conforming pages. Experiments show PQA mitigates trait interference and improves the reliability of LLM-based simulator evaluation under activity shifts. Our code is available at https://github.com/chaehyun1/PQA

Related papers