SimGuide: Typed Multi-Context User Representations for Preference-Conditioned Agent Planning
cs.AI
Submitted: 2026-05-14
Updated: 2026-08-29
Comments: Previous version had some results that were hard to verify or reproduce, so new experiments were done and many parts of the paper were changed to reflect that
Code: https://github.com/VersarAI/SimBench
License: http://creativecommons.org/licenses/by/4.0/
The gist: Agents that act on a user's behalf must plan differently for different users, and increasingly do so from some structured representation of user context and not from raw interaction history.
Terminology
Abstract
Agents that act on a user's behalf must plan differently for different users, and increasingly do so from some structured representation of user context and not from raw interaction history. How much that structure is worth, and which parts of it carry the value, is largely unmeasured. We introduce SimBench, 47 preference-conditioned planning tasks over 9 synthetic users represented as 28 typed, potentially conflicting context blocks, where the correct plan depends on which contexts are active and how their conflicts are resolved. Against it we evaluate SimGuide, a framework combining typed multi-context representation, explicit conflict arbitration, and optional procedural grounding of individual constraints. Across three models and six user-context representations, SimGuide's typed blocks with arbitration outperform retrieval over the same user's past decisions by +0.210, +0.205 and +0.144 Preference Adherence on Llama 3.3 70B, GPT-4o and Claude Sonnet 4.5 respectively (all p < 0.001); removing the arbitration instruction alone costs up to +0.209. Grounding each constraint with a worked example of past application helps only where the model has headroom: +0.094 on Llama 70B (p < 0.001), falling to +0.029 on GPT-4o and +0.002 on Claude, which already scores perfectly on 35 of 47 tasks without it. We report the benchmark's minimum detectable effect alongside its results. The benchmark ships with a provenance audit that re-derives every reported number from the prompt that produced it.
Sources
- PREDICT: Preference Reasoning by Evaluating Decomposed preferences Inferred from Candidate Trajectories
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Mistral 7B
- AgentBench: Evaluating LLMs as Agents
- GPT-4o System Card
- MemGPT: Towards LLMs as Operating Systems
- LLMs + Persona-Plug = Personalized LLMs
- Personalized Pieces: Efficient Personalized Large Language Models through Collaborative Efforts
- Instant Personalized Large Language Model Adaptation via Hypernetwork
- Personalized Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection