TRACER: Trajectory-Aligned Learning for Multi-Turn User Simulation
cs.AI
Submitted: 2026-09-23
Updated: 2026-09-28
Terminology
Sources
- MOA: Multi-Objective Alignment for Role-Playing Agents
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior Data
- Flipping the Dialogue: Training and Evaluating User Language Models
- LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
- UserBench: An Interactive Gym Environment for User-Centric Agents
- UserRL: Training Interactive User-Centric Agent via Reinforcement Learning
- Interactive Agents: Simulating Counselor-Client Psychological Counseling via Role-Playing LLM-to-LLM Interactions
- Qwen2.5 Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- HybridFlow: A Flexible and Efficient RLHF Framework
- RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents
- RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward
- HumanLM: Simulating Users with State Alignment Beats Response Imitation
- Qwen3 Technical Report
- Harmonizing Dense and Sparse Signals in Multi-turn RL: Dual-Horizon Credit Assignment for Industrial Sales Agents
- CPO: Addressing Reward Ambiguity in Role-playing Dialogue via Comparative Policy Optimization
- Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models
- UserLM-R1: Modeling Human Reasoning in User Language Models with Multi-Reward Reinforcement Learning
- LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection