Disclosure-Gated User Simulation for Companion-Agent Evaluation
cs.CL, cs.AI, cs.HC
Submitted: 2026-09-01
Updated: 2026-09-01
Code: https://github.com/liuyaox/CompanionBench
License: http://creativecommons.org/licenses/by/4.0/
The gist: Using a large language model to play the user is now standard in scalable evaluation.
Terminology
Abstract
Using a large language model to play the user is now standard in scalable evaluation. It has a repeatedly diagnosed failure: the simulated user is excessively cooperative, so a system under test can score by the sheer number of questions it asks rather than by making the user willing to speak. We answer with a disclosure gate conditioning information release on the companion agent's behaviour: its state is a ladder of five ordered gates, merged onto three observable depth layers. We specify, ablate, and audit it, and train a user simulator against that specification. Gating behaviour is learned from the training corpus's synthetic branch, while the real branch supplies how people speak and react; after training, the simulator need not be told at runtime which gate each item sits behind. The gate is a load-bearing component of the environment: on the English corpus of a published companion-agent benchmark (CompanionBench), once training no longer states per example which gate each item sits behind, the largest rank displacement across 12 systems under test exceeds the noise band set by re-running that environment under a new seed, while per-system scores show no detectable change. We state two acceptance criteria: a ranking must be order-preserving, and absolute scores must be scale-stable. Of the candidates we examine, only one passes both -- the simulator we release -- and its leaderboard correlates at 0.993 with the benchmark's original simulator. By contrast, prompting a frontier model as the simulator barely moves the ranking while shifting every score upward -- a shift invisible to anyone checking the ranking alone. The environment we specify is the one that benchmark already used. That publication describes the mechanism in about four hundred words, and we supply what it lacked: specification, ablations, human studies, negative controls, and downstream sensitivity analysis.
Sources
- The Empirically Grounded Adaptive Virtual Patient for Psychotherapy Training: Disclosure That Responds to Therapist Micro-Skills
- Simulated Customers Never Walk Away: Decision Fidelity of LLM User Simulators Measured Against Real Purchase Outcomes
- Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
- MedConceal: A Benchmark for Clinical Hidden-Concern Reasoning Under Partial Observability
- PSI-Bench: Interpretable and Clinically Meaningful Evaluation of Depression Patient Simulators
- Large Language Models Pass the Turing Test
- CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship
- Quantifying Variance in Evaluation Benchmarks
- Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants
- Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
- HumanLM: Simulating Users with State Alignment Beats Response Imitation
- Mind the Sim2Real Gap in User Simulation for Agentic Tasks
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering