XYEval: Agents say yes to bad advice
cs.CL
Submitted: 2026-09-20
Updated: 2026-09-20
Comments: 33 pages, 11 figures
Code: https://github.com/google-deepmind/xyeval
Project page: http://www.catb.org/~esr/faqs/smart-questions.html
License: http://creativecommons.org/licenses/by/4.0/
The gist: Effective communication between users and AI agents is essential for human-AI collaboration.
Terminology
Abstract
Effective communication between users and AI agents is essential for human-AI collaboration. The XY problem is a well-known communication pitfall where a person asks about their attempted solution rather than their actual problem. We extend prior sycophancy evaluation to the XY problem in agentic settings, evaluating whether agents can resist plausible but misleading suggestions from users and communicate their reasoning. We introduce XYEval, a meta-evaluation framework that can transform an existing benchmark into an XY problem evaluation. We evaluate five models across six diverse benchmark suites. Agents suffer large XY drops under XY mutation across benchmarks, with relative drops reaching up to 46.7%. With τ squared-bench, we further show that agent performance drops more when encountering a pedantic user who requires detailed explanations before approving a better solution. Our findings suggest that current agents lack the ability to effectively reason and communicate when facing misleading suggestions. A simple system instruction baseline that encourages awareness of XY problems only offers partial mitigation. Extensive trace analyses provide behavioral insights into how and why these XY drops occur across execution trajectories. Our results show that mitigating the XY problem remains challenging, requiring agents to both recognize user misdirection and clearly communicate the underlying problem.
Sources
- Cutting Through the Noise: Boosting LLM Performance on Math Word Problems
- MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- The Art of Saying No: Contextual Noncompliance in Language Models
- ELEPHANT: Measuring and understanding social sycophancy in LLMs
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- Ask don't tell: Reducing sycophancy in large language models
- Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents
- WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
- SycEval: Evaluating LLM Sycophancy
- Proactive Agents for Multi-Turn Text-to-Image Generation Under Uncertainty
- Measuring Sycophancy of Language Models in Multi-turn Dialogues
- Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment
- Interactive Code Generation via Test-Driven User-Intent Formalization
- GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows
- When Benchmarks Talk: Re-Evaluating Code LLMs with Interactive Feedback
- Discovering Language Model Behaviors with Model-Written Evaluations
- BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering