Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations

arXiv:2604.27093 · cs.CL, cs.AI · Submitted 2026-04-29 · Read on arXiv

cs.CL, cs.AI

Submitted: 2026-04-29

Updated: 2026-09-12

Code: https://github.com/tatsu-lab/alpaca_eval

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Terminology

Sources

Related papers