Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
cs.CL, cs.AI
Submitted: 2026-04-29
Updated: 2026-09-12
Code: https://github.com/tatsu-lab/alpaca_eval
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
- AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
- Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
- Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
- Accumulating Context Changes the Beliefs of Language Models
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
- WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
- Prioritization First, Principles Second: An Adaptive Interpretation of Helpful, Honest, and Harmless Principles
- LLMs Get Lost In Multi-Turn Conversation
- StructFlowBench: A Structured Flow Benchmark for Multi-turn Instruction Following
- OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment
- Learning Query-Specific Rubrics from Human Preferences for DeepResearch Report Generation
- DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
- Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks
- CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models
- Checklists Are Better Than Reward Models For Aligning Language Models
- Tradeoffs Between Alignment and Helpfulness in Language Models with Steering Methods
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering