Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents
cs.CL
Submitted: 2026-03-27
Updated: 2026-09-07
Comments: EMNLP 2026 camera-ready version
Code: https://github.com/nedwards99/ask-or-assume
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: As Large Language Model (LLM) agents are increasingly deployed in open-ended domains like software engineering, they frequently encounter underspecified instructions that lack crucial context.
Terminology
Abstract
As Large Language Model (LLM) agents are increasingly deployed in open-ended domains like software engineering, they frequently encounter underspecified instructions that lack crucial context. While human developers naturally resolve underspecification by asking clarifying questions, current agents are largely optimized for autonomous execution. In this work, we systematically evaluate the clarification-seeking abilities of LLM agents on an underspecified variant of SWE-bench Verified. We propose an uncertainty-aware multi-agent scaffold that decouples underspecification detection from code execution. Across both proprietary and open-weight frontier LLMs, our scaffold achieves a 69.40% task resolve rate, significantly outperforming a standard single-agent setup and closing the performance gap with agents operating on fully specified instructions. Furthermore, we find that the multi-agent system exhibits well-calibrated information-seeking behavior, conserving queries on simple tasks while proactively seeking information on more complex issues. These findings indicate that current models can be turned into proactive collaborators, where agents independently recognize when to ask questions to elicit missing information in real-world, underspecified tasks.
Sources
- Prompt Baking
- GPT-4o System Card
- Language Models (Mostly) Know What They Know
- Curiosity by Design: An LLM-based Coding Assistant Asking Clarification Questions
- ClarEval: A Benchmark for Evaluating Clarification Skills of Code Agents under Ambiguous Instructions
- Training Proactive and Personalized LLM Agents
- SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
- Asking What Matters: Reward-Driven Clarification for Software Engineering Tasks
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering