Clarify, Abstain or Answer? Strategising in Conversation with Belief-Augmented Generation
cs.CL, cs.AI, cs.LG
Submitted: 2026-05-25
Updated: 2026-09-07
Comments: Accepted at INLG 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) define a distribution over text, which can be viewed as a probabilistic representation of uncertainty: sampling K responses yields a belief state - responses a model
Terminology
Abstract
Large language models (LLMs) define a distribution over text, which can be viewed as a probabilistic representation of uncertainty: sampling K responses yields a belief state - responses a model deems plausible. Existing work exploits this representation for narrow tasks like either decoding or selective prediction, and often requires manual interventions, not controlling generation directly. We propose Belief-Augmented Generation (BAG): grounding LLMs in their own belief state via the prompt and letting them reason over these K samples to decide on and execute a conversational strategy: clarify, abstain, or answer. In a multi-turn ambiguous question answering (QA) setting, we find that LLMs by default rarely clarify or abstain, ignoring uncertainty about the input (aleatoric) or facts (epistemic). BAG improves QA accuracy across six models and yields strategy decisions more faithful to their belief state than prompt-only baselines. Disentangling when to clarify from when to abstain, however, remains challenging.
Sources
- Learning Steerable Clarification Policies with Collaborative Self-play
- Uncertainty in Natural Language Generation: From Theory to Applications
- Teaching Language Models to Faithfully Express their Uncertainty
- Olmo 3
- Reasoning about Intent for Ambiguous Requests
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering