CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement

arXiv:2608.11631 · cs.AI · Submitted 2026-08-12 · Read on arXiv

Kuangzhao Yang, Ziliang Zhao, Zhicheng Dou

Gaoling School of Artificial Intelligence, Renmin University of China

cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: 11 pages, 4 figures, and 3 tables

Code: https://github.com/ykun49365/CLAIM-final

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement Summary: This paper proposes CLAIM, an uncertainty-driven framework for active clarification

Terminology

Summary

CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement

Summary:

This paper proposes CLAIM, an uncertainty-driven framework for active clarification learning in open-domain human–LLM interactions. The authors address two fundamental challenges in clarification: determining when clarification is necessary and which aspect of the query should be clarified. Existing approaches rely on manually annotated data or preference alignment, which incurs high annotation costs and limits generalization. CLAIM eliminates the need for explicit human preference annotations by quantifying query uncertainty through the entropy induced by answer disagreements across multiple models.

The core intuition is that "When a user query is well-specified, different large language models tend to produce semantically consistent responses; in contrast, when the query is ambiguous or underspecified, the responses generated by different models often diverge substantially at the semantic level." This semantic divergence reflects uncertainty and provides a strong intrinsic signal for whether clarification is needed.

The framework consists of several stages:

  1. Entropy-driven Uncertainty Estimation: Given a user query q, CLAIM uses k1 different LLMs to independently generate candidate direct answers. These answers are semantically clustered, and the entropy E1(q) of the resulting answer distribution is computed as a measure of query uncertainty.

  2. Clarification Judgement: Combines entropy-based judgement (threshold τ = 0.45) with LLM-based semantic completeness judgement. When the two signals disagree, a conflict resolution model arbitrates the final decision.

  3. Clarifying Question Generation: For queries requiring clarification, multiple diverse candidate clarifying questions are generated with dimension labels, using history-based constraints to encourage diversity.

  4. Information Gain-based Selection: Each candidate question is evaluated by computing post-clarification entropy E2(q, cq, A) after simulating user answers, and the information gain IG(q, cq) = E1(q) − E2(q, cq, A) is used to select the optimal clarifying question.

  5. Training: A two-stage paradigm combining supervised fine-tuning (SFT) with group-relative policy optimization (GRPO) to learn a unified clarification decision model.

The authors distinguish between CLAIM-Agent (the offline multi-model data-construction pipeline) and CLAIM (the trained single-model policy used for online inference). The multi-model cost is paid only during offline synthetic data construction; after training, CLAIM performs one standard model inference per user query.

Experiments are conducted on three benchmarks: ClariLM-test (synthetic clarification dataset), IN3 (task-oriented interactive scenarios), and CLAMBER (general open-domain clarification). Results show that CLAIM achieves state-of-the-art or near-SOTA performance on most metrics across all datasets. Key findings include:

  • CLAIM matches or outperforms ClariLM using only 10k uncertainty-constructed training instances versus ClariLM's 120k supervised examples.

  • Both entropy-based and LLM-based judgement signals are complementary; relying on either alone is insufficient.

  • Information gain-based selection significantly improves clarification quality (CDA and CQSS).

  • GRPO further improves decision stability and clarification quality.

  • SFT-Full generalizes better across domains than SFT-IN3, which overfits to its specific domain.

  • LLM-as-a-Judge and human evaluations confirm that CLAIM improves user-perceived clarification behavior.

The paper concludes that CLAIM demonstrates the effectiveness of uncertainty-aware modeling for general-domain clarification, offering a low-cost and robust solution for proactive understanding in real-world open-domain interactions with LLMs. Future work includes extending CLAIM to multi-turn interaction with dialogue-state tracking and history-dependent uncertainty estimation.

Improvements for AI systems

Improvements to AI Systems:

  1. Uncertainty-Aware Query Disambiguation: Integrate CLAIM’s entropy-based uncertainty estimation into AI assistants (e.g., chatbots, search engines) to automatically detect ambiguous or underspecified user queries without needing human-labeled clarification data. The AI can then proactively ask a single, high-information clarifying question only when needed, reducing unnecessary back-and-forth.

  2. Low-Cost Clarification Policy Learning: Replace expensive preference-annotation pipelines with the offline multi-model disagreement signal (k1 LLMs) to generate synthetic training data. This enables smaller or domain-specific AI systems to learn clarification behavior with 10k instances instead of 120k, making it feasible for resource-constrained deployments.

  3. Information-Gain-Based Question Selection: Use the post-clarification entropy reduction (IG = E1 − E2) to rank candidate clarifying questions. The AI system can select the question that maximizes expected information gain, leading to faster and more accurate user intent resolution, especially in task-oriented scenarios (e.g., booking, troubleshooting).

  4. Complementary Judgement with Conflict Resolution: Implement a dual-signal decision mechanism (entropy threshold + LLM semantic completeness) with an arbitration model for disagreements. This improves robustness—e.g., avoiding unnecessary clarification when the query is semantically complete but entropy is high due to stylistic variation, or clarifying when entropy is low but the query is genuinely incomplete.

  5. Domain-Generalizable Clarification via SFT-Full: Train the clarification policy on a broad, multi-domain synthetic dataset (SFT-Full) rather than a narrow domain-specific one. This allows the AI system to transfer clarification skills across unseen domains (e.g., from tech support to medical advice) without retraining, reducing overfitting and improving real-world adaptability.

  6. Stable Decision-Making via GRPO: Apply group-relative policy optimization to fine-tune the clarification model, improving decision stability (fewer inconsistent clarification choices for similar queries) and overall clarification quality. This makes the AI’s proactive behavior more predictable and trustworthy for users.

  7. Unified Single-Model Inference: After offline multi-model data construction, deploy a single trained model (CLAIM) for online inference, requiring only one standard model call per query. This makes the improvement computationally efficient and scalable for production AI systems, with no extra latency or cost at inference time.

What the Improved AI System Can Do:

  • Proactively ask the right clarifying question at the right time, only when the query is genuinely ambiguous, reducing user frustration from irrelevant or excessive questions.

  • Handle open-domain queries (e.g., vague requests like “Tell me about the weather” or “Recommend a good book”) by generating diverse, dimension-labeled clarifying questions (e.g., location, genre, time) and selecting the most informative one.

  • Operate without human annotation for clarification, enabling rapid deployment in new languages or specialized domains where labeled data is scarce.

  • Improve task completion efficiency in interactive systems (e.g., virtual assistants, customer support bots) by reducing the number of turns needed to resolve user intent, while maintaining high accuracy.

  • Adapt to user feedback through the simulated answer entropy, allowing the system to refine its understanding even when the user’s initial query is vague or multi-intent.

  • Provide consistent and stable clarification behavior across similar queries, enhancing user trust and perceived intelligence.

Sources

Related papers