Asymptotic Risk Calibration for Selective Question Answering

arXiv:2608.12008 · cs.CL · Submitted 2026-08-12 · Read on arXiv

Shufan Lin, Sijin Dong

Zhangjiang University · Ibaraki University

cs.CL

Submitted: 2026-08-12

Updated: 2026-08-13

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: The paper proposes A-CRC-QA, a post-hoc calibration framework for uncertainty-aware selective question answering.

Terminology

Summary

The paper proposes A-CRC-QA, a post-hoc calibration framework for uncertainty-aware selective question answering. The method reformulates selection-conditioned error control as a linear expectation constraint and applies a monotonized empirical-risk calibration procedure inspired by conformal risk control. Since the resulting instance-wise loss is generally non-monotone with respect to the acceptance threshold, the framework targets asymptotic rather than finite-sample risk control. A-CRC-QA is model-agnostic, requires no additional training, and can be combined with different uncertainty estimators. Experiments on CoQA and MedMCQA demonstrate its applicability to both open-ended and closed-ended question answering, achieving a favorable trade-off between accepted-answer reliability and answer retention compared with uncalibrated and confidence-bound-based baselines.

The paper addresses the problem that large language models (LLMs) may generate fluent but incorrect answers, making uncertainty quantification important for reliable question answering. However, heuristic uncertainty scores cannot perfectly distinguish correct predictions from incorrect ones, and directly applying a fixed uncertainty threshold provides no statistical control over the error rate among accepted answers.

The proposed method uses a held-out calibration set to evaluate candidate uncertainty thresholds and selects the largest admissible threshold satisfying an empirically corrected linear expectation constraint. The loss is defined as L(lambda) = S(lambda)(E − alpha), where S(lambda) is the selection indicator, E is the error indicator, and alpha is the target risk level. The population risk is g(lambda) = E[S(lambda)E] − alphaE[S(lambda)], and the condition g(lambda) ≤ 0 is equivalent to SCER(lambda) ≤ alpha, where SCER is the selection-conditioned error rate.

A key difficulty is that the loss is not monotone with respect to the selection threshold. Increasing lambda until a previously accepted example is rejected changes the loss in opposite directions for correct and incorrect examples. The standard finite-sample result of Conformal Risk Control (CRC) requires the loss to be monotone, so it cannot be directly applied. Instead, the framework adopts the asymptotic treatment of general non-monotone losses, monotonizing the empirical population-level risk rather than assuming each instance-wise loss is monotone.

The monotonized empirical risk is defined as gbn↑(lambda) = sup t≥lambda gbn(t), which is non-increasing in lambda and serves as an upper envelope of the original empirical risk. The calibrated threshold is defined as lambdabn = inf lambda ∈ Λ: gbn↑(lambda) + gamman ≤ 0, where gamman is a non-negative calibration correction satisfying gamman → 0, with default choice gamman = (1−alpha)/(n+1). This selects the least conservative threshold satisfying the monotonized constraint, maximizing empirical retention within the feasible threshold family.

The paper provides an asymptotic risk guarantee (Theorem 1) stating that under exchangeability and standard regularity conditions, the calibrated decision rule provides asymptotic control of the error rate among accepted answers. The result is an asymptotic marginal guarantee over the joint randomness of the calibration set and a future test example, not a finite-sample or conditional-on-calibration guarantee.

Experiments are conducted on CoQA (open-ended conversational question answering) and MedMCQA (closed-ended medical multiple-choice question answering), using LLaMA-3.1-8B-Instruct and Qwen2.5-7B-Instruct. Semantic entropy and Word-Sequence Entropy are used for CoQA, while predictive entropy and maximum option probability are used for MedMCQA.

At a target accepted-answer error rate of 0.15, A-CRC-QA achieves average error rates of 0.143 and 0.138 on CoQA and MedMCQA, while retaining 46.2% and 58.7% of test answers, respectively. Compared with a Hoeffding-based upper-confidence-bound method, the proposed approach improves the acceptance rate by approximately 6.1 percentage points on CoQA and 7.4 percentage points on MedMCQA.

The main results show that uncalibrated methods (Fixed-50 and Empirical) frequently exceed the prescribed risk level, with violation rates ranging from 61% to 93%. Confidence-bound methods (UCB-HFD and UCB-CLP) maintain conservative test risks but have substantially lower acceptance rates. LEC-Direct achieves the highest retention among statistically calibrated methods but is more sensitive to local non-monotonic fluctuations, with higher violation rates. A-CRC-QA reduces the average violation rate from 25.5% to 12.5% on CoQA and from 19.0% to 8.0% on MedMCQA compared with LEC-Direct, at the cost of approximately 3.3 percentage points in acceptance rate.

Across target risk levels from 0.05 to 0.25, the empirical SCER remains below the target level in all settings. Increasing alpha permits the system to accept more answers and substantially improves power. Very small risk levels can be infeasible when the base model does not produce a sufficiently reliable low-uncertainty subset, with infeasibility rates of 16% on CoQA and 22% on MedMCQA at alpha = 0.05.

The robustness experiments show that all uncertainty estimators support average risk control after calibration, but uncertainty quality strongly affects retention. On CoQA, WSE achieves higher AUROC and accepts approximately four percentage points more answers than Semantic Entropy. On MedMCQA, PE consistently outperforms MSP.

The calibration-set size experiments show that small calibration sets yield conservative thresholds and relatively high variability. As the calibration size increases from 100 to 1,500 examples, the average SCER approaches the target level from below, the acceptance rate increases (from 30.8% to 48.0% on CoQA), and the violation rate decreases (from 26% to 12% on CoQA).

The ablation study isolates the effects of the two main components. Removing monotonization produces the largest acceptance rate but fails to maintain the target risk. Removing the correction leads to average SCER values close to or above alpha and substantially increases the violation rate. Monotonizing each instance-wise loss is extremely conservative, accepting fewer than 20% of predictions on both datasets. The full method provides a more useful compromise between retention and empirical stability.

The paper concludes that uncalibrated uncertainty thresholds do not reliably control the error rate among accepted answers, confidence-bound-based calibration provides strong finite-sample conservativeness but may reject many correct predictions, and monotonizing the empirical linear risk reduces sensitivity to isolated feasible thresholds and provides a practical middle ground between aggressive and conservative methods. The results should be interpreted according to the scope of the theory, targeting asymptotic marginal risk control rather than finite-sample guarantees.

Improvements for AI systems

Improvements to AI Systems:

  1. Implement a post-hoc calibration layer for selective QA: Integrate A-CRC-QA as a model-agnostic, training-free module that wraps any LLM’s uncertainty scores. This allows the system to dynamically choose an acceptance threshold on a held-out calibration set, ensuring the error rate among accepted answers stays at or below a user-specified target (e.g., 0.15) without retraining.

  2. Enable statistically guaranteed answer retention: The system can now trade off between reliability and coverage by tuning the target risk level α. For high-stakes applications, set α low (e.g., 0.05) to accept only highly certain answers; for general use, set α higher (e.g., 0.25) to retain more answers while still controlling error. This gives operators a principled knob for risk tolerance.

  3. Robust handling of non-monotone uncertainty losses: The monotonized empirical-risk procedure (using the upper envelope of the risk function) prevents the system from selecting overly aggressive thresholds that happen to look feasible due to local fluctuations. This reduces violation rates (e.g., from 25.5% to 12.5% on CoQA) while sacrificing only 3 percentage points in acceptance rate, making the system more stable across different calibration sets.

  4. Adaptive calibration with small data: The system works with calibration sets as small as 100 examples, though it becomes more conservative. This allows deployment in low-data regimes (e.g., new domains or languages) while still providing asymptotic risk control, with acceptance rates improving as more calibration data is added (e.g., from 30.8% to 48.0% on CoQA when increasing from 100 to 1,500 examples).

  5. Uncertainty-estimator-agnostic integration: The framework can be paired with any uncertainty estimator (e.g., semantic entropy, word-sequence entropy, predictive entropy, max softmax probability). This lets the system automatically select the best estimator for a given task—e.g., word-sequence entropy for open-ended QA (higher AUROC, +4% acceptance) and predictive entropy for closed-ended QA—without changing the calibration logic.

  6. Asymptotic marginal risk control for non-exchangeable or complex losses: The system provides a theoretical guarantee (Theorem 1) that, under exchangeability, the long-run error rate among accepted answers will not exceed α, even when instance-wise losses are non-monotone. This makes it suitable for production systems where finite-sample guarantees are infeasible but long-term reliability is required.

  7. Improved decision-making in mixed-format QA: The system handles both open-ended (e.g., conversational) and closed-ended (e.g., multiple-choice) QA tasks with the same calibration framework, enabling a unified uncertainty-aware interface across different model outputs and evaluation metrics.

  8. Practical fallback for infeasible risk levels: When the target α is too low (e.g., 0.05) and the model cannot produce a sufficiently reliable subset, the system detects infeasibility (16% on CoQA, 22% on MedMCQA) and can alert the operator to either raise α or improve the base model, preventing silent over-rejection of all answers.

  9. Reduced over-conservatism compared to finite-sample bounds: Unlike Hoeffding-based upper-confidence-bound methods, the system improves acceptance rates by 6–7 percentage points on CoQA and MedMCQA while maintaining the same target risk, leading to higher utility (more answered questions) without sacrificing reliability.

  10. Calibration-set size guidance: The system can recommend a minimum calibration size (e.g., ≥500 examples) to balance variance and retention, based on observed trends—helping practitioners allocate data for calibration in real-world deployments.

Sources

Related papers