Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

arXiv:2608.13430 · cs.CL, cs.AI · Submitted 2026-08-13 · Read on arXiv

Cohere Labs Community · Laboratoire Hubert Curien, UMR CNRS 5516, Saint-Étienne, France · School of Computer Science Engineering and Technology (SCSET), Bennett University, Greater Noida, India · School of Physical Therapy, Faculty of Health Sciences, Western University, London, Canada

cs.CL, cs.AI

Submitted: 2026-08-13

Updated: 2026-09-18

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper investigates how instruction tuning affects model confidence and the lexical diversity of generated answer rationales in question-answering tasks.

Terminology

Summary

The paper investigates how instruction tuning affects model confidence and the lexical diversity of generated answer rationales in question-answering tasks. The authors evaluate three matched base and instruction-tuned model pairs (Qwen2.5-7B, Mistral-7B-v0.3, Llama-3.1-8B) across three multiple-choice benchmarks (ARC-Easy, MMLU, CommonsenseQA). They measure likelihood-based uncertainty via choice entropy, verbalized confidence via a two-stage prompting protocol, and lexical diversity using Unique-2 (proportion of distinct bigrams) and 1-SelfBLEU (cross-rationale variability).

The main findings are as follows. First, instruction tuning consistently increases model confidence without corresponding improvements in predictive accuracy. Across all models and benchmarks, answer entropy decreases and verbalized confidence increases. For example, for Llama on ARC-Easy, accuracy remains unchanged at 82.2% while verbalized confidence increases from 49.2% to 90.4%, and choice entropy decreases from 0.188 to 0.173. For Mistral on CSQA, choice entropy decreases from 0.736 to 0.268, and verbalized confidence increases from 47.9% to 92.2%, while accuracy improves from 57.4% to 69.2%.

Second, the effect of instruction tuning on lexical diversity is heterogeneous. Cross-rationale diversity, measured as 1-SelfBLEU, consistently decreases across all models and benchmarks, with the largest decrease for Mistral on ARC-Easy (from 0.813 to 0.626). However, surface-level lexical diversity (Unique-2) varies in both direction and magnitude. For example, Unique-2 decreases for Mistral on ARC-Easy (0.701 to 0.673) and MMLU (0.731 to 0.674), but increases on CSQA (0.719 to 0.750). For Llama, Unique-2 decreases on ARC-Easy (0.704 to 0.697) and MMLU (0.721 to 0.702), but increases on CSQA (0.709 to 0.731).

Third, decreases in answer uncertainty do not consistently coincide with reduced lexical diversity. The directional analysis shows that for Qwen, lower uncertainty coincides with lower diversity for 61.8% of examples under Unique-2 and 69.3% under 1-SelfBLEU. For Mistral, lower uncertainty coincides with higher Unique-2 for 61.8% of examples but with lower 1-SelfBLEU for 77.6%. Llama exhibits a similar divergence between the two measures. The authors conclude that the direction of the association depends on the diversity measure and model.

Fourth, the diversity shifts persist after controlling for answer selection and rationale length. When restricting to examples where base and instruction-tuned models select the same answer and matching rationale length, the patterns remain. For Qwen, Unique-2 remains nearly unchanged (−0.001) while 1−SelfBLEU decreases by 0.036. For Mistral and Llama, Unique-2 increases significantly by 0.050 and 0.053, respectively, whereas 1−SelfBLEU decreases by 0.069 and 0.012.

Fifth, changes in lexical diversity are not consistently associated with calibration. The authors compute Expected Calibration Error (ECE) separately for likelihood-based and verbalized confidence. For example, for Qwen on ARC-Easy, verbalized-confidence ECE decreases from 35.3 to 22.8 after instruction tuning, while both diversity measures decrease. In contrast, for Llama on MMLU, both likelihood-based and verbalized ECE increase (0.5 to 5.9 and 16.6 to 23.7, respectively), while both diversity measures decrease.

The paper concludes that instruction tuning affects confidence and rationale diversity differently. The authors state: instruction tuning consistently alters answer confidence, despite limited changes in predictive accuracy and decreases in likelihood-based calibration and cross-rationale diversity consistently decreases, whereas surface-level lexical diversity varies in both direction and magnitude across models and benchmarks. They also note that these differences persist after controlling for answer selection and rationale length, confirming that confidence and rationale diversity capture distinct effects of instruction tuning. The findings motivate future work on uncertainty estimation that jointly considers predictive confidence and variation in generated rationales.

Improvements for AI systems

Improvements to AI Systems:

  1. Calibration-Aware Instruction Tuning: Modify the instruction-tuning objective to include a regularization term that penalizes overconfidence without accuracy gains. The improved system can maintain or reduce verbalized confidence when predictive accuracy plateaus, preventing misleadingly high confidence in deployed QA systems.

  2. Dual-Metric Uncertainty Estimation: Implement an uncertainty estimator that combines likelihood-based entropy with cross-rationale diversity (1-SelfBLEU). The improved system can flag low-confidence answers even when verbalized confidence is high, by detecting low diversity in generated rationales as a separate risk signal.

  3. Adaptive Rationale Generation Control: Add a decoding-time or fine-tuning mechanism that adjusts lexical diversity based on task and model pair. For models like Mistral where instruction tuning reduces cross-rationale diversity, the system can explicitly sample multiple rationales with higher temperature or add diversity-promoting penalties to preserve explanatory variation.

  4. Benchmark-Specific Confidence Thresholds: Train a meta-model that predicts per-benchmark calibration offsets (e.g., ARC-Easy vs. MMLU) after instruction tuning. The improved system can recalibrate verbalized confidence outputs dynamically, reducing ECE for tasks where instruction tuning worsens calibration (e.g., Llama on MMLU).

  5. Rationale-Length-Controlled Diversity Reporting: Integrate a post-hoc analysis module that separates diversity changes due to answer selection and rationale length from intrinsic generation style. The improved system can provide more interpretable uncertainty reports, distinguishing genuine diversity loss from artifacts, and alert users when diversity drops are not explained by these confounders.

  6. Confidence-Diversity Joint Training: Train a multi-task model that simultaneously optimizes for calibrated confidence and controlled lexical diversity. The improved system can generate rationales that are both confidently correct and sufficiently varied across plausible answers, reducing the risk of spurious uniformity in reasoning.

Abstract

Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhibit verbalized overconfidence. In question answering, verbalized model overconfidence may be associated with the consistency of the generated supporting rationales. In this paper, we study whether corresponding changes in the lexical diversity of generated answer rationales accompany changes in model confidence induced by instruction tuning. We evaluate three matched base and instruction-tuned models across question-answering benchmarks and find that instruction tuning consistently alters answer confidence, despite limited changes in predictive accuracy and decreases in likelihood-based calibration. Secondly, we observe a non-uniform effect of instruction tuning on rationale diversity: cross-rationale diversity consistently decreases, whereas surface-level lexical diversity varies in both direction and magnitude across models and benchmarks. Finally, we find that these differences persist after controlling for answer selection and rationale length, confirming that confidence and rationale diversity capture distinct effects of instruction tuning.

Sources

Related papers