Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs

arXiv:2607.06327 · cs.CL, cs.AI · Submitted 2026-07-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs".

Jane: The paper was written by Andrea Bacciu, Andrea Alfarano, Saab Mansour, Amin Mantrach and Marcello Federico from Amazon and INSAIT Institute, Sofia.

Tom: Stay tuned as we take you through the paper and discuss its implications.

The Summary and Key Findings: Tom: So, after showing us how they built this massive study, what were the most important takeaways from **Estimating Uncertainty from Reasoning: A Large-Scale Study of Multiand Crosslingual MCQA Performance in LLMs**?

Jane: The findings are really transformative, especially when looking at the influence of language on performance across those twenty-two languages.

Tom: It’s fascinating that if we prompt models to reason in English, it significantly boosts their performance on low-resource languages like Yoruba or Swahili.

Lu: This is a profound discovery because it suggests the core problem isn't understanding the input; the bottleneck is generation fidelity itself.

Meng: The practical implication here is that we can design AI systems whose reliability isn't compromised by language barriers, which would be a massive win for deployment.

Lalam: If an AI can reliably articulate its uncertainty in English, it provides a dependable anchor for trust, even if the user’s native language differs from the target output.

Tom: Another crucial finding relates to how model size dictates which method we should use for uncertainty estimation.

Jane: The study found that at smaller models, open-box probability methods showed great stability and consistency across all resource levels.

Lu: It’s interesting that as the scale increases, the dynamics of how we measure uncertainty shift completely and become much more complex in their behavior.

Meng: The data suggests that at larger scales, closed-box techniques like Self Verbalized become superior to open-box approaches such as Token Entropy.

Lalam: This is a profound shift in how we approach trust; it implies that as our systems grow, they gain genuine meta-cognitive abilities that smaller models simply don’t have.

Tom: It seems the model scale is a huge factor in deciding what kind of uncertainty measurement tool to select for us.

Jane: We are seeing a clear divide: the open-box methods work reliably when the model is small, but when they fail on low-resource languages, closed-box methods succeed.

Lu: This suggests that we are moving toward a future where complexity and scale unlock new levels of self-awareness in AI.

Meng: The engineering challenge here is to match the right tool to the right size of deployment environment.

Lalam: We're seeing how our systems gain genuine meta-cognition, allowing us to build AI that is trustworthy at every stage, regardless of its size or scale.

Actionable Improvements and Guidance: Tom: Now that we’ve seen the core results of **Estimating Uncertainty from Reasoning: A Large-Scale Study of Multiand Crosslingual MCQA Performance in LLMs**, let's talk about the concrete guidance the authors offer.

Jane: The study offers a practical guide for selecting an uncertainty method, depending on whether your system uses large or small models.

Tom: They found that at smaller scales, open-box probability methods like Token Entropy are quite robust, while at larger scales, closed-box self-verbalized methods become the best.

Lu: This is a critical piece of information because it changes how we think about the design of AI systems; it’s not a one-size solution problem.

Meng: The guidance suggests that we cannot use one fixed method for all applications; we need to choose the right tool based on our specific operational needs.

Lalam: This allows us to build AI systems that are reliable and transparent across different sizes, ensuring trust in whatever scale is deployed.

Tom: The authors also provide specific strategies for threshold selection—like using English-only calibration or pooled multilingual data—depending on the required safety level.

Jane: It’s clear that this study is forcing us to rethink how we calibrate these complex AI systems in a way that truly reflects the global nature of our users.

Lu: We are entering an era where understanding *how* an AI knows something is just as important as having the knowledge itself, and this work puts that into practice.

Meng: This forces us to be more disciplined in our engineering choices, making sure we select the right methodology for a specific deployment environment.

Lalam: We are excited to see this framework deployed in systems that serve people from all these diverse linguistic groups with a high degree of confidence.

Tom: The guidance on threshold selection is about managing risk and ensuring that we only let the most reliable outputs pass through the system.

Jane: It helps us balance the trade-off between letting many correct answers through and filtering out a few that are definitely unreliable.

Lu: This sophisticated approach provides a clear path toward achieving trustworthy AI at scale.

Meng: It allows us to make data-driven decisions about whether we prioritize safety or maximum coverage in our specific use case.

Lalam: We can now design systems that provide both high accuracy and verifiable confidence for every step of the process, thanks to this guidance.

The Global Picture and Conclusion: Tom: As we wrap up our discussion on **Estimating Uncertainty from Reasoning: A Large-Scale Study of Multiand Crosslingual MCQA Performance in LLMs**, what is the final word on this revolutionary work?

Jane: This study has provided us with a robust, scalable solution to a problem that has plagued AI development for years.

Tom: It's monumental because it proves that even if we are dealing with low-resource languages, the uncertainty signal can be reliably extracted using the right methods.

Lu: The brilliance here lies in recognizing the gap between generation and comprehension and closing it through thoughtful prompting strategies, which is a huge step for AI.

Meng: This means we can build deployment pipelines that handle uncertainty systematically across multiple languages without relying on luck or guesswork.

Lalam: This entire study is a major step toward making AI truly global, ensuring its reliability is tied to its capability in every culture and language.

Tom: It's not just about knowing the answer; it's about how we define and measure the certainty of the answer.

Jane: That’s right; it gives us confidence in a level of reliability that simply didn't exist before this research was published.

Lu: The ability to see how model scale impacts these results is so critical for understanding the long-term trajectory of AI development as we move forward.

Meng: We can finally make data-driven decisions about which models are suitable for different global applications, based on their inherent uncertainty profiles.

Lalam: This work allows us to shift from just asking "what's the answer?" to asking "how confident is the AI in the answer?", and that changes everything for AI.

Conclusion: Tom: So, wrapping up our look at "Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs," it really shows us where the field is headed.

Jane: We've seen that reliability isn't just about getting the right answer; it’s about knowing how sure you are when you give it to someone.

Lu: Knowing how these systems process information, step by step, offers a completely different way to build trust in AI tools.

Meng: The practical meaning for deployment is that we can't afford to treat all languages or all model sizes the same way anymore.

Lalam: What this research provides is a real blueprint for making these powerful models trustworthy across genuinely diverse global settings.

Tom: These findings push us past just evaluating capability and straight into assessing accountability, which is huge.

Jane: Honestly, it changes the whole conversation around how we certify that complex AI systems will perform when they matter most.

Lu: It's a massive step forward for building tools that people can actually depend on in their daily lives.

Meng: I think the biggest impact will be forcing developers to build in these uncertainty checks from day one, rather than adding them as an afterthought.

Lalam: Truly, this work elevates the standard for what we expect from large language models globally.

Tom: Thank you all for joining us today as we wrapped up our discussion on "Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs."

Jane: We'll have to take a quick break, but when we come back, we're going to look at how multimodal inputs are changing the entire landscape of content creation.

Amazon · INSAIT Institute, Sofia

cs.CL, cs.AI

Submitted: 2026-07-07

Updated: 2026-09-04

Comments: Accepted at Findings of EMNLP 2026

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 93/100

The gist: This study presents a large-scale evaluation of Uncertainty Estimation (UE) methods, addressing a critical gap in prior research that was "predominantly focused on English." By utilizing two

Key concepts

Low-Resource Languages
These are languages, such as Yoruba or Swahili, where there is limited available data. The study found that prompting models in English significantly boosts performance on these languages.
Open-Box vs. Closed-Box Methods
This refers to two types of uncertainty measurement tools. Open-box methods (like Token Entropy) are stable for smaller models, while closed-box techniques (like Self Verbalized) are superior for larger, more complex models.
Meta-Cognitive Abilities
In the context of AI, this refers to the system's ability to gain self-awareness. The discussion suggests that as model scale increases, AI gains genuine meta-cognitive abilities necessary for measuring its own uncertainty.
Crosslingual Performance
This measures how well an LLM performs across multiple languages. The study highlights that understanding uncertainty is crucial for building reliable, global AI systems regardless of the user's native language.

Terminology

Summary

This study presents a large-scale evaluation of Uncertainty Estimation (UE) methods, addressing a critical gap in prior research that was predominantly focused on English. By utilizing two human-curated Multiple-Choice Question Answer (MCQA) datasets and eliciting longform, language-rich reasoning, the paper provides comprehensive evidence for reliable guidance when deploying trustworthy, multilingual LLM systems.

How it works

The study employs a label-grounded framework that avoids model-based correctness metrics or automated ground truth matching (like ROUGE or BERTScore). Instead of relying on short, open-ended answers, the researchers prompt the LLM to generate a step-by-step reasoning explanation (averaging 150 words) before selecting an answer. This approach allows for the evaluation of continuous-valued uncertainty signals from long text generation while maintaining objective correctness against human-curated labels. The methodology involves testing nine UE methods across 22 languages, covering high-, mid-, and low-resource settings, using two parallel datasets: Global-MMLU and MMLU-ProX.

How it works

The study reveals a significant asymmetry between natural language understanding and generation regarding UE performance. When researchers prompt models to reason in English while the questions remain in low-resource languages, substantially improves UE performance. This suggests that the reliability bottleneck is found in the model's ability to generate coherent reasoning in specific languages, rather than its initial comprehension of those questions. Furthermore, this finding demonstrates that generation language matters more than question language, effectively closing the performance gap between high and low-resource settings.

How it works

The study conducted a detailed scaling analysis of how model size affects UE efficacy across the Qwen3 family (from 4B to 235B parameters). The results show a clear divergence in scaling behavior:

  • At smaller scales, open-box probability-based methods outperform alternatives and maintain consistent performance.

  • At larger scales, closed-box selfverbalized uncertainty becomes superior. This is because meta-cognitive abilities required for accurate self-assessment emerge primarily at larger model sizes. The closed-box Self Verbalized method achieved the best overall performance, reaching an average AUROC of 0.72 across all languages and models.

How it works

The researchers investigated the robustness of UE methods under complex, real-world scenarios:

  • Cross-Lingual Answers: When evaluation involves a question in one language but answer options drawn from different languages, UE performance remains stable, confirming that these methods do not rely solely on lexical matching.

  • Threshold Selection: The paper analyzed three calibration strategies for selective prediction (i.e., when to abstain):

  1. tEN: Optimized on English data only (most practical).

  2. tGLOBAL: Optimized on pooled multilingual data (moderate coverage loss).

  3. tLANG: Using language-specific thresholds (achieves the strongest practical performance).

How it works

The study concludes with actionable guidance for deploying uncertainty-aware systems in multilingual environments, emphasizing that while self-verbalized approaches are superior at large scales, sampling-based consistency methods degrade substantially on low-resource languages due to their reliance on surface-level lexical overlap.

Improvements for AI systems

The following improvements integrate the findings of the large-scale study to construct a robust, trustworthy AI system capable of reliable operation across diverse linguistic and computational environments.


Improvement: The system must be engineered to extract uncertainty signals not from the final predicted label (the Guess) but from the entire intermediate reasoning trace (the Chain-of-Thought explanation).

Specific Implementation: Apply UE methodologies—specifically those derived from semantic entropy and graph theory—to the sequence of token probabilities across a about.150 word reasoning trace, rather than relying on the final choice probability.

System Capability: The system can now detect subtle inconsistencies or lack of coherent internal logic before generating an answer, significantly reducing hallucinations that stem from superficial consistency checks.

Improvement: Abandon a monolithic one-size-fits-all approach to uncertainty estimation. The system must dynamically select the optimal UE method based on the underlying LLM's parameter scale and architecture.

Specific Implementation:

  • For Smaller Models (<1B parameters): Prioritize Open-Box Probability Methods (e.g., Token Entropy, Self-Certainty). These are computationally efficient and provide the most reliable signal in resource-constrained environments where larger models might fail to achieve meta-cognitive stability.

  • For Larger Models (235B parameters): Utilize Closed-Box Self-Verbalized Uncertainty. Leverage the model's inherent metacognitive capabilities (e.g, using a follow-up prompt to elicit a confidence score s in [0, 1]) to achieve superior AUROC performance, despite the increased computational cost.

System Capability: The system maintains high reliability across its entire operational spectrum, knowing exactly how to interpret uncertainty from different model sizes without requiring manual tuning or retraining.

Improvement: Recognize that in low-resource languages, the bottleneck for reliability is not understanding the question (comprehension), but generating a coherent, stable reasoning trace (generation).

Specific Implementation: When operating in a low-resource language environment, the system should be designed to utilize English Reasoning as a Diagnostic Baseline. If native language generation yields an unacceptably low certainty score, the the system can trigger an internal switch to generate and evaluate its reasoning in English. This leverages stronger generative capabilities while retaining the original question context.

System Capability: The AI achieves high-level performance parity across all 22 supported languages, ensuring that low-resource language users receive answers with a reliability level comparable to those using high-resource languages.

Improvement: Implement a sophisticated, multi-layered calibration strategy for selective prediction (abstention).

Specific Implementation: Instead of applying a single threshold derived from English data (t EN) across all languages, the system must use Language-Specific Calibration (t LANG). This involves tuning validation thresholds specific to the target language's characteristics and using pooled multilingual data (t GLOBAL) as a fallback for cross-lingual contexts.

System Capability: The system maximizes safety and accuracy simultaneously. It can achieve a significant reduction in false positives (FP) by abstaining when uncertainty is high, while maintaining maximal coverage of correct predictions (TP), even in complex, multi-language deployments.

Improvement: Design the UE pipeline to be robust against cross-lingual input/output complexity (e.g., a monolingual question with multiple foreign language answer options).

Specific Implementation: The system must confirm that its uncertainty estimation methods do not rely on simple lexical overlap between the reasoning trace and the final answer choice. The UE architecture must be verified to remain stable when considering the semantic distance between different language families (e.g., comparing a French reason trace against an Italian option).

System Capability: The system is inherently reliable in complex, real-world applications like Cross-Lingual Retrieval-Augmented Generation (RAG), where input and output languages may be mixed, ensuring that confidence scores reflect the model's internal certainty rather than surface-level linguistic coincidence.

Abstract

Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing research has predominantly focused on English. We present the first large-scale evaluation of UE methods across 22 languages, spanning high-, mid-, and low-resource settings. Using two human-curated Q&A datasets, we compare open and closed box UE methods (nine in total) across different model sizes and architectures while eliciting long-form reasoning, avoiding LLM-as-a-judge and embedding-based scoring, which can introduce evaluation noise. We report three main actionable findings. First, we find that prompting models to reason in English while keeping questions in low-resource languages substantially improves UE performance, suggesting that comprehension of low-resource languages is largely intact, and that the reliability bottleneck lies in generation rather than understanding. Second, prompting models to reason in English closes the UE performance gap between low and high-resource languages, demonstrating that generation language matters more than the question language. Third, the choice of UE method should depend on model scale: at smaller scales, open-box probability-based methods outperform alternatives; at larger scales, closed-box self-verbalized uncertainty becomes superior. Finally, we provide an analysis of threshold selection for selective prediction, offering guidance on calibrating abstention in multilingual settings.

Sources

Related papers