Auditing MCQA Benchmarks through Probability Landscapes
cs.CL
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: Accepted to EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: As Large Language Models rapidly advance, performance on standard multiple-choice question answering (MCQA) benchmarks is reaching saturation.
Terminology
Abstract
As Large Language Models rapidly advance, performance on standard multiple-choice question answering (MCQA) benchmarks is reaching saturation. While the community has responded by developing increasingly difficult datasets, validating question quality and filtering flawed items remains a labor-intensive process. To provide a scalable diagnostic approach, we propose a two-component probabilistic framework for auditing MCQA benchmarks using model output distributions. First, for benchmark-level analysis, we characterize the probability landscape using the top prediction probability (P top1) and normalized residual entropy (H norm), summarized globally by Mean Pairwise Distance (MPD). Second, for item-level diagnostics, we introduce noise injection to reduce meaningful distractor competition, enabling us to flag candidate items for targeted human review and categorize residual failure patterns. Across four MCQA benchmarks, our landscape analysis reveals benchmark-level differences in model confidence and residual option competition. Concurrently, our noise-injection method flags potentially actionable item-level issues, showing alignment with expert error annotations from MMLU-Redux. These results suggest that our probability-based framework provides a lightweight audit lens for comparing macro-level benchmark structure and prioritizing individual items for targeted human review.
Sources
- The Llama 3 Herd of Models
- When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
- Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
- The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Gemma 2: Improving Open Language Models at a Practical Size
- Qwen2.5 Technical Report
- Measuring Massive Multitask Language Understanding
- Better Distractions: Transformer-based Distractor Generation and Multiple Choice Question Filtering
- The Linear Representation Hypothesis and the Geometry of Large Language Models
- Humanity's Last Exam
- Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data
- Rethinking Generative Large Language Model Evaluation for Semantic Comprehension
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
- Large Language Models Are Not Robust Multiple Choice Selectors
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering