The Flaw of Averages: Measuring Benchmark-Level Distributional Robustness
cs.CL, cs.SE
Submitted: 2025-09-30
Updated: 2026-09-05
Comments: EMNLP 2026 Main
Code: https://github.com/EleutherAI/lm-evaluation-harness
License: http://creativecommons.org/licenses/by/4.0/
The gist: Benchmarks are central to measuring progress in language models, but aggregate scores can obscure substantial variation across subdomains, making models appear broadly competent despite concentrated
Terminology
Abstract
Benchmarks are central to measuring progress in language models, but aggregate scores can obscure substantial variation across subdomains, making models appear broadly competent despite concentrated strengths and weaknesses. We study this issue as benchmark-level distributional robustness: whether aggregate scores faithfully reflect performance across benchmark subdomains. We operationalize this notion with benchmark Harmony, an entropy-based measure of how uniformly model performance is distributed across subdomains. Measuring Harmony on 19 language model benchmarks across five model families, we find substantial variation in benchmark-level distributional robustness. Low-Harmony benchmarks are more likely to yield aggregate scores that overstate broad competence, whereas high-Harmony benchmarks provide more representative summaries of model capability. Rebalancing benchmarks by pruning overrepresented subdomains to increase Harmony substantially shifts aggregate scores for low-Harmony benchmarks, but leaves high-Harmony benchmarks comparatively stable. For example, while BoolQ remains comparatively stable as Harmony increases, PubMedQA, which evaluates performance in a medically consequential domain, exhibits substantial, often statistically significant, shifts in aggregate accuracy. Together, these findings show that aggregate scores can misrepresent broad competence when performance is unevenly distributed. We therefore recommend reporting benchmark Harmony alongside aggregate accuracy as a diagnostic of benchmark representativeness when interpreting claims about broad model competence.
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms
- Language Models are Few-Shot Learners
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- ARC Prize 2024: Technical Report
- RobustBench: a standardized adversarial robustness benchmark
- Underspecification Presents Challenges for Credibility in Modern Machine Learning
- Gemini: A Family of Highly Capable Multimodal Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- The Benchmark Lottery
- Investigating Data Contamination in Modern Benchmarks for Large Language Models
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- Understanding Dataset Difficulty with $\mathcal{V}$-Usable Information
- Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies
- Time Travel in LLMs: Tracing Data Contamination in Large Language Models
- MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Scaling Laws for Neural Language Models
- TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
- Towards General Text Embeddings with Multi-stage Contrastive Learning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering