Anytime-Valid LLM Leaderboards via Benchmark-weighted and Block-Factorized e-Processes
cs.LG
Submitted: 2026-09-26
Updated: 2026-09-26
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language model (LLM) leaderboards compare model capabilities by ranking models according to their mean performance on fixed benchmarks.
Terminology
Abstract
Large language model (LLM) leaderboards compare model capabilities by ranking models according to their mean performance on fixed benchmarks. However, variability in evaluation outcomes across runs may produce unsupported claims of model superiority on the benchmark, a risk compounded by leaderboard updates. In this paper, we propose BB-EDGE (Benchmark-Weighted and Block-Factorized e-processes for Directed Graph Evaluation), a principled framework that represents an LLM leaderboard as a directed graph whose edges certify pairwise mean-performance advantages, with anytime-valid family-wise error rate (FWER) control. Concretely, for each direction, BB-EDGE constructs an empirical-Bernstein e-process by factorizing evidence over protocol-defined blocks and assigning stakes proportional to the corresponding block weights, then applies direct e-Holm across these e-processes to certify directional advantages as edges. Theoretically, we characterize weight-proportional linear stakes under heterogeneous benchmark-average nulls and prove anytime FWER control under arbitrary within-block and cross-pair dependence. BB-EDGE further supports anytime-valid Top- k certification and simultaneous rank intervals. Extensive experiments on synthetic data and four real-world benchmarks demonstrate that BB-EDGE maintains anytime FWER control while achieving high efficiency.
Sources
- Phi-4 Technical Report
- Intern-S1: A Scientific Multimodal Foundation Model
- Lessons from the Trenches on Reproducible Evaluation of Language Models
- Falcon-H1R: Pushing the Reasoning Frontiers with a Hybrid Model for Efficient Test-Time Scaling
- evalci: A Python Library for Statistically Rigorous Comparison of Language Model Evaluations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- Training Verifiers to Solve Math Word Problems
- On the Stability of Prompt Ranking in Large Language Model Evaluation
- Do Repetitions Matter? Strengthening Reliability in LLM Evaluations
- The Llama 3 Herd of Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Pluralistic Leaderboards
- Family-wise Error Rate Control with E-values
- Efficient Sequential Evaluation of Large Language Models
- Mistral 7B
- MiniMax Sparse Attention
- LLM Evaluation as Tensor Completion: Low Rank Structure and Semiparametric Efficiency
- Low Rank for Rank: Uncertainty-Aware Task-Specific LLM Ranking under Sparse Pairwise Comparisons
- Quantifying Variance in Evaluation Benchmarks
- Prompt-Dependent Ranking of Large Language Models with Uncertainty Quantification
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks