Rank Confidence Sequences:Anytime-valid Leaderboards
stat.ME, cs.AI
Submitted: 2026-09-26
Updated: 2026-09-26
Code: https://github.com/HamedKhosravi99/rank-confidencesequences-supplement
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Sequential model confidence sets
- Prompt-Dependent Ranking of Large Language Models with Uncertainty Quantification
- Finite-Sample Valid Rank Confidence Sets for a Broad Class of Statistical and Machine Learning Models
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- Admissible online closed testing must employ e-values
- Family-wise Error Rate Control with E-values
- Efficient Sequential Evaluation of Large Language Models
- Resolution Diagnostics for Paired LLM Evaluation
- Low Rank for Rank: Uncertainty-Aware Task-Specific LLM Ranking under Sparse Pairwise Comparisons
- Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
- Quantifying Ranking Uncertainty in LLM Benchmarks
- Rank Intervals for Leaderboards: A Hierarchical Framework for Model Evaluation
- CELEUS: Certifiable and Efficient LLM Evaluation via E-Processes
Related papers
- Doubly robust inference via calibration
- Bayesian Empirical Bayes: Simultaneous Inference from Probabilistic Symmetries
- Flexible Nonparametric Inference for Causal Effects under the Front-Door Model
- Deployment of AI-Assisted Interventions: Capacity Constraints and Noisy Compliance
- A Survey on Archetypal Analysis
- Dynamic Spatial Bayesian Machine Learning Model: Applications to Intergenerational Economic Mobility and Geographic Income Inequality in the United States