How Sensitive Are LLM Leaderboard Claims to Hidden Model Selection?
stat.ML, cs.LG
Submitted: 2026-09-23
Updated: 2026-09-25
Terminology
Sources
- Flexible Inference for Winners with Conditional Validity
- The Ladder: A Reliable Leaderboard for Machine Learning Competitions
- Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints
- Climbing a shaky ladder: Better adaptive risk estimation
- Correlated Errors in Large Language Models
- Resolution Diagnostics for Paired LLM Evaluation
- Critical Values Robust to P-hacking
- Hidden Measurement Error in LLM Pipelines Distorts Annotation, Evaluation, and Benchmarking
- Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
- Improving Your Model Ranking on Chatbot Arena by Vote Rigging
- Inference conditional on selection: a review
- Efficient multi-prompt evaluation of LLMs
- Powerful rank verification for multivariate Gaussian data with any covariance structure
- Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking
- A Flexible Defense Against the Winner's Curse
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey