Efficient Sequential Evaluation of Large Language Models
Chia-Yu Hsu, Shubhanshu Shekhar
stat.ML, cs.LG, stat.ME
Submitted: 2026-07-19
Comments: This is a preliminary version; feedback is welcome
License: http://creativecommons.org/licenses/by/4.0/
The gist: We study the problem of sequentially evaluating a new large language model (LLM) on a fixed question set using historical performance data from prior LLMs.
Terminology
Abstract
We study the problem of sequentially evaluating a new large language model (LLM) on a fixed question set using historical performance data from prior LLMs. Our goal is to construct a confidence sequence (CS) for the model's capability on this question set and to design active querying rules that shrink the CS width as quickly as possible. For CS construction, we invert a family of test supermartingales and focus on two representative approaches: a reverse information projection (RIPr)-based approach and a testing-by-betting-based approach. We first study these approaches under an oracle setting, and demonstrate the oracle optimality of the RIPr-based construction. We then propose a growth-oriented querying rule that aims to maximize the worst-case one-step expected log-increment over the endpoints of the current CS. In practice, we build these test supermartingales and the querying rule on predictions of question-level correctness learned from historical data. We then analyze the shrinkage behavior of the resulting CSs and identify two key factors that slow the shrinkage rate of CSs: accumulated prediction mismatch and the spikiness of the querying distribution. Finally, motivated by this analysis, we propose several mixture querying rules that combine growth-oriented querying, prediction refinement, and uniform exploration, trying to mitigate the effects that slow the shrinkage rate. We provide experiments comparing different querying rules for the RIPr-based and testing-by-betting-based CSs across several synthetic testing datasets. Interestingly, we observe that the simplest querying rule, uniform sampling, can sometimes outperform more adaptive querying rules for both methods.
Sources
- Evaluating Large Language Models: A Comprehensive Survey
- Efficient Evaluation of LLM Performance with Statistical Guarantees
- Active Statistical Inference
- Time-Uniform Confidence Spheres for Means of Random Vectors
- On the near-optimality of betting confidence sets for bounded means
- Safe Testing
- Game-theoretic statistics and safe anytime-valid inference
- Exact Anytime-valid Confidence Intervals for Contingency Tables and Beyond
- Design-Based Confidence Sequences: A General Approach to Risk Mitigation in Online Experimentation
- Semiparametric Efficient Inference in Adaptive Experiments
- Anytime-valid off-policy inference for contextual bandits
- Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization
- Prediction-Powered Inference
- Revisiting Active Sequential Prediction-Powered Mean Estimation
- The numeraire e-variable and reverse information projection
- CollabEval: Statistically Efficient Collaborative Model Evaluation via Matrix Completion
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey