Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks
stat.ML, cs.LG, stat.AP
Submitted: 2025-01-08
Updated: 2026-09-11
Comments: LA-UR-24-25289; presented at the Workshop on Statistical Frontiers in LLMs and Foundation Models at NeurIPS 2024
Code: https://github.com/google-research/task_adaptation
Project page: https://google-research.github.io/task_adaptation/benchmark
License: http://creativecommons.org/licenses/by/4.0/
The gist: Modern artificial intelligence is supported by machine learning models (e.g., foundation models) that are pretrained on a massive data corpus and then adapted to solve a variety of downstream tasks.
Terminology
Abstract
Modern artificial intelligence is supported by machine learning models (e.g., foundation models) that are pretrained on a massive data corpus and then adapted to solve a variety of downstream tasks. To summarize performance across multiple tasks, evaluation metrics are often aggregated into a summary metric, e.g., average accuracy across 10 question-answering tasks. When aggregating evaluation metrics, it is useful to incorporate uncertainty in the aggregate metric in order to gain a more realistic understanding of model performance. Our objective in this work is to demonstrate how statistical methodology can be used for quantifying uncertainty in metrics that have been aggregated across multiple tasks. The methods we emphasize are bootstrapping, Bayesian hierarchical (i.e., multilevel) modeling, and the visualization of task weightings that consider standard errors. These techniques reveal insights such as the dominance of a specific model for certain types of tasks despite an overall poor performance. We use a popular ML benchmark, the Visual Task Adaptation Benchmark (VTAB), to demonstrate the usefulness of our approaches.
Sources
- GPT-4 Technical Report
- The Benchmark Lottery
- Uncertainty in Ranking
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- LLaMA: Open and Efficient Foundation Language Models
- A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey