Prediction-Powered Active Testing
Kianoosh Ashouritaklimi, Valentin Kilian, Daolang Huang, Tom Rainforth, François Caron
stat.ML, cs.LG
Submitted: 2026-07-09
Code: https://github.com/treforevans/uci_datasets
License: http://creativecommons.org/licenses/by/4.0/
The gist: Active testing provides a label--efficient approach to risk estimation by adaptively selecting which test points should be labelled.
Terminology
Abstract
Active testing provides a label--efficient approach to risk estimation by adaptively selecting which test points should be labelled. However, existing estimators fail to exploit the informative predictions of powerful black--box models, even though such predictions are increasingly available in settings where labels remain expensive. To address this, we propose Prediction--Powered Active Testing (PPAT), a novel label--efficient risk estimation framework that combines the unbiased LURE estimator with a prediction--powered control variate. Rather than using proxy predictions as biased pseudo--labels, PPAT uses them to residualise the loss, preserving unbiasedness while reducing variance. Beyond the estimator itself, PPAT also changes which points should be acquired: we derive oracle and practical surrogate--based acquisition rules tailored to reducing the variance of our estimator. Moreover, we establish asymptotic normality for PPAT, yielding asymptotically valid confidence intervals and thus a principled estimate of the uncertainty around our estimates. Across tabular regression and image--classification tasks, PPAT outperforms existing methods in risk estimation, while its confidence intervals attain the target coverage with substantially fewer labels and smaller widths.
Sources
- Understanding intermediate layers using linear classifier probes
- PPI++: Efficient Prediction-Powered Inference
- Label-Efficient Model Selection for Text Generation
- Scaling Up Active Testing to Large Language Models
- DiffDock: Diffusion Steps, Twists, and Turns for Molecular Docking
- On Statistical Bias In Active Learning: How and When To Fix It
- TabPFN-2.5: Advancing the State of the Art in Tabular Foundation Models
- Active Measurement: Efficient Estimation at Scale
- Adam: A Method for Stochastic Optimization
- DINOv2: Learning Robust Visual Features without Supervision
- tinyBenchmarks: evaluating LLMs with fewer examples
- SigmaDock: Untwisting Molecular Docking With Fragment-Based SE(3) Diffusion
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- Revisiting Active Sequential Prediction-Powered Mean Estimation
- Gemini: A Family of Highly Capable Multimodal Models
- LLaMA: Open and Efficient Foundation Language Models
- Sample Efficient Model Evaluation
- Active Statistical Inference
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey