Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation
stat.ML, cs.AI, cs.LG, stat.AP, stat.ME
Submitted: 2026-09-17
Updated: 2026-09-17
Comments: 15 pages of main text, 30 pages total, 4 figures
Code: https://github.com/shokawano/disagg-ai-eval
License: http://creativecommons.org/licenses/by/4.0/
The gist: Evaluating an AI system requires disaggregated assessment, as performance varies across domains such as benchmark task types or conversation types in deployed agents.
Terminology
Abstract
Evaluating an AI system requires disaggregated assessment, as performance varies across domains such as benchmark task types or conversation types in deployed agents. Exhaustive testing is expensive, so evaluation rests on a sample of labeled units. We treat the evaluation set as a finite population and seek accurate point and interval estimates of each domain mean. Direct estimators, including prediction-powered inference (PPI), use only a domain's own labels and are imprecise where labels are few. Small area estimation addresses this problem, and we build on it to develop an integrated workflow for estimation and validation. For estimation, we propose prediction-powered smoothing (PP-S), a Bayesian model fit to each domain's prediction-powered estimate, with an extension that borrows strength across a reporting taxonomy (PP-TS). For validation, we derive a new, approximately unbiased design-based cross-validation score for choosing among direct and smoothed estimators. We study a curated benchmark with verifiable grading and deployed agent traffic graded by humans, each with every outcome observed. In both, the proposed estimators improve on the direct estimators in point and interval estimation, with near-nominal coverage. At the same sampling budget, our score selects as well as an independent validation sample does and estimates the selected estimator's error far more accurately.
Sources
- PPI++: Efficient Prediction-Powered Inference
- Measuring the Prevalence of Policy Violating Content with ML Assisted Sampling and LLM Labeling
- Design-Based Cross-Validation for Comparing Small Area Estimators
- Prediction-Powered Inference Across Many Tasks for AI Evaluation & Social Science Research
- Stratified Prediction-Powered Inference for Hybrid Language Model Evaluation
- Precise Model Benchmarking with Only a Few Observations
- A Framework for Efficient Model Evaluation through Stratification, Sampling, and Estimation
- A structured regression approach for evaluating model performance across intersectional subgroups
- On Data Thinning for Model Validation in Small Area Estimation
- The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
- SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
- Prediction-Powered Adaptive Shrinkage Estimation
- Holistic Evaluation of Language Models
- tinyBenchmarks: evaluating LLMs with fewer examples
- Optimal Hold-Out Size in Cross-Validation
- PPI is the Difference Estimator: Recognizing the Survey Sampling Roots of Prediction-Powered Inference
- Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairness
- Introducing v0.5 of the AI Safety Benchmark from MLCommons
- Efficient Evaluation of LLM Performance with Statistical Guarantees
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey