Quantifying Protocol-Induced Uncertainty in Comparative Predictive-Model Evaluation: Evidence from Large-Scale Daily PM10 Forecasting
cs.LG
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: 29 pages, 8 figures, 7 tables
License: http://creativecommons.org/licenses/by/4.0/
The gist: Comparative studies of predictive models often end by ranking candidate models, yet these rankings depend on evaluation protocols whose influence is rarely treated as a source of uncertainty.
Terminology
Abstract
Comparative studies of predictive models often end by ranking candidate models, yet these rankings depend on evaluation protocols whose influence is rarely treated as a source of uncertainty. We formalize this problem as protocol-induced ranking uncertainty and introduce a framework that compares ranking displacement caused by switching protocols with displacement produced by conventional choices within a fixed protocol. We quantify these effects using the Protocol Sensitivity Score (PSS) and a full-refit intraprotocol reference. We validate the framework in a large-scale sequential prediction study of daily PM10. Static-split and rolling-origin evaluation are compared across 425 European background stations and 365 US EPA monitors. Switching protocols produces mean PSS values of 0.801 and 0.772 and changes the selected model at 35.3% and 31.5% of stations, respectively. In Europe, intraprotocol perturbations with identical scored targets produce PSS values of 0.072 and 0.230, with winner-swap rates of 0.8% and 4.9%. Between-protocol displacement is therefore substantially larger than the selected within-protocol references. Expanding the candidate set from three to nine models increases the between-protocol winner-swap rate to 60.2% in Europe. The pattern also persists under a frozen protocol applied to held-out background and non-background stations. These results show that model-selection conclusions can depend materially on legitimate evaluation choices. We recommend reporting ranking stability under a small set of defensible intraprotocol perturbations alongside claims of model superiority.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks