The geometry of AI validation: From structural blindness to reusable audits
cs.LG, math.ST, stat.ML, stat.TH
Submitted: 2026-08-21
Updated: 2026-09-11
Comments: 19 pages, 5 figures. Substantially revised and extended version; includes supplementary material. Code and data: https://github.com/rfitas-lab/geometry-of-ai-validation
Code: https://github.com/rfitas-lab/geometry-of-ai-validation
License: http://creativecommons.org/licenses/by/4.0/
The gist: AI systems increasingly search among candidate answers and deploy the highest-scoring one.
Terminology
Abstract
AI systems increasingly search among candidate answers and deploy the highest-scoring one. Increasing search changes which errors matter, so a precise evaluation at one computation budget can leave another budget unresolved. We connect this information gap to the cost of closing it. For independent best-of-n search, aggregate reliability measurements identify deployment only through the directions they observe; we derive an exact ambiguity frontier when only small search widths are audited. Retaining candidate ranks and truth labels enables a constructive alternative: one audit can estimate reliability across all widths up to N. With known score percentiles, the minimax worst-coordinate mean squared error scales as (1 + log N)/T + N/M, capped at a constant, for expected budgets of T truth labels and M candidate observations. Matching lower bounds allow adaptive label acquisition, establishing that the distinct label and candidate costs are intrinsic to this experiment. An explicit design attains this order; a complementary record-based procedure supplies simultaneous guarantees without a known score distribution. Retrospective mathematical-reasoning and code-generation analyses show why search-dependent validation matters. In held-out CodeRM pools, a shared audit reduces the 95th-percentile maximum error across 100 widths by 58% and 40% relative to uniform labeling. These results turn structural ambiguity into a quantitative prescription for reusable validation.
Sources
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- Don't Gamble, GAMBLe: An Analytical Framework for AI-Driven Research Systems
- The moment problem with bounded density
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility
- ROC-n-reroll: How verifier imperfection affects test-time scaling
- The Evaluation Blind Spot: A Stereological Theory of Benchmark Coverage for Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks