One Capability or Many? Structural and Predictive Tests of Benchmark Validity Disagree About Economic Benchmarks for Frontier AI
cs.LG, cs.CY, cs.SE
Submitted: 2026-08-29
Updated: 2026-09-25
Comments: 25 pages, 11 figures. Analysis plan deposited at https://doi.org/10.17605/OSF.IO/VD34J (retrospective deposit; see the note there). Code and data: https://github.com/louisyzhu/frontier-ai-economic-validity
Code: https://github.com/louisyzhu/frontier-ai-economic-validity
License: http://creativecommons.org/licenses/by/4.0/
The gist: Frontier-model leaderboards now rank systems based on economic benchmarks, tests of how well models carry out professional tasks from software engineering to banking workflows, and those rankings
Terminology
Abstract
Frontier-model leaderboards now rank systems based on economic benchmarks, tests of how well models carry out professional tasks from software engineering to banking workflows, and those rankings inform what organisations buy, what regulators scrutinise, and expectations of how work will change. Whether such benchmarks measure a capability distinct from general test-taking, or re-express the one axis along which every benchmark rises as models improve, is a question of construct validity that has not yet been studied. We test it on a hash-pinned leaderboard snapshot of 421 model configurations across twelve benchmarks, four of them economic, treating benchmarks as items and models as respondents in a latent-variable model with four hypotheses and their thresholds fixed before analysis. A single factor explains 74.5% of common variance and tracks model release date (R squared = 0.505), so the leading axis of capability is substantially a time trend; where prior work controls for scale, compute adds little once date is removed. Removing the date trend lowers that share by 14.9 points, and by 24.1 with one row per base model. Under the dimensionality rule fixed in advance the economic benchmarks form no distinct factor, yet a leave-one-benchmark-out test with factors re-estimated inside every fold shows that a multi-factor representation predicts held-out economic scores better than a single general index (pooled Delta-MSE 0.037, 95% bootstrap interval [0.019, 0.055]). Economic benchmarks therefore add incremental predictive information to a largely date-driven general factor, and the evidence does not support treating them as a distinct latent capability. Leaderboards remain a sound guide to overall progress, but most of the gap between models released months apart is calendar, so a small gap between contemporaneous models should be date-adjusted before being read as a capability difference.
Sources
- When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
- Lessons from the Trenches on Reproducible Evaluation of Language Models
- Revealing the structure of language model capabilities
- Quantifying construct validity in large language model evaluations
- Holistic Evaluation of Language Models
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
- Humanity's Last Exam
- Generalizing Verifiable Instruction Following
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations
- $\tau$-Knowledge: Evaluating Conversational Agents over Unstructured Knowledge
- SciCode: A Research Coding Benchmark Curated by Scientists
- APEX-Agents
- Toward an Evaluation Science for Generative AI Systems
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks