Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation
cs.AI
Submitted: 2026-08-01
Updated: 2026-09-11
Comments: 36 pages (review/double-spaced format), 2 figures, 6 tables. Submitted to Knowledge-Based Systems (Elsevier)
Code: https://github.com/williamcaban/experiment-measurement-without-validity
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- The Measurement Imbalance in Agentic AI Evaluation Undermines Industry Productivity Claims
- Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
- EvalCards: A Framework for Standardized Evaluation Reporting
- Counting on Consensus: Selecting the Right Inter-annotator Agreement Metric for NLP Annotation and Evaluation
- Establishing Best Practices for Building Rigorous Agentic Benchmarks
- Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations
- Agent-as-a-Judge: Evaluate Agents with Agents
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- GAIA: a benchmark for General AI Assistants
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Auditing Automated Evaluation, Error Propagation, and Runtime Mitigation in Tool-Using Language Agents
- AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents
- Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
- Annotation alignment: Comparing LLM and human annotations of conversational safety
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- Training language models to follow instructions with human feedback
- Dialectal Toxicity Detection: Evaluating LLM-as-a-Judge Consistency Across Language Varieties
- Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection