EnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise
cs.AI, stat.ML
Submitted: 2026-09-18
Updated: 2026-10-05
License: http://creativecommons.org/licenses/by/4.0/
The gist: Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically valuable tasks, yet most enterprise GenAI
Terminology
Abstract
Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically valuable tasks, yet most enterprise GenAI initiatives fail to show a measurable business effect and a large fraction of agentic projects are expected to be cancelled. We argue that this is substantially a measurement problem: public benchmarks answer "what can the model do?", whereas a deployment decision requires "is this workflow fit, reliable, safe and worth scaling - here, on our data, under our controls?". We present EnterpriseVal, a use-case-level evaluation system that closes this gap. It comprises (i) a formal specification of the use case and of the frozen socio-technical configuration under test, model, prompts, retrieval, tools, guardrails and human oversight, with an autonomy level and consequence tier that jointly set the required evaluation intensity; (ii) a metric catalogue spanning fidelity, utility, efficiency, reliability, assurance and oversight; (iii) a grading protocol that scales blinded expert judgement with calibrated LLM-as-judge scoring through prediction-powered inference; (iv) a two-tier threshold gate, stated as an executable algorithm, that maps metric vectors with confidence bounds to REJECT/CONDITIONAL/SCALE decisions; and (v) a value-and-risk model in which the reviewer catch rate is a measured parameter. We report a pilot across three workflows in a global bank. In credit-memo drafting, human-graded citation precision reached 88% and hallucination rate 1.6% for the best model against gates of 70% and 5%; in procedure transformation, analyst refinement effort fell from an estimated 27.4 to 2.9 hours per document. We separate established results, documented pilot evidence, the proposed system and open hypotheses, and specify the experiments required for full validation
Sources
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
- A Survey on LLM-as-a-Judge
- Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
- AI Agents That Matter
- Log analysis is necessary for credible evaluation of AI agents
- Measuring AI Ability to Complete Long Software Tasks
- Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems
- The Measurement Imbalance in Agentic AI Evaluation Undermines Industry Productivity Claims
- Measuring Agents in Production
- Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
- Towards a Science of AI Agent Reliability
- Towards a Standard, Enterprise-Relevant Agentic AI Benchmark: Lessons from 5.5 billion tokens' worth of agentic AI evaluations
- Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments
- Enterprise Large Language Model Evaluation Benchmark
- Agent-as-a-Judge: Evaluate Agents with Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection