Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
cs.CL, cs.AI, cs.DB
Submitted: 2026-10-01
Updated: 2026-10-01
Code: https://github.com/TextQLLabs/Argo-Bench
Project page: https://spider2-sql.github.io
Terminology
Sources
- Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software
- BEAVER: An Enterprise Benchmark for Text-to-SQL
- DABstep: Data Agent Benchmark for Multi-step Reasoning
- The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents
- Pervasive Annotation Errors Break Text-to-SQL Benchmarks and Leaderboards
- AvalancheBench: Evaluating Enterprise Data Agents Through Latent World Recovery
- Business Arena: Benchmarking LLM Agents in a Realistic Marketplace
- Data Agent: A Holistic Architecture for Orchestrating Data+AI Ecosystems
- ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering