ShowTellArena: Evaluating Business Workflow Understanding from Demonstrations
cs.AI
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 6 pages, 1 figure, 2 tables. Code and dataset linked in the paper
License: http://creativecommons.org/licenses/by/4.0/
The gist: We often teach a colleague by showing the work and explaining the decisions as we go.
Terminology
Abstract
We often teach a colleague by showing the work and explaining the decisions as we go. How can we check what an agent understood from the same lesson? We introduce ShowTellArena, a benchmark protocol and public dataset for comprehension after narrated business demonstrations. The v1.0 release contains 50 business workflow tasks, with recordings, screenshots, narration, fixture seeds, and 502 questions. Tasks span finance, hiring, procurement, customer decisions, inventory, and logistics. The protocol holds the business scenario and quiz fixed while allowing each product to capture the lesson through its own teaching interface. Questions test operational rules, boundaries, exceptions, and errors in proposed automations. We analyze 218 selected pilot attempts across 39 workflow cases, including 28 cases attempted by all three evaluated systems. These exploratory results expose both answer errors and failures to complete the teaching experience. We describe the release's verification gaps and the pilot's uneven coverage, exclusions, and grading provenance. The contribution is an inspectable dataset and assessment workflow that others can extend; the selected pilot is not a controlled product ranking.
Sources
- VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks
- WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection