RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents
cs.AI
Submitted: 2026-09-16
Updated: 2026-09-16
License: http://creativecommons.org/licenses/by/4.0/
The gist: AI agents are usually evaluated by whether they complete a task.
Terminology
Abstract
AI agents are usually evaluated by whether they complete a task. In interactive service settings, a successful agent can still frustrate users by asking repeated questions, performing redundant searches, or making avoidable revisions. We introduce RideWay, an efficiency-centered benchmark for ridehailing agents in a stateful tool-calling environment, together with Efficiency Utility, a success-gated metric that discounts successful trajectories for excess tool calls and user-facing turns relative to task-specific reference effort. Human paired preferences calibrate the relative penalties, reflecting an aggregate service-workflow trade-off: extra dialogue often creates visible friction, whereas extra tool use can sometimes verify constraints or preserve user intent. Across 58 tasks and 24 models, the fitted penalty for excess turns is about twice that for excess tool calls. On task-disjoint held-out preferences, Efficiency Utility achieves 78.7% accuracy overall: 90.6% when trajectories differ in turns, but chance-level accuracy when they differ solely in tool calls - the axis on which human annotators agree least. RideWay therefore makes interaction efficiency measurable alongside task success, while exposing the boundary of count-based tool-use evaluation.
Sources
- On Evaluation of Embodied Navigation Agents
- VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Evaluating Large Language Models Trained on Code
- AgentAtlas: Beyond Outcome Leaderboards for LLM Agents
- LLM Evaluators Recognize and Favor Their Own Generations
- PyBench: Evaluating LLM Agent on various real-world coding tasks
- T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection