E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation
cs.LG, cs.CL
Submitted: 2026-08-31
Updated: 2026-08-31
Code: https://github.com/QwenLM/E-CommerceBench
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
- $\texttt{YC-Bench}$: Benchmarking AI Agents for Long-Term Planning and Consistent Execution
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- UltraHorizon: Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios
- Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
- When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
- TERMS-Bench: Diagnosing LLM Negotiation Agents Beyond Deal Rate
- Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
- RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments
- EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks