Evaluating Open-Weight E-Commerce Agents with Environment-Grounded Verification
cs.LG
Submitted: 2026-09-14
Updated: 2026-09-14
License: http://creativecommons.org/licenses/by/4.0/
The gist: A shopping conversation has many routes to the same cart, and a task-success rate reduces all of them to one score.
Terminology
Abstract
A shopping conversation has many routes to the same cart, and a task-success rate reduces all of them to one score. We build a deterministic and reproducible e-commerce environment that precommits each trial's customer and trajectory parameters, including the persona, difficulty, target cart, and an item reveal schedule. A simulated consumer attempts to buy a target cart from the environment with assistance from the evaluated model. The environment guides the simulator's actions and records every assistant action alongside the environment state at that point. After the trial, these records allow the evaluator to assess individual parts of the conversation against the retained evidence. For example, the evaluator penalizes a search for failing to surface a target product only when the customer has already mentioned that product. We further use this evidence to apply different penalties to tool calls depending on how the assistant's actions compare with an expected tool-call set. Our environment also interacts with the simulator bidirectionally, reading its output to stop the trial when the simulator determines that the customer has become too frustrated and injecting directives in real time that specify when to explore, defer buying an item, or recall a previous exchange. This interaction creates an open-ended and verifiable simulation. Across eight open-weight agents from 20B to 35B parameters, with 160 trials per agent and 44 metrics, the resulting capability profiles distinguish under-action, over-purchase, unsupported product attributes, and poor search, all of which terminal success obscures.
Sources
- FacTool: Factuality Detection in Generative AI -- A Tool Augmented Framework for Multi-Task and Multi-Domain Scenarios
- Mind2Web: Towards a Generalist Agent for the Web
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- Log analysis is necessary for credible evaluation of AI agents
- AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
- AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
- WebMall -- A Multi-Shop Benchmark for Evaluating Web Agents
- COMPASS: Benchmarking Constrained Optimization in LLM Agents
- Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations
- ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents
- Large Language Models are not Fair Evaluators
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- WebArena: A Realistic Web Environment for Building Autonomous Agents
- Mind the Sim2Real Gap in User Simulation for Agentic Tasks
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks