ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons
Adrian Li, Kelong Mao, Yudong Guo, Heming Xia, Xinwei Yang, Lirui Luo, Jace Wong, Pu Yao, Sulong Xu, Simiu Gu
JD.com
cs.AI, cs.CL
Submitted: 2026-08-10
Updated: 2026-08-11
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: ComboShoppingBench is a benchmark introduced to evaluate large language model (LLM) agents on the task of "combo shopping," which involves constructing a basket of complementary items under joint
Terminology
Summary
ComboShoppingBench is a benchmark introduced to evaluate large language model (LLM) agents on the task of combo shopping,
which involves constructing a basket of complementary items under joint semantic and transactional constraints, such as item compatibility, availability, store-level requirements, delivery fees, coupons, and budgets. The paper notes that real-world shopping often requires building a basket of complementary items rather than retrieving a single product, and this arises in scenarios like device setup, meal preparation, event planning, and group takeout ordering.
The benchmark addresses two fundamental challenges: task synthesis must generate natural and diverse user requests with meaningful cross-item dependencies while ensuring every task has at least one feasible basket, and evaluation must reliably verify arbitrary agent-generated baskets since multiple valid solutions may exist. The central design principle is solution-first construction,
where an exploration agent first identifies and validates a purchasable basket in the environment, which is retained as a hidden witness for synthesizing the user request, budget, and transactional constraints. This ensures every task is solvable by design, while the witness is never used as a reference answer during evaluation.
The benchmark spans 291 tasks across three domains: 151 product-only tasks, 90 takeout or instant-retail tasks, and 50 mixed-domain tasks. The construction pipeline involves five stages: theme-brief curation, witness discovery, coupon-pack synthesis, budget instantiation, and query and rubric synthesis. For theme-brief curation, 16 procurement briefs are defined, covering scenarios such as a core product with required accessories, complementary product sets, individual meals, group meals, and cross-channel pairings, paired with 291 distinct themes. Witness discovery uses an exploration agent that searches the environment and assembles a candidate basket, which undergoes automated validation for domain match, item-store associations, ordering requirements, and well-formed price and fee information. Coupon-pack synthesis creates packs of five or six coupons per task, including threshold-reduction, direct-reduction, and capped percentage-discount coupons, with varied scopes, thresholds, caps, and mutual-exclusion rules, including plausible but suboptimal decoy coupons. Budget instantiation constructs a moderately widened interval around the optimal payable amount p*, rounded to natural currency values, with three modes: upper limit, approximate target, and explicit range. Query and rubric synthesis generates natural-language requests and task-specific semantic rubrics.
Evaluation uses a four-dimensional hybrid framework: semantic satisfaction, rule-based validation, response quality, and claim faithfulness. Semantic satisfaction uses an LLM judge to evaluate whether selected products fulfill the user's shopping intent against task-specific rubrics. Rule-based validation reconstructs a structured basket from the agent's response and deterministically verifies SKU validity, coupon legality and optimality, and budget compliance. Response quality assesses how effectively the solution is communicated, and claim faithfulness checks whether factual statements in the response are consistent with environment-recomputed results. Overall Success is defined as the intersection of all four dimensions.
The paper evaluates 11 proprietary and open-source agents under Think and No-think configurations, yielding 22 configurations. The strongest agent, GPT-5.5 with thinking enabled, achieves an Overall success rate of only 61.2%. The paper reports that high pass rates on individual dimensions do not imply reliable end-to-end performance, with GPT-5.5 (Think) exceeding 83% on both Semantic and Rule-based validation and 90% on response-related dimensions but still failing nearly 40% of tasks. Failure analyses reveal persistent limitations in compositional requirement satisfaction, coupon and budget optimization, and faithful reporting.
The failure analysis identifies several key bottlenecks. Agents struggle to choose mutually exclusive coupons, with the main bottleneck being selection among mutually exclusive alternatives rather than missing stackable coupons. Think reduces the rates of selecting a suboptimal exclusive coupon and skipping an exclusive group from 15% each to 8% and 7%, respectively, increasing the optimal-set rate from 68% to 83%. The difficulty of compositional shopping stems primarily from constraint accumulation, as task-level Semantic all-pass rates drop substantially as the number of semantic criteria increases, while criterion-level failure rates rise only slightly. Cross-item requirements have the highest semantic failure rate, followed by quantity requirements, while channel and store requirements are the easiest to satisfy.
The paper also examines whether thinking configurations help, finding that Think improves aggregate performance for most agents but is not uniformly positive across tasks. Think changes the solution trajectory, resolving some tasks that fail under No-think while causing a subset of previously successful tasks to fail. For example, Qwen3.6-27B loses more tasks than it recovers under Think, with its average calculator usage dropping from 4.21 to 1.19 calls, suggesting over-reliance on internal reasoning at the expense of necessary exact computation.
The reliability of LLM-based evaluators is assessed through a human-annotated reference set from 30 outputs, with two annotators independently labeling 700 rubric instances and an expert adjudicating disagreements. Three evaluators (Gemini-3.1-Pro, GPT-5.5, and Kimi-K2.6) achieve at least 97% overall agreement with the human reference, with agreement consistently high for Semantic (97.71–98.39%) and reaching 99.12% for Claim Faithfulness across all evaluators. Response Quality exhibits greater variation but remains above 91%.
The paper concludes that ComboShoppingBench highlights substantial room for improvement in reliable, constraint-aware combo shopping, with the strongest configuration achieving only 61.2% Overall success. The analysis identifies cross-item constraint accumulation and optimization over mutually exclusive coupons as key bottlenecks, and shows that Think configurations generally help but are not uniformly beneficial.
Improvements for AI systems
Improvements to AI Systems:
-
Implement explicit compositional constraint tracking and decomposition. The AI system should parse user requests into a structured dependency graph of semantic criteria (e.g., cross-item compatibility, quantity, channel, store) and maintain a live checklist during basket construction. It must verify each criterion incrementally, rather than relying on holistic reasoning, to prevent constraint accumulation failures.
-
Add a dedicated coupon-optimization module with exhaustive local search. The system should enumerate all mutually exclusive coupon groups, simulate each combination against the candidate basket, and select the globally optimal set (including decoy rejection). This replaces heuristic coupon selection with deterministic optimization, addressing the 15–17% failure rate on exclusive-coupon choices.
-
Integrate a mandatory external calculator for all arithmetic operations. The system must route price, fee, discount, and budget calculations to a deterministic tool (e.g., Python interpreter) instead of internal reasoning. This prevents over-reliance on latent arithmetic, which caused Qwen3.6-27B to lose performance when internal reasoning replaced calculator calls.
-
Enable dual-mode reasoning with automatic fallback. The system should run both Think and No-think trajectories in parallel, compare intermediate solution validity (e.g., via rule-based checks), and select the trajectory with higher verified constraint satisfaction. This mitigates cases where thinking degrades performance (e.g., by skipping exact computation) while retaining its benefits for complex tasks.
-
Build a self-verification loop using environment-recomputed facts. After generating a basket, the system must re-fetch live prices, fees, and coupon applicability from the environment, then re-run all rule-based validations (SKU validity, coupon legality, budget compliance) before finalizing the response. Any discrepancy triggers a correction pass, ensuring claim faithfulness.
-
Add cross-item compatibility reasoning via structured embeddings. The system should encode item attributes (e.g., device ports, meal dietary tags, event type) into a compatibility matrix and use graph-based constraint satisfaction to propose only mutually compatible items. This directly targets the highest semantic failure rate (cross-item requirements).
-
Implement rubric-aware response generation. The system should generate its final answer by explicitly mapping each semantic rubric criterion to a stated product or action, and then self-check that every rubric item is addressed. This improves Response Quality and Semantic all-pass rates by forcing explicit coverage.
What the Improved AI System Can Do:
-
Achieve >80% Overall success on ComboShoppingBench by reducing compositional constraint failures (from 40% to <20%) and eliminating coupon-optimization errors (from 15–17% to <5%).
-
Reliably handle multi-item, multi-constraint shopping tasks across product-only, takeout, and mixed domains, including group meals and device setups, with deterministic verification of every transactional rule.
-
Optimize baskets under coupon mutual-exclusion and budget intervals by exhaustively searching coupon combinations and using exact arithmetic, ensuring the chosen basket is provably cost-minimal or within the stated range.
-
Maintain high performance even when thinking is enabled or disabled by automatically selecting the better trajectory, avoiding regressions seen in smaller models.
-
Produce faithful, verifiable responses where every claim (prices, fees, savings) matches environment recomputation, eliminating hallucinated discounts or availability.
-
Scale to arbitrary new procurement themes by generalizing the structured constraint graph and compatibility embeddings, without requiring task-specific fine-tuning.
Abstract
Real-world shopping often requires constructing a basket of complementary items rather than retrieving a single product. Such combo-shopping tasks arise in device setup, meal preparation, event planning, and group takeout ordering, requiring joint reasoning about item compatibility, availability, store-level requirements, delivery fees, coupons, and budgets. Evaluation is challenging because multiple baskets may satisfy the same request, making exact-match metrics unsuitable, whereas semantic evaluation alone cannot detect infeasible orders, invalid coupon combinations, or incorrect payments. We introduce ComboShoppingBench, an agentic shopping benchmark for open-ended yet verifiable basket construction in a simulated commerce and takeout environment. During task synthesis, an exploration agent constructs a feasible and semantically coherent basket of purchasable products; this witness guides the generation of coupons, budget constraints, user queries, and aligned evaluation rubrics. During evaluation, LLM judges assess semantic satisfaction, response quality, and claim faithfulness, while deterministic validation checks product-ID validity, budget compliance, and coupon optimality. Experiments with diverse LLM agents demonstrate that even strong agents struggle on ComboShoppingBench, highlighting substantial room for improvement in reliable, constraint-aware combo shopping.
Sources
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
- DeepShop: A Benchmark for Deep Research Shopping Agents
- ShoppingComp: Are LLMs Really Ready for Your Shopping Cart?
- WebMall -- A Multi-Shop Benchmark for Evaluating Web Agents
- Shopping Companion: Benchmarking and Training LLM Agents for Long-Horizon Preference-Grounded E-Commerce Tasks
- EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent
- Mind2Web: Towards a Generalist Agent for the Web
- WebArena: A Realistic Web Environment for Building Autonomous Agents
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
- WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
- ShopGym: An Integrated Framework for Realistic Simulation and Scalable Benchmarking of E-Commerce Web Agents
- A Survey on Bundle Recommendation: Methods, Applications, and Challenges
- BRIDGE: Bundle Recommendation via Instruction-Driven Generation
- EpicCBR: Item-Relation-Enhanced Dual-Scenario Contrastive Learning for Cold-Start Bundle Recommendation
- Time-Interval-Aware Disentangled Expert Modeling for Next-Basket Recommendation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection