Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy
Ravi Teja Chunduri, Srikaran Reddy Boya, Deep Narayan Mishra, Ajay Kumar B, Karthik Kumaran, Pranay Kona
Walmart Global Tech
cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: 8 pages. Accepted in the Main Conference of IEEE ICMLA 2026
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 85/100
Terminology
Summary
Summary
This paper presents a scalable, context-aware Multi-Agent Framework designed to automate the construction of Lines and Ladders
pricing taxonomies for large-scale retail catalogs. The framework addresses the challenge of maintaining price consistency and executing an Every Day Low Price strategy for global retailers with catalogs spanning millions of active items, where manual governance of price relationships is infeasible. Inconsistent pricing across item variants distorts customer value perception and cannibalizes sales.
The core problem is not merely the volume of items, but the operational bottleneck created by inconsistent and non-standardized catalog data. The paper identifies two common gaps in unmanaged catalogs: Variant Inconsistency (e.g., identical bikes priced at 238 vs 199 solely due to color) and Data Discrepancy (e.g., a 0.1-inch typo in a mirror's dimensions creating a 7.50 gap).
The framework decomposes pricing governance into Similarity and Variance discovery via a 3-agent architecture. The methodology involves scoping the catalog by ProductType (e.g., Candles
) as a granular anchor, then using two specialized LLM agents for dual-pronged attribute discovery:
-
Agent 1 (Similarity Attributes): Identifies base features common to identically priced items, defining the core value proposition. It uses density-based sampling to find the K most frequent (Brand, MSRP) tuples, generates corpora of 15 item descriptions per tuple, and iteratively extracts consistent core attributes using persistent memory. The prompt strategy enforces quantifiable-only attributes, consistency, evidence tiering (explicit, inferred, speculative), and standardization.
-
Agent 2 (Variance Attributes): Identifies premium features that justify price stratification. It uses stratified sampling of the top 5 brands and their top 5 MSRP points (up to 25 groups), ingests the entire stratified dataset as a single input, and performs comparative analysis between low- and high-priced groups to identify differentiating features. The agent acts as a
Pricing Strategist
and classifies attributes as Type 1 (Differentiating) or Type 2 (Foundational).
Parameter selection was tuned to optimize the Pareto frontier between signal-to-noise ratio and LLM token costs. The proposed thresholds (30 groups / 15 items for Agent 1, and 25 groups / 5 items for Agent 2) achieved an F1 score of 0.83 for Lines and 0.82 for Ladders, outperforming larger configurations (e.g., 80/30 achieved 0.77 Lines F1 but scaled runtime to 155 minutes vs 65 minutes for proposed thresholds).
A Synthesis Agent performs a mapping function Φ to produce a canonical schema Spt = Φ(Asim ∪ Avar), executing three sequential operations: Attribute Consolidation & Semantic Resolution (eliminating redundancy, resolving synonymy, enforcing standardized naming), Attribute-Unit Decoupling (separating numerical values from units to resolve UOM inconsistencies), and Schema Definition (defining data types, defaults, and extraction prompts).
For Multi-Modal Attribute Extraction, the framework populates the schema for every item using a multi-modal function fLLM that maps raw data (text + images) to feature vectors. This capability is critical because key specifications often appear only on packaging images. The system uses a single generalized agent to extract arbitrary attributes defined by the dynamic schema, resolving the scalability bottleneck of traditional NER models. Initial experiments with smaller open-weight models (Nemotron VLM 2B/8B/12B) failed to achieve production accuracy, necessitating reliance on frontier models.
Data Standardization and Normalization employs a post-processing pipeline Ψ including: Numerical Standardization (UOM conversion with a strict 5% numerical tolerance), Categorical Semantic Clustering (LLM-based clustering to resolve semantic fragmentation like Soy
vs. Soy Blend
), and Brand Normalization (hybrid function reconciling extracted brand with catalog master using Jaccard Similarity and brand frequency).
Adaptive Taxonomy Construction is scoped at the brand level. Feature selection applies heuristics: dropping attributes with >65% nullity or >70% cardinality, and applying 5% tolerance numerical binning. Hierarchical grouping logic generates Lines by grouping items by the complete set of categorical and numerical attributes (each unique combination forms a Line with a persistent UUID), and Ladders by relaxing constraints (excluding numerical attributes) to group Lines by categorical attributes only, linking items differing only by size or pack quantity.
A Human-in-the-Loop (HITL) feedback mechanism captures merchant corrections as labeled signals, updating a ProductType-specific contextual memory Pcontext that drives future pipeline refinements. This allows agents to learn strategic alignment, schema refinement, and business-specific naming conventions.
Deployment at scale uses a distributed compute cluster with a refresh orchestrator supporting full refresh and delta refresh modes. The architecture decomposes the pipeline by Department → ProductType for independent failure domains. A multi-model routing strategy uses: a proprietary enterprise-managed frontier LLM (comparable to >100B parameter thinking
models) for reasoning-heavy tasks, a cost-optimized multi-modal model (comparable to 7B-13B VLMs) for high-throughput extraction, a SQL-based semantic cache for synonym mappings, and hierarchical clustering logic. The system processed 1,700 ProductTypes (1 million items) in 30 minutes using 10 worker nodes, with LLM inference dominating spend (4K tokens/item).
Experimental evaluation includes:
-
Architecture Ablation: A single-agent LLM suffered severe cognitive overload (0.64 Lines F1), a 2-agent system improved to 0.77 F1, but the 3-agent system (with Synthesis Agent) achieved 0.83 F1, strictly necessary for production (>80% accuracy).
-
Quantitative Manual Evaluation: Domain experts evaluated 1000 randomly sampled items across 11 diverse General Merchandise ProductTypes, achieving 80.0% weighted average accuracy. Qualitative feedback revealed three error categories: Technical Granularity (missing specific technical variants like
Ultrasonic
vs.Evaporative
), Aesthetic & Subjective Nuance (intangible qualities likesilhouette elegance
), and Data Completeness & Complexity (sparse descriptions or complex bundle combinations). In most failure cases, grouping logic was directionally correct but missed a single differentiating attribute. -
Comparative Evaluation against Merchant Reference: In Food & Consumables across 14 ProductTypes, the system achieved average Line Precision of 98% and Line Recall of 75% (Ladders: 92% Precision, 81% Recall). This high-precision/low-recall profile is an intentional guardrail—high-recall groupings risk applying price changes to unrelated items. The recall gap stems from merchants frequently over-grouping items based on promotional strategies or broad
flavor families
rather than strict attribute equivalence. -
Post-Launch Business Impact: The fully launched production system processes tens of thousands of active items, generating structured Lines and Ladders for >90% of the previously un-managed catalog for the first time. End-to-end telemetry over a 13-week period indicates an 86.1% reduction in time-on-task for existing workflows, calculated via: New Hours = Old Hours × Remaining Workload (0.625) × Remaining Time (0.222), resulting in workflows requiring only 13.9% of original hours.
Limitations include LLM non-determinism despite temperature 0, sensitivity to sampling bias, compute costs restricting batch processing, reliance on specific LLM families requiring re-validation upon model updates, and the need for future empirical benchmarking against non-LLM baselines. Future directions include refining anchoring to brand level, longitudinal studies on HITL convergence, statistical significance testing, cross-category transfer learning for niche categories (5%), and teacher-student distillation into Small Language Models for real-time inference.
The paper concludes that multi-agent architectures are highly viable for industrial-scale governance, offering a framework broadly applicable to other heterogeneous domains like industrial supply chains and online marketplaces.
Improvements for AI systems
Improvements to AI Systems:
-
Cognitive Load Decomposition via Role-Specialized Agents: Instead of a single monolithic LLM handling complex classification tasks, decompose the problem into distinct agents with narrow, well-defined roles (e.g., similarity discovery, variance discovery, schema synthesis). This reduces cognitive overload and improves F1 scores from 0.64 to 0.83 in production settings. The improved system can handle large-scale data governance tasks with higher accuracy by preventing context dilution.
-
Dynamic Schema Generation with Attribute-Unit Decoupling: Implement a synthesis layer that separates numerical values from their units (e.g., "10
vs.
inches") and resolves synonymy across attributes. This enables the AI to automatically generate a canonical, machine-readable schema for any new domain without manual annotation, making it adaptable to heterogeneous catalogs, supply chains, or marketplaces. -
Evidence-Tiered Prompting with Persistent Memory: Use a prompting strategy that requires the LLM to classify extracted attributes into explicit, inferred, or speculative evidence tiers, while maintaining a persistent memory of previously validated attributes. This improves consistency and reduces hallucination in attribute extraction, enabling the system to produce reliable, standardized taxonomies from noisy, unstructured data.
-
Multi-Modal Extraction with Dynamic Schema Following: Train or configure a single generalized vision-language agent to extract arbitrary attributes defined by a dynamic schema (rather than fixed NER labels). This allows the system to read specifications from packaging images and text, overcoming the scalability bottleneck of traditional NER models. The improved system can process millions of items with mixed media inputs (text + images) at production accuracy.
-
Hybrid Normalization Pipeline with Tolerance-Based Validation: Implement a post-processing layer that combines LLM-based semantic clustering (e.g., resolving
Soy
vs.Soy Blend
) with rule-based numerical standardization (e.g., unit conversion with a strict 5% tolerance) and brand reconciliation using Jaccard similarity. This yields a robust, error-tolerant normalization system that can clean inconsistent enterprise data at scale. -
Adaptive Hierarchical Grouping with Guardrails: Use a two-tier grouping logic—first grouping by all attributes (Lines), then relaxing numerical constraints (Ladders)—with heuristics like dropping attributes with >65% nullity or >70% cardinality. This enables the AI to automatically generate a high-precision, low-recall taxonomy that avoids risky price changes on unrelated items, a critical safety feature for pricing automation.
-
Human-in-the-Loop Contextual Memory: Capture merchant corrections as labeled signals and feed them back into a ProductType-specific memory that updates future pipeline runs. This allows the AI system to learn business-specific naming conventions, strategic alignment, and schema refinements over time, improving accuracy and reducing manual oversight in recurring workflows.
-
Multi-Model Routing for Cost-Performance Optimization: Implement a routing strategy that sends reasoning-heavy tasks (e.g., attribute discovery) to a large frontier model, while delegating high-throughput extraction tasks to a smaller, cost-optimized vision-language model. This reduces token spend (4K tokens/item) and enables processing of 1 million items in 30 minutes on a 10-node cluster, making the system economically viable for industrial-scale deployment.
-
Stratified Sampling for Balanced Comparative Analysis: Use density-based and stratified sampling (e.g., top 5 brands × top 5 MSRP points) to generate balanced corpora for comparative analysis. This improves the signal-to-noise ratio in attribute discovery and prevents sampling bias, enabling the AI to identify true differentiating features rather than artifacts of skewed data.
-
Failure-Aware Error Categorization: Incorporate a feedback loop that classifies errors into technical granularity, aesthetic/subjective nuance, or data completeness issues. This allows the AI to adjust its extraction prompts and grouping logic dynamically, improving performance on edge cases like sparse descriptions or complex bundle combinations.
What the Improved AI System Can Do:
-
Automatically generate and maintain consistent pricing taxonomies (Lines and Ladders) for catalogs with millions of items, reducing manual governance time by 86.1%.
-
Extract and standardize attributes from heterogeneous, noisy data (text + images) across diverse product categories with >80% accuracy, even for previously unmanaged items.
-
Learn from human corrections and improve over time, adapting to business-specific naming and strategic pricing rules.
-
Operate at industrial scale (1,700 product types, 1 million items in 30 minutes) with distributed compute and cost-efficient model routing.
-
Provide high-precision, low-recall groupings that minimize risk of erroneous price changes, ensuring safe deployment in live retail environments.
-
Transfer to other heterogeneous domains (e.g., industrial supply chains, online marketplaces) by dynamically generating schemas and taxonomies without manual re-engineering.
Sources
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection