Learning from Online User Feedback for Shopping Agents

arXiv:2608.11604 · cs.AI · Submitted 2026-08-12 · Read on arXiv

Haobo Zhang, Kelong Mao, Sulong Xu, Simiu Gu, Zhicheng Dou

Renmin University of China · JD.com

cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: LOFA, a framework that enables shopping agents to learn directly from real online interaction logs without human annotation.

Terminology

Summary

LOFA, a framework that enables shopping agents to learn directly from real online interaction logs without human annotation. LOFA combines reinforcement learning over verifiable purchase outcomes with feedback-aware on-policy distillation, which identifies users’ in-dialogue directives and converts them into dense token-level supervision. These complementary objectives capture both collaborative behavioral patterns and user-specific preferences. Extensive experiments on real-world e-commerce logs demonstrate that LOFA consistently improves recommendation quality, response helpfulness, and user-satisfaction alignment over strong baselines, highlighting the effectiveness of learning shopping agents from real online user feedback.

The paper identifies two complementary forms of online user feedback: explicit behavioral feedback (e.g., purchase outcomes) and in-dialogue directive feedback (e.g., user criticisms, preference refinements, and corrective feedback in subsequent conversational turns). Existing methods primarily rely on explicit behavioral feedback while overlooking in-dialogue directive feedback. Purchase signals are inherently sparse and delayed, and in-dialogue directive feedback is heterogeneous, sparse, and noisy, making it difficult to automatically transform into reliable learning signals.

LOFA systematically transforms heterogeneous online feedback into effective training signals. For explicit behavioral feedback, LOFA automatically identifies interaction sessions associated with purchase outcomes and formulates them as training data, then optimizes the agent through reinforcement learning (GRPO) with purchase-based rewards. For implicit conversational feedback, LOFA leverages an LLM-based feedback mining pipeline to identify directive signals, categorizing them into four actionable types: Explicit Criticism, Implicit Deduction, Pure Negative Feedback, and Comparative Preference (Topic Shift/Abandonment is discarded). LOFA then introduces an on-policy distillation framework where subsequent user feedback is incorporated as additional guidance to construct an enhanced teacher model, providing dense token-level supervision via reverse KL divergence.

The training pipeline is sequential: the agent is first optimized with GRPO using purchase signals, then refined through on-policy distillation using directive feedback. This order (RL→OPD) outperforms the reversed order (OPD→RL), suggesting that learning reliable outcome-level supervision before fine-grained directive supervision is more effective.

Experiments are conducted on real-world interaction logs from Jingyan, a deployed shopping agent on JD.com. The behavioral-feedback dataset (JD-Search) contains 7,001 sessions with 6,704 users and 125,694 candidates, while the directive-feedback dataset (JD-conv) contains 8,249 turn-level samples across 6,121 sessions. The backbone model is Qwen3-8B, with DeepSeek-V3.2 used as the judge model for feedback mining and evaluation.

Main results show that +RL (behavioral feedback only) substantially improves recommendation quality (NDCG, Recall, MAP), while +OPD (directive feedback only) substantially improves the Success Rate in addressing user concerns. Combining both objectives achieves the best overall performance. RL→OPD consistently outperforms OPD→RL. Both RL→OPD and OPD→RL substantially outperform the Base model and the inference-time Self-Reflection baseline. NoThink and +SFT perform substantially worse than all feedback-learning variants.

Ablation studies show that removing the recommendation reward causes the largest performance degradation in ranking metrics, confirming it is the primary learning signal. Removing the format reward causes a smaller decline. Removing any directive-feedback category reduces Success Rate, with Implicit Deduction causing the largest degradation and Pure Negative Feedback also noticeably degrading performance. Recommendation ranking remains largely unchanged across directive-feedback ablations.

Generalization experiments on Qwen3-4B show that LOFA consistently improves both recommendation quality and directive-feedback resolution across backbone scales, demonstrating that the framework is backbone-agnostic.

Case studies illustrate that LOFA correctly identifies hard constraints (e.g., same-day delivery) and resolves user dissatisfaction, and that it ranks the user's actual purchased item first by understanding gift intent, whereas the base model collapses gift into homogeneous skincare recommendations.

Improvements for AI systems

Improvements to AI systems:

  1. Dual-signal learning architecture: Implement a training pipeline that first optimizes an AI agent via reinforcement learning on verifiable outcome rewards (e.g., purchase completion), then refines it via on-policy distillation using in-dialogue user directives (criticisms, preference changes, corrective feedback). This sequential order (outcome-first, directive-second) yields better performance than the reverse.

  2. Automated feedback mining and categorization: Add a module that uses a large language model to parse user conversational turns and classify them into four actionable types—Explicit Criticism, Implicit Deduction, Pure Negative Feedback, and Comparative Preference—while discarding irrelevant topic shifts. This converts noisy, heterogeneous user complaints into structured, dense token-level supervision signals.

  3. Token-level dense supervision via reverse KL divergence: Enhance the AI agent’s policy by distilling from an improved teacher model that incorporates subsequent user feedback as additional context. This provides fine-grained, per-token guidance, improving the agent’s ability to resolve specific user concerns rather than only optimizing for final outcomes.

  4. Hybrid reward shaping: Combine sparse purchase-based rewards with format-compliance rewards during reinforcement learning. This ensures the agent not only recommends the right item but also produces well-structured, helpful responses, preventing degradation in conversational quality.

  5. Preference-aware ranking with hard-constraint handling: Train the agent to identify and prioritize hard constraints (e.g., same-day delivery) and nuanced intents (e.g., gift implying a specific product category) from user dialogue, enabling it to rank the user’s actual desired item first, even when the base model would collapse the intent into generic recommendations.

What the improved AI system can do:

  • Learn from real-world, unlabeled interaction logs without human annotation, automatically extracting both behavioral outcomes (purchases) and conversational directives.

  • Resolve user dissatisfaction in real time by detecting criticisms and preference shifts, then adjusting recommendations or responses mid-dialogue to align with user-specific needs.

  • Outperform strong baselines (including self-reflection and supervised fine-tuning) in both recommendation quality (NDCG, Recall, MAP) and user-satisfaction alignment (Success Rate).

  • Generalize across backbone model sizes (e.g., from 8B to 4B parameters) without performance collapse, making the framework scalable and model-agnostic.

  • Handle sparse and delayed feedback by using dense token-level supervision from immediate user directives, compensating for the rarity of purchase signals.

  • Improve conversational helpfulness by generating responses that address both the explicit query and implicit corrective feedback, leading to higher user satisfaction in deployed e-commerce or assistant settings.

Sources

Related papers