Learning from Online User Feedback for Shopping Agents
Haobo Zhang, Kelong Mao, Sulong Xu, Simiu Gu, Zhicheng Dou
Renmin University of China · JD.com
cs.AI
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: LOFA, a framework that enables shopping agents to learn directly from real online interaction logs without human annotation.
Terminology
Summary
LOFA, a framework that enables shopping agents to learn directly from real online interaction logs without human annotation. LOFA combines reinforcement learning over verifiable purchase outcomes with feedback-aware on-policy distillation, which identifies users’ in-dialogue directives and converts them into dense token-level supervision. These complementary objectives capture both collaborative behavioral patterns and user-specific preferences. Extensive experiments on real-world e-commerce logs demonstrate that LOFA consistently improves recommendation quality, response helpfulness, and user-satisfaction alignment over strong baselines, highlighting the effectiveness of learning shopping agents from real online user feedback.
The paper identifies two complementary forms of online user feedback: explicit behavioral feedback (e.g., purchase outcomes) and in-dialogue directive feedback (e.g., user criticisms, preference refinements, and corrective feedback in subsequent conversational turns). Existing methods primarily rely on explicit behavioral feedback while overlooking in-dialogue directive feedback. Purchase signals are inherently sparse and delayed, and in-dialogue directive feedback is heterogeneous, sparse, and noisy, making it difficult to automatically transform into reliable learning signals.
LOFA systematically transforms heterogeneous online feedback into effective training signals. For explicit behavioral feedback, LOFA automatically identifies interaction sessions associated with purchase outcomes and formulates them as training data, then optimizes the agent through reinforcement learning (GRPO) with purchase-based rewards. For implicit conversational feedback, LOFA leverages an LLM-based feedback mining pipeline to identify directive signals, categorizing them into four actionable types: Explicit Criticism, Implicit Deduction, Pure Negative Feedback, and Comparative Preference (Topic Shift/Abandonment is discarded). LOFA then introduces an on-policy distillation framework where subsequent user feedback is incorporated as additional guidance to construct an enhanced teacher model, providing dense token-level supervision via reverse KL divergence.
The training pipeline is sequential: the agent is first optimized with GRPO using purchase signals, then refined through on-policy distillation using directive feedback. This order (RL→OPD) outperforms the reversed order (OPD→RL), suggesting that learning reliable outcome-level supervision before fine-grained directive supervision is more effective.
Experiments are conducted on real-world interaction logs from Jingyan, a deployed shopping agent on JD.com. The behavioral-feedback dataset (JD-Search) contains 7,001 sessions with 6,704 users and 125,694 candidates, while the directive-feedback dataset (JD-conv) contains 8,249 turn-level samples across 6,121 sessions. The backbone model is Qwen3-8B, with DeepSeek-V3.2 used as the judge model for feedback mining and evaluation.
Main results show that +RL (behavioral feedback only) substantially improves recommendation quality (NDCG, Recall, MAP), while +OPD (directive feedback only) substantially improves the Success Rate in addressing user concerns. Combining both objectives achieves the best overall performance. RL→OPD consistently outperforms OPD→RL. Both RL→OPD and OPD→RL substantially outperform the Base model and the inference-time Self-Reflection baseline. NoThink and +SFT perform substantially worse than all feedback-learning variants.
Ablation studies show that removing the recommendation reward causes the largest performance degradation in ranking metrics, confirming it is the primary learning signal. Removing the format reward causes a smaller decline. Removing any directive-feedback category reduces Success Rate, with Implicit Deduction causing the largest degradation and Pure Negative Feedback also noticeably degrading performance. Recommendation ranking remains largely unchanged across directive-feedback ablations.
Generalization experiments on Qwen3-4B show that LOFA consistently improves both recommendation quality and directive-feedback resolution across backbone scales, demonstrating that the framework is backbone-agnostic.
Case studies illustrate that LOFA correctly identifies hard constraints (e.g., same-day delivery) and resolves user dissatisfaction, and that it ranks the user's actual purchased item first by understanding gift
intent, whereas the base model collapses gift
into homogeneous skincare recommendations.
Improvements for AI systems
Improvements to AI systems:
-
Dual-signal learning architecture: Implement a training pipeline that first optimizes an AI agent via reinforcement learning on verifiable outcome rewards (e.g., purchase completion), then refines it via on-policy distillation using in-dialogue user directives (criticisms, preference changes, corrective feedback). This sequential order (outcome-first, directive-second) yields better performance than the reverse.
-
Automated feedback mining and categorization: Add a module that uses a large language model to parse user conversational turns and classify them into four actionable types—Explicit Criticism, Implicit Deduction, Pure Negative Feedback, and Comparative Preference—while discarding irrelevant topic shifts. This converts noisy, heterogeneous user complaints into structured, dense token-level supervision signals.
-
Token-level dense supervision via reverse KL divergence: Enhance the AI agent’s policy by distilling from an improved teacher model that incorporates subsequent user feedback as additional context. This provides fine-grained, per-token guidance, improving the agent’s ability to resolve specific user concerns rather than only optimizing for final outcomes.
-
Hybrid reward shaping: Combine sparse purchase-based rewards with format-compliance rewards during reinforcement learning. This ensures the agent not only recommends the right item but also produces well-structured, helpful responses, preventing degradation in conversational quality.
-
Preference-aware ranking with hard-constraint handling: Train the agent to identify and prioritize hard constraints (e.g., same-day delivery) and nuanced intents (e.g.,
gift
implying a specific product category) from user dialogue, enabling it to rank the user’s actual desired item first, even when the base model would collapse the intent into generic recommendations.
What the improved AI system can do:
-
Learn from real-world, unlabeled interaction logs without human annotation, automatically extracting both behavioral outcomes (purchases) and conversational directives.
-
Resolve user dissatisfaction in real time by detecting criticisms and preference shifts, then adjusting recommendations or responses mid-dialogue to align with user-specific needs.
-
Outperform strong baselines (including self-reflection and supervised fine-tuning) in both recommendation quality (NDCG, Recall, MAP) and user-satisfaction alignment (Success Rate).
-
Generalize across backbone model sizes (e.g., from 8B to 4B parameters) without performance collapse, making the framework scalable and model-agnostic.
-
Handle sparse and delayed feedback by using dense token-level supervision from immediate user directives, compensating for the rarity of purchase signals.
-
Improve conversational helpfulness by generating responses that address both the explicit query and implicit corrective feedback, leading to higher user satisfaction in deployed e-commerce or assistant settings.
Sources
- BPR: Bayesian Personalized Ranking from Implicit Feedback
- LLaSA: Large Language and E-Commerce Shopping Assistant
- R$^2$ec: Towards Large Recommender Models with Reasoning
- Reason-to-Recommend: Using Interaction-of-Thought Reasoning to Enhance LLM Recommendation
- A Multi-Agent Conversational Recommender System
- Reason4Rec: Deliberative User Preference Alignment of Large Language Models for Recommendation
- RecThinker: An Agentic Framework for Tool-Augmented Reasoning in Recommendation
- SimUSER: Simulating User Behavior with Large Language Models for Recommender System Evaluation
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- A Survey on Knowledge Distillation of Large Language Models
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
- Qwen3 Technical Report
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection