AURA: Agentic Diagnosis and Refinement for Production Recommender Systems at Scale
cs.IR, cs.AI, cs.LG
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: 14 pages, 1 figure, 6 tables. Accepted at GenAIECommerce'26: The Third Workshop on Agentic and Generative AI for E-Commerce, co-located with RecSys 2026, September 28, 2026, Minneapolis, MN, USA
License: http://creativecommons.org/licenses/by/4.0/
The gist: How and why does a recommender system fail the users it serves? Oftentimes, practitioners are left to improve their algorithms based on a combination of feedback from stakeholder teams, domain
Terminology
Abstract
How and why does a recommender system fail the users it serves? Oftentimes, practitioners are left to improve their algorithms based on a combination of feedback from stakeholder teams, domain expertise, and insights from data analyses. Yet the nuances of how and where recommendations perform well or poorly for end users are difficult to discern from aggregate quantitative metrics. Whereas these metrics provide a high-level and incomplete picture, further granularity into the quality of recommendations and their patterns requires reasoning with domain understanding and objectivity, at scale. We contemplate this complex conundrum and describe a method and implementation that uses the latest AI agentic advances to provide actionable diagnoses and improvements for production recommender systems. We present AURA (Agentic Understanding and Refinement of recommender Algorithms), an end-to-end agentic system that performs qualitative evaluation at scale and can then generate improvements to our algorithms at the code level. Specialized agents read production engagement logs, from thousands of sessions to millions, and surface patterns and examples of how the recommender fails real users. The next step uses those diagnoses and context about the recommender's own code, data, and training pipeline to propose and implement refinements grounded in that codebase. We report the system design, initial tests on production data from two large consumer platforms at a major media-streaming company, safeguards, operational learnings, and early results toward a self-improving recommender system. Finally, the diagnostic gap AURA closes is not specific to streaming. The architecture is built to transfer: every domain-specific element enters through the configuration layer that already ported it between our two platforms. We map it concretely to e-commerce and online-retail recommendation.
Sources
- Simpson's Paradox in Recommender Fairness: Reconciling differences between per-user and aggregated evaluations
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- No-Human in the Loop: Agentic Evaluation at Scale for Recommendation
- Large Language Models are Zero-Shot Rankers for Recommender Systems
- Large Language Models are not Fair Evaluators
- Style Over Substance: Evaluation Biases for Large Language Models
- Human Feedback is not Gold Standard
- Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation
- No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
- Beyond NDCG: behavioral testing of recommender systems with RecList
- On Generative Agents in Recommendation
- User Behavior Simulation with Large Language Model based Agents
- Rethinking the Evaluation for Conversational Recommendation in the Era of Large Language Models
- RecAI: Leveraging Large Language Models for Next-Generation Recommender Systems
- Self-Evolving Recommendation System: End-To-End Autonomous Model Optimization With LLM Agents
- Self-EvolveRec: Self-Evolving Recommender Systems with LLM-based Directional Feedback
- MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
- AIDE: AI-Driven Exploration in the Space of Code
- Self-Refine: Iterative Refinement with Self-Feedback
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG