CORAL: An LLM-Native Harness for Production Recommender Systems
cs.CL
Submitted: 2026-09-02
Updated: 2026-09-02
Comments: Accepted by RecSys '26 OARS Workshop
License: http://creativecommons.org/licenses/by/4.0/
The gist: Production recommender systems shape what billions of people see, and sustaining their performance requires continual optimization: as content, user behavior, and upstream models shift, the choices
Terminology
Abstract
Production recommender systems shape what billions of people see, and sustaining their performance requires continual optimization: as content, user behavior, and upstream models shift, the choices governing retrieval, ranking, and serving must be revisited. Traditionally, human engineers test such changes through online experiments--a slow, reactive process limited by engineering effort, leaving parts of the system unrevised as conditions change. Although large language models have been applied to ranking, user modeling, and offline model development, few systems place an agent in a continual closed loop that acts on a live recommender and learns from the measured effects of its decisions. We present CORAL (Constraint-Optimized Recommender via an Agentic Loop), an LLM-native harness that closes this loop: each cycle, the agent observes operating signals, reasons over a memory of past decisions and outcomes, and invokes tools--including a numerical optimizer that keeps changes within a fixed operating budget--to reconfigure the recommender, with measured outcomes informing the next cycle. We formulate this as a partially observed, non-stationary, constrained optimization problem in which the policy improves in context, without parameter updates, from its prior actions. Across two large-scale social platforms, evaluated with A/B experiments, the same harness improves engagement at no additional serving cost on one and reduces serving cost without degrading engagement on the other, spanning the engagement-efficiency frontier. Performance improves as the loop iterates, suggesting that a single agentic loop can automate continual optimization work traditionally performed by human algorithm engineers under explicit guardrails.
Sources
- MemRec: Collaborative Memory-Augmented Agentic Recommender System
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Interactive Recommendation Agent with Active User Commands
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
- Rethinking Recommendation Paradigms: From Pipelines to Agentic Recommender Systems
- Recommender AI Agent: Integrating Large Language Models for Interactive Recommendations
- AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems
- MACRec: a Multi-Agent Collaboration Framework for Recommendation
- RecNet: Self-Evolving Preference Propagation for Agentic Recommender Systems
- MemOS: A Memory OS for AI System
- From Atom to Community: Structured and Evolving Agent Memory for User Behavior Modeling
- NOVA: A Verification-Aware Agent Harness for Architecture Evolution in Industrial Recommender Systems
- Deep Learning Recommendation Model for Personalization and Recommendation Systems
- AMEM4Rec: Leveraging Cross-User Similarity for Memory Evolution in Agentic LLM Recommenders
- Deep Research for Recommender Systems
- MemGPT: Towards LLMs as Operating Systems
- A Survey on LLM-powered Agents for Recommender Systems
- Agentic Recommender System with Hierarchical Belief-State Memory
- User Behavior Simulation with Large Language Model based Agents
- RecMind: Large Language Model Powered Agent For Recommendation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering