Netflix Artwork Personalization via LLM Post-training
cs.IR, cs.AI
Submitted: 2026-01-06
Updated: 2026-08-26
Comments: Pluralistic Alignment @ ICML 2026 Workshop; 6 pages
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Large language models (LLMs) have demonstrated success in various applications of user recommendation and personalization across e-commerce and entertainment.
Terminology
Abstract
Large language models (LLMs) have demonstrated success in various applications of user recommendation and personalization across e-commerce and entertainment. On many entertainment platforms such as Netflix, users typically interact with a wide range of titles, each represented by an artwork. Since users have diverse preferences, an artwork that appeals to one type of user may not resonate with another with different preferences. Given this user heterogeneity, our work explores the novel problem of personalized artwork recommendations according to diverse user preferences. Similar to the multi-dimensional nature of users' tastes, titles contain different themes and tones that may appeal to different viewers. For example, the same title might feature both heartfelt family drama and intense action scenes. Users who prefer romantic content may like the artwork emphasizing emotional warmth between the characters, while those who prefer action thrillers may find high-intensity action scenes more intriguing. Rather than a one-size-fits-all approach, we conduct post-training of pre-trained LLMs to make personalized artwork recommendations, selecting the most preferred visual representation of a title for each user and thereby improving user satisfaction and engagement. Our experimental results with Llama 3.1 8B models (trained on a dataset of 110K data points and evaluated on 5K held-out user-title pairs) show that the post-trained LLMs achieve 3-5% improvements over the Netflix production model, suggesting a promising direction for granular personalized recommendations using LLMs.
Sources
- Flamingo: a Visual Language Model for Few-Shot Learning
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
- Bootstrapping Language Models with DPO Implicit Rewards
- Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective
- Otter: A Multi-Modal Model with In-Context Instruction Tuning
- Chat-REC: Towards Interactive and Explainable LLMs-Augmented Recommender System
- DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training
- How Can Recommender Systems Benefit from Large Language Models: A Survey
- MiniLLM: On-Policy Distillation of Large Language Models
- Understanding Reference Policies in Direct Preference Optimization
- Measuring Mathematical Problem Solving With the MATH Dataset
- LLM Post-Training: A Deep Dive into Reasoning Large Language Models
- RLPF: Reinforcement Learning from Prediction Feedback for User Summarization with LLMs
- ClipCap: CLIP Prefix for Image Captioning
- A Survey on Knowledge Distillation of Large Language Models
- Training language models to follow instructions with human feedback
- Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs
- STaR: Bootstrapping Reasoning With Reasoning
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG