Contextual Scalarisation Thompson Sampling for multi-objective decisions in public media
cs.IR, cs.LG
Submitted: 2026-05-29
Updated: 2026-09-13
Comments: 15 pages, 3 figures, 3 tables. Submitted-manuscript version of a paper published at ICPR 2026 (LNCS vol. 16824, Springer). v2 adds the publisher acknowledgement and the DOI of the Version of Record, and corrects bibliography metadata; no other changes
Journal ref: Pattern Recognition. ICPR 2026. Lecture Notes in Computer Science, vol. 16824, pp. 357-372. Springer, Cham (2027)
DOI: 10.1007/978-3-032-31927-2_24
License: http://creativecommons.org/licenses/by/4.0/
The gist: Recommender systems may operate under multiple, competing objectives.
Terminology
Abstract
Recommender systems may operate under multiple, competing objectives. For example, audience reach, cultural values, public service mandate, and operational constraints must be balanced in editorial decisions of public service media. Existing approaches relying on fixed combinations of objectives or Pareto-based optimisation do not adapt to changing priorities across situations. In this paper, we propose Contextual Scalarisation Thompson Sampler (CSTS), a multi-objective contextual bandit method that learns to weight objectives as a function of the observed context. We evaluate CSTS on real programming data from Radio Télévision Suisse, the Swiss national broadcaster, showing improved contextual relevance and better alignment with expert curation practices compared to fixed weight and standard contextual bandit approaches.
Sources
- Thompson Sampling for Contextual Bandits with Linear Payoffs
- Neural Collaborative Filtering
- Pareto-based Multi-Objective Recommender System with Forgetting Curve
- A Tutorial on Thompson Sampling
- HyperBandit: Contextual Bandit with Hypernewtork for Time-Varying User Preferences in Streaming Recommendation
- Contextual-Bandit Based Personalized Recommendation with Time-Varying User Interests
- Scalable Neural Contextual Bandit for Recommender Systems
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG