FedSlate:A Federated Deep Reinforcement Learning Recommender System

arXiv:2409.14872 · cs.IR, cs.AI · Submitted 2025-04-28 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "FedSlate:A Federated Deep Reinforcement Learning Recommender System".

Jane: The paper was written by Yongxin Deng, Xihe Qiu, Xiaoyu Tan and Yaochu Jin from Shanghai University of Engineering Science and INFLY TECH (Shanghai) Co., Ltd. and Westlake University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Welcome back to the show, everyone! Today we're diving into a paper that's got a mouthful of a title: "FedSlate: A Federated Deep Reinforcement Learning Recommender System." Jane, I have to say, just reading that title out loud made me want to grab a coffee.

Jane: Ha! I know exactly what you mean, Tom. But behind that dense title is a really clever idea. So, the authors — Yongxin Deng, Xihe Qiu, Xiaoyu Tan, and Yaochu Jin — they're tackling a problem that's been bugging recommender systems for a while. You know how Netflix suggests what to watch, or Spotify suggests what to listen to?

Tom: Sure, I'm a binge-watcher, I'm a victim of these systems daily.

Jane: Well, these systems usually learn from your behavior on that one platform. But the researchers here are saying, look, your behavior on Netflix might be influenced by what you saw on Spotify or on a shopping app. They're all connected in your head.

Tom: So they're trying to connect the dots between platforms without actually sharing your private data between them. That sounds like a tall order.

Jane: Exactly. And that's where the "Federated" part comes in. It's a way to train a shared model without ever moving your raw data. Each platform keeps its own data, but they can still learn from each other.

Tom: And the "Deep Reinforcement Learning" part? That's the other half of the puzzle.

Jane: Right. So instead of just predicting what you'll click on next, they're using reinforcement learning to think long-term. It's not just about the immediate click; it's about keeping you engaged and satisfied over weeks and months. They use a specific algorithm called SlateQ to handle the fact that they recommend a whole slate of items at once, not just one.

Tom: So it's like a chess player thinking several moves ahead, but for recommending videos and products.

Jane: Precisely. And the really clever bit is how they combine these two ideas. The paper is from a group that includes folks at Shanghai University of Engineering Science and Westlake University. They're not just theorizing; they built a simulation to test it.

Tom: And we're going to get into all that in a bit. But first, Jane, what's the big "why" here? Why should we care if my shopping app knows I like sci-fi movies?

Jane: Because it could make recommendations so much better. Imagine you're looking for a new book on your e-reader, and it knows you've been watching a lot of space documentaries on a video platform. It might suggest a hard sci-fi novel you wouldn't have found otherwise. That's the promise, and it all happens without your reading history ever leaving your e-reader.

Tom: Okay, I'm hooked. Let's dig into how they actually pulled this off. That's up next.

Paper Summary: Tom: Alright, we're back with "FedSlate: A Federated Deep Reinforcement Learning Recommender System." Jane, we've set the stage, but what's the core setup here? How does this thing actually work?

Jane: So imagine two platforms, let's call them Platform A and Platform B. They're both recommending content to the same user. The key assumption is that what happens on Platform A affects how the user feels and acts on Platform B. They call this "cross-platform" influence.

Tom: So my mood after watching a sad movie on one app might affect whether I buy that comfort food on the other app.

Jane: Exactly that. Now, the problem is, in their setup, Platform A gets direct feedback from the user—like, did they click, did they stay, did they engage? But Platform B doesn't get that feedback at all. It's like recommending things in the dark.

Tom: That sounds like a terrible position for Platform B to be in. How does it ever learn?

Jane: That's the magic of the federated part. They have a central server, but it doesn't see any user data. Each platform has its own local "Q-network" that calculates a value for each item. These values are sent to the central server. The server has a global network that takes these values and combines them to produce a new, "federated" value for each platform.

Tom: So the server is like a blind chef who only gets the ingredients' flavor profiles, not the actual ingredients, and then decides the final recipe.

Jane: That's a great analogy. The server sends these new, combined values back to the platforms. Then, Platform A, which has the reward, uses that to update its own network and the global network. Platform B, which has no reward, just uses the global network's output to make its recommendations.

Tom: And they tested this in a simulation called RecSim, right? I remember that from the paper.

Jane: Yes, they built a "Choc vs. Kale" scenario. Chocolate is the fun, engaging content that's bad for you long-term, and Kale is the boring, healthy content that's good for you. The agent has to balance recommending both to maximize the user's long-term satisfaction.

Tom: And the results? Did Platform B, the one in the dark, actually learn something useful?

Jane: That's the headline result. Platform B, using FedSlate, consistently outperformed a random recommendation strategy. It actually learned to recommend good content even though it never saw a single reward signal. It was learning indirectly, through the information shared by Platform A.

Tom: That's wild. It's like learning to cook by only watching someone else taste the food. So it's not just about making a good platform better; it's about making a platform with no data at all become competent. That's a huge deal. What about the platform that does have the data? Did it get worse?

Jane: That's the trade-off they found. Platform A, which had direct feedback, learned faster with FedSlate, but its final, optimal reward was slightly lower than if it had just trained alone. It's a trade-off between speed and peak performance, but the benefit of bringing Platform B up to speed seems to outweigh that cost.

Tom: Okay, so we have a system that works in a simple two-platform world. But what happens when you throw more platforms into the mix? Let's talk about that next.

Improvements and Extensions: Tom: We're back with "FedSlate: A Federated Deep Reinforcement Learning Recommender System." So, Jane, we've seen it work with two platforms. But the real world has dozens of apps fighting for our attention. Does this thing scale?

Jane: Great question, and the authors actually tested that. They ran experiments with five platforms in two different configurations. In the first, all five platforms had access to user feedback. In the second, two of the five platforms were "blind," just like Platform B in the earlier experiment.

Tom: And what happened? Did the whole thing collapse under the weight of more participants?

Jane: Quite the opposite. In the first configuration, all five platforms learned faster and more robustly than if they'd been working alone. They were sharing information and everyone benefited. It was a clear win.

Tom: And the second configuration, with the blind platforms?

Jane: This is where it gets really interesting. The two blind platforms, which had no direct feedback, achieved performance comparable to the original SlateQ algorithm that had full access to feedback. They got to the same level of recommendation quality just by being part of the federation.

Tom: So they're getting the benefits of a well-trained system without ever seeing the reward signal. That's like a student passing an exam by only studying their classmates' notes, never the textbook.

Jane: Exactly. And there's another improvement they made. In the basic FedSlate, the reward is sparse. It only comes from one platform. But what if the rewards are too sparse, and the algorithm struggles to learn? They created an extended version where both platforms have access to their own rewards, and they use that to train more effectively.

Tom: So it's a more flexible framework. You can plug in different configurations depending on what data each platform has.

Jane: Right. And they compared this extended version against a baseline called FedQ. FedQ is a simpler federated approach, and it struggled in this long-term scenario. It kept recommending the "chocolate" items for short-term gain, which hurt the user's long-term satisfaction. FedSlate, on the other hand, learned to balance the two and ended up with a much better long-term outcome.

Tom: So it's not just about sharing data; it's about sharing the right kind of information to optimize for the long haul. That's a real step forward. But I'm curious about the practical side. What does this mean for a real engineer trying to build this? Let's bring in Meng for that.

Meng: Hey Tom, Jane. I've been listening, and the results are compelling. But the paper also mentions communication costs. Every time a platform sends its Q-values to the server, that's bandwidth. In a real system with millions of users, that could get expensive fast.

Jane: That's a really good point, Meng. The paper does acknowledge that as a limitation. The benefits of federated learning come with the overhead of constant communication. It's a trade-off that any real-world deployment would have to carefully measure.

Meng: And the gains are uneven, too. The blind platforms get a huge boost, but the platform that provides the feedback sees a slight dip in its own peak performance. That could be a hard sell for a company that's already doing well.

Tom: So it's a classic "rising tide lifts all boats" scenario, but some boats get lifted a lot more than others. That's a conversation for the business folks, not just the engineers. But the potential here is undeniable. Let's get Lu's take on the bigger picture.

Conclusion: Tom: We're wrapping up our discussion on "FedSlate: A Federated Deep Reinforcement Learning Recommender System." Lu, you've been quiet. What's your big-picture take on this?

Lu: I think the most exciting implication is that this could fundamentally change how we think about user modeling. Right now, every app has a siloed, incomplete picture of who we are. FedSlate offers a way to build a more holistic model of user intent and long-term satisfaction without ever centralizing the data. That's a paradigm shift.

Jane: It really is. And it's not just about better recommendations. It's about creating a system that respects privacy by design. The central server never sees raw user data, only abstracted Q-values. That's a huge step toward trustworthy AI.

Meng: And from a practical standpoint, it gives a real path forward for smaller platforms. A new app with no user data could join a federation and immediately benefit from the collective knowledge of larger partners, without having to buy or steal that data.

Tom: So it's a win for privacy, a win for new businesses, and a win for users who get better recommendations. The trade-offs around communication costs and uneven benefits are real, but the direction is clear.

Lalam: And I'd add that this has a cultural dimension. By enabling platforms to collaborate without compromising individual privacy, we're fostering an ecosystem where services can be more attuned to our genuine, long-term interests rather than just chasing immediate clicks. It moves us from a culture of short-term engagement to one of sustained, meaningful interaction.

Tom: That's a beautiful way to put it, Lalam. So, to sum up: "FedSlate: A Federated Deep Reinforcement Learning Recommender System" shows us a way to train recommendation agents across platforms, sharing knowledge without sharing data. It helps platforms with no feedback learn, and it speeds up learning for everyone else, all while keeping user privacy intact.

Jane: And it does this by cleverly combining reinforcement learning for long-term value with federated learning for privacy. It's a smart, practical solution to a very modern problem.

Tom: Well said. We've covered the title, the method, the results, and the implications. We'll be back soon with another paper, but for now, thanks for listening, and keep your recommendations thoughtful.

Jane: And your data private. See you next time!

Yongxin Deng, Xihe Qiu, Xiaoyu Tan, Yaochu Jin

Shanghai University of Engineering Science · INFLY TECH (Shanghai) Co., Ltd. · Westlake University

cs.IR, cs.AI

Submitted: 2025-04-28

Updated: 2026-08-12

Journal ref: IEEE Transactions on Emerging Topics in Computational Intelligence, vol. 9, no. 6, pp. 4202-4216, (2025)

DOI: 10.1109/TETCI.2025.3573250

Code: https://github.com/TianYaDY/FedSlate

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 49/100

Terminology

Summary

Summary

This paper introduces FedSlate, a federated deep reinforcement learning recommendation algorithm designed to optimize long-term user engagement across multiple platforms while addressing privacy and legal constraints. The authors state: We propose FedSlate, a federated reinforcement learning recommendation algorithm that effectively utilizes information that is prohibited from being shared at a legal level. The core motivation is that existing reinforcement learning-based recommendation systems do not fully exploit the relevance of individual user behavior across different platforms, and while aggregating data centrally is a potential solution, this approach raises economic and legal concerns, including increased communication costs and potential threats to user privacy.

The paper extends the application scope of recommendation systems from single-user single-platform to single-user multi-platform by introducing federated learning. The algorithm employs the SlateQ algorithm to assist FedSlate in learning users’ long-term behavior and evaluating the value of recommended content. The authors use RecSim to construct a simulation environment for evaluation and compare performance with state-of-the-art benchmark models.

The problem is formalized as an augmented Markov Decision Process (MDP) for the single-user, multi-platform context, where platforms employ recommendation systems to curate slates of content. The model includes discrete recommendation agents for each platform (agent α for platform A and agent β for platform B), with corresponding user states, actions, transition probabilities, rewards, and Q-functions. Four key assumptions are made: A1 states A user selects only one item at a time (or may choose not to select, represented as ⊥ for a null item); A2 states Transitions depend solely on the selection and while a user engages with a platform, the states of other platforms remain frozen; A3 states There is interconnectedness between the user’s behaviors on different platforms, and the impact of a single platform on the user is 'cross-platform'; A4 states Only the output values of Qα and Qβ are shared for learning the joint policy πf∗ed. Other information, including transitions Dα = ⟨sα, Aα, s′α, rα ⟩ and Dβ = ⟨sβ, Aβ ⟩, is locally visible only.

The FedSlate algorithm incorporates both local and global models. Unlike adaptive personalized federated learning (APFL), FedSlate does not blend local and global models proportionally but instead adopts a Q-value sharing approach where the local models generate Q-values that are passed as inputs to the central server. The central server then computes the global Q-values used for content recommendation selection. The algorithm is divided into stages: first, dedicated agents on each platform calculate local Q-values based on observed user states and candidate recommendation content states. They transmit these values to the central server. Second, the central server collects the received local Q-values as inputs to the global Q-network and calculates the corresponding global Q-values for each local agent. Third, the central server distributes the Q-values to the respective agents, and the local agents make policy selections based on the received Q-values. Each agent is unaware of the Q-network parameters of other agents.

The architecture includes basic Q networks (Qα and Qβ) for each agent, which output tensors of the same size as the number of candidate documents. A federated agent (agent f ed) receives these outputs and uses a multi-layer perceptron (MLP) module, denoted as Qf, to derive Q-values for slate selection. The federated Q-values are defined as: Qfα (·; θα, θβ, θf) = M LP ([Qα (sα; θα)Qβ (sβ; θβ); θf) and Qfβ (·; θα, θβ, θf) = M LP ([Qβ (sβ; θβ)Qα (sα; θα); θf), where θf represents the parameters of the MLP and [··] denotes concatenation. During training, the parameters of the Q network for agents on platforms where the user is not present are fixed to ensure stability.

The learning process uses Huber loss functions for agents α and β, with the loss for agent α defined as Lα (θα, θf) = (1/B) Σ L(Yα − Qα f (·, Cβ; θα, θf)) and for agent β as Lβ (θβ, θf) = (1/B) Σ L(Yα − Qβ f (·, Cα; θβ, θf)), where B is the batch size and Yα = r + γ max A′α ∈Aα Σ j∈A′α P (js′α, A′α)Qα f (s′α, j) is the target Q-value. The updates of both Qα and Qβ depend on rα since agent β has no access to rβ.

The algorithm is divided into acting and learning phases. In the acting phase, agent f ed initiates inquiry requests to agents α and β, which calculate Qα and Qβ based on their current states and send them to agent f ed. Agent f ed computes Qfα and Qfβ and sends them back, and the agents construct slates using a greedy method based on the received Q-values. In the learning phase, agents α and β are assigned random indices IDs of size B, calculate Qα, Q′α, and Qβ, and transmit these values to agent f ed for computation of Qα f and Q′α f. Agent f ed returns these to agent α for derivation of Yα and updating networks Qα and Qf. After updates, Yα is conveyed to agent β, and agent f ed calculates Qfβ using updated Qα and initial Qβ, dispatching it to agent β, which updates networks Qβ and Qf. The learning protocol mandates a single update for each local network and two updates for the global network.

The paper also presents an extended version of the algorithm to address sparse team rewards, which duplicates Algorithm 1 to replace Algorithm 2 and makes minor modifications to Algorithm 3, requiring platform B to have access to rewards.

Experiments were conducted using RecSim with a Choc vs. Kale recommendation scenario, where chocolate represents interesting but not conducive content for long-term satisfaction, and kale represents less exciting but beneficial content. The goal is to maximize long-term user satisfaction by balancing these elements. The user model includes features of net kale exposure (nket) and satisfaction (satt), related through a sigmoid function: satt = σ(τ · nket), where τ is a user-specific sensitivity parameter. Users select items based on the Kaleness scale, with probability p ∼ e1−kaleness(i). The net kale exposure evolves as nket+1 = β · nket + 2(ki − 1/2) + N (0, η), where β is a user-specific memory discount, ki is the kaleness of the selected item, and η is noise standard deviation. User engagement si follows a log-normal distribution: si ∼ logN (ki µk + (1 − ki)µc, ki σk + (1 − ki)σc).

The evaluation environment includes two document models representing Platform A and Platform B, integrated with a user model. Platform A has access to user feedback and includes the user's previous 5 engagements in the state, resulting in a tensor of size [1 + 5 × n + N], where n is the slate size. Platform B lacks access to user feedback and has a state tensor of size [1 + N]. Actions for both agents involve recommending a slate of content, represented as an integer tensor of size [n]. Rewards are cumulative user engagement, with the environment providing engages on both platforms, but agent β cannot access engages on Platform B.

Two evaluation criteria are defined. For agent α, Episodes to Reach Optimal Reward (ETROR) measures learning time consumption, comparing the number of episodes for the baseline method (SlateQ) versus FedSlate. For agent β, the optimal reward achieved with FL is compared against the average rewards obtained by a random recommendation method.

Experimental results show that FedSlate enhances training velocity for agent α, with M′2 consistently less than or equal to M′1 across different settings. For example, with N=100, n=10, SlateQ's ETROR is 3400 while FedSlate's is 2340; with N=500, n=10, SlateQ's ETROR is 2020 while FedSlate's is 1700. However, FedSlate may compromise the agent's learning of an optimal local policy by potentially reducing the optimal reward, attributed to its design focus on optimizing the aggregate long-term benefit for users across platforms rather than maximizing the immediate value for individual users on a single platform. In the N=10, n=3 scenario, neither algorithm converged, performing worse than random recommendations due to overfitting in an overly simplistic environment.

For agent β, FedSlate consistently outperforms random recommendations in various settings, with optimal rewards of 1115.281, 1102.269, and 1129.96 for N=10/n=3, N=100/n=10, and N=500/n=10 respectively, compared to mean rewards of 948.294, 944.982, and 939.979 for random recommendations. This demonstrates that FedSlate enables entities without feedback data to benefit from user feedback.

An ablation experiment was conducted to investigate whether the global network tends to discard Q-values that do not originate from itself. The global network was reduced to a single hidden layer network without activation functions, transforming it into a simple linear model. The results show that the ablated algorithm's rewards diverge completely throughout training, with no significant improvement compared to random recommendation algorithms, demonstrating the effectiveness of the FedSlate framework in terms of vertical federated learning.

The extended FedSlate algorithm was evaluated in sparse team reward scenarios. In environments with sparse team rewards, FedSlate agents required extended learning periods and showed diminished policy effectiveness. The extended algorithm demonstrated superior performance in both ETROR and Optimal Reward metrics, with agent β showing particularly marked improvements. For example, in the sparse environment, FedSlate-Alpha's ETROR was 3540 while FedSlate(exp)-Alpha's was 2280, and FedSlate-Beta's ETROR was 2940 while FedSlate(exp)-Beta's was 780. The extended algorithm also achieved higher optimal rewards (1091.515 for Alpha and 1108.855 for Beta) compared to the original (1081.422 and 1042.926 respectively). The FedQ algorithm, which lacks cross-platform LTV consideration, performed barely above random recommendation benchmarks and exhibited significant instability.

Extended experiments across multiple platforms were conducted under two configurations. In the first configuration with five platforms all having access to user feedback, all platforms demonstrated accelerated and more robust learning of recommendation strategies, with consistently superior recommendation performance compared to baseline methods. In the second configuration with two platforms lacking feedback access, those platforms achieved performance comparable to the original SlateQ algorithm through indirect benefits from the other platforms, demonstrating the framework's scalability and sustained effectiveness with increased federation size.

The paper concludes that FedSlate effectively resolves the challenges of cross-platform learning in recommendation systems without necessitating the exchange of private data between platforms. The main contributions are summarized as: (1) proposing FedSlate, a novel algorithm integrating SlateQ methodology with federated learning principles; (2) designing an innovative update mechanism with sequential updating protocol involving multiple updates of the global network while maintaining single updates for local networks; (3) demonstrating FedSlate's effectiveness in developing recommendation strategies using data not directly accessible to local agents.

The limitations acknowledged include higher communication costs introduced by FL, which may negate some performance enhancements, and uneven distribution of gains among participants, with entities lacking access to user feedback benefiting disproportionately. Future research will investigate cost-benefit trade-off thresholds, promote fairness in the FL process, and incorporate advanced privacy-preserving technologies such as secure aggregation, differential privacy, and encrypted data.

Improvements for AI systems

Based on the FedSlate paper, here are the specific improvements I can make to AI systems and the resulting capabilities:

  • Improvement: Implement a federated Q-value sharing architecture where local agents compute item-level Q-values from their own user states and transmit only these values (not raw user data) to a central server. The server uses a global MLP network to compute fused Q-values for each agent.

  • Capability: The system can learn recommendation policies that account for user behavior correlations across multiple platforms (e.g., a user's activity on a social network influencing their e-commerce purchases) without violating GDPR/CCPA/PDPA regulations, as no raw user data leaves individual platforms.

  • Improvement: Implement the paper's update mechanism where the global federated Q-network is updated twice per training iteration, while each local Q-network is updated only once. The local network for the platform without feedback is updated using target Q-values derived from the feedback-available platform's rewards.

  • Capability: The system achieves faster convergence (e.g., 2340 episodes vs. 3400 for baseline SlateQ in N=100, n=10 settings) and can train platforms that have no direct access to user feedback (e.g., a platform that cannot track user engagement) by leveraging rewards from other platforms in the federation.

  • Improvement: Apply SlateQ's decomposition formula (Q(s,A) = Σ P(is,A) Q(s,i)) within the federated framework, using greedy optimization for slate construction that accounts for item position effects (Eq. 8).

  • Capability: The system can handle recommendation slates with large candidate sets (e.g., 500 items) and slate sizes (e.g., 10 items) without exponential action-space explosion, making it computationally tractable for real-world e-commerce or content platforms.

  • Improvement: Implement the extended FedSlate variant (Algorithm 4) that duplicates the feedback-available agent's update process for the feedback-deprived agent, allowing both agents to use their own rewards when team rewards are too sparse.

  • Capability: In scenarios where user behavior on one platform does not strongly influence another (sparse cross-platform correlation), the system still learns effective policies. The extended version achieves ETROR of 780 episodes for the feedback-deprived agent vs. 2940 for the basic version, and improves optimal reward from 1042.93 to 1108.86.

  • Improvement: Use a multi-layer perceptron (MLP) with nonlinear activation (e.g., Mish) for the global network, rather than a linear model, to ensure the global network genuinely fuses local Q-values rather than merely passing through one agent's values.

  • Capability: The system avoids the failure mode where the global network degenerates into a single-platform policy. In ablation tests, a linear global network caused complete divergence (no ETROR achieved), while the full MLP-based FedSlate converged reliably, demonstrating robust cross-platform learning.

  1. For feedback-available platforms: 31% faster training convergence (ETROR reduced from 3400 to 2340 episodes) while maintaining near-optimal reward (1106.57 vs. 1115.07 baseline, a 0.8% trade-off for cross-platform optimization).

  2. For feedback-deprived platforms: The system achieves optimal rewards of 1102.27-1129.96 (vs. random recommendation means of 939.98-948.29), representing a 16-20% improvement, enabling these platforms to provide personalized recommendations they previously could not.

  3. For multi-platform scalability: With 5 platforms, the system shows that feedback-deprived platforms achieve performance parity with the original SlateQ algorithm (which requires full feedback), and all platforms benefit from accelerated learning and higher mean rewards (e.g., 1177.76-1196.26 vs. baseline 1155.97-1185.04).

  4. For privacy-sensitive deployments: The system operates with a central server that only sees Q-values (tensors of size [batch, candidate items]), never raw user states, transitions, or rewards, making it compliant with strict data protection regulations while still enabling collaborative learning.

Sources

Related papers