FedCGR: Federated Cross-Domain Generative Recommendation

arXiv:2608.10929 · cs.AI · Submitted 2026-08-11 · Read on arXiv

Zhuodong Liu, Hugen Lv, Xiangyu Li, Bohan Guo, Peiyu Hu

Beijing Jiaotong University · Shanghai Jiao Tong University · University of Malaya · Xi'an Jiaotong-Liverpool University

cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: Accepted at CIKM 2026. 10 pages, 5 figures, 6 tables

Code: https://github.com/ZhuodLiu/FedCGR

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: FedCGR: Federated Cross-Domain Generative Recommendation proposes a federated cross-domain recommendation (CDR) framework that formulates federated CDR "as generation over a stable semantic item

Terminology

Summary

FedCGR: Federated Cross-Domain Generative Recommendation proposes a federated cross-domain recommendation (CDR) framework that formulates federated CDR as generation over a stable semantic item language. The paper addresses the fundamental tension in federated CDR: "the behavioral anchors that make cross-domain item spaces alignable—overlapping users, co-occurrence graphs, or shared interaction signals—are precisely the signals that become sparse, unavailable, or privacy-sensitive under federated isolation."

The method represents items as discrete semantic ID (SID) sequences derived from public item-side metadata, which creates a shared item language: cross-domain item alignment becomes a property of the shared discrete vocabulary rather than an additional objective built on private behavioral data. Directly federating SID-based generators introduces two design constraints: first, the SID tokenizer must remain fixed to preserve cross-client token consistency, which creates a semantic-only bottleneck because local collaborative filtering (CF) signals cannot be globally shared or aligned; second, standard federated averaging can cause negative transfer under domain heterogeneity.

To overcome these constraints, FedCGR keeps the item language stable and makes adaptation explicit through two coupled mechanisms. First, a reliability-aware semantic interface injects local CF evidence: each interacted item is represented by its SID token embedding plus a residual CF signal whose strength is controlled by item-level reliability and a client-local gate. The item-level reliability score is computed from local interaction frequency as rho v,i = log(1 + n v,i) / (epsilon + maxv ′ ∈ Vi log(1 + n v ′,i)). The input representation is constructed as ht = ∑Ll=1 Elsid(c lvt) + Epos(t) + rho vt,i · gi · Ai(ecf vt,i), where Elsid is the trainable SID token embedding table at level l, Epos is the positional embedding, and gi = sigma(alphai) is a client-local gate. A training-only dense regularization uses a confidence-weighted local InfoNCE loss to encourage the shared encoder to preserve reliable local collaborative evidence.

Second, a prototype-personalized generator trains shared generator parameters [that] are aggregated more strongly from behaviorally related domains, while domain embeddings, gates, private experts, and local CF statistics stay on each client. The generator uses a shared-private expert block based on Mixture-of-Experts architecture with Ns shared experts and one client-local private expert. Each client computes a domain prototype from encoder memories, updated with exponential moving average and normalized. The server then constructs a personalized aggregate for each target domain using weights omegaij(t+1) = n j exp(cos(pi(t+1), p j(t+1))/taua) / ∑k=1K nk exp(cos(pi(t+1), pk(t+1))/taua). The local training objective includes Li = Lsid + lambdad Ldense + lambdaa Laux + (mu/2)‖Wi,sh − W¯ i,sh‖22 with a proximal term applied only to shared parameters.

Experiments on six Amazon cross-domain scenarios (FK, GB, GS, FKB, GBS, GKBS) show that FedCGR consistently outperforms federated generative baselines and achieves competitive performance against strong sequential and federated CDR methods under both full-ranking and sampled evaluation protocols. Under full-ranking evaluation, FedCGR is the strongest federated generative method across all 22 domain-metric cells, with the most pronounced improvements on the most heterogeneous scenario GKBS. Under the 999-neg sampled protocol, FedCGR achieves the best results across all six scenarios on both metrics, improving H@10 from 0.180 to 0.186 (+3.3%) on GKBS and from 0.170 to 0.208 (+22.4%) on GS compared to FedDCSR.

Ablation studies reveal a cross-over pattern in the division of labor between the two design principles: "On the high-affinity FK scenario, removing local CF evidence (−12.5%) is more harmful than replacing personalized aggregation with FedAvg (−8.3%)... On the heterogeneous GKBS scenario, the pattern reverses: FedAvg (−14.1%) is more harmful than CF removal (−13.0%). Cold-start analysis shows the stable SID vocabulary provides the main cold-start gain; the reliability-aware CF residual adds an increasing benefit as user activity grows. Sensitivity analysis finds a clear optimum around 0.10 for both the dense loss weight lambdad and aggregation temperature taua, while performance is relatively stable across n shared ∈ 2, 3, 4 and L ∈ 3, 4. Diagnostic statistics confirm that domains with higher average interactions per item... learn higher gate values and item confidence, while sparser domains such as Sports... learn lower gates, relying more on the SID-only representation."

Efficiency analysis shows that despite an 18% higher per-round communication (5.1M shared parameters vs. 4.3M), faster convergence reduces the total budget by 34–45%: on GBS, FedCGR reaches the 0.09 N@10 target with 288 MB cumulative upload versus 435 MB for TIGER+FedAvg and 365 MB for TIGER+FedProx; on GKBS, FedCGR crosses the 0.07 target at 288 MB versus 522 MB and 365 MB respectively.

Improvements for AI systems

Improvements to AI Systems Based on FedCGR:

  1. Cross-Domain Recommendation with Privacy-Preserving Alignment
  • The improved system can recommend items across disparate domains (e.g., books → movies, groceries → sports) without sharing raw user behavior or interaction graphs. It uses a fixed, public semantic ID (SID) vocabulary derived from item metadata (e.g., titles, categories) to create a shared item language, enabling zero-shot alignment across domains even when users do not overlap.

  • This eliminates the need for co-occurrence graphs or shared user identifiers, which are typically sparse or privacy-violating in federated settings.

  1. Adaptive Fusion of Local Collaborative Signals with Global Semantics
  • The system dynamically balances global semantic knowledge (from SID sequences) with local collaborative filtering (CF) evidence per client. It computes an item-level reliability score (based on local interaction frequency) and a client-local gate to control how much CF residual is injected into the item representation.

  • This allows the model to rely more on semantic-only representations for sparse domains (e.g., Sports) and more on CF signals for dense domains (e.g., Fashion), improving recommendation accuracy without sharing raw interaction data.

  1. Personalized Federated Aggregation via Domain Prototypes
  • Instead of naive FedAvg, the system aggregates shared generator parameters using a weighted scheme based on cosine similarity between domain prototypes (computed from local encoder memories). This prevents negative transfer when domains are heterogeneous (e.g., Groceries vs. Kindle Store).

  • The improved system can automatically identify behaviorally related domains and give them higher aggregation weights, while keeping domain-specific embeddings, gates, and private experts local—preserving personalization and reducing communication overhead.

  1. Robust Cold-Start Handling via Stable Semantic Vocabulary
  • For new users or items with few interactions, the system leverages the fixed SID tokenizer to generate recommendations purely from item metadata, providing strong cold-start performance. As user activity grows, the reliability-aware CF residual incrementally boosts accuracy.

  • This enables the system to serve new domains or items immediately, without waiting for collaborative data to accumulate.

  1. Efficient Federated Training with Reduced Total Communication
  • Despite a slightly higher per-round communication cost (due to shared expert parameters), the system converges faster (34–45% less total upload budget) because personalized aggregation reduces oscillation and speeds up optimization.

  • This makes the system practical for bandwidth-constrained federated deployments, achieving target accuracy with significantly lower cumulative data transfer (e.g., 288 MB vs. 435 MB for baselines).

  1. Controllable Trade-off Between Global and Local Knowledge
  • The system exposes hyperparameters (e.g., dense loss weight λ d, aggregation temperature τ a) with a clear optimum, allowing operators to tune the balance between preserving local CF evidence and leveraging global semantic structure.

  • It also supports a proximal term on shared parameters to prevent catastrophic forgetting during local updates, making training stable across heterogeneous clients.

  1. Scalable Mixture-of-Experts Architecture for Domain Heterogeneity
  • The generator uses shared experts (trained globally) plus a client-local private expert, enabling the model to capture both universal patterns and domain-specific quirks. This improves performance on diverse domains without increasing server-side model size proportionally.

  • The system can be extended to new domains by simply adding a local private expert and updating the domain prototype, without retraining the entire shared architecture.

What the improved AI system can do:

  • Deploy federated recommendation across multiple e-commerce categories (e.g., books, movies, groceries) with no shared user data, achieving state-of-the-art accuracy under strict privacy constraints.

  • Provide instant, high-quality recommendations for cold-start items and users using only public metadata, then seamlessly refine them as local interaction data accumulates.

  • Operate efficiently on edge devices with limited bandwidth, reducing total communication cost by up to 45% while maintaining or exceeding baseline accuracy.

  • Automatically adapt to domain heterogeneity—e.g., prioritizing semantic alignment for sparse domains and CF signals for dense ones—without manual tuning per domain.

  • Support continuous learning in federated settings where domains evolve or new clients join, thanks to prototype-based aggregation and fixed SID vocabulary.

Abstract

Cross-domain recommendation (CDR) transfers preference knowledge across related domains, but federated deployment makes cross-domain alignment difficult because the behavioral anchors that align item spaces, such as overlapping users and shared interaction signals, are often sparse, unavailable, or privacy-sensitive across clients. To address this tension, we revisit federated CDR as generation over a stable semantic item language. By representing items as discrete semantic ID (SID) sequences derived from public item-side metadata, cross-domain item alignment is induced by a shared vocabulary rather than by exchanging private interactions or aligning domain-specific embeddings. Directly federating SID-based generators, however, introduces two design constraints: the SID tokenizer must remain fixed to preserve cross-client token consistency, which creates a semantic-only bottleneck because local collaborative filtering (CF) signals cannot be globally shared or aligned; meanwhile, standard federated averaging can cause negative transfer under domain heterogeneity. To overcome these constraints, we propose FedCGR, a federated generative CDR framework that keeps the item language stable and makes adaptation explicit. FedCGR injects local CF evidence through a reliability-aware semantic interface and trains a prototype-personalized generator that selectively aggregates shared parameters according to domain relatedness while keeping domain-specific quantities local. Experiments on six Amazon cross-domain scenarios show that FedCGR consistently outperforms federated generative baselines and achieves competitive performance against strong sequential and federated CDR methods under both full-ranking and sampled evaluation protocols.

Sources

Related papers