FSGR: Mitigating Token Frequency Bias for Fair SID-Based Generative Recommendation

arXiv:2608.12845 · cs.IR, cs.AI, cs.LG · Submitted 2026-08-13 · Read on arXiv

Yuchen Zheng, Sihan Xu, Jingwen Yang, Xiangrui Cai, Haiwei Zhang, Xiaojie Yuan

Nankai University

cs.IR, cs.AI, cs.LG

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 75/100

The gist: FSGR: Mitigating Token Frequency Bias for Fair SID-Based Generative Recommendation proposes a fairness optimization framework to address a previously overlooked fairness issue in Semantic ID

Terminology

Summary

FSGR: Mitigating Token Frequency Bias for Fair SID-Based Generative Recommendation proposes a fairness optimization framework to address a previously overlooked fairness issue in Semantic ID (SID)-based generative recommendation, termed Token Frequency Bias, where high-frequency SID tokens are systematically over-predicted while low-frequency SID tokens are under-predicted. This bias "originates from the combined effects of imbalanced semantic codebooks during SID construction, and popularity bias together with the maximum likelihood estimation objective during recommendation training, resulting in unfair exposure across item categories."

The paper identifies two complementary sources of this bias: "The first arises during SID construction, where the semantic codebook is inherently imbalanced. Tokens representing broad semantic categories are repeatedly assigned to a large number of items, resulting in higher usage frequencies than other tokens, whereas tokens corresponding to less populated categories are rarely utilized. The second arises during recommendation training: On the one hand, popularity bias in the training data causes SIDs of popular items to appear much more frequently than those of long-tail items. On the other hand, the maximum likelihood estimation (MLE) objective amplifies high-frequency training signals, assigning higher prediction probabilities to frequent SID tokens while suppressing infrequent tokens."

To address this, FSGR consists of two main components. First, during SID construction, it introduces Balanced Semantic Quantization (BSQ) which includes OT-based Assignment Optimization (OTA) and Dual-Criteria Re-anchor (DCR) mechanism. OTA models the quantization process within each mini-batch as an Optimal Transport (OT) problem, aiming to minimize transport cost while promoting uniform marginal distributions for codebook utilization, using the Sinkhorn-Knopp algorithm and minimizing KL divergence between the soft assignment and the optimal transport plan. DCR periodically identifies dead codewords and re-initializes them to either under-covered or high-density regions according to two complementary criteria: Strategy A (OT-Cost Void Detection Re-anchor) fills geometric voids in the latent feature space by re-anchoring dead codes to samples with high transport costs, and Strategy B (Density Aware Demand Re-anchor) focuses on overcrowded regions by re-anchoring dead codes to representative instances from densely populated semantic regions.

Second, during recommendation training, FSGR adopts a two-stage training strategy. The first stage is Semantic Alignment Pre-training, where the model is optimized via standard cross-entropy to learn semantic correspondence between user behaviors and SID sequences. The second stage is Hierarchical Frequency Calibration (HFC), which computes the logarithmic frequency prior bl = −log(fl + ϵ) and calibrates logits as ẑl = zl + τl bl, where τl is the layer-specific calibration temperature. HFC uses a layer-aware temperature schedule where the calibration temperature increases progressively across SID layers according to τl = l/L, so deeper layers receive stronger frequency calibration, while lower layers remain relatively stable, aligning with the hierarchical semantics of SID tokens.

Experiments were conducted on three Amazon datasets (Luxury Beauty, Industrial and Scientific, Software) with three backbone models (TIGER, Llama3.1-8B, Qwen3-8B). Results show that FSGR mitigates token frequency bias and delivers an average Gini fairness improvement of over 20% while maintaining competitive recommendation accuracy. The ablation study confirms that BSQ and HFC complement each other in mitigating token frequency bias, and removing either module degrades fairness. Codebook utilization analysis shows FSGR achieves the highest Coverage and the lowest Gini on all datasets, with Coverage approaching 100% on Industrial. The layer-wise temperature analysis demonstrates that the proposed HFC achieves the best overall trade-off between recommendation accuracy and fairness across all three datasets compared to reverse or fixed temperature assignments.

Improvements for AI systems

Improvements to AI Systems:

  1. Fairness-Aware Generative Recommenders: Integrate FSGR’s two-stage training (semantic alignment pre-training + hierarchical frequency calibration) into any SID-based generative recommendation model. The improved system will reduce systematic over-prediction of popular or broad-category items, ensuring long-tail and niche items receive fair exposure without sacrificing ranking accuracy.

  2. Balanced Semantic Codebook Construction: Apply Balanced Semantic Quantization (BSQ) with OT-based assignment optimization and dual-criteria re-anchoring to any vector-quantized embedding pipeline. The improved system will eliminate dead or under-utilized codewords, achieve near-100% codebook coverage, and prevent geometric or density-based gaps in latent space, leading to more expressive and unbiased item representations.

  3. Hierarchical Logit Calibration for Sequential Models: Implement the layer-aware temperature schedule (τl = l/L) with logarithmic frequency priors for any hierarchical token prediction task (e.g., product IDs, category trees, or structured sequences). The improved system will apply stronger calibration at deeper semantic layers, correcting frequency bias where it matters most while preserving low-level token stability, thus improving fairness in multi-level generative outputs.

  4. Optimal Transport for Mini-Batch Quantization: Replace heuristic codebook assignment with Sinkhorn-Knopp-based optimal transport in any self-supervised or contrastive learning framework that uses discrete latent codes. The improved system will promote uniform code utilization, reduce mode collapse, and enhance representation diversity across batches, benefiting downstream tasks like clustering, retrieval, or generation.

  5. Bias-Aware Pre-training + Fine-tuning Pipeline: Adopt the two-stage strategy (standard cross-entropy pre-training, then frequency-calibrated fine-tuning) for any generative model trained on imbalanced sequential data. The improved system will first learn semantic structure, then explicitly counteract token frequency bias during fine-tuning, leading to more robust and equitable predictions in domains like e-commerce, news, or social media recommendation.

  6. Dead-Code Re-initialization Heuristics: Use the dual-criteria re-anchor mechanism (OT-cost void detection + density-aware demand) to periodically refresh unused or collapsed units in any neural codebook or embedding layer. The improved system will dynamically adapt to shifting data distributions, preventing permanent loss of representational capacity and improving long-term model stability and fairness.

Abstract

Semantic ID (SID)-based generative recommendation has recently achieved remarkable success. However, existing methods suffer from a previously overlooked fairness issue, which we term Token Frequency Bias, where high-frequency SID tokens are systematically over-predicted while low-frequency SID tokens are under-predicted. This bias originates from the combined effects of imbalanced semantic codebooks during SID construction, and popularity bias together with the maximum likelihood estimation objective during recommendation training, resulting in unfair exposure across item categories. Existing SID methods mainly focus on improving codebook quality and overlook the impact of token frequency imbalance on downstream recommendation fairness, while LLM debiasing methods often yield suboptimal results when directly applied to SID-based recommendation, due to the hierarchical semantics of SID tokens. To address this issue, we propose FSGR, a fairness optimization framework for SID-based generative recommendation. During SID construction, FSGR employs OT-based Assignment Optimization and Dual-Criteria Re-anchor mechanism to form a more balanced SID representation space. During recommendation training, it adopts a two-stage training strategy and introduces Hierarchical Frequency Calibration for layer-specific fairness fine-tuning. Experiments on three public datasets with three backbone models demonstrate that FSGR mitigates token frequency bias and delivers an average Gini fairness improvement of over 20% while maintaining competitive recommendation accuracy.

Sources

Related papers