Generalised Balanced Softmax: A Finite-Data Perspective on Logit Adjustment for Long-Tailed Recognition
cs.LG, cs.CV
Submitted: 2026-07-24
Updated: 2026-09-20
Comments: 26 pages, 8 figures. Revised author version of the CBS work published in Pattern Recognition. Title and terminology updated to Generalised Balanced Softmax (GBS); theoretical presentation and relationship to existing logit-adjustment methods clarified. Underlying loss function and reported experimental results unchanged
Journal ref: Pattern Recognition (2026), Article 114961
DOI: 10.1016/j.patcog.2026.114961
Code: https://github.com/YihangZhu/Class-Balanced-Softmax
License: http://creativecommons.org/licenses/by/4.0/
The gist: Models trained on long-tailed data using standard softmax tend to exhibit higher training error and a larger generalisation gap for classes with fewer training samples.
Terminology
Abstract
Models trained on long-tailed data using standard softmax tend to exhibit higher training error and a larger generalisation gap for classes with fewer training samples. We characterise this class-wise disparity as the preference issue and quantify it using a new metric, the model imbalance level I. To understand this issue, we analyse how imbalanced training data adversely affects class-wise gradients under standard softmax training. This paper then develops a finite-data Generalised Balanced Softmax (GBS) framework for analysing and mitigating the preference issue. The framework uses the training-time logit adjustment z nc+β N c, which is algebraically identical to the training-time logit-adjusted loss of Menon et al. (2021) when τ=β. The case β=1 also coincides with Balanced Softmax and with the unit adjustment supported by the Fisher-consistency argument under the true data distribution, corresponding to an idealised infinite-data setting. Building on this existing loss family, this paper uses a heuristic power-law assumption to motivate the adjustable coefficient and studies how β affects trained models. Across the evaluated long-tailed benchmarks, β=1 does not attain the highest average testing recall on most datasets, showing that a different coefficient can be preferable when training on finite data. The selected values of β reduce I and improve average testing recall relative to the β=1 reference, while retaining negligible computational overhead and compatibility with existing representation-learning frameworks.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks