Pre-train to Gain: Robust Learning Without Clean Labels
cs.LG, cs.AI, cs.NE
Submitted: 2025-11-25
Updated: 2026-09-16
Comments: 10 pages, 8 figures
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Training deep networks with noisy labels leads to poor generalization and degraded accuracy due to overfitting to label noise.
Terminology
Abstract
Training deep networks with noisy labels leads to poor generalization and degraded accuracy due to overfitting to label noise. Existing approaches for learning with noisy labels often rely on the availability of a clean subset of data. By pre-training a feature extractor on the target dataset without labels using in-domain self-supervised learning (SSL), followed by standard supervised training on the same noisy dataset, we can train a more noise robust model without requiring a subset with clean labels. We evaluate both contrastive and non-contrastive SSL pre-training methods across datasets with synthetic and real-world label noise, demonstrating the broad applicability of our approach across large-scale datasets, diverse downstream tasks, and model architectures. Across all noise rates, in-domain self-supervised pre-training consistently improves classification accuracy and downstream label-error detection (F1 and Balanced Accuracy) compared with supervised training from scratch. The performance gap widens as the noise rate increases, demonstrating improved robustness. Notably, our approach achieves comparable results to ImageNet and DinoV2 pre-trained models at low noise levels, while substantially outperforming them under high noise conditions.
Sources
- Combating noisy labels by agreement: A joint training method with co-regularization
- Barlow Twins: Self-Supervised Learning via Redundancy Reduction
- Understanding deep learning requires rethinking generalization
- Contrast to Divide: Self-Supervised Pre-Training for Learning with Noisy Labels
- A Simple Framework for Contrastive Learning of Visual Representations
- Exploring Simple Siamese Representation Learning
- Bootstrap your own latent: A new approach to self-supervised Learning
- Co-teaching: Robust Training of Deep Neural Networks with Extremely Noisy Labels
- Deep Residual Learning for Image Recognition
- Momentum Contrast for Unsupervised Visual Representation Learning
- Adam: A Method for Stochastic Optimization
- DivideMix: Learning with Noisy Labels as Semi-supervised Learning
- Early-Learning Regularization Prevents Memorization of Noisy Labels
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks