Triple Expert Learning from Noisy Labels for Semi-Supervised Vision Foundation Model Adaptation
Xuanyu Liu, Zheng Fang, Hongyang He, Yundi Hong, Daizong Liu
University of Warwick · Manifolda.AI · Wuhan University
cs.CV, cs.AI
Submitted: 2026-08-10
Updated: 2026-08-11
Comments: Accepted for publication at the British Machine Vision Conference (BMVC) 2026. Official list of accepted papers: https://bmvc2026.bmva.org/programme/accepted_papers/
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: TriNoL, a Triple-expert learning framework from Noisy Labels for semi-supervised vision foundation model (VFM) adaptation, addresses the problem that "pseudo-labels have mixed reliability, and a
Terminology
Summary
TriNoL, a Triple-expert learning framework from Noisy Labels for semi-supervised vision foundation model (VFM) adaptation, addresses the problem that pseudo-labels have mixed reliability, and a single LoRA adapter must absorb reliable, ambiguous, and noisy gradients in the same low-rank space,
which can make VFM adaptation sensitive to pseudo-label noise.
The paper identifies that pseudo-label reliability should control the adaptation pathway, not only the loss weight.
TriNoL "routes unlabeled samples into three confidence regions and assigns them to three LoRA experts: a Positive Expert for high-confidence pseudo-labels, an Alignment Expert for medium-confidence ambiguous samples, and a Negative Expert for low-confidence noisy samples. The VFM backbone remains frozen, and only the LoRA experts and classifier head are updated.
By separating different pseudo-label reliability regions into specialized adaptation paths, TriNoL improves robustness to noisy supervision while keeping the training cost low."
The method decomposes pseudo-label supervision into three confidence regions using two thresholds, τ+ and τ−, where "High-confidence samples train the Positive Expert with hard pseudo-label supervision. Medium-confidence samples train the Alignment Expert with softer pseudo-label alignment. Low-confidence samples train the Negative Expert with a noise-aware objective that suppresses over-commitment to unreliable pseudo-labels." The Positive Expert uses cross-entropy on hard pseudo-labels, the Alignment Expert uses KL divergence to match strong-view predictions to detached weak-view probabilities, and the Negative Expert uses a loss that discourages assigning high probability to the current top pseudo-label: Lneg = −(1/Bu−) Σ log(1 − p−i,ŷi + ε). The final objective combines supervised loss with weighted expert losses: L = Lsup + λpos Lpos + λalign Lalign + λneg Lneg.
Theoretical analysis shows that a single LoRA adapter directly absorbs ambiguous and noisy pseudo-label gradients,
leading to gradient deviation, while TriNoL reduces the contamination entering the main reliable adaptation path
by routing different reliability regions into different experts. The paper proves that the Positive Expert learns a low-rank subspace closer to the reliable target subspace than a single adapter, and that the Alignment and Negative Experts preserve uncertainty and suppress unreliable top-label commitment, respectively. It also shows that reducing pseudo-label bias improves local optimization stability.
Experimental results on CIFAR-100, FOOD-101, Semi-Aves, and ImageNet show that TriNoL achieves the best results on CIFAR-100 N4, FOOD-101 N2/N4/N10, Semi-Aves with both in-distribution and mixed OOD unlabeled data, and ImageNet 1%.
For example, "on FOOD-101 N2, TriNoL improves over V-PET from 87.58% to 88.42%, and over FineSSL from 87.27% to 88.42%. On Semi-Aves with mixed in-distribution and OOD unlabeled data, TriNoL improves over FineSSL from 61.30% to 61.82%. Efficiency analysis shows that
TriNoL outperforms the rank-matched Single LoRA rank 24 baseline. Both methods have 0.962M trainable parameters and the same relative LoRA cost of 3.0×, but TriNoL improves the accuracy from 87.73% to 88.42%, demonstrating that
the gain comes from confidence-aware expert specialization, where reliable, ambiguous, and unreliable pseudo-labels are routed to different adaptation paths."
Ablation studies confirm that using only the Positive Expert already improves performance by separating high-confidence pseudo-labels from the mixed unlabeled supervision,
and Adding the Alignment Expert further improves the results... Adding the Negative Expert also brings gains, with clearer improvements under the mixed OOD setting of Semi-Aves.
Routing strategy ablation shows that Random routing only brings marginal improvement over Single LoRA... Confidence routing with a shared hard pseudo-label loss performs better than random routing, but it is still weaker than TriNoL,
indicating that routing samples by confidence is useful, but each confidence region also requires a suitable learning objective.
Sensitivity analysis finds that thresholds (τ−, τ+) = (0.3, 0.7) and loss weights λa = 1.0, λn = 0.1 give the best balance. Expert routing dynamics show that TriNoL does not statically split data into fixed groups. Instead, it dynamically transfers samples from uncertain or noisy regions to the reliable pseudo-label region as the model becomes more confident.
Noise robustness analysis shows that under injected pseudo-label corruption, TriNoL degrades more slowly than Single LoRA.
The paper concludes that "TriNoL treats pseudo-labels as mixed-reliability supervision and routes unlabeled samples into three LoRA experts according to confidence... This design separates different pseudo-label gradients into different adaptation paths and reduces interference in the low-rank update space. Limitations include dependence on
confidence-based routing, which may be affected by miscalibrated predictions or strong domain shift, slightly higher parameter and training cost than a single rank-8 LoRA adapter, and the work
mainly studies image classification," with extension to dense prediction, open-vocabulary recognition, and multimodal adaptation left for future work.
Improvements for AI systems
Improvements to AI systems:
-
Confidence-aware expert routing for semi-supervised learning: Implement a multi-expert architecture where unlabeled data is dynamically routed into separate adaptation modules (e.g., LoRA experts) based on pseudo-label confidence thresholds. This prevents a single low-rank adapter from absorbing conflicting gradients from reliable, ambiguous, and noisy pseudo-labels, improving robustness to label noise.
-
Specialized loss functions per confidence region: Use distinct objectives for each expert—hard cross-entropy for high-confidence samples, KL divergence for medium-confidence (soft alignment to weak-view predictions), and a noise-averse loss (e.g., suppressing top-pseudo-label probability) for low-confidence samples. This reduces over-commitment to unreliable labels and preserves uncertainty.
-
Dynamic sample re-routing over training: Allow samples to transition between experts as model confidence evolves, rather than static grouping. This enables the system to progressively move uncertain samples into the reliable region, improving adaptation stability and final accuracy.
-
Theoretical gradient-deviation reduction: By separating pseudo-label gradients into different low-rank subspaces, the system provably reduces contamination in the main adaptation path, leading to better local optimization and faster convergence under noisy supervision.
What the improved AI system can do:
-
Achieve higher accuracy in semi-supervised image classification with noisy or out-of-distribution unlabeled data (e.g., +0.84% over prior SOTA on FOOD-101 N2, +0.52% on Semi-Aves with mixed OOD).
-
Maintain robust performance under injected pseudo-label corruption, degrading more gracefully than single-adapter baselines.
-
Operate with low training cost (only LoRA experts and classifier head updated, backbone frozen) while matching or exceeding rank-matched single-adapter efficiency.
-
Dynamically adapt to changing confidence distributions, improving reliability in real-world scenarios with domain shift or miscalibrated model outputs.
Abstract
Semi-supervised adaptation of vision foundation models (VFMs) commonly freezes the pretrained backbone and updates lightweight modules such as LoRA. However, pseudo-labels have mixed reliability, and a single LoRA adapter must absorb reliable, ambiguous, and noisy gradients in the same low-rank space. This can make VFM adaptation sensitive to pseudo-label noise. We propose TriNoL, a Triple-expert learning framework from Noisy Labels for semi-supervised VFM adaptation. TriNoL routes unlabeled samples into three confidence regions and assigns them to three LoRA experts: a Positive Expert for high-confidence pseudo-labels, an Alignment Expert for medium-confidence ambiguous samples, and a Negative Expert for low-confidence noisy samples. The VFM backbone remains frozen, and only the LoRA experts and classifier head are updated. By separating different pseudo-label reliability regions into specialized adaptation paths, TriNoL improves robustness to noisy supervision while keeping the training cost low.
Sources
- SoftMatch: Addressing the Quantity-Quality Trade-off in Semi-supervised Learning
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Erasing the Bias: Fine-Tuning Foundation Models for Semi-Supervised Learning
- Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey
- Revisiting Chain-of-Thought Reasoning under Limited Supervision: Semi-supervised Chain-of-Thought Learning
- Semi-Supervised Vision-Language-Action Model
- DivideMix: Learning with Noisy Labels as Semi-supervised Learning
- Unlabeled Data vs. Pre-trained Knowledge: Rethinking SSL in the Era of Large Models
- DINOv2: Learning Robust Visual Features without Supervision
- FreeMatch: Self-adaptive Thresholding for Semi-supervised Learning
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models