CLEAR: Class-wise Expert Aggregation with Structured Sampling for Long-Tailed Classification
Gawon Lim
University of Illinois Urbana-Champaign
cs.CV, cs.AI, cs.LG
Submitted: 2026-08-11
Updated: 2026-08-13
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: CLEAR (Class-wise reLiability-aware Expert Aggregation for long-tailed Recognition) is a modular ensemble framework for long-tailed classification that addresses the challenge of uneven prediction
Terminology
Summary
CLEAR (Class-wise reLiability-aware Expert Aggregation for long-tailed Recognition) is a modular ensemble framework for long-tailed classification that addresses the challenge of uneven prediction reliability across frequent and underrepresented classes. The paper states: Long-tailed classification poses a reliability challenge because models trained on imbalanced data are unevenly reliable across frequent and underrepresented classes.
The central problem is that the central challenge in long-tailed recognition is not merely low average accuracy, but uneven prediction reliability across different regions of the label space.
The framework has two key components: "First, structured sampling generates multiple sub-training sets with different imbalance levels while preserving all classes in every subset, producing experts specialized under different class-distribution regimes. Second, class-wise trust-weighted aggregation assigns each expert a separate reliability score for each class. The trust score is
computed as a smoothed estimate of class-wise precision, measuring how reliably an expert's predictions for each class are supported by reference data, using
a Beta prior, which yields a closed-form estimate and stabilizes precision when only a few predictions are available. The final prediction is obtained
through a class-wise generalized product-of-experts aggregation, allowing each class to place greater weight on experts estimated to be more reliable for it."
The method defines class-wise reliability as θm,c = P(correct ŷm = c), representing the probability that expert m is correct when it predicts class c.
With Beta-prior smoothing, the trust score is qm,c = (α0 + nm,c) / (α0 + β0 + Nm,c), where nm,c is the number of correct predictions and Nm,c is the total number of predictions for class c. These scores are normalized across experts via softmax: wm,c = exp(τ qm,c) / Σj exp(τ qj,c), where τ is a sharpness parameter. The aggregation in log-space is Sc(x) = Σm wm,c log pm(cx), with the final prediction ŷ = arg maxc pCLEAR(cx).
For expert generation, CLEAR uses threshold-based clipping of the original training data: classes with instance counts larger than a clipping threshold are downsampled, while classes with equal or fewer samples are fully retained.
The sub-training set at stage i is Si = ∪c Sample(Dc, min(nc, Ti)), with Exponential Decay Clipping (EDC) defining Ti = Cmax δ(i−1), where Cmax is the maximum class frequency and δ ∈ (0,1) controls the decay rate. This produces a sequence of sub-training sets that gradually moves from the original long-tailed distribution toward more balanced distributions while retaining all classes at every stage.
CLEAR is modular and does not replace existing long-tailed training objectives, but can be combined with components such as Balanced Softmax, Balanced Contrastive Learning, and post-hoc Logit Adjustment.
The main contributions are: Class-wise expert reliability is identified as an underexplored principle for long-tailed ensemble learning, showing why global expert weighting is insufficient under severe class imbalance
; CLEAR is introduced as a modular framework that combines structured expert generation with class-wise reliability-aware aggregation
; and CLEAR is evaluated across CIFAR-100-LT, ImageNet-LT, and Places-LT, showing competitive overall accuracy and particularly strong few-shot performance across multiple backbone architectures.
Experimental results show that on CIFAR-100-LT, CLEAR (BSM/BCL/LA) achieves a Few-shot accuracy of 41.25%, outperforming strong recent baselines such as SADE (33.9%), ConCutMix (35.8%), and MGS (37.2%) under the standard 200-epoch training schedule.
On ImageNet-LT, CLEAR (BSM/BCL/LA) achieves an overall accuracy of 58.43%,
and CLEAR (BSM/LA) attains 43.48% accuracy on the Few-shot split, remaining competitive with SADE (43.5%).
On Places-LT, the fully integrated CLEAR (BSM/BCL/LA) model achieves 42.15% overall accuracy
and reaching 40.82% with BSM/BCL/LA
on Few-shot, which compares favorably with strong ensemble and contrastive learning methods, including SADE+RL (38.7%) and GPaCo+ConCutMix (34.9%).
Ablation studies show that most gains are obtained in the early ensemble stages
and 5–8 stages provide a reasonable trade-off between accuracy and inference cost.
Regarding the clipping strategy, EDC achieves the best overall accuracy, while uniform interval selection performs competitively with the same number of experts.
The decay rate δ controls the trade-off between expert diversity and expert reliability,
with larger values of δ generally improve both Macro Accuracy and Few-shot Accuracy.
The sharpness parameter τ acts as a stage-quality amplifier in the class-wise ensemble,
being most beneficial when the structured sampling process creates heterogeneous experts, but overly sharp weighting can be harmful when class-wise reliability estimates are noisy.
For BCL weight, a moderate BCL weight improves Few-shot Accuracy, while overly large values reduce Macro Accuracy,
with the best Few-shot Accuracy obtained around λBCL = 0.4–0.5.
For logit adjustment, small values such as αLA = 0.1–0.2 provide a conservative trade-off, while larger values are useful only when tail-class accuracy is prioritized.
The main limitation is the additional computational cost introduced by training and aggregating multiple experts.
Another limitation is that for classes with nc ≤ Tm... no separate held-out samples remain, and trust is therefore estimated from training-side (in-bag) predictions,
which may lead to optimistic trust estimates, particularly for rare classes.
Future work will explore expert pruning, early stopping of uninformative stages, or parameter-efficient adaptation methods such as LoRA,
as well as automatic expert selection, more diverse expert generation strategies, and noise-resilient class-wise trust estimation.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems:
Improvement: Replace global expert weighting with per-class trust scores computed via Beta-prior smoothed precision. Each expert gets a separate reliability weight for each class, normalized across experts via softmax with a sharpness parameter.
Improved capability: The system can now dynamically trust different experts for different classes—e.g., trusting an expert trained on balanced data for rare classes while trusting a long-tailed expert for frequent classes—leading to more accurate predictions on underrepresented categories without sacrificing performance on common ones.
Improvement: Generate multiple sub-training sets by clipping class frequencies with a geometric decay schedule (Ti = Cmax·δ(i−1)), preserving all classes at every stage while shifting from original imbalance toward balanced distributions.
Improvement: Combine CLEAR with Balanced Softmax, Balanced Contrastive Learning, and post-hoc Logit Adjustment without replacing them. Use tuned hyperparameters (e.g., λBCL ≈ 0.4–0.5, αLA ≈ 0.1–0.2) for optimal few-shot and macro accuracy.
Improvement: Aggregate expert predictions via Sc(x) = Σm wm,c·log(pm(cx)), where weights are class-specific and sharpened by τ. This allows each class to be decided by the most reliable experts for that class.
Improvement: Compute trust scores as qm,c = (α0 + nm,c)/(α0 + β0 + Nm,c), which stabilizes precision estimates when only a few predictions are available—critical for rare classes.
Improvement: Use 5–8 ensemble stages as a trade-off sweet spot, with early stages providing most gains. Optionally, implement early stopping of uninformative stages or expert pruning to reduce computational overhead.
Improvement: Detect when classes have nc ≤ Tm (no held-out samples for trust estimation) and apply correction factors or use out-of-bag estimates from other experts to reduce optimistic bias in trust scores.
Improvement: Automatically tune τ (sharpness) and δ (decay rate) based on the heterogeneity of experts and noise in reliability estimates—larger δ for diversity, moderate τ to avoid over-sharpening when estimates are noisy.
What the improved AI system can do overall:
It can perform long-tailed classification with state-of-the-art few-shot and macro accuracy (e.g., 41.25% few-shot on CIFAR-100-LT, 58.43% overall on ImageNet-LT, 42.15% overall on Places-LT), while being modular enough to integrate with existing training objectives, and computationally manageable via stage pruning and early stopping—making it suitable for production systems where rare classes matter (e.g., medical diagnosis, fraud detection, rare species identification).
Sources
- Aligned Contrastive Loss for Long-Tailed Recognition
- Enhanced Long-Tailed Recognition with Contrastive CutMix Augmentation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models