CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers

arXiv:2608.12773 · cs.CV, cs.LG, eess.IV · Submitted 2026-08-13 · Read on arXiv

Ebenezer Tarubinga

Ebenworks Systems

cs.CV, cs.LG, eess.IV

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: Submitted to IEEE TPAMI. 22 pages, 11 figures, 17 tables

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: CW-BASS v2 is a saturation-aware pseudo-label selection method for semi-supervised semantic segmentation (SSSS) under foundation-model teachers.

Terminology

Summary

CW-BASS v2 is a saturation-aware pseudo-label selection method for semi-supervised semantic segmentation (SSSS) under foundation-model teachers. The paper's central premise is that "no single threshold rule is right across confidence regimes, the same adaptive filtering that helps an under-confident teacher hurts a saturated one, so CW-BASS v2 does not commit to one rule. It reads the teacher's confidence regime with a one-pass diagnostic and deploys whichever rule the regime calls for, strict filtering when the teacher saturates, adaptive filtering when it does not."

The method addresses a regime shift caused by self-supervised foundation encoders like DINOv2. With a DINOv2 teacher, confidence saturates, so the filtering that helped a weak teacher can hurt a strong one. The paper shows that "a ResNet-50 teacher spreads its pixel confidence across [0.3, 1], giving a threshold a wide operating range, whereas the DINOv2 teacher collapses almost all of its mass into [0.95, 1] (98% of pixels exceed 0.95, versus 53% for the ResNet)."

CW-BASS v2 has three main constructs:

  1. Held-out calibration: The labeled set is split into a training slice and a small calibration slice (α=5%). Per-class pseudo-label noise is estimated only on the held-out slice, which is proven to be unbiased (Proposition 1), unlike in-batch estimators that are downward-biased because they measure error on data the student has been trained to fit.

  2. A self-adaptive confidence floor: This floor scales with the teacher's mean confidence and is proven (Theorem 1, Corollary 1) to pin retention to a fixed quantile bounded away from 1, whereas the bare dynamic threshold provably collapses to full retention.

  3. A one-pass saturation gate: The method measures one statistic on a held-out slice, the reliability of the teacher's confident set, πkept = Pr[correct c≥τ], and deploys strict filtering when it is at least as accurate as the confidence demanded (πkept ≥τ), and its adaptive floor otherwise.

The gate criterion is "a calibration test read off the teacher itself: does the retained set match the confidence it required? When it does (a strong, well-calibrated teacher), hard-training on the confident set is already near-optimal and relaxing the cutoff only admits the error-enriched band; when it does not (a confidently-miscalibrated teacher), hard-thresholding trusts wrong labels at full weight and the floor's softened, confidence-weighted loss is preferable."

The paper reports that across six DINOv2 teachers it makes the correct strict-vs-floor call blind. On the saturated benchmarks, CW-BASS v2 "recovers the UniMatch V2 operating point on the saturated benchmarks by selecting strict (Pascal VOC 1/8 87.4 against its reported 87.9; Cityscapes within 0.5), and improves on it where the confident set is unreliable (πkept ≈89%, ADE20K), where the floor edges ahead (+1.5, single seed)."

The paper identifies two failure modes of confidence-adaptive selection that are latent at ResNet strength and dominant at foundation-model strength:

  1. The overconfidence of in-batch noise estimation: "Any method that estimates pseudo-label noise from the labeled training pixels measures error on data the student has been trained to fit. The estimate is biased toward zero and grows more so as training proceeds, dragging any noise-driven threshold downward independently of the true pseudo-label quality."

  2. The collapse of the dynamic threshold: "Under DINOv2 the teacher's confidence saturates, and we show that the CW-BASS dynamic threshold is upper-bounded by a small constant in this regime. Once nearly all confident pixels exceed it, retention drifts to 1 and training floods with whatever residual noise remains."

The measured failure chain is: "Confidence saturation (98% of pixels ≥ 0.95, ECE 0.007) ⇒ dynamic-range collapse (the dynamic cutoff pinned in [0.300, 0.331], directly beneath the analytic ceiling τ0σ(β/2) ≈ 0.34 of Corollary 1) ⇒ mask flooding (retention crosses 0.95 by epoch 7 and 0.99 by epoch 8, reaching 1.000; 99.9% of teacher errors admitted vs. 63% for strict) ⇒ early peak (best EMA at epochs 4–20 for the adaptive rules vs. epoch 42 for strict) ⇒ decline (the per-class evaluation model itself sheds 6.14 mIoU, 84.07→77.93)."

The paper emphasizes that "the gate is principled because the failure it avoids is measured, not assumed: on a reliable, saturated teacher the confidence distribution's dynamic range collapses (98% of Pascal pixels ≥0.95), so an adaptive cutoff floods the retention mask and self-training decays into confirmation bias."

The paper also reports that CW-BASS v2's own self-adaptive floor, run unconditionally on a saturated teacher, is the worst variant we test (82.32 vs. strict's 87.40), which is exactly why the method does not run it there.

The contributions are summarized as:

  1. A saturation-aware selection method with an explicit gate

  2. An unbiased noise diagnostic: held-out calibration

  3. A stability-guaranteed confidence floor

  4. The mechanism that decides the gate

The paper concludes with four cheap checks for practitioners: 1) Measure confident-set reliability, not just saturation... 2) Match the batch... 3) Report trajectories and final-vs-best, not best checkpoints... 4) Check the best-vs-final gap, on the EMA teacher itself.

The practical message is "not 'always strict' but a decision rule CW-BASS v2 embodies: measure whether the confident set earns the confidence you demand, filter strictly when it does, soften when it does not. The methodological counterpart is shorter still: measure before you adapt."

Improvements for AI systems

Improvements to AI Systems:

  1. Saturation-Aware Pseudo-Label Selection for Semi-Supervised Learning: Replace fixed or single-rule confidence thresholds with a two-mode selector that first measures confident-set reliability (πkept) on a held-out calibration slice. If πkept ≥ τ, use strict filtering; otherwise, use a self-adaptive confidence floor with confidence-weighted loss. This prevents performance collapse when teacher models (e.g., DINOv2) saturate confidence, and prevents noise flooding when they are miscalibrated.

  2. Unbiased Noise Estimation via Held-Out Calibration: Instead of estimating pseudo-label noise from in-batch training pixels (which is downward-biased due to student overfitting), split the labeled set into a training slice and a small calibration slice (α=5%). Compute per-class pseudo-label noise only on the held-out slice. This yields unbiased estimates (Proposition 1) and prevents thresholds from drifting too low as training progresses.

  3. Stability-Guaranteed Confidence Floor with Quantile Pinning: When using adaptive filtering, scale the confidence floor with the teacher's mean confidence and pin retention to a fixed quantile bounded away from 1 (Theorem 1, Corollary 1). This prevents the dynamic threshold from collapsing to full retention under saturated confidence distributions, which otherwise floods the training mask with errors.

  4. One-Pass Saturation Gate for Model Selection: Automatically decide between strict and adaptive filtering using a single diagnostic statistic (πkept) computed on a held-out slice. This gate is robust across different teacher strengths (e.g., ResNet-50 vs. DINOv2) and does not require retraining or multiple runs. It correctly identifies when hard training is near-optimal (reliable, saturated teacher) versus when softened loss is needed (confidently-miscalibrated teacher).

  5. Early-Stopping and Checkpoint Selection Based on Final-vs-Best Gap: Monitor the gap between the best EMA checkpoint and the final EMA checkpoint. If the gap is large (e.g., >2 mIoU), it indicates confirmation-bias decay from mask flooding. Use this signal to trigger early stopping or to revert to the best checkpoint, avoiding the 6+ mIoU drop observed in the paper.

  6. Confidence-Regime Diagnosis for Foundation-Model Teachers: Add a pre-training diagnostic that measures the teacher's confidence distribution (e.g., % pixels ≥0.95, ECE). If saturation is detected (e.g., >90% pixels above 0.95), automatically switch to strict filtering by default, unless the held-out reliability check shows πkept < τ, in which case use the adaptive floor. This prevents the worst-case scenario where adaptive filtering runs unconditionally on a saturated teacher (82.32 vs. 87.40 mIoU).

What the Improved AI System Can Do:

  • Achieve higher final accuracy in semi-supervised semantic segmentation with foundation-model teachers, recovering near-optimal operating points (e.g., 87.4 vs. 87.9 on Pascal VOC 1/8) without manual tuning.

  • Avoid catastrophic self-training decay by preventing mask flooding (retention >0.95) and the resulting confirmation bias, maintaining performance over longer training runs.

  • Adapt automatically to teacher quality—whether weak (ResNet) or strong (DINOv2)—without requiring prior knowledge of the teacher's calibration, making it a drop-in replacement for existing pseudo-label selection methods.

  • Provide reliable uncertainty estimates for pseudo-labels, enabling better downstream decision-making in low-label regimes (e.g., 1/8, 1/16 labeled data).

  • Reduce computational waste by avoiding long, unproductive training phases (e.g., epochs 4–20 where adaptive rules peak early but then decline), and by using a one-pass gate instead of expensive grid searches over thresholds.

Abstract

Semi-supervised semantic segmentation has long turned on one question, which pseudo-labels to trust, and a generation of selection rules, dynamic thresholds, per-class curricula, soft confidence weights, answered it for the noisy, under-confident ResNet teachers of their day. Self-supervised foundation encoders change the regime: with a DINOv2 teacher, confidence saturates, so the filtering that helped a weak teacher can hurt a strong one. We propose CW-BASS v2, a saturation-aware pseudo-label selection method that reads the teacher's confidence regime rather than committing to one rule. It pairs held-out calibration, an unbiased per-class noise estimate, with a self-adaptive confidence floor that provably bounds retention away from 1, and combines them in a one-pass gate: measure the reliability of the teacher's confident set, pi kept = Pr[correct c >= tau], on a held-out slice, and filter strictly when it meets the confidence demanded (pi kept >= tau), falling back to the adaptive floor otherwise. The boundary is the pre-existing operating threshold, not a value tuned to mIoU, and across six DINOv2 teachers it makes the correct strict-vs-floor call blind. CW-BASS v2 thus recovers the UniMatch V2 operating point on the saturated benchmarks by selecting strict (Pascal VOC 1/8 87.4 against its reported 87.9; Cityscapes within 0.5), and improves on it where the confident set is unreliable (pi kept 89%, ADE20K), where the floor edges ahead (+1.5 mIoU, single seed). The gate is principled because the failure it avoids is measured, not assumed: on a reliable, saturated teacher the confidence distribution's dynamic range collapses (98% of Pascal pixels >= 0.95), so an adaptive cutoff floods the retention mask and self-training decays into confirmation bias.

Sources

Related papers