PseudoMapLabeler: Confidence-Aware Pseudo-Label Generation for Semi-Supervised Online Mapping
Chikao Tsuchiya, Dhaval Bhanderi, David Ilstrup, Hsinmin Cheng, Christopher Ostafew
Nissan Advanced Technology Center - Silicon Valley, Nissan North America
cs.CV, cs.AI
Submitted: 2026-08-12
Updated: 2026-08-14
Comments: 17 pages, 4 figures, Accepted at ECCV 2026 DriveX Workshop on Foundation Models for Autonomous Driving
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: PseudoMapLabeler: Confidence-Aware Pseudo-Label Generation for Semi-Supervised Online Mapping Abstract Summary: The paper addresses the critical challenge of labeled training data scarcity in
Terminology
Summary
PseudoMapLabeler: Confidence-Aware Pseudo-Label Generation for Semi-Supervised Online Mapping
Abstract Summary: The paper addresses the critical challenge of labeled training data scarcity in deploying online HD map construction systems to real-world scenarios. The authors propose a teacher-student semi-supervised learning (SSL) framework that generates high-quality pseudo-labels from unlabeled data through confidence-aware map refinement. The approach first trains a teacher model on limited labeled data, then leverages Beta-distribution-based confidence maps to assess the reliability of predicted map elements across temporal observations. Unlike conventional filtering methods that discard entire elements, the paper introduces a spatial clipping technique that selectively preserves high-confidence regions while removing unreliable segments. The refined map elements serve as map priors that improve the teacher model's prediction accuracy on unlabeled data in a second pass. These enhanced predictions become pseudo-labels for training a student model from scratch, followed by fine-tuning on the original labeled data. Experimental results on the nuScenes dataset demonstrate that the teacher-student framework with refined pseudo-labels improves performance by +6.1 mAP under a low-label regime compared to training on labeled data alone.
Introduction Summary: The paper notes that online HD map construction is critical for autonomous driving, enabling real-time perception of road topology and semantic elements from onboard sensors. Recent advances in deep learning-based approaches have demonstrated impressive performance on benchmark datasets such as nuScenes and Argoverse. However, a fundamental challenge remains: the scarcity of labeled training data, as annotating HD maps requires extensive manual effort by domain experts, making it prohibitively expensive to collect labeled data for every new environment. Semi-supervised learning offers a promising solution by leveraging abundant unlabeled data alongside limited labeled samples, but the key challenge lies in generating high-quality pseudo-labels from model predictions on unlabeled data. Naive approaches that directly use raw predictions as pseudo-labels often introduce significant noise, leading to error accumulation during training. For online HD map construction, this problem is particularly acute due to the temporal nature of predictions: while individual frame predictions may contain errors, aggregating predictions across multiple temporal observations can provide more reliable estimates. However, simple temporal aggregation without confidence assessment fails to distinguish between consistent high-quality predictions and noisy outliers.
Key Contributions:
-
Beta-Distribution Confidence Map: The paper proposes a Beta-distribution based evidence aggregation scheme that summarizes spatiotemporal confidence over a BEV grid and yields a per-cell reliability score for map refinement. Unlike deterministic approaches, this method accounts for both the number of observations and the prediction confidence at each spatial location, providing an interpretable way to assess reliability.
-
Spatial Clipping for Prior Map Generation: The paper proposes a novel spatial clipping technique that selectively preserves high-confidence regions of map elements while removing unreliable segments. To handle class-specific confidence distributions, percentile-based adaptive thresholding is employed that ensures balanced prior map generation across different map element types. This approach maximizes the utilization of unlabeled data by retaining partial information from predictions, rather than discarding entire elements as in conventional filtering methods.
-
Model-Agnostic Pseudo-Labeling Pipeline: The pseudo-label generation and refinement procedure is architecture-agnostic: it operates on predicted vectorized map elements and does not rely on model-specific internals. While the teacher-student framework is instantiated with Uni-PrevPredMap in experiments for controlled evaluation, the refined pseudo-labels can be used to train other online vectorized mapping architectures. The paper validates this on a second architecture (MapTR), observing consistent gains.
Method Details:
-
Problem Formulation: The paper assumes access to a limited labeled dataset DL and a larger unlabeled dataset DU where NU ≫ NL. Each sequence F consists of frames with corresponding ego-vehicle poses. Ground truth map elements consist of vectorized polylines with semantic class labels c ∈ divider, ped crossing, boundary.
-
Teacher Model: Built upon PrevPredMap and Uni-PrevPredMap, which are state-of-the-art online vectorized mapping methods that leverage temporal priors. The model predicts map elements for each frame, where each element consists of a polyline represented by a sequence of 2D points, along with a semantic class label and a prediction confidence score.
-
Temporal Accumulation: Map elements are transformed from ego-vehicle coordinates to a scene-local coordinate system using ego-vehicle poses. For nuScenes, sequences are divided into 20-second scenes, which serve as the processing unit.
-
Probability Calibration: Temperature scaling is applied to address overconfident or poorly calibrated probability estimates. The temperature parameter T is optimized on a held-out validation set by minimizing the negative log-likelihood of binary cross-entropy. The optimal temperature parameter is obtained as T* = 0.696.
-
Beta-Distribution Confidence Map: For each grid cell (resolution δ = 0.5m) and class, the method tracks the number of times the cell was observed and the weighted detection count based on calibrated prediction confidence scores. The detection probability is modeled using a Beta distribution with shape parameters α and β updated based on observations. The confidence at each grid cell is computed as the posterior mean of the Beta distribution. The formulation naturally accounts for both the frequency of detections and prediction confidence scores, provides higher confidence when multiple consistent observations are available, gracefully handles sparse observations through the Bayesian prior, and produces confidence estimates that reflect detection frequency and prediction scores.
-
Confidence-Based Map Refinement: The spatial clipping method selectively preserves high-confidence portions of map elements while discarding unreliable segments. For each accumulated map element represented as a polyline, confidence values are sampled at each point using bilinear interpolation. Continuous segments where confidence exceeds a threshold are identified, with additional filtering to remove segments that are too short or have too few points. For polygon-type elements (e.g., pedestrian crossings), a stricter criterion is adopted: the polygon is retained only if all points satisfy the confidence threshold. Percentile-based adaptive thresholding is used to handle class-specific confidence distributions, setting the threshold as the percentile of confidence values for each class.
-
Teacher-Student Framework: The teacher model is trained using only the labeled dataset DL. The trained teacher model is applied to the unlabeled dataset to obtain initial predictions, which are calibrated and aggregated across temporal sequences. Beta-distribution-based confidence maps are constructed, and spatial clipping with percentile-based thresholding is applied to extract high-confidence map element segments. The teacher model is then applied again to the unlabeled dataset with the refined temporal priors as additional input, generating enhanced predictions that serve as pseudo-labels. Finally, a student model is trained from scratch using the pseudo-labeled dataset, followed by fine-tuning on the original labeled dataset.
Experimental Setup: Experiments are conducted on the nuScenes dataset using the geographically disjoint split proposed by StreamMapNet to prevent data leakage. The training set is split into a labeled set DL (115 scenes, 4,636 samples, 16.5%) and a pseudo-unlabeled set DU (581 scenes, 23,332 samples, 83.5%). Both teacher and student models are based on UPPM, with the global map prior component removed entirely. Temporal priors are cleared at scene boundaries to ensure independent processing of each 20-second scene.
Main Results:
-
Teacher Model Training: The teacher model trained on labeled data alone achieved 21.5 mAP on the validation set.
-
Map Prior Generation: The effect of percentile threshold on pseudo-label quality was evaluated. The mAP peaks at p = 20, while rasterized Dice/IoU peak at p = 30 with only a marginal mAP reduction. The paper chose p = 30 as a balanced operating point for student training.
-
Student Model Training: The results show that Ours (Clipping) achieves a +6.1 mAP improvement over the Baseline (21.5 to 27.6 mAP), validating the effectiveness of leveraging unlabeled data through confidence-aware pseudo-labeling. Ours (Clipping) outperforms Ours (Filtering) by +2.8 mAP (27.6 vs 24.8 mAP), confirming the superiority of spatial clipping over conventional element-wise filtering. Notably, pedestrian crossing AP improves from 5.4 to 9.3, suggesting that segment-level refinement particularly benefits structurally complex classes.
-
Cross-Architecture Generalization: The MapTR baseline achieves 16.5 mAP, and training with Clipping-based pseudo-labels improves this to 21.5 mAP (+5.0 mAP). The Clipping method also outperforms Filtering method by +1.4 mAP, supporting the claim that the refinement-and-pseudo-labeling pipeline is model-agnostic.
Conclusion: The paper presents a teacher-student framework for semi-supervised online HD map construction that addresses labeled data scarcity. The approach leverages Beta-distribution-based confidence maps to assess the reliability of temporally accumulated predictions and introduces spatial clipping to selectively preserve high-confidence regions for pseudo-label generation. The refined map elements serve as map priors that improve prediction accuracy in a second pass, resulting in high-quality pseudo-labels for student model training. The baseline uses a fully temporal-prior-enabled UPPM, so the +6.1 mAP gain conservatively reflects genuine SSL benefit.
Limitations: The paper acknowledges several limitations: (1) the temporal accumulation process requires accurate ego-pose estimation and sensor calibration; errors in localization can lead to misaligned map elements and degraded confidence maps; (2) the method assumes that the teacher model's predictions, even if noisy, contain sufficient signal for confidence-based refinement; in scenarios with extremely limited labeled data or severe domain shift, the initial teacher quality may be insufficient; (3) an ≈10 mAP gap remains between refined pseudo-labels and the GT Prior upper bound, indicating headroom attributable to limited teacher quality at 16.5% labeled data; (4) evaluation is only on nuScenes, with extending to additional datasets and cross-region settings as future work.
Improvements for AI systems
Improvements to AI Systems:
- Confidence-Aware Temporal Aggregation Module
-
Integrate Beta-distribution-based confidence scoring into any multi-frame perception system (e.g., object detection, lane detection, or semantic segmentation) to replace naive averaging or voting.
-
The system can now produce per-cell reliability maps that explicitly account for observation frequency and prediction confidence, enabling robust fusion of noisy temporal data without requiring hand-tuned thresholds.
- Spatial Clipping for Partial Label Retention
-
Replace binary element-level filtering (keep/discard) with segment-level spatial clipping that preserves high-confidence sub-regions of predicted structures.
-
The improved system can salvage partially correct predictions (e.g., a lane boundary with a noisy segment) by retaining only reliable portions, maximizing usable information from unlabeled data and reducing label noise in downstream training.
- Adaptive Percentile Thresholding for Class Imbalance
-
Implement class-specific, percentile-based confidence thresholds instead of fixed global thresholds.
-
The system can now handle heterogeneous map element types (e.g., dividers vs. pedestrian crossings) with different confidence distributions, preventing over-pruning of rare or complex classes and improving recall for structurally intricate objects.
- Two-Pass Teacher-Student Refinement with Prior Feedback
-
Use refined high-confidence map elements as explicit priors to re-query the teacher model on the same unlabeled data, generating enhanced pseudo-labels.
-
The improved system can iteratively correct its own errors by conditioning on spatial priors, leading to higher-quality pseudo-labels that capture global context and reduce hallucinated or misplaced elements.
- Model-Agnostic Pseudo-Label Generation Pipeline
-
Decouple the pseudo-labeling process from any specific neural architecture (e.g., MapTR, PrevPredMap) by operating solely on vectorized outputs and confidence scores.
-
The improved system can be plugged into any online mapping or vectorized perception framework, enabling semi-supervised learning gains across diverse model families without architectural changes.
- Calibrated Confidence Estimation via Temperature Scaling
-
Apply temperature scaling to model outputs before confidence aggregation to correct overconfidence or miscalibration.
-
The improved system yields more reliable confidence scores, which directly improves the quality of Beta-distribution updates, spatial clipping decisions, and final pseudo-label accuracy—especially under distribution shift or limited labeled data.
- Bayesian Prior for Sparse Observation Handling
-
Use Beta-distribution priors to model detection confidence in sparsely observed grid cells.
-
The improved system gracefully degrades in regions with few temporal observations, avoiding overconfident or unstable pseudo-labels, and provides uncertainty estimates that can be used for active learning or safe deployment.
- Geographically Disjoint Data Splitting for Realistic Evaluation
-
Adopt the paper’s evaluation protocol (geographically disjoint train/val splits) to prevent data leakage in SSL benchmarks.
-
The improved system can be validated more realistically, ensuring that reported gains reflect genuine generalization to new environments rather than memorization of overlapping scenes.
What the Improved AI System Can Do:
-
Self-Train on Unlabeled Data with Minimal Noise: Automatically generate high-quality pseudo-labels from raw sensor streams, even with only 10–20% labeled data, by leveraging temporal consistency and confidence-aware refinement.
-
Handle Complex Map Elements Robustly: Preserve partial information from noisy predictions (e.g., a partially occluded crosswalk) rather than discarding the entire element, improving recall for rare or structurally complex classes.
-
Adapt to New Environments Without Re-annotation: Deploy to new cities or regions using only unlabeled driving logs, with the system automatically refining its own predictions via temporal aggregation and spatial clipping.
-
Integrate into Any Online Mapping Stack: Provide a drop-in pseudo-labeling module for existing vectorized mapping models (e.g., MapTR, StreamMapNet) without requiring model-specific modifications.
-
Provide Uncertainty-Aware Outputs: Output per-cell confidence maps that reflect both observation count and prediction reliability, enabling downstream planners to weight map elements by certainty or trigger human review in low-confidence zones.
-
Achieve Consistent Performance Gains: Improve mAP by +5–6 points in low-label regimes (e.g., from 21.5 to 27.6 mAP on nuScenes) and generalize across architectures (e.g., +5.0 mAP for MapTR), making SSL practical for real-world autonomous driving deployment.
Abstract
A critical challenge in deploying online HD map construction systems to real-world scenarios is the scarcity of labeled training data, which limits model generalization in diverse environments. To address this limitation, we propose a teacher-student semi-supervised learning (SSL) framework that generates high-quality pseudo-labels from unlabeled data through confidence-aware map refinement. Our approach first trains a teacher model on limited labeled data, then leverages Beta-distribution-based confidence maps to assess the reliability of predicted map elements across temporal observations. Unlike conventional filtering methods that discard entire elements, we introduce a spatial clipping technique that selectively preserves high-confidence regions while removing unreliable segments. The refined map elements serve as map priors that improve the teacher model's prediction accuracy on unlabeled data in a second pass. These enhanced predictions become pseudo-labels for training a student model from scratch, followed by fine-tuning on the original labeled data. Experimental results on the nuScenes dataset demonstrate that our teacher-student framework with refined pseudo-labels improves performance by +6.1 mAP under a low-label regime compared to training on labeled data alone, offering a practical solution to the labeled data scarcity problem in online HD map construction.
Sources
- VMA: Divide-and-Conquer Vectorized Map Annotation System for Large-Scale Driving Scene
- Improved Regularization of Convolutional Neural Networks with Cutout
- A Review of Pseudo-Labeling for Computer Vision
- Uni-PrevPredMap: Extending PrevPredMap to a Unified Framework of Prior-Informed Modeling for Online Vectorized HD Map Construction
- Pseudo-Label Noise Suppression Techniques for Semi-Supervised Semantic Segmentation
- Semi-Supervised Regression with Heteroscedastic Pseudo-Labels
- Using Language and Road Manuals to Inform Map Reconstruction for Autonomous Driving
- Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models