Learning from Distributed Eyes: Leveraging Collaborative Perception for Automated Model Adaptation
cs.CV, cs.LG
Submitted: 2026-09-16
Updated: 2026-09-16
Comments: 9 pages, 3 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: In autonomous driving, perception models often struggle to generalize to new environments due to domain shifts.
Terminology
Abstract
In autonomous driving, perception models often struggle to generalize to new environments due to domain shifts. While unsupervised model adaptation offers a feasible solution without labor-intensive manual labeling, existing methods that rely solely on the ego-vehicle's data often lead to inferior pseudo-labeling performance. To address this critical issue, we propose LDE, Learning from Distributed ``Eyes", a novel framework that transforms collaborative perception (CP) into a source of high-quality supervision for model adaptation. This pseudo-labeling approach is hyperparameter-insensitive and relatively reliable, assuming CP often outperforms single-agent's perception. However, naively implementing this approach encounters (1) the communication bottleneck of sharing rich features under time and bandwidth constraints, (2) the view discrepancy between the CP view and the learner's Field of View (FoV), and (3) the unreliability even in CP-generated labels. To address these issues, we design an adaptation-oriented feature sharing mechanism that selectively transmits the most critical information for adaptation, an FoV filtering method that meticulously eliminates mismatched labels, and a curriculum learning strategy to progressively exploit pseudo labels. Extensive experiments on 3D object detection tasks demonstrate that LDE consistently outperforms both the pre-trained models and state-of-the-art unsupervised adaptation methods.
Sources
- Continual Test-Time Adaptation for Object Detection with Adaptive Monitoring and Randomized Restoration
- Birdcast: Interest-aware BEV Multicasting for Infrastructure-assisted Collaborative Perception
- Update the Unseen Only: Minimizing AoI for Collaborative Perception through Online Learning
- Revisiting Batch Normalization For Practical Domain Adaptation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models