GT-Space: Enhancing Heterogeneous Collaborative Perception with Ground Truth Feature Space
cs.LG, cs.AI, cs.MA
Submitted: 2026-03-13
Updated: 2026-09-12
Comments: Accepted by ICLR 2026
Code: https://github.com/KingScar/GT-Space
License: http://creativecommons.org/licenses/by/4.0/
The gist: In autonomous driving, multi-agent collaborative perception enhances sensing capabilities by enabling agents to share perceptual data.
Terminology
Abstract
In autonomous driving, multi-agent collaborative perception enhances sensing capabilities by enabling agents to share perceptual data. A key challenge lies in handling heterogeneous features from agents equipped with different sensing modalities or model architectures, which complicates data fusion. Existing approaches often require retraining encoders or designing interpreter modules for pairwise feature alignment, but these solutions are not scalable in practice. To address this, we propose GT-Space, a flexible and scalable collaborative perception framework for heterogeneous agents. GT-Space constructs a common feature space from ground-truth labels, providing a unified reference for feature alignment. With this shared space, agents only need a single adapter module to project their features, eliminating the need for pairwise interactions with other agents. Furthermore, we design a fusion network trained with contrastive losses across diverse modality combinations. Extensive experiments on simulation datasets (OPV2V and V2XSet) and a real-world dataset (RCooper) demonstrate that GT-Space consistently outperforms baselines in detection accuracy while delivering robust performance. Our code will be released at https://github.com/KingScar/GT-Space.
Sources
- QUEST: Query Stream for Practical Cooperative Perception
- STAMP: Scalable Task And Model-agnostic Collaborative Perception
- ActFormer: Scalable Collaborative Perception via Active Queries
- Adam: A Method for Stochastic Optimization
- An Extensible Framework for Open Heterogeneous Collaborative Perception
- AB3DMOT: A Baseline for 3D Multi-Object Tracking and New Evaluation Metrics
- HM-ViT: Hetero-modal Vehicle-to-Vehicle Cooperative perception with vision transformer
- CoBEVT: Cooperative Bird's Eye View Semantic Segmentation with Sparse Transformers
- Meta-Transformer: A Unified Framework for Multimodal Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks