Rethinking Multi-Branch and Cross-Backbone Fusion for Vehicle Re-Identification under Foundation-Model Pretraining
cs.CV, cs.LG
Submitted: 2026-07-24
Updated: 2026-09-22
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multi-branch architectures and CNN-Transformer fusion are widely believed to improve vehicle re-identification (Re-ID) by combining complementary representations.
Terminology
Abstract
Multi-branch architectures and CNN-Transformer fusion are widely believed to improve vehicle re-identification (Re-ID) by combining complementary representations. We revisit this for a DINOv3-pretrained backbone. A single DINOv3-pretrained ConvNeXt with a tuned recipe reaches 88.19 mAP on VeRi-Wild Small and 77.47 on Large from visual cues alone, within the combined evaluation and optimization noise of the strongest protocol-verified metadata-dependent multi-branch baseline, and 92.38/83.68 with training-free re-ranking. Using this baseline and retrieval-level branch diagnostics, we ask whether representational diversity still pays at this scale. In our runs, it does not. Across both benchmarks and every converged configuration, concatenating multiple heads over a shared backbone moves the best single head by under one mAP point in either direction while costing four times the embedding dimension; 99.7% of the concatenation's variance lies in 512 principal components, so the heads not only duplicate one another but each occupies a quarter of its nominal 2048 dimensions. Pushing diversity to its architectural limit, CNN versus Transformer, we grant fusion every advantage through an asymmetric frozen-anchor scheme. Every Transformer configuration still lands 13-15 mAP below the ConvNeXt backbone, and a paired per-query bootstrap bounds the fusion gain at +0.11 mAP (95% CI) even for the most favourable snapshot we obtained. One strong backbone with the right recipe and re-ranking is the efficiency frontier. All results use single-seed training and one foundation-model family; differences of this size are therefore reported as bounds rather than orderings, and we list falsifiers.
Sources
- A Comprehensive Survey on Deep-Learning-based Vehicle Re-Identification: Models, Data Sets and Challenges
- DINOv3
- ResNet strikes back: An improved training procedure in timm
- Deep Ensembles: A Loss Landscape Perspective
- CLIP-SENet: CLIP-based Semantic Enhancement Network for Vehicle Re-identification
- In Defense of the Triplet Loss for Person Re-Identification
- Parameter Efficient Fine-tuning of Self-supervised ViTs without Catastrophic Forgetting
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models