Diverse via bounded Agreement: Geometric Regularization for Multimodal Fusion
cs.CV, cs.LG
Submitted: 2026-01-29
Updated: 2026-08-28
Terminology
Sources
- VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning
- A Closer Look at Multimodal Representation Collapse
- Microsoft COCO Captions: Data Collection and Evaluation Server
- X-Fi: A Modality-Invariant Foundation Model for Multimodal Human Sensing
- What to align in multimodal contrastive learning?
- Balancing Multi-modal Sensor Learning via Multi-objective Optimization
- Deep Incomplete Multi-view Learning via Cyclic Permutation of VAEs
- Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models
- Learning Multimodal VAEs through Mutual Supervision
- See-Saw Modality Balance: See Gradient, and Sew Impaired Vision-Language Balance to Mitigate Dominant Modality Bias
- Generalized Multimodal ELBO
- MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance
- Explaining and Mitigating the Modality Gap in Contrastive Multimodal Learning
- Decipher the Modality Gap in Multimodal Contrastive Learning: From Convergent Representations to Pairwise Alignment
- Towards Uniformity and Alignment for Multimodal Representation Learning
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models