SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport
cs.LG, cs.AI
Submitted: 2026-02-26
Updated: 2026-09-21
Comments: ICML 2026
Code: https://github.com/ExplainableML/SOTAlign
License: http://creativecommons.org/licenses/by/4.0/
The gist: The Platonic Representation Hypothesis posits that neural networks trained on different modalities converge toward a shared statistical model of the world.
Terminology
Abstract
The Platonic Representation Hypothesis posits that neural networks trained on different modalities converge toward a shared statistical model of the world. Recent work exploits this convergence by aligning frozen pretrained vision and language models with lightweight alignment layers, but typically relies on contrastive losses and millions of paired samples. In this work, we ask whether meaningful alignment can be achieved with substantially less supervision. We introduce a semi-supervised setting in which pretrained unimodal encoders are aligned using a small number of image-text pairs together with large amounts of unpaired data. To address this challenge, we propose SOTAlign, a two-stage framework that first recovers a coarse shared geometry from limited paired data using a linear teacher, and then refines the alignment on unpaired samples via an optimal-transport-based divergence that transfers relational structure without overconstraining the target space. SOTAlign effectively leverages unpaired images and text, learning robust joint embeddings across datasets and encoder pairs, and significantly outperforming supervised and semi-supervised baselines. Code is available at https://github.com/ExplainableML/SOTAlign.
Sources
- Reliability of CKA as a Similarity Measure in Deep Learning
- Ranking via Sinkhorn Propagation
- Llama-Embed-Nemotron-8B: A Universal Text Embedding Model for Multilingual and Cross-Lingual Tasks
- Learning Permutations with Sinkhorn Policy Gradient
- MMTEB: Massive Multilingual Text Embedding Benchmark
- LLM2CLIP: Powerful Language Model Unlocks Richer Cross-Modality Representation
- The quest for the GRAph Level autoEncoder (GRALE)
- Linear Spherical Sliced Optimal Transport: A Fast Metric for Comparing Spherical Data
- Fine-Grained Visual Classification of Aircraft
- DINOv3
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Representation Learning with Contrastive Predictive Coding
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks