Hypersolid: Emergent Vision Representations via Short-Range Repulsion
cs.CV, cs.AI, cs.LG
Submitted: 2026-01-29
Updated: 2026-09-28
Comments: 17 pages, 16 figures
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: A central problem in self-supervised learning is preventing representation collapse.
Terminology
Abstract
A central problem in self-supervised learning is preventing representation collapse. Most methods avoid it through global mechanisms, such as contrastive expansion, variance constraints, decorrelating dimensions, or enforcing certain output distributions. In this work, we study a different design: short-range repulsion. We introduce Hypersolid, a self-supervised objective that combines view alignment with local collision avoidance. Our method induces a latent geometry of compact, semantically aligned neighborhoods with low anisotropy. This geometry is especially effective for unsupervised clustering and fine-grained separation, although it comes at the cost of weaker transferability.
Sources
- LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
- Contrastive Self-Supervised Learning As Neural Manifold Packing
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models