Spatial Matryoshka Training for Multi-Granularity Visual Document Retrieval
cs.AI, cs.IR
Submitted: 2026-08-30
Updated: 2026-08-30
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multi-modal late-interaction retrievers achieve strong retrieval on visually rich documents by representing each page as per patch embeddings and matching at the token level.
Terminology
Abstract
Multi-modal late-interaction retrievers achieve strong retrieval on visually rich documents by representing each page as per patch embeddings and matching at the token level. However, this approach incurs high storage costs. Existing compression methods typically fix a single compression level at indexing time, limiting flexibility. We present ColSNAP (Spatial Nested Average Pooling)1, a training method that generates a nested hierarchy of compression levels directly from a backbone's patch grid. By spatially pooling patch embeddings into pro- gressively coarser tiers and training all tiers simultaneously, a single model learns to support retrieval at multiple compression levels without architectural changes. Crucially, a single encoding pass yields every tier, enabling the accuracy-storage trade-off to be configured at indexing time to match avail- able storage budgets, rather than being fixed during training. We demonstrate that models trained using ColSNAP maintain near full-resolution retrieval performance under substantial compression and that ColSNAP transfers effectively across multiple late-interaction backbones, and achieves most of its improvements via a lightweight adaptation stage applied to a pre-trained retriever.
Sources
- PaliGemma: A versatile 3B VLM for transfer
- Matryoshka Multimodal Models
- ReinPool: Reinforcement Learning Pooling Multi-Vector Embeddings for Retrieval System
- Reducing the Footprint of Multi-Vector Retrieval with Minimal Performance Impact via Token Pooling
- Nemotron ColEmbed V2: Top-Performing Late Interaction Embedding Models for Visual Document Retrieval
- MUVERA: Multi-Vector Retrieval via Fixed Dimensional Encodings
- MMTEB: Massive Multilingual Text Embedding Benchmark
- ColPali: Efficient Document Retrieval with Vision Language Models
- jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval
- LoRA: Low-Rank Adaptation of Large Language Models
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
- Matryoshka Representation Learning
- Rethinking the Role of Token Retrieval in Multi-Vector Retrieval
- Structural Anchor Pruning: Training-Free Multi-Vector Compression for Visual Document Retrieval
- ViDoRe V3: A Comprehensive Evaluation of Retrieval Augmented Generation in Complex Real-World Scenarios
- Decoupled Weight Decay Regularization
- Unifying Multimodal Retrieval via Document Screenshot Embedding
- Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings
- ViDoRe Benchmark V2: Raising the Bar for Visual Retrieval
- VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection