NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval
cs.IR, cs.CV, cs.LG
Submitted: 2026-03-13
Updated: 2026-08-29
Comments: Accepted to EMNLP 2026 (Main Track)
License: http://creativecommons.org/licenses/by/4.0/
The gist: Vision-Language Model (VLM) based retrievers have advanced visual document retrieval (VDR) to impressive quality.
Terminology
Abstract
Vision-Language Model (VLM) based retrievers have advanced visual document retrieval (VDR) to impressive quality. They require the same multi-billion parameter encoder for both document indexing and query encoding, incurring high latency and GPU dependence even for plain-text queries. We observe that this design is unnecessarily symmetric: documents are visually complex and demand strong visual understanding, whereas queries are just short text strings. NanoVDR exploits this query--document asymmetry by decoupling the two encoding paths: a frozen 2B VLM teacher indexes documents offline, while a distilled text-only student as small as 69M parameters encodes queries at inference. The key design choice is the distillation objective. Through systematic comparison of six objectives across three backbones and 22 ViDoRe benchmark datasets, we find that pointwise cosine alignment on query text consistently outperforms ranking-based and contrastive alternatives, while requiring only pre-cached teacher query embeddings and no document processing during training. Furthermore, we identify cross-lingual transfer as the primary performance bottleneck, and resolve it cheaply by augmenting training data with machine-translated queries. The resulting NanoVDR-S-Multi (DistilBERT, 69M) retains 95.1% of teacher quality and outperforms DSE-Qwen2 (2B) on v2 and v3 with 32 times fewer parameters and 50 times lower CPU query latency, at a total training cost under 13 GPU-hours.
Sources
- Qwen3-VL Technical Report
- Distilling the Knowledge in a Neural Network
- Improving Efficient Neural Ranking Models with Cross-Architecture Knowledge Distillation
- ViDoRe V3: A Comprehensive Evaluation of Retrieval Augmented Generation in Complex Real-World Scenarios
- ViDoRe Benchmark V2: Raising the Bar for Visual Retrieval
- SERVAL: Surprisingly Effective Zero-Shot Visual Document Retrieval Powered by Large Vision and Language Models
- Representation Learning with Contrastive Predictive Coding
- Dense Passage Retrieval for Open-Domain Question Answering
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- Jina CLIP: Your CLIP Model Is Also Your Text Retriever
- Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking
- ModernVBERT: Towards Smaller Visual Document Retrievers
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference
- VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG