MIDR: Enrichment-Augmented Indexing for Multimodal Document Retrieval
cs.IR, cs.AI, cs.CL, cs.CV, cs.LG
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: To appear in Proceedings of EMNLP 2026
Code: https://github.com/Sygil-Dev/whoosh-reloaded
License: http://creativecommons.org/licenses/by/4.0/
The gist: Retrieval over visually rich documents has a representation problem: important content often lives in tables, charts, figures, and layout relations that plain OCR linearizes, corrupts, or omits.
Terminology
Abstract
Retrieval over visually rich documents has a representation problem: important content often lives in tables, charts, figures, and layout relations that plain OCR linearizes, corrupts, or omits. ColPali-family visual retrievers address this with patch-level multi-vector indexes and late-interaction scoring, keeping image-derived retrieval on the query-time serving path. We introduce MIDR (Multimodal Indexing for Document Retrieval), a training-free framework for enrichment-augmented indexing that shifts multimodal reasoning to index time. During ingestion, a multimodal LLM converts rendered pages into verified textual fields that are indexed with BM25F and optionally fused with dense retrieval, enabling text-centric serving over multimodally grounded evidence. On ViDoRe V3, MIDR Hybrid achieves 0.6219 average nDCG across five English domains, a 23.0% relative gain over BM25, remaining competitive with ColQwen2.5. On two French-document domains, enrichment bridges English queries and French page text, lifting BM25 from 0.1532 to 0.5448 nDCG and outperforming ColQwen2.5. Across all seven domains, MIDR leads ColQwen2.5 on four while using approximately 9x smaller index memory and approximately 2x lower query latency. These results establish index-time multimodal reasoning as a compelling accuracy-deployment alternative to serving-time visual late interaction.
Sources
- IndexRAG: Index-Time Reasoning for Multi-Hop Retrieval-Augmented Generation
- The Faiss library
- A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models
- EnrichIndex: Using LLMs to Enrich Retrieval Indices Offline
- ColPali: Efficient Document Retrieval with Vision Language Models
- M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding
- Zero-shot Multimodal Document Retrieval via Cross-modal Question Generation
- MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding
- Long-Context Long-Form Question Answering for Legal Domain
- Doc2Query++: Topic-Coverage based Document Expansion and its Application to Dense Retrieval via Dual-Index Fusion
- Nemotron ColEmbed V2: Top-Performing Late Interaction Embedding Models for Visual Document Retrieval
- Document Expansion by Query Prediction
- FLASH-MAXSIM: IO-Aware Fused Kernels for Late-Interaction Retrieval
- Qwen3.5-Omni Technical Report
- A Comprehensive Survey on Long Context Language Modeling
- ViDoRe V3: A Comprehensive Evaluation of Retrieval Augmented Generation in Complex Real-World Scenarios
- TileMaxSim: IO-Aware GPU MaxSim Scoring with Dimension Tiling and Fused Product Quantization
- Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization
- EmbeddingGemma: Powerful and Lightweight Text Representations
- C-Pack: Packed Resources For General Chinese Embeddings
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG