ViSAR: Training-Free Adaptive- k Retrieval for Visual Document Question Answering
cs.IR, cs.AI, cs.CL, cs.CV
Submitted: 2026-09-02
Updated: 2026-09-03
Comments: 13 pages, 5 figures, 4 tables
Code: https://github.com/adrienmialland/ViSAR
License: http://creativecommons.org/licenses/by/4.0/
The gist: Document Visual Question Answering (DocVQA) often leverages Retrieval-Augmented Generation (RAG), where late-interaction encoders are commonly used to identify document pages relevant to a user
Terminology
Abstract
Document Visual Question Answering (DocVQA) often leverages Retrieval-Augmented Generation (RAG), where late-interaction encoders are commonly used to identify document pages relevant to a user query, before answer generation by a Large Vision-Language Model (LVLM). Existing approaches typically retrieve a fixed top- k number of pages regardless of query complexity, which increases LVLM latency and may degrade answer accuracy. We introduce ViSAR (Visual Semantic Activation Retrieval), a training-free adaptive- k retrieval method for late-interaction visual document retrieval. ViSAR operates directly in the embedding space to construct a query-conditioned page-level similarity matrix that highlights query-relevant semantics and dynamically determines the number of pages to retrieve. Across multiple encoders and LVLMs, ViSAR retrieves compact, query-adapted page sets that reduce RAG latency by up to 58.7%, while maintaining or improving answer accuracy compared with fixed top- k and adaptive retrieval heuristics. Furthermore, we show that the similarity matrix structure correlates with answer accuracy, suggesting future directions for retrieval quality-aware document understanding.
Sources
- M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding
- MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding
- ScreenAI: A Vision-Language Model for UI and Infographics Understanding
- Qwen3-VL Technical Report
- Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends
- ModernVBERT: Towards Smaller Visual Document Retrievers
- SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval
- Context-Aware Sentence/Passage Term Importance Estimation For First Stage Retrieval
- Enhancing Retrieval-Augmented Generation with Topic-Enriched Embeddings: A Hybrid Approach Integrating Traditional NLP Techniques
- ViDoRe V3: A Comprehensive Evaluation of Retrieval Augmented Generation in Complex Real-World Scenarios
- Cluster-based Adaptive Retrieval: Dynamic Context Selection for RAG Applications
- Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook
- MHier-RAG: Multi-Modal RAG for Visual-Rich Document Question-Answering via Hierarchical and Multi-Granularity Reasoning
- No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG