Less Is More: Graph-free Multimodal RAG via Multi-signal Late Fusion
cs.CL
Submitted: 2026-09-16
Updated: 2026-09-16
Code: https://github.com/hk-zh/TrioRAG
License: http://creativecommons.org/licenses/by/4.0/
The gist: Graph-based retrieval-augmented generation (RAG) is widely used for multimodal, cross-document question answering.
Terminology
Abstract
Graph-based retrieval-augmented generation (RAG) is widely used for multimodal, cross-document question answering. However, building corpus-level graphs is expensive, slow to query, and difficult to maintain. We present TrioRAG, a graph-free multimodal framework that integrates evidence from three complementary signals: the question, the anchor image, and a VLM-enhanced query generated from both. Each signal retrieves independently over a shared multi-vector index of page text and page images, and the results are combined through late fusion. Further, we introduce AutoQA, a multimodal automotive benchmark whose questions are grounded in noisy, web-sourced images rather than clean document-sourced figures. Its questions require reasoning across manuals. We position it as a model-curated testbed rather than a human-validated gold standard. Across three benchmarks, TrioRAG matches or outperforms graph-based systems while reducing total cost and accelerating per-query inference by 1.6-2.3 times. By construction, AutoQA grounds its questions in out-of-corpus web images. In this setting image retrieval reaches only 19.3% document-level recall, while text-derived signals, especially the VLM-enhanced query, keep retrieval robust.
Sources
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- RAG-Anything: All-in-One RAG Framework
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
- VAT-KG: Knowledge-Intensive Multimodal Knowledge Graph Dataset for Retrieval-Augmented Generation
- RAG-Fusion: a New Take on Retrieval-Augmented Generation
- Beyond Text: Optimizing RAG with Multimodal Inputs for Industrial Applications
- MMGraphRAG: Bridging Vision and Language with Interpretable Multimodal Knowledge Graphs
- MinerU: An Open-Source Solution for Precise Document Content Extraction
- Evaluating Knowledge Graph Based Retrieval Augmented Generation Methods under Knowledge Incompleteness
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering