PlanSightRAG: A Visual-First Multimodal RAG for Automating Question Answering and Compliance Checking for Civil Standard Plans
cs.IR, cs.CL, cs.CV
Submitted: 2026-08-26
Updated: 2026-08-26
Terminology
Sources
- ColPali: Efficient Document Retrieval with Vision Language Models
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Gemini: A Family of Highly Capable Multimodal Models
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- GPT-4 Technical Report
- The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
- HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Nemotron ColEmbed V2: Top-Performing Late Interaction Embedding Models for Visual Document Retrieval
- VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG