Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting
cs.CL, cs.CV, cs.IR, cs.MM
Submitted: 2026-06-30
Updated: 2026-06-30
Comments: Accepted to ECCV 2026. The datasets and code are available in https://github.com/VAN-QIAN/ECCV26-ARA
DOI: 10.1007/978-3-032-37574-2_1
Code: https://github.com/VAN-QIAN/ECCV26-ARA
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GPT-4 Technical Report
- Qwen2.5-VL Technical Report
- M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
- Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
- A Comprehensive Survey of Knowledge-Based Vision Question Answering Systems: The Lifecycle of Knowledge in Visual Reasoning Task
- The Llama 3 Herd of Models
- Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
- CoRe-MMRAG: Cross-Source Knowledge Reconciliation for Multimodal RAG
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Towards General Continuous Memory for Vision-Language Models
- Qwen3 Technical Report
- BERTScore: Evaluating Text Generation with BERT
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering