When Does an Image Determine the Answer? Benchmarking Visual Answerability across Charts and Scenes
cs.CV
Submitted: 2026-09-28
Updated: 2026-09-28
Terminology
Sources
- Towards Causal VQA: Revealing and Reducing Spurious Correlations by Invariant and Covariant Semantic Editing
- VQA: Visual Question Answering
- Analyzing the Behavior of Visual Question Answering Models
- Don't Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering
- Neural Module Networks
- A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification
- Learn then Test: Calibrating Predictive Algorithms to Achieve Risk Control
- CLOSURE: Assessing Systematic Generalization of CLEVR Models
- Scene Text Visual Question Answering
- Querying Incomplete Numerical Data: Between Certain and Possible Answers
- Underspecification Presents Challenges for Credibility in Modern Machine Learning
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- BLINK: Multimodal Large Language Models Can See but Not Perceive
- Datasheets for Datasets
- Selective Classification for Deep Neural Networks
- SelectiveNet: A Deep Neural Network with an Integrated Reject Option
- Shortcut Learning in Deep Neural Networks
- Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
- HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
- Detecting and Preventing Hallucinations in Large Vision Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models