MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA
cs.CV, cs.AI
Submitted: 2026-08-27
Updated: 2026-08-27
Comments: Accepted by CVPR 2026
Code: https://github.com/NUST-Machine-Intelligence-Laboratory/MedFG
License: http://creativecommons.org/licenses/by/4.0/
The gist: Medical Visual Question Answering (Med-VQA) holds significant promise for clinical decision support, yet faces challenges due to limited annotated data and the high computational demands of existing
Terminology
Abstract
Medical Visual Question Answering (Med-VQA) holds significant promise for clinical decision support, yet faces challenges due to limited annotated data and the high computational demands of existing large vision-language models. We propose MedFG-VQA, a lightweight framework that leverages a memory bank to augment DCT-based low-frequency features and employs graph-enhanced cross-attention for effective visual-textual alignment. Specifically, our approach features two key components: Frequency-Memory Fusion (FMF), which enhances low-frequency features by retrieving from a learnable memory bank built on DCT decomposition, and Graph-Aware Cross-Attention (GACA), which aligns visual-textual features via cross-attention and refines them through graph-convolutional aggregation. To address data scarcity, we construct SynMed-VQA, a large-scale synthetic dataset comprising over 2 million question-answer pairs across 9 imaging modalities and 10 major organs, generated with GPT-4o. Extensive experiments on SynMed-VQA and three other standard biomedical VQA benchmarks demonstrate that MedFG-VQA achieves competitive or superior performance compared to much larger models while maintaining significantly lower computational costs, highlighting its efficiency and potential for clinical deployment.
Sources
- GPT-4 Technical Report
- Lung and Colon Cancer Histopathological Image Dataset (LC25000)
- HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
- EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
- LLaDA-MedV: Exploring Large Language Diffusion Models for Biomedical Image Understanding
- PathVQA: 30000+ Questions for Medical Visual Question Answering
- Gemma 3 Technical Report
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Qwen3 Technical Report
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models