LiteMedCoT-VL: Parameter-Efficient Adaptation for Medical Visual Question Answering
cs.CV, cs.AI, q-bio.QM
Submitted: 2026-05-10
Updated: 2026-09-18
Comments: Accepted at NLPCC 2026 (The 15th CCF International Conference on Natural Language Processing and Chinese Computing), Springer proceedings. 17 pages, 5 figures
Code: https://github.com/R4nzer/LiteMedCoT-VL
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- A Comprehensive Survey on Knowledge Distillation
- Step-CoT: Stepwise Visual Chain-of-Thought for Medical Visual Question Answering
- CheXthought: A global multimodal dataset of clinical chain-of-thought reasoning and visual attention for chest X-ray interpretation
- Improving Medical VQA through Trajectory-Aware Process Supervision
- CARE: Towards Clinical Accountability in Multi-Modal Medical Reasoning with an Evidence-Grounded Agentic Framework
- Elicit and Enhance: Advancing Multimodal Reasoning in Medical Scenarios
- PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
- Qwen3-VL Technical Report
- TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation
- PathVQA: 30000+ Questions for Medical Visual Question Answering
- Knowledge Distillation Based on Transformed Teacher Matching
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- SmolVLM: Redefining small and efficient multimodal models
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- MIRAGE: The Illusion of Visual Understanding
- MedVR: Annotation-Free Medical Visual Reasoning via Agentic Reinforcement Learning
- Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
- Dual Causal Inference: Integrating Backdoor Adjustment and Instrumental Variable Learning for Medical VQA
- Beyond Accuracy: Evaluating Visual Grounding In Multimodal Medical Reasoning
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models