RAU: Reference-based Anatomical Understanding with Vision Language Models
cs.CV, cs.AI
Submitted: 2025-09-26
Updated: 2026-09-08
Comments: ECCV 2026
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Qwen Technical Report
- R-LLaVA: Improving Med-VQA Understanding through Visual Region of Interest
- How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition
- A Survey on In-context Learning
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Label-Efficient Deep Learning in Medical Image Analysis: Challenges and Future Directions
- ECHOPulse: ECG controlled echocardio-grams video generation
- Less Could Be Better: Parameter-efficient Fine-tuning Advances Medical Vision Foundation Models
- DeepSeek-V3 Technical Report
- A Survey on Hallucination in Large Vision-Language Models
- SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data
- DINOv2: Learning Robust Visual Features without Supervision
- SAM 2: Segment Anything in Images and Videos
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
- Consensus, dissensus and synergy between clinicians and specialist foundation models in radiology report generation
- Foundation Models in Medical Imaging: A Review and Outlook
- SVQA-R1: Reinforcing Spatial Reasoning in MLLMs via View-Consistent Reward Optimization
- SAMed-2: Selective Memory Enhanced Medical Segment Anything Model
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models