Resolution Meets Reduction: Efficient Visual Context for 3D Radiology Report Generation
Jonathan Suprijadi, Raphael Stock, Moritz Langenberg, David Zimmerer, Kim-Celine Kahl, Stefan Denner, Yannick Kirchhoff, Karol Gotkowski, Maximilian Rokuss, Jeremias Traub, Tassilo Wald, Constantin Ulrich, Klaus Maier-Hein
cs.CV, cs.AI
Submitted: 2026-08-09
Updated: 2026-08-11
Code: https://github.com/salaniz/pycocoevalcap
License: http://creativecommons.org/licenses/by/4.0/
The gist: Vision-language models offer a promising path toward automating radiology report generation, but applying them to full 3D CT volumes poses substantial computational challenges.
Terminology
Abstract
Vision-language models offer a promising path toward automating radiology report generation, but applying them to full 3D CT volumes poses substantial computational challenges. Modern foundation vision encoders (VEs) can produce tens of thousands of vision tokens per scan, making the visual sequence passed to the large language model (LLM) a primary computational bottleneck. Vision-to-language projectors can compress this sequence to reduce computation, but may discard clinically relevant detail; conversely, effective compression can accommodate higher-resolution inputs while keeping the downstream token count fixed. How this vision-token budget should be allocated across input field of view, spatial resolution, and vision-to-language projection therefore remains an open design question. We systematically evaluate four heterogeneous VEs (CNN- and ViT-based), five token-reducing projectors at up to 64x compression alongside a non-reducing MLP projector baseline, and five instruction-tuned LLMs (1.7B--4B) on two large-scale CT report datasets (CT-RATE and Merlin). At matched LLM token budgets, anatomy-guided region of interest cropping is the most consistent strategy, improving clinical macro F1 in 19 of 20 settings by +3.7 points on average for the 3D ViT Primus encoder and +1.1 for the slice-based 2D ViT Curia encoder. Increasing input resolution further is strongly projector-dependent: the PerceiverResampler, paired with higher-resolution Curia features, yields the strongest configuration in the resolution study on both datasets. Our best configurations achieve state-of-the-art clinical macro F1 on the test sets, reaching 49.5 on CT-RATE and 49.0 on Merlin. Code and models will be published upon publication.
Sources
- MAIRA-2: Grounded Radiology Report Generation
- Scaling medical imaging report generation with multimodal reinforcement learning
- Comprehensive language-image pre-training for 3D medical image understanding
- M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models
- MedGemma 1.5 Technical Report
- Jolia: Concept-Level Vision-Language Alignment for 3D CT Contrastive Learning
- Curia: A Multi-Modal Foundation Model for Radiology
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference
- gpt-oss-120b & gpt-oss-20b Model Card
- Beyond the Embedding Bottleneck: Adaptive Retrieval-Augmented 3D CT Report Generation
- Qwen3-VL Technical Report
- Vision Foundation Models for Computed Tomography
- Qwen3 Technical Report
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Gemma 3 Technical Report
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Decoupled Weight Decay Regularization
- Curriculum-Driven 3D CT Report Generation via Language-Free Visual Grafting and Zone-Constrained Compression
- U-VLM: Hierarchical Vision Language Modeling for Report Generation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models