From Alignment to Synthesis: Contrastive Volumetric Grounding for Text-to-CT Generation
cs.CV, cs.AI
Submitted: 2025-05-31
Updated: 2026-09-16
Comments: Accepted at BMVC 2026
Code: https://github.com/danielemolino/Text2CT
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Generating semantically controllable 3D CT volumes from radiology reports requires more than a rich text encoder, it requires vision-language alignment grounded in volumetric space.
Terminology
Abstract
Generating semantically controllable 3D CT volumes from radiology reports requires more than a rich text encoder, it requires vision-language alignment grounded in volumetric space. Existing Text-to-CT approaches condition generation on encoders pretrained with language only or 2D vision-language objectives, providing conditioning signals that are linguistically expressive but volumetrically blind. We argue this is a structural limitation: the quality of 3D vision-language alignment, not the richness of the text encoder, is the primary bottleneck for semantic controllability in volumetric diffusion models. To address this, we propose a generation-oriented 3D-CLIP encoder trained with structured hard negatives that operate exclusively at the text level. This design increases contrastive difficulty without any additional 3D memory cost, overcoming the small-batch constraints inherent to volumetric encoders. The resulting encoder conditions a fully end-to-end latent diffusion model that operates directly in 3D latent space, eliminating the spatial artifacts and cross-slice inconsistencies introduced by super-resolution pipelines. Through systematic ablations, we establish a clear empirical link between grounding quality and downstream generative controllability. Evaluated on CT-RATE across 18 pathological conditions, our method achieves state-of-the-art performance on both image fidelity and factual correctness, while requiring less inference time and GPU memory than all competing methods. Code is at https://github.com/danielemolino/Text2CT.
Sources
- Radiology Report Conditional 3D CT Generation with Multi Encoder Latent diffusion Model
- Med3D: Transfer Learning for 3D Medical Image Analysis
- Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography
- Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical Imaging
- Classifier-Free Diffusion Guidance
- ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission
- Auto-Encoding Variational Bayes
- Decoupled Weight Decay Regularization
- Representation Learning with Contrastive Predictive Coding
- Contrastive Learning with Hard Negative Samples
- Denoising Diffusion Implicit Models
- BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs
- CLIP in Medical Imaging: A Survey
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models