What Makes a Good Medical Image Tokenizer? Rethinking Reconstruction and Generation in Medical Image Tokenization
cs.CV, cs.AI
Submitted: 2026-09-21
Updated: 2026-09-21
License: http://creativecommons.org/licenses/by/4.0/
The gist: Latent diffusion models now dominate medical image generation, and every such pipeline rests on a tokenizer that compresses images into the latent codes for image generation to operate on.
Terminology
Abstract
Latent diffusion models now dominate medical image generation, and every such pipeline rests on a tokenizer that compresses images into the latent codes for image generation to operate on. Thereby, the tokenizer choice bounds every downstream task from reconstruction fidelity and generation quality to the representations available for downstream analysis. Yet, medical imaging pipelines routinely utilize tokenizers from natural imaging on the hypothesis that their behavior carries over. However, this is an assumption never tested in the medical imaging regime, where datasets are orders of magnitude smaller and images exhibit far lower inter-sample variance. We present a systematic evaluation of medical image tokenizers evaluating thirty configurations across ten model families on twelve datasets at three compression factors, spanning reconstruction, generation, latent geometry, downstream classification, and memorization. We find that (1) performance on image reconstruction and generation strongly correlate, unlike prior reports on natural images; (2) modern tokenizers use nearly all of their codebook entries, but still leave most of the latent space unused; (3) training-set memorization is mild and is further suppressed by stronger latent space compression; and (4) discrete quantization can largely preserve downstream classification, with lookup-free schemes being the main exception.
Sources
- TumorFlow: Physics-Guided Longitudinal MRI Synthesis of Glioblastoma Growth
- RoentGen: Vision-Language Foundation Model for Chest X-ray Generation
- Latent Interpolation Learning Using Diffusion Models for Cardiac Volume Reconstruction
- Clinically Relevant Latent Space Embedding of Cancer Histopathology Slides through Variational Autoencoder Based Image Compression
- Unified Medical Image Tokenizer for Autoregressive Synthesis and Understanding
- Scalable next-scale autoregression for medical image generation across anatomical regions
- Making Reconstruction FID Predictive of Diffusion Generation FID
- Missing Fine Details in Images: Last Seen in High Frequencies
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models