Tokenizer-Generator Coupling in Medical Image Generation
cs.CV, cs.LG
Submitted: 2026-08-07
Updated: 2026-09-27
Code: https://github.com/liamchalcroft/medtokenizers
License: http://creativecommons.org/licenses/by/4.0/
The gist: Latent medical image generators usually treat the tokenizer as fixed preprocessing.
Terminology
Abstract
Latent medical image generators usually treat the tokenizer as fixed preprocessing. We test whether this separation is valid in a controlled ChestMNIST study at 64x64, crossing discrete tokenizers, generator families, and sampler settings under a shared latent grid, with continuous-latent reference cells. In this controlled setting, rankings depend jointly on the tokenizer, generator, and sampler: the best quantizer changes with the generator, and validation-based sampler selection changes the apparent generator ranking. We retrain the vocabulary-1024 interaction block at three seeds and the interaction survives (6 of 9 pairwise quantizer comparisons exceed three seed standard deviations), and we scope the wider single-seed grid accordingly. Reconstruction PSNR alone is not a reliable selection criterion; we instead introduce a generator-free statistic, neighbour-conditional predictive gain, that separates the quantizer families by downstream generation quality (rank-AUC 1.00) where reconstruction PSNR and marginal token entropy do not. On LFQ-1024, retuning D3PM and SE-D3PM (selected on a held-out validation split) moves them from default FID-192 0.44/0.41 to 0.09/0.10 at lower NFE, replicated across seeds; the continuous references were not given an equivalent sampler sweep. We report FID-192 as an internal ranking metric; it ranks consistently with standard FID-2048 (Spearman 0.80) and with a label-free classifier two-sample test (0.78). We interpret these results through a rate-distortion-modelability framing, where modelability is conditional on the generator, sampler, and inference budget. All experiments are at 64x64 on low-resolution medical-style images, unconditional, and evaluated with non-clinical FID-based metrics, and we scope every claim to that setting. Code: https://github.com/liamchalcroft/medtokenizers and https://github.com/liamchalcroft/medlatents.
Sources
- How Faithful is your Synthetic Data? Sample-level Metrics for Evaluating and Auditing Generative Models
- Factorized Visual Tokenization and Generation
- CheXGenBench: A Unified Benchmark For Fidelity, Privacy and Utility of Synthetic Chest Radiographs
- Bayesian Flow Networks
- Fr'echet Radiomic Distance (FRD): A Versatile Metric for Comparing Medical Imaging Datasets
- VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
- Unified Medical Image Tokenizer for Autoregressive Synthesis and Understanding
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations
- Generative AI for Medical Imaging: extending the MONAI Framework
- Evaluating and Improving the Effectiveness of Synthetic Chest X-Rays for Medical Image Analysis
- Image Tokenizer Needs Post-Training
- When Worse is Better: Navigating the compression-generation tradeoff in visual tokenization
- Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
- Recent Advances in Autoencoder-Based Representation Learning
- MedVAE: Efficient Automated Interpretation of Medical Images with Large-Scale Generalizable Autoencoders
- Feature Extraction for Generative Medical Imaging Evaluation: New Evidence Against an Evolving Trend
- TokBench: Evaluating Your Visual Tokenizer before Visual Generation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models