VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation

arXiv:2505.13439 · cs.CV, cs.AI, cs.LG · Submitted 2025-05-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation".

Jane: Autoregressive (AR) models for image generation rely critically on visual tokenizers (VT), and this paper introduces VTBench,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, let’s talk about the title and who wrote this piece. The paper is titled "VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation," and it lists Huawei Lin, Tong Geng, Zhaozhuo Xu, Weijie Zhao from Rochester Institute of Technology and Stevens Institute of Technology.

Jane: That’s a solid team behind it; having researchers from different institutions usually means a broad perspective on the problem. The authors are trying to establish a standard for how we measure these visual tokenizers in image generation pipelines.

Lu: The title itself is very direct about what they are doing: evaluating VTs for AR image generation, which gets right to the heart of the current research bottleneck.

Meng: I’m looking at the authors and thinking about the scope; it suggests they didn't just look at one specific type of tokenizer but tried to cover a wide variety of them in their testing.

Lalam: It’s interesting that they focused on isolating the tokenizer rather than just evaluating the final image, which shows a commitment to deep analysis.

The paper's summary: Tom: Moving on to what this paper is actually summarizing, it basically says that current discrete VTs are struggling compared to continuous variational autoencoders because they often result in distorted reconstructions and lose fine-grained textures and text integrity.

Jane: That’s a strong statement, Tom. It sounds like the core issue is that the way we turn continuous image data into discrete tokens isn't preserving all the necessary visual information for high-quality outputs.

Lu: They point out that existing benchmarks are flawed because they evaluate overall generation quality without separating what the tokenizer contributes to that result, which makes it hard to diagnose why things aren't working well.

Meng: So, if we look at the summary, the paper is basically arguing that we need a more focused way to test these components rather than just checking the final output score.

Lalam: I agree; this focus on isolating the component is exactly what’s needed when trying to build more robust and reliable AI systems.

The paper's improvements: Tom: Now, the paper proposes VTBench as their solution, which is a comprehensive benchmark designed to systematically evaluate VTs across three main tasks: Image Reconstruction, Detail Preservation, and Text Preservation.

Jane: That’s a multi-faceted approach; they aren't just checking one thing; they are testing how well the tokenizer handles different kinds of visual information under specific conditions.

Lu: The scenarios they designed are quite thorough, covering everything from imageNet inputs to high resolution and varying resolutions, which shows a good attempt at covering real-world usage.

Meng: I see the structure of VTBench is very systematic because it tests these three distinct areas—reconstruction, detail retention, and text preservation—which should give us a clear picture of where the weaknesses lie.

Lalam: The idea of using specific metrics for each task, like Character Error Rate for text and PSNR or SSIM for reconstruction, makes this framework much more useful than just one general quality score.

Conclusion: Tom: To wrap up, the paper suggests that continuous VAEs produce superior visual representations because they manage spatial structure and semantic detail better than discrete VTs, which is a big finding in their VTBench evaluation.

Jane: And they emphasize that for future AR models to keep up with things like GPT-4o, we really need a visual tokenizer that is resolution-flexible, semantically robust, and reusable across different inputs <ref:2505.13439#pg1>.

Lu: The implication here is a clear direction: the next generation of VTs needs to be built not just for one fixed size but for flexibility and deep semantic understanding.

Meng: From an engineering standpoint, this means our design goals should shift toward architectures that can handle those variable resolutions without needing massive retraining for every single input size.

Lalam: I think the most impactful part is realizing that we need a reusable VT that understands both the visual appearance and the underlying semantic context to handle complex text accurately.

Tom: So, in closing, "VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation" gives us a very clear roadmap for where research needs to go concerning tokenization quality. That’s what we’ve been discussing today.

Huawei Lin, Tong Geng, Zhaozhuo Xu, Weijie Zhao

Rochester Institute of Technology · University of Rochester · Stevens Institute of Technology

cs.CV, cs.AI, cs.LG

Submitted: 2025-05-19

Updated: 2026-10-01

Code: https://github.com/huawei-lin/VTBench

Importance score: 82/100

The gist: Autoregressive (AR) models for image generation rely critically on visual tokenizers (VT), and this paper introduces VTBench, a comprehensive benchmark designed to systematically evaluate VTs across

Key concepts

Visual Tokenizer (VT)
A component within an AR model that converts images into a discrete sequence of tokens. These tokens are the fundamental units the model uses to process visual information, and their quality directly impacts the final image generation.
Image Reconstruction
The task of testing how accurately a VT can rebuild an original image from its tokenized representation. This assesses the VT's ability to capture overall spatial structure and semantic content when converting tokens back into pixels.
Detail Preservation
Measuring a VT's success in retaining high-frequency visual information, such as fine textures, edges, and small objects. This task specifically tests how well the tokenizer handles the intricate details of an image.
Text Preservation
Evaluating a VT's ability to accurately reproduce text content under different conditions like complex writing or multiple languages. This ensures that the tokenizer does not corrupt or distort textual information during processing.

Terminology

Summary

Autoregressive (AR) models for image generation rely critically on visual tokenizers (VT), and this paper introduces VTBench, a comprehensive benchmark designed to systematically evaluate VTs across three core tasks—Image Reconstruction, Detail Preservation, and Text Preservation—to diagnose the performance gap between discrete VTs and continuous variational autoencoders. The findings reveal that while continuous VAEs produce superior visual representations by retaining spatial structure and semantic detail, existing discrete VTs often lead to distorted reconstructions with loss of fine-grained textures and failures in preserving text integrity.

The Problem Addressed

Current benchmarks focus on end-to-end generation quality without isolating the VT performance, leaving critical issues unaddressed. The paper highlights three main problems: (1) Lack of VT-Specific Evaluation, as the performance of the VT often determines the upper bound of AR model quality; (2) Benchmark Misalignment, where existing benchmarks evaluate overall image generation rather than isolating the tokenizer's contribution; and (3) Inadequate Evaluation Metrics, as common metrics like FID are insufficient to capture fine-grained failures such as high-frequency detail loss or incorrect text reconstruction. The core issue is that current visual tokenizers often fail to preserve fine-grained details and semantic integrity during the quantization process.

VTBench Framework

VTBench is a comprehensive benchmark systematically evaluating VTs across three core tasks: (1) Image Reconstruction, (2) Detail Preservation, and (3) Text Preservation. This framework provides a multi-faceted framework for assessing visual tokenizers by covering diverse evaluation aspects, including high-resolution inputs, multilingual text scenarios, and varying-resolution conditions. The evaluation metrics are designed to stress different aspects of tokenization quality. For Image Reconstruction, metrics include PSNR, SSIM, LPIPS (Learned Perceptual Image Patch Similarity), and FID.

Evaluation Scenarios for Task 1: Image Reconstruction

Task 1 evaluates the fundamental ability of a VT to reconstruct an image from its tokenized representation across three settings: (1) ImageNet (model-specific input size), (2) High Resolution (1024 × 1024 inputs), and (3) Varying Resolution. The paper notes that most VTs are limited to model-specific input sizes and fail to generalize to arbitrary resolutions, unlike continuous VAEs which naturally support flexible image dimensions.

Evaluation Scenarios for Task 2: Detail Preservation

Task 2 focuses on measuring how well VTs retain high-frequency visual information crucial for perceptual fidelity, testing textures, facial features, edges, and small objects. This task assesses how well VTs retain high-frequency information using a dataset of patterned and texture-rich images. Results show that while continuous VAEs lead in all metrics, discrete VTs struggle to preserve local textures and structural integrity.

Evaluation Scenarios for Task 3: Text Preservation

Task 3 evaluates how well VTs preserve text content under varying complexity and linguistic diversity across three scenarios: (1) Movie Posters with clean, short English text (Easy), (2) ArXiv Abstracts containing dense, long-form academic writing (Hard), and (3) Multilingual Text rendered in non-Latin scripts including Chinese, Hindi, Japanese, and Korean. For text preservation metrics, the paper utilizes Optical Character Recognition (OCR) via Gemma 3 to compute Character Error Rate (CER) and Word Error Rate (WER).

Key Findings and GPT-4o Insights

The experiments reveal that continuous VAEs produce superior visual representations compared to discrete VTs, particularly in retaining spatial structure and semantic detail. Furthermore, the analysis of GPT-4o image generation suggests it may employ an autoregressive backbone. The paper hypothesizes that GPT-4o might use a residual next-scale VAE (RVAE) [41, 14] or a diffusion-based encoder-decoder [49], suggesting its VT must be capable of encoding not only fine-grained appearance but also semantic structure and spatial context. The paper concludes that the need for a resolution-flexible, semantically robust, and reusable VT is urgent to keep pace with modern LLMs.

Contributions

The main contributions include: (1) Introducing VTBench, a high-quality benchmark designed specifically for evaluating VTs in AR image generation; (2) Designing three tasks—Image Reconstruction, Detail Preservation, and Text Preservation—to provide a multi-faceted framework; (3) Conducting extensive experiments on VTs used in SOTA AR models to uncover key findings regarding the limitations of discrete tokenization; and (4) Providing an open-source codebase and dataset to foster further research. The paper also provides additional qualitative results illustrating reconstruction quality, detail retention, and text preservation across various VTs.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed the VTBench paper, which systematically evaluates Visual Tokenizers (VTs) in Autoregressive (AR) image generation models. The core finding is that discrete VTs significantly underperform continuous VAEs in reconstruction fidelity, detail preservation, and text accuracy.

Based on this scientific evidence, here are the specific improvements I would implement across the AI systems:


  1. Acknowledge and Integrate a Continuous Tokenization Backbone (VAEs) for High-Fidelity Synthesis:

  2. Develop Resolution-Agnostic Tokenization Mechanisms for Generalization:

  3. Implement Task-Specific Metrics to Diagnose Tokenizer Failure Modes:

  4. Design Semantic/Linguistic Alignment Layers in the VT for Text Preservation:

  5. Acknowledge and Integrate a Continuous Tokenization Backbone (VAEs) for High-Fidelity Synthesis:

The primary improvement is shifting away from purely discrete tokenization schemes (like standard VQ-VAE) toward continuous latent spaces, similar to those used in diffusion models (e.g., SD3.5L or FLUX.1).

  • This involves replacing discrete quantization steps with a continuous latent space representation that allows the AR model to operate on high-dimensional, smooth features rather than rigid index lookups.

  • The improved AI system can achieve significantly higher PSNR and SSIM scores in the Image Reconstruction task (Task 1), leading to images with superior overall visual fidelity.

  1. Develop Resolution-Agnostic Tokenization Mechanisms for Generalization:

The current limitations show that most discrete VTs are constrained by fixed input sizes, failing on High Resolution and Varying Resolution inputs.

  • Implement architectures like the Residual Next-Scale VAE (RVAE) or advanced BSQ variants that inherently support hierarchical tokenization across multiple spatial scales.

  • The improved AI system will be capable of synthesizing and reconstructing images with arbitrary resolutions (e.g., 1024x1024 or even larger), ensuring robustness in real-world applications where input dimensions are not predefined.

  1. Implement Task-Specific Metrics to Diagnose Tokenizer Failure Modes:

Instead of relying solely on end-to-end metrics like FID, the system should be evaluated using the VTBench framework—specifically isolating performance across Reconstruction, Detail Preservation, and Text Preservation tasks.

  • By measuring Character Error Rate (CER) and Word Error Rate (WER) in Task 3 (Text Preservation), researchers can pinpoint exactly where a tokenizer fails—whether it's losing fine textual structure in dense academic layouts or failing to preserve multilingual scripts.

  • By analyzing LPIPS and SSIM scores in Task 2 (Detail Preservation), we can diagnose whether the quantization process is causing loss of high-frequency textures or structural integrity.

  • The improved AI system will provide a diagnostic report indicating whether its current VT design is weak in spatial structure, fine-grained textures, or symbolic fidelity.

  1. Design Semantic/Linguistic Alignment Layers in the VT for Text Preservation:

To address the failure of discrete VTs to handle text accurately, the tokenizer must be augmented to prioritize symbolic preservation.

  • This involves incorporating mechanisms (similar to those hypothesized for GPT-4o) that allow the tokenization process to encode semantic relationships between visual tokens and linguistic priors.

  • The improved AI system can perform high-accuracy Optical Character Recognition (OCR) and subsequent text editing/insertion tasks, even on complex, dense, or multilingual text inputs, leading to accurate document understanding and multimodal reasoning capabilities.

The overall improved AI system will be a highly robust Autoregressive Image Generation model that exhibits:

  • Near-diffusion level visual fidelity.

  • Universal resolution support (from small thumbnails to high-resolution outputs).

  • Superior accuracy in reconstructing complex text, including multilingual and densely laid out academic documents.

Sources

Related papers