PerCoV2: Ultra-Low Bit-Rate Perceptual Image Compression via Query-Based 1D Multimodal Image Tokens

arXiv:2503.09368 · cs.CV, eess.IV · Submitted 2025-03-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "PerCoV2: Ultra-Low Bit-Rate Perceptual Image Compression via Query-Based 1D Multimodal Image Tokens".

Tom: The paper introduces PerCoV2,

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So we’re looking at the paper now called "PerCoV2: Ultra-Low Bit-Rate Perceptual Image Compression via Query-Based 1D Multimodal Image Tokens <ref:2503.09368#pg0,Ultra-Low Bit-Rate Perceptual Image Compression>." It sounds like they're proposing a system that uses Stable Diffusion three for image compression, specifically targeting situations where you have really tight bandwidth or storage limits.

Jane: Exactly, it’s focused on making images smaller without sacrificing the visual quality you need for streaming or mobile use right now. They built this new system on top of the Stable Diffusion three architecture <ref:2503.09368#pg0>.

Lu: What stands out is how they tackle the compression part by explicitly modeling that discrete hyper-latent image distribution, which they call their dedicated entropy model <ref:2503.09368#pg1>. That seems like a direct way to boost efficiency compared to just using standard coding methods.

Meng: From an engineering standpoint, that sounds complicated, but if it really makes the bits go further for a given quality level, that’s practical. We need to see how much real savings they can actually get in the field.

Tom: The paper points out they compared their approach against autoregressive methods like VAR and MaskGIT for entropy modeling, and their method seems to do better across the board on the MSCOCO-30k benchmark as well as the Kodak dataset <ref:2503.09368#pg1>.

Jane: They’re showing that PerCoV2 can achieve higher image fidelity at even lower bit-rates while still keeping the perceptual quality competitive with other strong models.

Lu: That’s a key finding because it confirms that better autoencoder reconstruction ability doesn't automatically mean better overall generation performance, which is something the paper notes in relation to other work <ref:2503.09368#pg2>.

Meng: So they’re not just tweaking the decoder; they’ve changed how the compression itself is learned by integrating this dedicated entropy model into their learning objective. That sounds like a solid architectural improvement for efficiency.

Tom: Right, and I saw them list some specific bit-rate improvements on page one, showing things like a six point six two times saving over proprietary LDM performance at zero point zero two zero three eight bpp for the kodim10 model <ref:2503.09368#pg1>.

Jane: It’s impressive that they show such a wide range of improvements, from very aggressive settings down to those lower bit-rates where the gains are most noticeable.

Lu: They also highlight that PerCoV2 features a hybrid generation mode specifically designed for further bit-rate savings, which means you can switch between high-fidelity generation and aggressive compression modes depending on what you need <ref:2503.09368#pg1>.

Meng: I wonder how flexible that switching is in practice. Can a user just toggle a setting and get those savings without needing to re-tune anything?

Tom: That’s what they’re aiming for, making it a more flexible deployment strategy where the model can adapt to different resource constraints <ref:2503.09368#pg1>.

Jane: And they've shown that their method is built solely on public components, which is always a big plus for open research and community adoption.

Lu: That’s a major point because building something on public parts makes it much more accessible than relying on proprietary models.

Meng: So if we look at the overall results, they are delivering more faithful reconstructions while preserving high perceptual quality compared to PerCo, MS-ILLM, DiffC and DiffEIC models <ref:2503.09368#pg2>.

Tom: That’s the comparison that really matters here—they outperformed those strong baselines on the large-scale MSCOCO-30k benchmark <ref:2503.09368#pg1>.

Jane: So, in short, PerCoV2 is presenting a novel and open ultra-low bitrate perceptual image compression system based on the Stable Diffusion three architecture that significantly extends prior research by improving entropy coding efficiency <ref:2503.09368#pg0,a novel and open ultra-low bitrate perceptual image compression system>.

Lu: And they did this by conducting a comprehensive comparison of recent autoregressive entropy modeling techniques like VAR and MaskGIT to show the benefits of their approach in the ultra-low bit-range <ref:2503.09368#pg2>.

Meng: Before we wrap up, I want to ask about where this actually lands for deployment. If someone is using a mobile device right now, what’s the real impact of these bit-rate savings?

Tom: Well, the main thing is that they achieve considerable improvements in entropy coding efficiency compared to previous methods like uniform coding <ref:2503.09368#pg2>, which translates directly into better compression ratios at those extreme bit-rates.

Jane: And the paper also shows improved semantic preservation metrics, specifically mIoU scores, which means the reconstructions align better with ground-truth labels when you use PerCoV2 over strong baselines like MS-ILLM <ref:2503.09368#pg1>.

The paper's summary: Tom: So, basically, PerCoV2 is taking Stable Diffusion three and turning it into something that can compress images down to incredibly tiny sizes while keeping the picture looking good for what you’re doing right now.

Jane: That's right. It’s not just about making a file smaller; it’s about finding a way to squeeze a lot of visual information into very few bits without losing the actual look of the scene.

Lu: The big thing they did was modeling that discrete hyper-latent image distribution explicitly, which is like giving the compression system a specific rulebook for how those latent images are structured.

Meng: I wonder what that means for us on the ground. If this compression method is really good at those extreme bitrates we’re talking about, does it mean AI images can actually be used in real-time mobile apps without massive data costs?

Tom: That's the practical angle there. The paper shows they achieved higher image fidelity even when the bit-rate is extremely low, like zero point zero three bits per pixel <ref:2503.09368#pg1>. That’s a huge jump compared to what we see in other compression methods right now.

Jane: It means for apps that need fast loading, like streaming video or mobile previews, we could get much better quality images using less bandwidth than we thought possible before.

Lu: And they proved that this works across different big datasets like MSCOCO-30k and Kodak <ref:2503.09368#pg0>. That’s not just a lab result on one small image; it holds up when you test it against complex, real-world scenarios.

Meng: So the caveat is they say it gets less effective if you push the bit-rate higher than that ultra-low range, which makes sense because they seem really tuned for those constrained settings.

Tom: Right. It’s a specialized tool. They aren't trying to beat every compression system at high quality; they're focused on doing the absolute best thing possible at those very low bandwidth limits where most of the problems happen.

Jane: And what they showed about semantic preservation—using metrics like mIoU scores—is important because it suggests that even when you compress aggressively, the essential stuff, like where objects are located, stays accurate.

Lu: That’s a strong point. It's not just blurring things up; the underlying structure is still there in a way that helps with understanding the scene.

Meng: I see what you mean. If we're using AI to understand environments—say for robotics or augmented reality—having that structural accuracy preserved at low bitrates could be really useful for on-device processing.

Tom: Exactly. It moves the goalposts a bit on what’s achievable in terms of compression efficiency versus perceptual quality, especially when you look at how they compared it to systems like PerCo and MS-ILLM.

Jane: So, the main implication is that we have a new way to think about image compression by integrating this explicit entropy modeling directly into the generative process itself.

Lu: And they’re opening up the use of Stable Diffusion three for these kinds of highly specialized, low-bitrate tasks, making it more accessible than those proprietary systems.

Meng: It’s interesting that they also mentioned a hybrid mode where you can switch between high-fidelity generation and this aggressive compression mode. That gives users flexibility depending on their exact needs.

Tom: That flexibility is key for deployment; you don't have to choose one setting if the situation changes, right?

Jane: It really does. So, moving forward, we need to think about how this level of efficiency can actually integrate into the larger ecosystem of generative AI tools and applications we use every day.

The paper's improvements: Tom: So, PerCoV2 isn't just about being a compression system; it’s about fundamentally changing how we use Stable Diffusion three for image tasks by adding this new entropy modeling layer on top.

Jane: It’s taking a powerful generative model and giving it a way to speak the language of extreme low bit-rates more efficiently than before.

Lu: They introduced these discrete entropy models, like MIM and VAR, which let the system predict what kind of image tokens are coming next in a structured way.

Meng: That structured prediction is what gives them that efficiency boost you mentioned earlier—it’s not just guessing the next bit; it’s modeling the distribution itself.

Tom: Exactly. They showed that by explicitly modeling those discrete distributions, they can achieve better compression ratios, specifically saving about thirteen point three percent over baseline methods at those super low settings.

Jane: That translates directly into smaller files for AI content, which is something everyone in the streaming and mobile space needs to hear about right now.

Lu: They also found that this works well when combined with a hybrid generation and compression mode, letting the model switch between high-quality output and aggressive compression modes depending on what you need at that moment.

Meng: That hybrid feature is interesting from a deployment standpoint; it means the system is adaptable to different resource constraints without needing a completely separate model for every single task.

Tom: Right. It’s not just one fixed setting; it’s a flexible strategy for balancing quality and size in real-time applications.

Jane: And they made sure to show that this improved reconstruction quality holds up across very different visual datasets, like MSCOCO-30k and the Kodak dataset <ref:2503.09368#pg0>.

Lu: That’s important because it means this isn't just a fluke on one type of image; it’s a more robust method for handling varied visual data.

Meng: So, we have a system that is both highly efficient at the edges and still maintains good perceptual quality on large benchmarks, which is exactly what we need to see in production environments.

Tom: It proves that you can build complex generative tools on top of existing architectures and layer new compression techniques to get specialized performance for specific use cases.

Jane: And they also focused on semantic preservation, showing that the reconstructions are better aligned with the actual labels when using PerCoV2 compared to other strong baselines like MS-ILLM <ref:2503.09368#pg1>.

Lu: That’s a big win because it shows that the compression isn't just making things look visually acceptable; it’s actually preserving the underlying structure needed for accurate object recognition and scene understanding.

Meng: I see what you mean. If we’re using AI to build three dee scenes or understand complex visual data, that structural accuracy is crucial for how well the AI performs its task.

Tom: So, this paper gives us a concrete way to improve the efficiency of generative models when bandwidth is severely limited, and it does so by deeply customizing the entropy coding process itself.

Jane: It’s a solid piece of research that shows how to combine cutting-edge generative models with better compression for real-world efficiency gains in bandwidth-limited scenarios.

Lu: We’re looking forward to seeing what they do next with those VAR and MaskGIT comparisons, which should give us more insight into the best way to handle those discrete distributions <ref:2503.09368#pg2>.

Conclusion: Tom: So we're wrapping up on PerCoV2: Ultra-Low Bit-Rate Perceptual Image Compression via Query-Based 1D Multimodal Image Tokens <ref:2503.09368#pg0,Ultra-Low Bit-Rate Perceptual Image Compression>. It’s a system that really pushes the boundaries on how we can compress images while keeping them looking good for constrained applications.

Jane: Right, it’s a system that uses Stable Diffusion three to create an efficient way to squeeze a lot of visual information into very few bits without losing the actual look of the scene.

Lu: The main implication is that we have a new way to think about image compression by integrating this explicit entropy modeling directly into the generative process itself, making it more open and accessible.

Meng: From an engineering view, this means we can expect smaller file sizes for AI-generated content in mobile apps down the line, provided these methods scale up well in terms of real-world usage.

Lalam: I think this work is important because it makes the visual world more accessible and less dependent on massive data pipelines for every application.

Tom: And they proved that this works across different big datasets like MSCOCO-30k, showing it’s a robust method for handling varied visual data while keeping perceptual quality high <ref:2503.09368#pg0>.

Jane: It really confirms that we can build powerful generative tools on top of existing architectures and layer new compression techniques to get specialized performance for specific use cases.

Lu: The future work they hinted at involves training those MIM and VAR models using standard cross-entropy loss, which should help refine their performance even further in those extreme bit-rate scenarios.

Meng: I wonder how practical it will be to deploy these complex flow matching objectives in a production environment without needing massive compute resources for the initial setup.

Tom: That’s a fair question, and they did address that by optimizing the objective in the latent space rather than pixel space for training <ref:2503.09368#pg1>.

Jane: They also use classifier-free guidance during inference to generate the final image reconstruction, which helps guide quality up without needing extra complex classifiers running alongside it.

Lu: It’s about building a system that is both efficient in training and effective during the actual generation process, which is a tough balancing act when dealing with these generative models.

Meng: So to wrap up on the technical side, they are using two types of masked image transformers—MIM and VAR—to explore discrete entropy modeling, but they’re demonstrating the benefits of their unified approach for compression and generation in the ultra-low bit-range.

Tom: And that brings us to the end of our discussion on PerCoV2: Ultra-Low Bit-Rate Perceptual Image Compression via Query-Based 1D Multimodal Image Tokens <ref:2503.09368#pg0,Ultra-Low Bit-Rate Perceptual Image Compression>. It’s a system that really pushes the boundaries on how we can compress images while keeping them looking good for constrained applications.

Jane: It's a solid piece of research that shows how to combine cutting-edge generative models with better entropy modeling for real-world efficiency gains in bandwidth-limited scenarios.

Lu: We’re looking forward to seeing what they do next with the VAR and MaskGIT comparisons, which should give us more insight into the best way to handle those discrete distributions <ref:2503.09368#pg2>.

Meng: For practical purposes, this means we can expect smaller file sizes for AI-generated content in mobile apps down the line, provided these methods scale up well.

Lalam: This work means that we can start thinking about visual media that is truly accessible and lightweight for everyday users across the board.

Nikolai Korber, Eduard Kromer, Andreas Siebert, Sascha Hauke, Daniel Mueller-Gritschneder, Bjorn Schuller

Technical University of Munich · University of Applied Sciences Landshut

cs.CV, eess.IV

Submitted: 2025-03-12

Updated: 2026-10-05

Comments: Major revision of the previous version. The initial draft corresponds to the PerCoV1++ variant described in Section 4 and illustrated in Figure 3. Code and pre-trained models will be released upon publication at https://github.com/nikolai10/PerCoV2

Code: https://github.com/Nikolai10/PerCoV2

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 90/100

The gist: The paper introduces PerCoV2, a novel and open ultra-low bitrate perceptual image compression system built upon Stable Diffusion 3 that enhances entropy coding efficiency by explicitly modeling the

Key concepts

Stable Diffusion 3 Architecture
PerCoV2 is fundamentally based on the Stable Diffusion 3 architecture, which includes components like a Latent Diffusion Model (LDM) encoder/decoder and various text encoders. This foundation provides the generative power and latent space structure necessary for image processing in this compression system.
Discrete Hyper-Latent Image Distribution Modeling
The system explicitly models the discrete distribution of hyper-latent images. This modeling is crucial for enhancing entropy coding efficiency, allowing PerCoV2 to better represent and compress the image data at very low bitrates by understanding how the latent features are distributed.
Hierarchical Masked Image Modeling (MIM/VAR)
PerCoV2 uses two autoregressive methods, Masked Image Model (MIM) and Visual Autoregressive Model (VAR), to model image formation. These models work either implicitly or explicitly across multiple scales, helping the system capture complex visual details for better compression.

Terminology

Summary

The paper introduces PerCoV2, a novel and open ultra-low bitrate perceptual image compression system built upon Stable Diffusion 3 that enhances entropy coding efficiency by explicitly modeling the discrete hyper-latent image distribution. This work matters because it extends prior research to the Stable Diffusion 3 ecosystem and demonstrates superior performance at extreme bit-rates, outperforming strong baselines on large benchmarks like MSCOCO-30k while maintaining competitive perceptual quality.

The gist: PerCoV2 achieves higher image fidelity at even lower bit-rates while maintaining competitive perceptual quality, see Figs. 2 and 5.<ref:2503.09368#pg2>

How it works

PerCoV2 is a novel and open ultra-low bitrate perceptual image compression system based on the Stable Diffusion 3 architecture [18] <ref:2503.09368#pg5>. It retains the core design principles of PerCo [11] but introduces two notable differences: i) replacing the proprietary LDM with an open alternative based on Stable Diffusion 3 [18] <ref:2503.09368#pg5>, and ii) enhancing entropy coding efficiency by explicitly modeling the discrete hyper-latent image distribution (Sec. 4.3) <ref:2503.09368#pg6>.

The system consists of several components:

  1. Stable Diffusion 3, which includes an LDM encoder and decoder, text encoders (CLIP-G/14 [52], CLIP-L/14 [52], and T5 XXL [53]), and a latent flow model [18, 43] <ref:2503.09368#pg7>.

  2. Feature extractors such as an image captioning model (e.g., BLIP 2 [39] or Molmo [15]) and a hyper-encoder <ref:2503.09368#pg7>.

  3. A discrete entropy model, denoted as VAR/MIM <ref:2503.09368#pg7>.

Encoding Process

During encoding, PerCoV2 extracts side information to better adapt the flow model for compression, represented as z = (zl, zg) where zl and zg correspond to local and global features respectively <ref:2503.09368#pg6>. The local features zl are vector-quantized (VQ) hyperlatent representations defined as zl = H(E(x)) <ref:2503.09368#pg6>, and the encoder E maps the input image x to a latent representation y of shape H/8 × W/8 × 16, which is then processed by the hyper-encoder H to yield zl with shape h × w × 320 <ref:2503.09368#pg6>. The global features zg correspond to image captions generated by a pre-trained large language model <ref:2503.09368#pg6>. Both zl and zg are then losslessly compressed using arithmetic coding and Lempel-Ziv coding <ref:2503.09368#pg6>.

Hierarchical Masked Image Modeling

PerCoV2 explores two types of masked image transformers for discrete entropy modeling: the masked image model (MIM) [12] <ref:2503.09368#pg6>, which models the image formation process in an implicit manner, and the visual autoregressive model (VAR) [62], which uses an explicit multi-scale representation <ref:2503.09368#pg6>. Both MIM and VAR are autoregressive methods that model the image formation process in either an implicit or explicit hierarchical manner <ref:2503.09368#pg6>.

The joint probability distribution is factorized as p(q) = Y K k=1 p(qk Ck), where q = (q1, q2,..., qK) denotes the sequence of token subsets/maps, and Ck represents the context used to predict qk <ref:2503.09368#pg6>. For MIM, Ck is the context derived from previously predicted token subsets <ref:2503.09368#pg6>, whereas for VAR, Ck corresponds to the previously generated token maps <ref:2503.09368#pg6>.

Optimization and Training

To train PerCoV2, a two-stage training protocol is used. In the first stage, PerCoV2 is optimized using the conditional flow matching objective Eq. (9), extended by z = (zl, zg) <ref:2503.09368#pg5>. This objective is formulated in the latent space of the auto-encoder rather than in the pixel space <ref:2503.09368#pg5>.

The second stage involves training MIM/VAR based on previously learned hyper-encoder representations using standard cross-entropy loss for optimization <ref:2503.09368#pg5>. During inference, classifier-free guidance [11, 27] is applied to generate the final image reconstruction <ref:2503.09368#pg5>.

Experimental Evaluation

PerCoV2 was empirically evaluated on the MSCOCO30k and Kodak datasets at a resolution of 512 × 512 <ref:2503.09368#pg6>. Performance is quantified using metrics such as FID [26] and KID [7] to quantify perception, PSNR, MS-SSIM [67], LPIPS [74] to quantify distortion, CLIP-score [25] to measure global alignment between reconstructed images and ground-truth captions, and mean intersection over union (mIoU) to assess semantic preservation <ref:2503.09368#pg6>.

The main results show that PerCoV2 considerably improves all metrics at ultralow to extreme bit-rates (0.003 − 0.03 bpp) while maintaining competitive perceptual quality <ref:2503.09368#pg8>. However, at higher bitrates, they become less effective, e.g., compared to PerCo (SD) [34] <ref:2503.09368#pg8>. The method consistently achieves more faithful reconstructions while preserving perceptual quality <ref:2503.09368#pg13> and outperforms strong baselines on the large-scale MSCOCO-30k benchmark <ref:2503.09368#pg6>.

REFERENCES

[1] Eirikur Agustsson, Michael Tschannen, Fabian Mentzer, Radu Timofte, and Luc Van Gool. Generative adversarial networks for extreme learned image compression. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019. 1, 3<ref:2503.09368#pg9>

[2] Michael Samuel Albergo and Eric Vanden-Eijnden. Building normalizing flows with stochastic interpolants. In The Eleventh International Conference on Learning Representations, 2023. 3, 4<ref:2503.09368#pg5>

[3] P. Astolfi, M. Careil, M. Hall, O. Manas, M. Muckley, J˜ Verbeek, A. R. Soriano, and M. Drozdzal. Consistencydiversity-realism pareto fronts of conditional image generative models. arXiv: 2406.10429, 2024. 2<ref:2503.09368#pg9>

[4] Tom Bachard, Tom Bordin, and Thomas Maugey. CoCliCo: Extremely low bitrate image compression based on CLIP semantic and tiny color map. In PCS 2024 - Picture Coding Symposium, 2024. 3<ref:2503.09368#pg12>

[5] Johannes Balle, Valero Laparra, and Eero P. Simoncelli. End-to-end optimized image compression. In International Conference on Learning Representations, 2017. 3<ref:2503.09368#pg6>

[6] Johannes Balle, David Minnen, Saurabh Singh, Sung Jin Hwang, and Nick Johnston. Variational image compression with a scale hyperprior. In International Conference on Learning Representations, 2018. 3<ref:2503.09368#pg6>

[7] Mikołaj Binkowski, Dougal J. Sutherland, Michael Arbel, ´and Arthur Gretton. Demystifying MMD GANs. In International Conference on Learning Representations, 2018. 6<ref:2503.09368#pg8>

[8] Yochai Blau and Tomer Michaeli. Rethinking Lossy Compression: The Rate-Distortion-Perception Tradeoff. In Proceedings of the 36th International Conference on Machine Learning, 2019. 1, 3<ref:2503.09368#pg9>

[9] Rishi Bommasani et al. On the opportunities and risks of foundation models. arXiv: 2108.07258, 2021. 1, 3<ref:2503.09368#pg5>

[10] Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. Cocostuff: Thing and stuff classes in context. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 6<ref:2503.09368#pg11>

[11] Marlene Careil, Matthew J. Muckley, Jakob Verbeek, and Stephane Lathuili´ ere. Towards image compression with perfect realism at ultra-low bitrates. In The Twelfth International Conference on Learning Representations, 2024. 1, 2, 3, 4, 5, 6, 8<ref:2503.09368#pg9>

[12] Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T. Freeman. Maskgit: Masked generative image transformer. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2, 3, 4, 5, 6<ref:2503.09368#pg6>

[13] Zhe Chen, Yuchen Duan, Wenhai Wang, Junjun He, Tong Lu, Jifeng Dai, and Yu Qiao. Vision transformer adapter for dense predictions. In The Eleventh International Conference on Learning Representations, 2023.

Improvements for AI systems

  1. Improved ultra-low bitrate perceptual image compression via PerCoV2, which achieves higher image fidelity at even lower bit-rates while maintaining competitive perceptual quality. This allows for significantly reduced bandwidth requirements in applications such as real-time streaming or mobile device storage.

  2. Enhanced entropy coding efficiency by explicitly modeling the discrete hyper-latent image distribution using a novel implicit hierarchical masked image model, which considerably improve[s] entropy coding efficiency compared to previous methods like uniform coding. This leads to better compression ratios at extreme bit-rates, such as achieving 13.33% savings over baseline for extreme-low settings.

  3. Hybrid generation and compression mode implementation, which allows for further bit-rate savings by utilizing both encoding and generative capabilities simultaneously. This system enables a more flexible deployment strategy where the model can switch between high-fidelity generation and aggressive compression modes to meet varying resource constraints.

  4. Improved semantic preservation metrics, as evidenced by the mIoU scores in Figure 4, allowing the AI system to produce reconstructions that better align with ground-truth labels when using PerCoV2 over strong baselines like MS-ILLM. This is crucial for tasks requiring accurate object localization and scene understanding.

Abstract

Despite recent progress in learned image compression, current image codecs still struggle to maintain realistic reconstructions at low bit-rates, often producing structured artifacts such as grid patterns or repetitive textures, even when trained with perceptual or adversarial losses. We introduce PerCoV2, an ultra-low bit-rate perceptual image compression system that unifies semantic tokenization, flow-based generation, and learned entropy modeling within a single framework. Building on the fully open flow-based SANA architecture, PerCoV2 introduces a novel resolution-adaptive 1D query-based tokenizer that produces compact semantic image tokens with a dual role in flow matching: providing a data-dependent reconstruction prior for initialization and a conditioning signal for flow-based refinement. By explicitly decoupling semantic representation from perceptual generation, our dual representation simplifies the flow-based learning objective, leading to more stable optimization and improved perceptual compression performance. PerCoV2 further introduces a dedicated 1D masked entropy model to improve rate efficiency and optional decoder-side multimodal enhancement via a vision-language model (Molmo) without increasing the transmitted bit budget. On MSCOCO-30k, PerCoV2 achieves state-of-the-art statistical fidelity, measured by FID and KID, across ultra-low and extreme bit-rates (0.0015-0.025 bpp). When trained solely on the general-purpose SA-1B dataset, PerCoV2 further demonstrates strong zero-shot generalization to widely adopted high-resolution benchmarks, including DIV2K and CLIC 2020, achieving competitive statistical fidelity with the current leading method, AEIC-ME. Finally, we introduce PerCoV2-distilled, a practical single-step variant derived from multi-step flow matching that accelerates decoding by 5.37x over PerCoV1, while preserving perceptual compression performance.

Sources

Related papers