Gen2-IC: Bridging Generative Models and Image Codecs through Latent Transport

arXiv:2606.21030 · eess.IV, cs.CV · Submitted 2026-06-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Gen2-IC: Bridging Generative Models and Image Codecs through Latent Transport".

Tom: Gen2-IC proposes FlowCodec, a streamlined framework for generative image compression that decouples efficient latent compression from one-step latent transport.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: Jane, I've been looking over the details of this paper now, "Gen2-IC: Bridging Generative Models and Image Codecs through Latent Transport," and it really seems to tackle a core issue in image compression. It proposes a new way to integrate those big generative models directly into codecs without needing tons of extra setup.

Jane: Oh, Tom, that's what caught my eye too; it sounds like they are trying to make the process much more direct for using these powerful generative priors in compression systems. The title itself suggests a bridge between two different worlds—generative models and image codecs—which is pretty interesting.

Lu: I think the key insight here, as I see it, is framing this as a latent transport problem where you're moving the noisy latent distribution towards the clean one using a pretrained prior to refine it in just one step. It moves us away from needing complex conditioning signals or auxiliary networks that we usually have to add on for these kinds of tasks.

Meng: From an engineering standpoint, I'm curious about how this decoupling works practically; if you're removing those auxiliary networks, what exactly replaces them to guide the transport? We need a solid mechanism there so it doesn't just become a black box operation.

Lalam: If I had to pick the most impactful aspect of this paper right now, it would be how it allows us to use models like MMDiT or Qwen-Image directly in ultra-low bitrate codecs without needing massive retraining for each specific compression task. This could drastically improve the accessibility of high-quality visual generation methods.

Tom: Exactly, Lalam; that flexibility to leverage existing strong priors is what makes this framework so appealing when we think about the future of image synthesis and storage. So, to recap, it’s a new framework called FlowCodec that uses two stages—latent compression and latent transport—to plug generative models into codecs simply by transporting the noisy latent toward the clean one in a single step.

Jane: That single-step refinement idea is what really simplifies things conceptually; instead of needing several iterative steps guided by external information, they are aiming for a much more streamlined decoding process. It sounds like they are prioritizing simplicity and efficiency in how these powerful models interact with the compression pipeline.

Title and authors: Lu: The paper highlights that this design removes the need for additional conditioning signals or auxiliary branches by leveraging stronger priors, which simplifies the entire structure significantly. This suggests that when you have a very good generative model, you don't need to complicate the codec architecture to make it work well at low bitrates.

Meng: That simplicity is attractive from a deployment perspective; less complexity means fewer moving parts and potentially faster encoding speeds, which is what we need for real-time applications. But I wonder how robust this single-path design is when the latent compression stage has already introduced significant information loss due to quantization and bottlenecks.

Lalam: The paper suggests that the stronger priors inherently make the decoder more tolerant to that information loss, which sounds like a really smart way to handle the trade-off between compression ratio and image quality in this new context. This points toward better semantic preservation overall.

Tom: Right, so we've established it’s about using those strong priors to simplify the architecture and enable flexible multi-bitrate control without extra overhead. The authors are essentially casting the generative model integration as a latent transport problem to achieve that clean mapping between data and latent distributions.

Jane: That framing as a transport problem really helps explain *why* they can decouple the stages; it gives them a mathematical structure for moving from the distorted noise distribution to the clean data distribution in one go. It connects the compression side and the generation side in a way that seems very elegant.

Lu: Building on that, I noticed they discuss how this allows for lightweight adaptation of massive generative models using techniques like LoRA without needing full retraining, which is a huge practical win for integrating these large models into production systems quickly.

Meng: That’s the engineering angle I’m focusing on; if we can adapt a model with just low-rank adapters, the cost of deploying this compression method isn't ballooning with every new generative prior we want to test out. That scalability is what matters for us.

Lalam: And from a cultural perspective, if image compression becomes this efficient and high-quality at very low bitrates, it could democratize access to incredibly detailed visual content in various applications, which is something I find really inspiring about the potential of this research.

Title and authors: Tom: So we’ve seen how they frame the core idea of FlowCodec—decoupling compression and transport—and how that leads to a method that supports flexible multi-bitrate control using just a refinement scalar, beta rate. That sounds like a practical feature for users and developers alike.

Jane: It really seems like the authors are showing us how to take the powerful generative knowledge we already have and apply it much more directly into the compression pipeline, rather than treating it as an afterthought that requires extra conditioning layers.

Lu: The paper emphasizes that this approach works across different generative priors, testing things like SD-two point one and FLUX.one-dev, which suggests the transport mechanism itself is quite general and doesn't require custom tuning for every single model we might want to use later on.

Meng: That generality is promising; it means we don’t have to reinvent the wheel for every new model release; we just plug in the prior and fine-tune a tiny adapter. It makes integrating state-of-the-art generative capabilities much more modular.

Lalam: I think what’s really compelling is that even when they test with complex models, like Qwen-Image, the noisy latent already provides enough conditioning signal for the transport to work well, meaning prompt conditioning doesn't have to be a heavy burden on top of this system.

Tom: So we’ve seen the technical setup—the two decoupled stages and the latent transport formulation—and how that translates into practical benefits like superior perceptual quality at ultra-low bitrates and flexible control via beta rate. The authors are really pushing for coupling generation and compression at the mechanism level, which is a significant design choice.

Jane: It sounds like the core contribution is showing that you can achieve state-of-the-art perceptual metrics in the ultra-low bitrate regime while keeping the architecture clean by relying on this transport mechanism instead of adding heavy conditioning networks.

Lu: The paper’s discussion on how a stronger decoder prior allows for a simpler codec that is more tolerant to information loss at low bitrates really drives home the intuition behind why this structure is effective.

Meng: From an implementation standpoint, the training strategy where Stage one optimizes rate-distortion objectives before moving to image-level objectives seems very sensible; it prevents instability during the initial latent mapping phase.

Title and authors: Lalam: And I think looking at the results, specifically how it performs on metrics like LPIPS and DISTS compared to established diffusion methods, shows a tangible improvement in what users perceive as quality.

Tom: So we’ve covered the core idea of Gen2-IC—the Latent Transport framework—and its practical implications for achieving high fidelity at very low bitrates using existing generative models with minimal architectural additions. We're heading toward the conclusion now, but I want to make sure we wrap up with some final thoughts on what this means for the industry.

Jane: It seems like the main implication is that we can start treating powerful generative models not just as tools for image *generation*, but as integral components of efficient image *compression* systems.

Lu: And I think the design offers a clear path forward by demonstrating how to leverage existing priors in a way that makes compression inherently more coupled with generation, which is a very deep structural idea.

Meng: Practically, this means we can expect compression techniques to become much more adaptable as we integrate ever-stronger generative AI models into our workflows without needing complete overhauls of the codec itself.

Lalam: I'm excited because if this framework proves robust across different priors, it sets a new standard for how we think about combining these powerful generative capabilities into practical, efficient applications.

Tom: Alright team, so that wraps up our discussion on "Gen2-IC: Bridging Generative Models and Image Codecs through Latent Transport." It’s clear this paper offers a streamlined framework that makes integrating large generative models into ultra-low-bitrate compression more feasible by framing it as a latent transport problem.

Jane: It was fascinating seeing how they managed to decouple the stages so cleanly, allowing the latent transport to do the heavy lifting of refinement in one step.

Lu: The coupling at the mechanism level is definitely a key takeaway that shows how generation and compression can be inherently linked rather than bolted together.

Meng: From an engineering viewpoint, this modular approach with lightweight LoRA adaptation for the prior seems like a very practical path for rapid deployment across different generative models.

Lalam: I think the potential here is really about making high-quality visual content accessible and usable in ways we haven't thought of before due to this new efficiency.

The paper's summary: Tom: So, we've seen how FlowCodec uses that two-stage setup to map clean latents to noisy ones and then transports them toward the clean image in one step using a pretrained prior, right?

Jane: Exactly! Think of it like this: the first stage compresses the image into a slightly fuzzy version efficiently, and then the second stage uses that powerful generative knowledge to instantly polish that fuzzy version back into something really high quality.

Lu: That's the beauty of decoupling; they’re not trying to force a single model to do both compression and refinement all at once, which would be incredibly inefficient. They treat it like a problem of moving data through a specific space where the generative prior knows exactly how to guide that movement.

Meng: From what I've seen in terms of implementation, the authors are being very smart by keeping the VAE encoder and decoder frozen while only training those latent codec parts and then using lightweight LoRA adapters for the transport step. That makes sense for rapid integration into existing pipelines without massive retraining cycles.

Lalam: This whole idea speaks to how we can start thinking about generative models as a resource that can be piped directly into compression workflows, which could really reshape how we think about storing and retrieving visual information across different applications.

Tom: Right, Lalam! The summary really boils down to this: they’ve created a unified framework where the generative model’s knowledge is leveraged in a very direct way—through latent transport—to achieve excellent perceptual quality even at extremely low bitrates.

Jane: And what's so exciting about that is the flexibility it gives us; instead of being locked into one specific compression setting, you can just tweak that refinement strength parameter to get different levels of quality for the same file size.

Lu: The generality they show with different priors like SD-two point one and Qwen-Image is huge; it means this technique isn't tied to one model anymore, it’s a general method for using strong generative knowledge in compression.

Meng: I'm thinking about the practical impact on deployment; if we can use LoRA adapters to tune these models for different bitrates quickly, it opens up possibilities for highly customized compression solutions tailored precisely to the needs of a specific platform.

Lalam: I think this moves us toward a future where image quality isn't just about maximizing bits per pixel, but about how much semantic information we can retain while keeping the file size tiny and the visual experience top-tier.

Tom: That’s a fantastic way to put it, Lalam; it shifts the focus from just technical metrics to the actual user experience of what they see on screen. So, we're looking at a system that is both technically sophisticated and incredibly practical for real-world use.

Jane: And remember, even though this paper focuses on combining generation and compression, it also shows how robust these latent transport mechanisms are across a wide variety of generative architectures.

Lu: Indeed, the way they structure the training—optimizing the codec first before introducing the image-level objectives—is a very disciplined approach that leads to stable results for this kind of complex coupling.

Meng: It’s interesting how they address prompt conditioning by showing that with a good noisy latent already present, you don't need heavy text guidance to get the refinement right, which simplifies the overall input pipeline.

Lalam: I see this as a cultural shift; if we can make high-fidelity visual content accessible and usable in such an efficient way, it could really democratize access to incredibly detailed visual storytelling across all media.

Tom: So, we’ve seen how FlowCodec uses that latent transport problem to bridge the gap between generative models and image codecs, offering a scalable path for ultra-low bitrate quality that's adaptable and powerful. What aspect of this architecture do you think will be the most influential for future research?

The paper's improvements: Tom: So, we've heard how FlowCodec uses that latent transport problem to bridge generative models and codecs in one go, but now we're talking about what this paper actually proposes as improvements for the technology itself.

Jane: Right! The authors are suggesting ways to make the whole system more adaptable and powerful by decoupling things even further, which I think is really smart for making these tools more versatile.

Lu: They emphasize that this framework isn't just about compression; it’s about creating a new way to integrate generative priors so they become an inherent part of the compression process rather than just an external layer we slap on top.

Meng: I see them proposing that the latent transport mechanism is designed to be reused across different generative backbones, meaning you can swap out the image generator later without rebuilding your entire codec pipeline from scratch. That modularity is key for a startup environment.

Lalam: This suggests a future where compression becomes less about squeezing data and more about intelligently guiding the latent space using rich generative knowledge, which could lead to much richer and more meaningful visual representations in AI applications.

Tom: Exactly! They’re pushing this idea that we can achieve better results by focusing on how the latent transport interacts with the prior, rather than just optimizing a single end-to-end model. It’s a different kind of optimization entirely.

Jane: And they suggest that the training strategy itself is improved by keeping those initial latent compression stages separate from the final image objectives, which makes the learning process much smoother and more stable for complex tasks like this.

Lu: That staged training approach addresses a major hurdle in coupling these domains—it tackles the rate-distortion problem first before worrying about image fidelity, which seems to be a very structured way to handle these trade-offs.

Meng: From an engineering viewpoint, that separation means we can test and tune the compression efficiency independently of how well the final image reconstruction looks initially, which cuts down on iterative debugging time significantly.

Lalam: I think this points toward a cultural shift in how we build these systems; instead of building monolithic solutions, we're moving toward building systems where different components—the compressor, the generator, and the transport—can be specialized and swapped out based on the specific task at hand.

Tom: It’s really about making these generative priors reusable assets rather than just fixed components in a single model structure. This opens up a lot of avenues for customization down the line.

Jane: And they also hint at how this structure might allow for more intuitive control over the final output quality through parameters that don't require deep understanding of diffusion sampling schedules.

Lu: The paper’s focus on the transport formulation itself is quite deep; it suggests a theoretical foundation for *why* this single-step refinement works, which is what makes it so compelling from a purely research standpoint.

Meng: I’m looking forward to seeing how scalable these LoRA adapters can truly be in practice, because if we can deploy this structure easily across many different generative models, the deployment barrier drops dramatically for everyone.

Lalam: For me, the most impactful vision is that this framework could allow us to create systems where visual content is synthesized and compressed on demand with unprecedented fidelity and efficiency.

Tom: So, to summarize these improvements: they’re proposing a more modular system that uses latent transport as the core mechanism for refinement, allowing for flexible adaptation and better stability during training. This really shows how the theoretical structure directly translates into practical engineering advantages.

Conclusion: Tom: So, we've walked through the technical details of Gen2-IC: Bridging Generative Models and Image Codecs through Latent Transport, and it’s clear that this framework is pushing us to think about how generative models and compression can work together directly.

Jane: That’s right! The main takeaway is seeing a really unified way to handle the process—using latent transport as the bridge between a compressed latent space and a high-fidelity generative prior in just one refined step.

Lu: I think the implication here is that we aren't limited to using compression techniques that are designed in isolation; we can start treating powerful generative models as integral components of image processing pipelines.

Meng: From an engineering standpoint, this means we have a much more flexible toolset for building new compression solutions because the core logic relies on a transport problem rather than rigid architectures.

Lalam: I feel that this could significantly improve how we think about visual data; it suggests that the latent structure itself is something we can leverage to create vastly more accessible and semantically rich visual experiences.

Tom: Absolutely, Lalam! The potential here is huge because it moves us toward systems where quality and efficiency are handled by a single, cohesive mechanism instead of separate, disconnected steps.

Jane: It really simplifies the concept for listeners by showing how we can use the existing knowledge in large generative models to enhance compression without needing to add extra conditioning networks that bog things down.

Lu: And the generality shown across different priors is what I find most exciting; it suggests this transport approach is a universal method applicable to many different types of generative models.

Meng: That modularity is what makes it attractive for production; if we can swap in a new generative prior with just a small adapter, we don't have to rebuild the whole system, which saves immense development time.

Lalam: I really see this as a cultural shift where the focus moves from simply achieving high numbers to creating systems that are inherently more intelligent and capable of handling complex visual tasks efficiently.

Tom: So, to wrap up, Gen2-IC: Bridging Generative Models and Image Codecs through Latent Transport gives us a practical blueprint for integrating massive generative capabilities into ultra-low bitrate compression using a streamlined latent transport mechanism.

Jane: It’s been a fascinating deep dive into how we can make these powerful models work in new ways, and I think the clarity on the decoupled stages really helps explain the mechanics.

Lu: The theoretical underpinning of framing this as a latent transport problem is definitely something worth exploring further for even more creative applications in generative modeling.

Meng: I'm just looking forward to seeing how these LoRA adapters perform when we start scaling up the number of different priors we can plug into this system.

Lalam: I think the future impact lies in creating visual tools that are inherently smarter, where compression and generation work together seamlessly to deliver top-tier results efficiently.

Tom: Well said, Lalam! That’s a lot of exciting potential for the industry right now. We’ll keep an eye on these developments as we explore other fascinating papers on arXiv next time.

Yinhuan Huang, Hao Cao Pu chen Wenqi Guo Zhijin Qin

Tsinghua University

eess.IV, cs.CV

Submitted: 2026-06-19

Updated: 2026-09-29

Comments: Substantially revised and expanded version of FlowCodec, renamed Gen2-IC; generalized methodological formulation and expanded technical and evaluation details

Code: https://github.com/black-forestlabs/flux

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 79/100

The gist: Gen2-IC proposes FlowCodec, a streamlined framework for generative image compression that decouples efficient latent compression from one-step latent transport.

Key concepts

FlowCodec
A new framework that uses two stages: latent compression followed by latent transport. It decouples efficient latent compression from one-step latent transport, allowing generative models to be plugged into codecs without needing complex auxiliary networks.
Latent Transport Problem
Framing the process as moving a noisy latent distribution toward a clean one using a pretrained prior in just one step. This approach replaces the need for complex conditioning signals or auxiliary networks, simplifying the structure of how generative models interact with compression systems.
Generative Priors
Strong knowledge from pre-trained generative models, such as MMDiT or Qwen-Image, is leveraged directly in the transport mechanism. This allows for high-quality refinement at ultra-low bitrates and makes the technique general across different generative architectures.
LoRA Adaptation
Using low-rank adapters to lightweight adapt massive generative models without full retraining. This modular approach allows for quick tuning of models for different bitrates, making integration into production systems scalable and cost-effective.

Terminology

Summary

Gen2-IC proposes FlowCodec, a streamlined framework for generative image compression that decouples efficient latent compression from one-step latent transport. This method is significant because it integrates powerful, pretrained large-scale generative priors directly into ultralow-bitrate codecs without requiring additional conditioning signals or auxiliary networks. By framing the process as a latent transport problem, FlowCodec allows for flexible multi-bitrate control and lightweight adaptation of massive generative models, enabling compression to benefit more directly from rapid advances in generative modeling.

FlowCodec Framework and Stages

FlowCodec decomposes the pipeline into two decoupled stages: (1) Latent Compression, which maps clean latents to bitrate-constrained noisy latents; and (2) Latent Transport, which leverages the pretrained prior to refine the noisy latents toward the clean ones in a single step. This design removes the need for additional conditioning signals and auxiliary networks while naturally supporting lightweight adaptation, one-step decoding, and flexible multi-bitrate control.

Latent Compression Mechanism

The first stage builds an efficient latent codec that supports multiple bitrates. It maps the clean latent representation to a noisy latent that is already close to the clean one. This stage is optimized independently to minimize the discrepancy between noisy and clean latents, providing an informative initialization for the subsequent stage. Key components include:

  1. Adopting depthwise convolutions for parameter efficiency and employing a 4-step autoregressive entropy model with quadtree partitioning for efficient entropy estimation.

  2. Introducing learnable per-channel scaling vectors to support multiple bitrates within a single model, avoiding separate models for different bitrate levels.

Latent Transport Mechanism

The second stage leverages the pretrained flow-matching prior (e.g., MMDiT) to refine the noisy latents toward clean ones in one step. This is achieved by injecting the noisy latents near the terminal step of the text-to-image sampling trajectory, realizing one-step refinement. The transport is formulated as:

(6)

ˆl1 = ˆl0 + βrate ∆t· vθ

where βrate controls the refinement strength. This mechanism reuses the pretrained generative backbone without architectural modification and activates its prior with only lightweight adaptation, enabling multi-bitrate decoding by modulating βrate.

Training Strategy

The training employs a decoupled, stage-wise procedure. The VAE encoder and decoder are kept frozen throughout training. The latent codec components (analysis/synthesis transforms and entropy model) are optimized first using a latent-space rate-distortion objective:

  1. Stage 1 optimizes the latent codec to minimize L(1) = λrate MSE(l, ˆl0) + R(y).

  2. Stage 2 decodes the noisy latent into a coarse reconstruction (xˆ0 = Ed(ˆl0)) and optimizes an image-level objective: L LC = λrate L1(x, xˆ0) + Lper(x, xˆ0) + Ladv(x, xˆ0) + R(y).

Finally, after Latent Compression is frozen, only lightweight LoRA adapters on the pretrained MMDiT backbone are finetuned to optimize the latent transport objective: LLT = L1(x, xˆ1) + Lper(x, xˆ1) + Ladv(x, xˆ1).

Experimental Validation and Performance

Experiments demonstrate that FlowCodec achieves superior perceptual quality in the ultra-low-bitrate regime. The Qwen-image variant significantly outperforms existing methods in terms of LPIPS and DISTS, while both variants deliver higher PSNR and clearly faster encoding than existing one-step diffusion-based methods. Ablation studies confirm that LoRA is crucial for effectively activating the generative prior, showing that performance is insensitive to the LoRA rank in our setting. Furthermore, the study shows that injecting noisy latents near the terminal step of the trajectory (near-terminal injection) consistently improves DISTS, supporting the near-terminal design. User studies confirm that FlowCodec-Qwen receives the highest top-1 preference, validating its perceptual gains for human observers.

Generality and Robustness

The framework demonstrates generality across different generative priors, including SD-2.1, SANA-1.5, FLUX.1-dev, and Qwen-Image-2512. Results show that stronger priors better preserve text, identity, and high-level visual content, suggesting that the design is robust for integrating ever-stronger generative models into compression systems. The analysis also confirms that prompt conditioning does not yield significant performance gains because the noisy latent already provides a strong conditioning signal for one-step refinement. The overall result is a streamlined framework that offers a "practical path for using modern generative priors in ultralow-bitrate image compression.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed the FlowCodec paper by Yinhuan Huang et al. The core contribution of FlowCodec is integrating powerful, large-scale generative priors (like Qwen-image or FLUX) into ultra-low-bitrate image compression via a decoupled Latent Compression and Latent Transport stage.

Here are the specific improvements to AI systems that can be made using the principles of FlowCodec:


The fundamental improvement is a new class of generative image codecs capable of achieving state-of-the-art perceptual quality at bitrates significantly lower than current diffusion-based methods, without relying on extensive auxiliary conditioning or retraining.

Specific improvements and capabilities include:

  1. A new generation of ultra-low bitrate image compression codecs that maintain high fidelity for complex visual elements (faces, textures, and text) by leveraging pretrained large-scale generative models as a single-step refinement mechanism.

  2. The ability to achieve significant bitrate savings (up to 73.7% relative to MS-ILLM) while simultaneously achieving superior perceptual quality metrics like LPIPS and DISTS in the ultra-low bitrate regime (below 0.05 bpp).

  3. The capability for flexible, multi-bitrate control using a single set of parameters (the refinement strength scalar, βrate), eliminating the need for separate models or complex bit-to-timestep mappings for different quality levels.

  4. Enhanced semantic preservation in compressed images, demonstrated by superior performance on metrics like OCR CER/WER and ArcFace face ROI similarity compared to traditional codecs and many diffusion baselines.

Specific operational improvements to the AI system:

  1. A new generative image compression pipeline that utilizes a frozen Variational Autoencoder (VAE) for latent extraction, followed by a learnable latent codec for compression, and finally refined by a pretrained flow-matching prior via one-step transport.

  2. The latent transport mechanism injects noisy latents near the terminal step of the generative trajectory to steer them toward the clean image manifold in a single refinement step (Eq. 6), effectively realizing one-step refinement.

  3. The system supports lightweight adaptation of large generative backbones (e.g., Qwen-image-2512 or FLUX) using low-rank LoRA adapters, requiring only a negligible fraction of the backbone's parameters (<0.54%), enabling rapid deployment and flexible multi-bitrate operation without substantial retraining costs.

  4. The latent codec component is designed to be optimized independently for rate-distortion objectives (Stage 1: Latent Compression) before image-level objectives are introduced (Stage 2), allowing for stable convergence during training.

  5. The system can be configured with different prompt modes (empty, generic, or compression-oriented prompts) to fine-tune the transport process, although the paper suggests that a strong noisy latent provides sufficient task alignment without explicit text guidance.

Abstract

Diffusion-based image compression has achieved strong perceptual quality at ultra-low bitrates. However, existing codecs are often tied to specific backbones and specialized components, making diverse, rapidly evolving generative models difficult to reuse. This raises a natural question: Can modern generative foundation models be connected to image compression through a simple and extensible interface? Two insights guide our design: stronger generative priors make a simpler codec interface viable, and generation and compression can be intrinsically linked through latent transport. We therefore propose Gen2-IC with two stages: (1) Latent Compression maps clean image latents to entropy-constrained latents; and (2) Latent Transport refines them with one near-terminal update based on the pretrained model. Gen2-IC requires neither auxiliary conditioning signals nor task-specific backbone modifications. With lightweight adaptation and no distillation, it supports fast encoding and one-step decoding across multiple bitrates. We validate Gen2-IC on SD-2.1, SANA-1.5, FLUX.1-dev, and Qwen-Image-2512, spanning U-Net and Transformer architectures as well as diffusion and flow-matching formulations. With stronger priors, Gen2-IC delivers gains below 0.05 bpp: the Qwen variant leads diffusion-based generative codecs in reconstruction fidelity (PSNR), perceptual similarity (LPIPS and DISTS), and recognizer-based semantic fidelity (OCR CER/WER and face-ROI similarity) across four benchmarks.

Sources

Related papers