HDR Video Generation via Latent Alignment with Logarithmic Encoding

arXiv:2604.11788 · cs.CV · Submitted 2026-04-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "HDR Video Generation via Latent Alignment with Logarithmic Encoding".

Jane: HDR generation from SDR input is challenging for generative models because HDR data's linear space and heavy-tailed distributions mismatch their training data,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, we’re looking at the paper "HDR Video Generation via Latent Alignment with Logarithmic Encoding" today, and we're starting by talking about what that title actually tells us about the work.

Jane: It really points to a specific technique they used—the logarithmic encoding—as the key mechanism for making this HDR video generation possible.

Lu: That encoding is essentially their bridge; it’s what lets them treat the complex, unbounded nature of HDR imagery as something that fits into a structure existing AI already understands.

Meng: From an engineering perspective, I’m thinking about how much effort goes into designing that specific mapping function and ensuring it doesn't introduce unforeseen artifacts when scaling up.

Lalam: What excites me is the implication for democratizing high-quality visual creation; if this approach is effective, it means anyone with a standard SDR camera could produce stunning HDR video without needing specialized, massive datasets.

Tom: I totally agree on the accessibility angle, Lalam; and beyond just making things easier to use, this approach fundamentally changes how we think about adapting powerful AI systems for new data regimes.

Jane: It’s like they're providing a standardized language for HDR data that existing models can read without needing a complete overhaul of their internal workings.

Lu: And that is where the core insight lies: observing that this logarithmic mapping places the HDR imagery into a distribution that is naturally aligned with the latent space of these models.

Meng: But I still have my questions about how stable this alignment holds up when you’re pushing it to generate something really complex like cinematic footage; the engineering reality of those KL divergences is something we need to consider.

Lalam: If they can prove it works well across out-of-distribution benchmarks, it opens the door for applications in everything from enhanced photography to immersive virtual reality experiences where realistic lighting is crucial.

Tom: So we're looking at a method that uses existing model knowledge and a clever mathematical transformation to unlock high-fidelity HDR generation without needing a massive retraining effort.

Jane: That’s the core idea, Tom—leveraging what the AI already knows instead of trying to teach it everything new about HDR lighting from zero.

Lu: And then they add that training strategy involving camera-mimicking degradations and exposure shifts, which is what really forces the model to recover those missing details intelligently during the fine-tuning phase.

Meng: That targeted corruption sounds like a smart way to train for robustness, but I’m curious if that extensive augmentation pipeline significantly slows down the training iteration compared to simpler data enhancements.

Lalam: The implication for culture is huge because this could mean high-quality, physically accurate visual content becomes much more readily available, which can really elevate the standard of digital media we consume and create.

Tom: So while the alignment technique is elegant, I’m really focused on how they manage that practical pipeline for real-time performance during inference.

Jane: It seems like a very clever way to handle the distribution mismatch by treating HDR content as an already known format for the AI, which is a huge conceptual step forward.

The paper's summary: Tom: Now that we’ve talked about the title and authors, let’s get into the actual substance of this paper, "HDR Video Generation via Latent Alignment with Logarithmic Encoding." They explain exactly how this system works in simple terms.

Jane: They are showing that HDR generation isn't as hard as it looks if you use a specific encoding to align the input data with the model’s existing understanding.

Lu: Basically, they demonstrate that because of the logarithmic encoding, the HDR imagery gets mapped into a distribution that is naturally compatible with the latent space of a pretrained Variational Autoencoder.

Meng: So, it's not about training a new encoder for HDR; it’s about finding a transformation to treat HDR content as familiar SDR data.

Lalam: If this works, it means we can bypass the massive data collection hurdles that usually stop us from creating high-qualityHDR video from standard inputs.

Tom: Exactly, Lalam; they're showing that we can leverage what the AI already knows instead of trying to teach it everything new about HDR lighting from zero.

Jane: It’s a very practical demonstration of how representation learning can be adapted to solve real-world physical problems like HDR synthesis.

Lu: They also integrate a differentiable transformation specifically to map those infinite scene-linear radiance values into the VAE’s expected input range, which they then test rigorously with pixel and latent space divergence measurements.

Meng: That rigorous testing sounds necessary before you try to scale this up; we need to know exactly how much computational overhead that alignment process adds to the overall generation time.

Lalam: If they can achieve these high perceptual metrics, it really validates the entire approach and suggests a future where AI-generated visual media can meet professional standards without needing incredibly complex manual input.

Tom: So we’re seeing a refinement process here—taking the initial alignment and layering on specific, realistic training challenges to really push the quality boundary in this paper.

Jane: It seems like they are moving from just getting a technically aligned latent space to actually training that space to produce visually rich, physically accurate output.

The paper's improvements: Tom: Let’s look at the specific steps they took to improve the results in this study, "HDR Video Generation via Latent Alignment with Logarithmic Encoding." These are the key modifications they made to get better quality video.

Jane: They focused on introducing specific training tricks on top of the core alignment mechanism, showing that just aligning the space wasn't enough to get high-quality results on its own.

Lu: They suggest applying a deliberate corruption pipeline—things like contrast clipping, compression artifacts, and selective blurring—to the SDR reference video itself to force the model to learn how to fill in those missing details.

Meng: I see the value in that targeted corruption; it pushes the model past simple pixel matching toward inferring physically plausible content from its learned priors, which is a strong step for realism.

Lalam: If they can achieve superior detail recovery through these methods, it means the resulting HDR videos won't just look bright; they’ll have that rich texture and depth that makes them truly professional-looking.

Tom: It’s definitely about making sure the model learns to be resilient against real-world capture issues rather than just being a perfect copier of the input.

Jane: And they also introduce applying exposure shifts jointly to both streams during training, which helps the AI become robust across a wider range of brightness levels while still maintaining physical correspondence.

Lu: That joint approach is important because it links the mathematical alignment directly to the physics of lighting, ensuring that whatever they generate looks like it could actually exist in a real-world scene.

Meng: From an engineering standpoint, I need to know if implementing this whole augmentation pipeline adds significant computational overhead to the training process compared to standard data augmentation methods.

Lalam: If they can achieve these high perceptual metrics, it really validates the entire approach and suggests a future where AI-generated visual media can meet professional standards without needing incredibly complex manual input.

Tom: So we’ve seen how this paper on "HDR Video Generation via Latent Alignment with Logarithmic Encoding" uses a clever mapping to align HDR data with existing model latent spaces, and the team is excited about its potential for efficiency and accessibility.

Jane: It’s a very warm conclusion, Tom; this method treats HDR content as familiar data by using a logarithmic encoding to fit it into a structure existing AI already understands.

Conclusion: Tom: So, we’ve covered the alignment, the training strategy involving degradations and why those specific training steps matter for detail recovery in this paper before we wrap up our discussion on "HDR Video Generation via Latent Alignment with Logarithmic Encoding."

Jane: We’ve also discussed how the LogC3 encoding maps linear radiance into a manageable latent space for existing models, which is a very neat conceptual step forward in this research.

Lu: The entire framework of HDR Video Generation via Latent Alignment with Logarithmic Encoding is built on the idea that we can adapt strong priors to solve a distribution mismatch.

Meng: I'm still focused on the practical deployment: if we have to run this full augmentation pipeline during inference, we need to make sure that doesn't introduce unacceptable latency for real-time use cases.

Lalam: This work really shows that sophisticated generative techniques can be simplified by smartly leveraging what we already have in place; it’s a great example of intelligent adaptation for improving the culture of visual content creation.

Tom: That’s the essence of what we were discussing today with "HDR Video Generation via Latent Alignment with Logarithmic Encoding," showing how simple representation alignment can unlock the latent capabilities of pretrained models.

Jane: It’s a method that treats HDR content as familiar data by mapping it into a distribution the model already understands through that specific logarithmic encoding.

Lu: The paper shows how versatile representation learning can be when you find the right mathematical encoding to bridge those different data domains.

Meng: My focus remains on the engineering hurdles for deployment; we need to see if this elegant alignment translates into a fast and reliable system for actual use.

Lalam: This work really suggests that sophisticated generative techniques can be simplified by smartly leveraging what we already have in place; it’s a great example of intelligent adaptation for improving the culture of visual content creation.

Lightricks Co. · Gear Productions Co. · Tel Aviv University

cs.CV

Submitted: 2026-04-13

Updated: 2026-10-07

Importance score: 82/100

The gist: HDR generation from SDR input is challenging for generative models because HDR data's linear space and heavy-tailed distributions mismatch their training data, but this work demonstrates that

Key concepts

LogC3 Encoding
This is a specific logarithmic compression transform used to map unbounded scene-linear radiance values into a bounded range suitable for the pretrained Variational Autoencoder (VAE). It acts as the primary alignment mechanism, transforming raw HDR data into a distribution that closely matches what the model was originally trained on, making it easier for the model to process.
Latent Alignment
The core idea is to force the high-dynamic-range (HDR) data manifold to align with the latent space learned by a pretrained generative model. This alignment is achieved through both pixel-space and latent-space divergence minimization, ensuring that the HDR content occupies a region where the model has learned robust representations, enabling successful adaptation.
Camera-Mimicking Degradations
During training, specific corruptions like contrast clipping and compression artifacts are deliberately applied only to the SDR reference video. This forces the model to learn how to reconstruct details from incomplete or degraded information rather than relying solely on direct pixel matching, which improves its ability to infer missing content in HDR scenes.
DiT Conditioning
The Diffusion Transformer (DiT) backbone is conditioned using the AVControl framework, which efficiently incorporates the SDR reference video into the generation process. This allows the model to leverage its learned temporal and structural priors from the video diffusion model while adapting to the specific lighting conditions of an SDR input.

Terminology

Summary

HDR generation from SDR input is challenging for generative models because HDR data's linear space and heavy-tailed distributions mismatch their training data, but this work demonstrates that high-quality HDR video can be achieved by leveraging pretrained video models through a simple representation alignment and targeted training strategy.

The gist

HDR generation can be achieved in a much simpler way by leveraging the strong visual priors already captured by pretrained generative models, specifically observing that a logarithmic encoding widely used in cinematic pipelines maps HDR imagery into a distribution that is naturally aligned with the latent space of these models, enabling direct adaptation via lightweight fine-tuning without retraining an encoder.

Methodology and Core Idea

The central idea of LumiVid is to align the HDR manifold with the model’s latent distribution via a fixed, camera-inspired LogC3 encoding. This transformation maps unbounded scene-linear radiance into the specific range for which the pretrained Variational Autoencoder (VAE) was originally optimized, effectively treating HDR content as familiar SDR data. To resolve the statistical mismatch between scene-linear radiance and the VAE’s learned SDR manifold, a differentiable transformation T is integrated to map unbounded radiance x ∈ [0,∞) into the VAE’s expected input range [−1, 1]. This alignment is confirmed by minimizing Kullback–Leibler (KL) divergence across two stages: pixel-space divergence (KLpx) and latent-space divergence (KLlat).

Training Strategy for Detail Recovery

To recover details not directly observable in the input, a key training strategy based on camera-mimicking degradations is introduced. This involves applying augmentations such as contrast clipping, compression artifacts, and selective blurring to the SDR reference video only. This deliberate corruption of extreme luminance regions prevents the model from relying on direct pixel reconstruction and instead encourages it to infer missing content from its learned priors. Furthermore, exposure shifts are applied jointly to both SDR and HDR streams to teach robustness across diverse brightness levels while preserving physical correspondence.

Inference Pipeline

The inference pipeline for LumiVid is designed for efficiency, requiring only a minimal adaptation mechanism. The process consists of three components: (1) a LogC3 compression transform that maps HDR values into the VAE’s expected input range, (2) the AVControl framework [4] for efficient conditioning of the Diffusion Transformer (DiT) on the SDR reference, and (3) a training pipeline incorporating realistic SDR degradation. At inference, an SDR input is VAE-encoded and processed through the LoRA-conditioned DiT. The resulting latents are then VAE-decoded and decompressed via the Inverse LogC3 transform to recover the scene-linear radiance as a float16 EXR file.

Results and Evaluation

The framework was evaluated on out-of-distribution benchmarks, including ARRI Cinema Footage and UPIQ, comparing it against state-of-the-art baselines like X2HDR image diffusion and HDRTVNet deterministic CNN reconstruction. LumiVid demonstrated superior performance across metrics: on UPIQ, it achieved 30.05 dB PU21-PSNR and a JOD of 8.22, substantially outperforming baselines. Crucially, for video benchmarks on ARRI footage, LumiVid was the only generative method with both high quality and temporal coherence, achieving a JOD of 7.86 and an F2F-PSNR of 45.63 dB, highlighting its ability to inherit temporal coherence from the native video diffusion backbone. Ablation studies confirmed that LogC3 provided the best alignment (lowest KL divergence) for both pixel and latent spaces, and that the full augmentation pipeline was essential for achieving peak perceptual metrics.

Conclusion

LumiVid successfully generates temporally coherent HDR video from a single SDR input by aligning the HDR manifold with a pretrained model’s latent distribution via LogC3 encoding and forcing reconstruction through camera-mimicking degradations. This approach unlocks the latent capabilities of pretrained models without requiring retraining, achieving high-quality float16 HDR video with minimal adaptation. Future directions include scaling training data, incorporating perceptual HDR metrics as training objectives, and exploring text-to-HDR-video generation in a single pass.

Improvements for AI systems

Based on the provided research paper, here are specific improvements that can be made to current AI systems, focusing on leveraging the core insights of LumiVid:

  1. Enhance Generative Video Synthesis with High Dynamic Range (HDR) Fidelity:

  2. Enable Direct Adaptation of Pretrained Models for New Data Regimes:

  3. Improve Reconstruction Quality in Extreme Lighting and Low-Data Environments:


  1. Enhance Generative Video Synthesis with High Dynamic Range (HDR) Fidelity:

The improved AI system can generate photorealistic, physically accurate video content directly in the HDR domain from standard SDR inputs. This goes beyond simple tone mapping by synthesizing details that are physically plausible but not directly observable in the input (e.g., recovering fine textures in deep shadows or accurately reproducing specular highlights).

  1. Enable Direct Adaptation of Pretrained Models for New Data Regimes:

The improved system can take existing, highly capable video diffusion models (which are typically trained on SDR data) and adapt them to handle HDR content with minimal retraining. By leveraging a fixed, camera-inspired logarithmic encoding (LogC3), the model treats HDR as a familiar representation within its existing latent space, allowing for high-quality adaptation via lightweight fine-tuning (LoRA adapters) rather than requiring the costly training of entirely new encoders or representations.

  1. Improve Reconstruction Quality in Extreme Lighting and Low-Data Environments:

The system can robustly reconstruct high dynamic range details even when the input SDR reference is degraded by realistic camera artifacts (like MP4 compression, contrast clipping, or selective blurring). By incorporating a training strategy based on these camera-mimicking degradations and utilizing an AVControl conditioning mechanism, the model is forced to infer missing radiance from its learned visual priors. This results in superior detail recovery compared to methods that rely solely on direct pixel reconstruction.

Sources

Related papers