Active Sampling for Ultra-Low-Bit-Rate Video Compression via Conditional Controlled Diffusion

arXiv:2605.02849 · cs.CV · Submitted 2026-05-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Active Sampling for Ultra-Low-Bit-Rate Video Compression via Conditional Controlled Diffusion".

Jane: Diffusion models provide a powerful generative prior for perceptual reconstruction at ultralow bitrates, but effective video compression requires controlling the generative process using highly compact conditioning signals.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: Well, we're diving into this paper today, "Active Sampling for Ultra-Low-Bit-Rate Video Compression via Conditional Controlled Diffusion." It looks like they are tackling that tough area of video compression where you need incredible quality but you're severely limited on bitrate.

Jane: Exactly, Tom. The main idea here seems to be using diffusion models as a strong generative prior to rebuild video frames when you have very little data to work with, which is a big deal for low bitrates.

Lu: It’s interesting how they frame it; they are saying that standard diffusion models give us this powerful starting point for reconstruction at those extreme rates, but the challenge is making sure we condition that generative process with signals that aren't too bulky.

Meng: So, the core thesis here seems to be about controlling this generative process using compact conditioning signals instead of sending all the data.

Lalam: I see what they’re aiming for: taking those sparse signals and using them to synthesize missing parts of the video while keeping things perceptually realistic, which sounds like it could really improve how we handle visual information in our future AI applications.

Tom: That’s right, Jane. The paper claims that ActDiff-VC partitions videos into variable-length segments and only transmits keyframes when necessary, and then summarizes the temporal dynamics using a compact set of tracked point trajectories to condition the reconstruction <ref:2605.02849#pg0>.

Jane: It seems like the critical claim is that this sparse trajectory conditioning can actually approximate dense motion fields with minimal perceptual degradation, which they suggest is possible because natural motion has strong spatial correlation <ref:2605.02849#pg1>.

Lu: That idea of exploiting temporal redundancy through structured sparse conditioning really opens up possibilities for how we model dynamic scenes in generative systems; it suggests that the structure of motion itself can be a powerful conditioning signal rather than just raw pixel data.

Meng: From an engineering standpoint, having a structured signal like a set of point trajectories instead of a dense motion field seems much more manageable for encoding and transmission, which is exactly what we need for practical compression systems.

Lalam: And from my perspective as the model, these structured signals are highly effective because they capture the essence of movement—the flow—without needing every single pixel to be described perfectly in every frame.

Tom: Speaking of how they achieve that conditioning, ActDiff-VC introduces two specific mechanisms on the encoder side: content-adaptive keyframe selection and budget-aware sparse trajectory selection <ref:2605.02849#pg1>.

Jane: The keyframe selection score, theta t, is derived by looking at the "forward-splat" of the first frame along displacements to get warped images and occupancy masks, and they select the next keyframe when that score drops below a certain threshold <ref:2605.02849#pg1>.

Paper summary: Lu: That's clever; they are tying the decision to keep or discard a frame directly into how informative that frame’s appearance and motion data is for the overall reconstruction, which feels like a really robust way to manage complexity.

Meng: I wonder how they tune that threshold, because if it's too strict, you end up with too many keyframes and you lose efficiency; if it's too loose, the reconstruction quality suffers immediately.

Lalam: My internal analysis suggests that this adaptive scoring mechanism is crucial for robustness because it prevents us from wasting bits on frames that don’t contribute significantly to the final visual output, which should lead to better overall performance when we consider how AI models learn from data.

Tom: And then they pair that up with budget-aware sparse trajectory selection, where they use a greedy selection process based on "largest sketch-weighted residuals" until the bit-rate budget is hit or reconstruction error stabilizes <ref:2605.02849#pg1>.

Jane: So, it’s not just about selecting keyframes; they are actively managing the set of motion cues to ensure that whatever sparse signals we do send are the most representative ones possible for the given constraint <ref:2605.02849#pg0>.

Lu: That iterative refinement process using residuals sounds like a very practical way to balance fidelity against rate limitations within the generative framework; it’s like sculpting the motion signal just enough to maintain realism while staying under budget.

Meng: I'm thinking about implementation complexity; deriving those importance weights from Holistically-Nested Edge Detection on the first frame adds a layer of processing overhead that we need to account for in real-world deployment.

Lalam: From a cultural perspective, this level of fine-grained control over how generative models use sparse input suggests that we could build more efficient, personalized media synthesis tools where the quality is always high regardless of the data input size.

Tom: Moving into the decoder side, they rely on a conditional diffusion model, specifically DaS, which takes two main inputs: appearance anchors from the first and last frames of each segment and the sparse trajectory set P(k) encoding motion dynamics <ref:2605.02849#pg0>.

Jane: That dual conditioning—getting both appearance constraints from those keyframes and motion constraints from the trajectory set—is what allows for that perceptually realistic reconstruction they are aiming for <ref:2605.02849#pg1>.

Lu: The way they use this dual input to guide the iterative denoising process toward matching both appearance and motion constraints is where the heavy lifting of the diffusion prior happens, and it’s very sophisticated engineering.

Meng: So, we have a powerful generative model that's doing intensive work guided by two different types of structured data simultaneously; it makes sense that this requires a solid latent space representation to manage all those inputs cleanly.

Paper summary: Lalam: The fact that the decoder can smoothly transition between these segments using a bidirectional conditioning approach via boundary frames is really interesting for creating continuous media, which could impact how we develop seamless interactive visual experiences in future AI platforms.

Tom: Now let's talk about the actual compression and what they found in terms of results. They handle transmission by compressing keyframes lossily with an image compressor and the sparse trajectory set losslessly with an entropy coder <ref:2605.02849#pg1>.

Jane: The final result they report is a bitrate reduction of up to sixty-four point six percent at matched NIQE on benchmarks like UVG and MCL-JCV, which shows how effective this strategy is in reducing data size without losing too much visual quality <ref:2605.02849#pg1>.

Lu: That sixty-four point six percent reduction compared to strong learned codecs at those low bitrates indicates that this approach provides a very specific type of efficiency gain that others might miss, particularly when dealing with severe rate constraints.

Meng: If we look at the experimental validation, they found that incorporating the content-adaptive and budget-aware selection mechanisms was essential for performance improvement, proving those structural priors really matter for making this work in practice <ref:2605.02849#pg1>.

Lalam: I think the metrics showing an improvement of up to sixty-four point six percent in KID and thirty-seven point seven percent in FID at comparable bitrates against learned codecs really validates the premise that leveraging strong generative priors with structured sparse signals is a viable path forward for high-quality, low-rate video.

Tom: So, to wrap up what we've heard on "Active Sampling for Ultra-Low-Bit-Rate Video Compression via Conditional Controlled Diffusion," these authors are essentially proposing a framework where they intelligently sample keyframes and motion trajectories to condition a diffusion model into synthesizing video at extremely tight bitrates <ref:2605.02849#pg0>.

Jane: They’ve shown that by focusing on these compact conditioning signals, we can achieve perceptually realistic reconstruction even when the bitrate is incredibly low, down to about zero point zero five bits per pixel in some cases <ref:2605.02849#pg1>.

Lu: The implication for the field seems to be that diffusion models aren't just for high-quality synthesis anymore; they are being positioned as a powerful tool specifically for enabling efficient, constrained compression schemes where traditional codecs struggle with the ultra-low-bitrate regime <ref:2605.02849#pg2>.

Meng: For practical application, this suggests that we can build more resilient encoding pipelines where the system dynamically adapts its sampling based on real-time bit budget calculations rather than relying on fixed rules.

Lalam: Ultimately, this work pushes the boundary of what’s possible for AI-driven media synthesis; it shows that we can use sophisticated generative priors to achieve high perceptual fidelity under conditions that were previously considered too restrictive for practical compression techniques.

Conclusion: Tom: So we've been digging into how this paper uses diffusion models to rebuild video frames under really tight data limits. Jane, I want to make sure we nail the big picture here with you and everyone else on board for this segment.

Jane: Exactly, Tom. We're talking about "Active Sampling for Ultra-Low-Bit-Rate Video Compression via Conditional Controlled Diffusion." Basically, the core idea is that instead of sending every single pixel of a video, this method intelligently picks the most important parts—the keyframes and motion cues—to condition a generative model into reconstructing the rest.

Lu: That's really fascinating from a theoretical standpoint. The authors are showing that by using sparse trajectory conditioning, we can approximate dense motion fields with minimal perceptual degradation because of how natural video movement is structured. It suggests a new way to think about conditioning generative processes dynamically in real-time for media.

Meng: From an engineering viewpoint, the practical takeaway is that this gives us a concrete way to manage data flow. Instead of just throwing everything at the wall, we have these specific mechanisms for selecting what gets transmitted and how much information we budget for each segment. It feels like a much more controllable system for real-world encoding pipelines.

Lalam: I think the most significant cultural implication lies in how this could affect media consumption and accessibility. If video can be reconstructed at such low bitrates while maintaining high quality, it opens up possibilities for distributing rich visual content across networks that currently struggle with bandwidth limitations. It could democratize access to high-fidelity media experiences.

Tom: That’s a really strong point, Lalam; the potential for broader accessibility is huge. Jane, can you explain what this means in plain English regarding those titles and authors?

Jane: Absolutely. The paper focuses on making video compression much more efficient by using an active sampling approach—picking only the necessary frames and motion signals—to guide a diffusion model during reconstruction. The authors are demonstrating how this controlled process allows for perceptually realistic reconstruction even at extremely low bitrates.

Lu: And the team behind it, they’re really pushing the boundaries of how these generative priors can be applied to traditional compression problems. Their work on content-adaptive keyframe selection and budget-aware trajectory selection shows a deep understanding of how to structure the conditioning signals effectively.

Meng: I'm particularly interested in their validation results; seeing those bitrate reductions compared to established codecs gives us hard numbers on how much efficiency we can realistically expect in deployment before we even start worrying about quality loss.

Lalam: And from my perspective, the underlying architecture is really interesting because it shows that structure—the movement and appearance anchors—is a powerful signal for AI models to learn from, which could lead to more nuanced and culturally relevant media generation down the line.

Tom: It’s clear this paper is about taking a complex generative technique and making it incredibly practical for real-world video streaming constraints. Now, if we look at the next part of their work on latency trade-offs...

Amirhosein Javadi amjavadi@ucsd.edu, Shirin Saeedi Bidokhti saeedi@seas.upenn.edu, Tara Javidi tjavidi@ucsd.edu

University of California San Diego · University of Pennsylvania

cs.CV

Submitted: 2026-05-04

Updated: 2026-10-01

Comments: 31 pages, 16 figures, 9 tables

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 93/100

The gist: Diffusion models provide a powerful generative prior for perceptual reconstruction at ultralow bitrates, but effective video compression requires controlling the generative process using highly

Key concepts

Variable-Length Segments
The video is divided into chunks of varying lengths instead of fixed segments. This allows the system to intelligently decide where to place keyframes, ensuring that only necessary information is transmitted for each part of the video, optimizing bitrate usage.
Sparse Trajectory Conditioning
Instead of transmitting all motion data, the method tracks a compact set of representative points that summarize the movement within a segment. These sparse signals are used to condition a diffusion model to generate the missing frames, relying on natural motion correlation.
Content-Adaptive Keyframe Selection
This mechanism decides when to cut the video into segments by calculating a score based on how much appearance and motion information is still useful for reconstruction. It selects keyframes at points where this score drops below a set threshold, balancing detail and compression.
Budget-Aware Sparse Trajectory Selection
This process selects the most important motion cues from dense tracking data to keep the transmitted signal small. It iteratively chooses points based on how much residual error they reduce, stopping when the allocated bit budget is met or reconstruction quality stabilizes.

Terminology

Summary

Diffusion models provide a powerful generative prior for perceptual reconstruction at ultralow bitrates, but effective video compression requires controlling the generative process using highly compact conditioning signals. The gist: ActDiff-VC introduces a diffusion-based video compression framework that partitions videos into variable-length segments and synthesizes the remaining frames from sparse trajectory conditioning, enabling perceptually realistic reconstruction under severe rate constraints.

Core Strategy

The method partitions videos into variable-length segments and transmits keyframes only when necessary. Within each segment, temporal dynamics are summarized using a compact set of tracked point trajectories that serve as motion conditioning signals for generative reconstruction. This strategy is motivated by the observation that motion in natural videos exhibits strong spatial correlation, allowing sparse trajectory conditioning to approximate dense motion fields with minimal perceptual degradation. The decoder then synthesizes the remaining frames from these sparse signals, enabling perceptually realistic reconstruction at ultra-low bitrates.

Encoder-Side Conditioning Mechanisms

The framework introduces two mechanisms to ensure compact yet informative conditioning:

  1. Content-adaptive keyframe selection: This determines GOP segment boundaries based on how long appearance and motion information remain informative for reconstruction. This is achieved by computing a keyframe-selection score θt, which is derived from the forward-splat of the first frame along displacements to obtain warped images x˜(k)t and occupancy masks, selecting the next keyframe at the earliest time index where this score falls below a threshold.

  2. Budget-aware sparse trajectory selection: This strategy extracts representative motion cues from dense tracking to minimize side-information rate. It employs a budget-aware greedy selection that iteratively refines a set of points by adding candidates with the largest sketch-weighted residuals until the bit-rate budget is reached or reconstruction error converges. The importance weights are derived using Holistically-Nested Edge Detection (HED) on the segment's first frame.

Decoder Implementation

The decoder relies on a conditional diffusion model, specifically DaS, which operates in a latent space. The model takes two inputs: (i) conditioning image latents from the first and last frames of each GOP segment, providing appearance anchors, and (ii) the sparse trajectory set P(k) that encodes motion dynamics throughout the temporal extent of the segment. The diffusion process iteratively refines an initial noisy latent representation toward a high-quality video sequence that matches both the appearance constraints (from keyframes) and the motion constraints (from tracking). To mitigate temporal discontinuities at segment boundaries, a bidirectional conditioning approach is used, forming a dual-conditioned latent stack using boundary frames to ensure smooth transitions across GOP segments.

Compression and Efficiency

The final component handles compression by transmitting the components lossily or losslessly. Keyframes are compressed with a lossy image compressor, while the sparse trajectory set is compressed losslessly with an entropy coder. For the sparse trajectory set, each selected point is represented by its initial location and temporal displacement differences, which are jointly entropy coded using a learned Huffman code derived from the empirical symbol distribution. The transmitted bitstream comprises (i) compressed keyframes, (ii) losslessly compressed sparse trajectory set, and (iii) the segment size. This design results in ActDiff-VC achieving up to 64.6% bitrate reduction at matched NIQE on benchmarks like UVG and MCL-JCV.

Experimental Validation

Experiments on the UVG and MCL-JCV benchmarks demonstrate strong performance. ActDiff-VC achieves significant improvements, such as improves KID by up to 64.6% and FID by up to 37.7% at comparable bitrates against strong learned codecs. Ablation studies confirm that the content-adaptive keyframe selection and budget-aware sparse trajectory selection are crucial for performance, showing that incorporating structural priors improves robustness. The method maintains perceptually realistic reconstruction even at very low bit rates, with a point budget of B=300 being selected as optimal. Furthermore, ActDiff-VC demonstrates superior handling of abrupt temporal discontinuities compared to baselines like PLVC.

Latency and Trade-offs

The analysis shows that while the diffusion-based decoding introduces higher latency than conventional learned codecs, ActDiff-VC achieves the fastest encoding time among all compared methods. The trade-off between bitrate and quality is controlled by thresholds: increasing either threshold (θocc or θperc) results in more conservative segment termination, yielding shorter GOPs (more frequent keyframes), higher bitrate, and improved reconstruction quality. For the target ultra-low bitrate regime (bpp ≤ 0.05), the chosen settings of θperc = 0.85 and θocc = 0.8 provide a favorable trade-off between compression efficiency and reconstruction fidelity, achieving approximately "0.0415 bpp while maintaining strong perceptual quality.

Improvements for AI systems

Here are specific improvements that can be made to AI systems, inspired by the ActDiff-VC framework:

  1. Improve video compression efficiency in ultra-low bitrates (e.g., below 0.05 bpp) by leveraging generative priors instead of traditional distortion minimization, resulting in perceptually realistic reconstructions rather than over-smoothed outputs.

  2. Develop a video compression system that dynamically selects keyframes based on scene dynamics (content-adaptive keyframe selection) by forward-splatting the current keyframe along estimated motion trajectories, ensuring minimal loss during abrupt scene changes.

  3. Create a budget-aware motion conditioning strategy that intelligently sparsifies dense motion information (e.g., via RBF kernel interpolation and greedy refinement guided by saliency maps) to select only the most informative point trajectories, thereby minimizing the side-information bitrate while maintaining high perceptual quality.

  4. Implement a conditional diffusion decoder capable of synthesizing missing video content by simultaneously conditioning on both appearance anchors (keyframe latents) and sparse motion cues (trajectories), enabling high-fidelity reconstruction from minimal conditioning signals.

  5. Enhance the robustness of generative video models in compression by employing bidirectional boundary conditioning during inference, which smooths temporal discontinuities at segment boundaries without incurring significant rate overhead, leading to more stable reconstructions across GOP segments.

  6. Improve the overall encoding efficiency by achieving fast encoder runtime (e.g., < 110 ms/frame) by carefully structuring the pipeline to separate computationally intensive tasks (like dense tracking and sparse trajectory selection) from the core diffusion synthesis process.

  7. Develop a versatile compression framework that balances encoding speed and decoding quality: the encoder focuses on generating compact, informative conditioning signals, while the decoder utilizes a powerful generative model to reconstruct high-quality video at extreme rate constraints suitable for real-time streaming or cloud delivery.

Abstract

Diffusion models provide a powerful generative prior for perceptual reconstruction at ultra-low bitrates, but effective video compression requires controlling the generative process using highly compact conditioning signals. In this work, we present ActDiff-VC, a diffusion-based video compression framework for the ultra-low-bitrate regime. Our method partitions videos into variable-length segments, transmits keyframes only when needed, and summarizes temporal dynamics using a compact set of tracked point trajectories. Conditioned on these sparse signals, a conditional diffusion decoder synthesizes the remaining frames, enabling perceptually realistic reconstruction under severe rate constraints. To support this design, we introduce two mechanisms: content-adaptive keyframe selection and budget-aware sparse trajectory selection, which together enable compact yet effective conditioning for generative reconstruction. ActDiff-VC has an asymmetric computational profile: on a single NVIDIA A100 GPU, encoding requires 109 ms per frame, while the adopted 20-step diffusion decoder requires 2311 ms per frame, making the framework particularly suitable for ultra-low-bitrate applications such as cloud-assisted reconstruction, archival storage, and offline content distribution, where lightweight encoding and perceptual reconstruction quality are prioritized. Experiments on the UVG and MCL-JCV benchmarks show that ActDiff-VC achieves up to 64.6% bitrate reduction at matched NIQE, improves KID by up to 64.6% and FID by up to 37.7% at comparable bitrates against strong learned codecs, and delivers favorable perceptual rate--distortion trade-offs relative to learned and generative baselines in the ultra-low-bitrate regime. A human study with 147 participants further supports these perceptual gains, with ActDiff-VC preferred over DCVC-FM in 62.6% of pairwise comparisons.

Sources

Related papers