Accelerating Video Inverse Problem Solvers with Autoregressive Diffusion Models

arXiv:2605.20624 · cs.CV, cs.AI, cs.LG · Submitted 2026-05-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Accelerating Video Inverse Problem Solvers with Autoregressive Diffusion Models".

Tom: This paper introduces Autoregressive Video Inverse problem Solver (AVIS) and its highly accelerated variant, AVIS Flash, which leverage autoregressive video diffusion models to restore videos in a streaming manner.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: Okay, so we're looking at the title and authors of "Accelerating Video Inverse Problem Solvers with Autoregressive Diffusion Models," and it’s clear this work is aimed directly at solving the practical problems we discussed earlier. The authors are focusing on using autoregressive diffusion models to fix video inverse problems in a way that allows for streaming restoration.

Jane: I agree, Tom. It sounds like they are taking existing, powerful diffusion priors and restructuring how they apply them so that videos can be fixed piece by piece rather than all at once, which is a major conceptual shift.

Lu: They are proposing a specific framework called AVIS to handle this streaming aspect, and the idea is that by initializing the reverse diffusion process with an estimate that already matches the measurements, they can cut down on how many steps are needed for sampling.

Meng: So if I'm hearing you right, they’re trying to make sure that when we start restoring a video chunk-by-chunk, we don't waste time redoing things because we started from a completely wrong place?

Lalam: Precisely. By using that measurement-consistent estimate right at the beginning, they are setting the trajectory for the diffusion process much more accurately than just starting with random noise.

The paper's summary: Tom: Now let's talk about what they actually propose in the AVIS framework. Essentially, their core idea is that instead of restoring every frame simultaneously, AVIS restores videos in a streaming manner, which naturally removes the initial latency problem because you start generating output immediately.

Jane: That makes sense; if you can see frame one quickly, it helps with user experience right away. They achieve this by initializing the reverse diffusion process with a measurement-consistent estimate to significantly reduce the sampling steps required for restoration.

Lu: The paper draws inspiration from existing techniques, like CCDF, which shows that starting from a coarse estimate really cuts down on the computational load needed for the rest of the restoration process.

Meng: But they aren't just using that initialization trick; they are also enforcing measurement updates during every video chunk during this streaming process to keep things consistent as they go.

Lalam: That continuous enforcement of consistency is what keeps the quality high throughout the whole sequence, not just at the very beginning. It’s a steady correction mechanism for every part of the video being generated.

The paper's improvements: Tom: So we've covered how AVIS works, but let’s look at what they added next with AVIS Flash. They introduce AVIS Flash to push the efficiency even further by changing where that measurement consistency enforcement happens during the streaming process.

Jane: That’s where things get really interesting for throughput; instead of applying measurement guidance to every single chunk in AVIS, AVIS Flash only enforces measurement consistency on the very first video chunk.

Lu: That simplification is clever because they observe that subsequent chunks can then be generated through autoregressive propagation from that one corrected prefix, which completely bypasses the need for those iterative VAE passes for every other part of the sequence.

Meng: Bypassing iterative passes sounds like a huge practical win for performance. So what’s the tangible result of this change in where they enforce consistency?

Lalam: The paper shows that this approach substantially boosts throughput, going from zero point seven one FPS to one point one eight FPS with AVIS, and then AVIS Flash jumps that even higher to five point nine one FPS on a single RTX four thousand ninety GPU. That’s a significant speed increase for video generation tasks, right?

Conclusion: Tom: Alright, so we've covered the initial setup of AVIS and how AVIS Flash dramatically improves throughput by shifting the consistency enforcement strategy, leading to much faster results than what was previously achievable with non-autoregressive solvers. This paper on Accelerating Video Inverse Problem Solvers with Autoregressive Diffusion Models really shows a viable path toward making these models practical for video applications.

Jane: I think the main implication is that we can move past those high initial latency problems by adopting streaming methods and smarter initialization, which makes zero-shot video restoration much more accessible in real-time settings.

Lu: The potential here is huge; imagine using this framework to create interactive tools where users see restored video frames instantly as they watch, rather than waiting for a long render time.

Meng: From a practical perspective, the AVIS Flash results suggest that we can actually deploy these solvers on consumer hardware and get usable frame rates that matter for applications like live content moderation or editing.

Lalam: It really strengthens the idea that autoregressive modeling provides a better temporal prior because it builds context sequentially, which is something non-autoregressive methods struggle with.

Tom: So we’re wrapping up our discussion on AVIS Flash and its impact on video inverse problem solving, which is a really solid piece of work in making diffusion models more useful for video tasks.

Taesung Kwon, Jonghyun Park, Hyungjin Chung

KAIST · EverEx

cs.CV, cs.AI, cs.LG

Submitted: 2026-05-20

Updated: 2026-09-29

Comments: NeurIPS 2026, Project page: https://avis-project.github.io/

Code: https://github.com/LAION-AI/aesthetic-predictor

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 85/100

The gist: This paper introduces Autoregressive Video Inverse problem Solver (AVIS) and its highly accelerated variant, AVIS Flash, which leverage autoregressive video diffusion models to restore videos in a

Key concepts

Autoregressive Diffusion Models
These models are used to restore videos in a streaming manner. They build context sequentially, which is beneficial for video because it allows the model to process information frame by frame or chunk by chunk, addressing latency issues.
AVIS Framework
This framework restores videos in a streaming fashion instead of all at once. It initializes the reverse diffusion process with an estimate that matches measurements to reduce the number of steps needed for restoration.
AVIS Flash
This is an improvement over AVIS that boosts efficiency by only enforcing measurement consistency on the very first video chunk. Subsequent chunks are generated through autoregressive propagation from this corrected prefix, bypassing iterative VAE passes.

Terminology

Summary

This paper introduces Autoregressive Video Inverse problem Solver (AVIS) and its highly accelerated variant, AVIS Flash, which leverage autoregressive video diffusion models to restore videos in a streaming manner. This approach addresses the critical limitations of existing non-autoregressive solvers—namely high initial latency and low throughput caused by multiple VAE passes—by initializing the reverse diffusion process with a measurement-consistent estimate. The work demonstrates that AVIS drastically reduces initial latency from 114s to 4s and increases throughput from 0.71 to 1.18 FPS, while AVIS Flash substantially boosts throughput to 5.91 FPS on a single RTX 4090 GPU, paving the way toward real-time deployment of diffusion-based video inverse problem solvers.

AVIS Framework and Core Mechanism

The AVIS framework leverages autoregressive video diffusion models to restore videos in a streaming manner, naturally eliminating latency bottlenecks. Specifically, AVIS initializes reverse diffusion with a measurement-consistent estimate, reducing the required sampling steps. This is achieved by drawing inspiration from CCDF [27], which demonstrates that initializing the reverse diffusion process from a coarse estimate significantly reduces the required sampling steps. Unlike CCDF, AVIS does not rely on an auxiliary pretrained restorer; instead, it obtains an initial restoration by directly minimizing the measurement consistency objective. AVIS enforces measurement consistency for every video chunk during the streaming process.

AVIS Flash Acceleration Strategy

To push efficiency further, the authors introduce AVIS Flash. Unlike AVIS, which applies iterative measurement guidance to every chunk, AVIS Flash enforces measurement consistency solely on the first video chunk. The authors observe that this is sufficient to restore the entire sequence because subsequent chunks are generated through autoregressive propagation from this corrected prefix. This process completely bypass[es] explicit measurement guidance, which eliminates iterative VAE passes for subsequent chunks, thereby accelerating restoration by 5× while preserving zero-shot performance.

Measurement Consistency Enforcement

Measurement consistency is enforced during the reverse diffusion process of the n-th noisy video chunk at each denoising step to guide the sampling trajectory toward the posterior. The conventional approach involves computing gradients of a measurement consistency term, but AVIS employs Decomposed Diffusion Sampling (DDS) [31] using conjugate gradient (CG) to efficiently enforce consistency. This is done by:

  1. Obtaining the clean latent estimate from the noisy state using Eq. (6).

  2. Decoding this estimate into pixel space: xˆn 0t = D(zˆn 0t).

  3. Solving a proximal optimization problem via CG to find an updated pixel-space estimate x˜n 0t, minimizing the objective: x˜n 0t:= arg min x γ 2∥y − A(x)∥ squared + 1/2∥x − xˆn 0t∥ squared.

  4. Re-encoding the updated pixel-space estimate back into the latent space: z˜n 0t = E(x˜n 0t).

Experimental Validation and Results

The methods were evaluated on 100 high-resolution videos across five restoration tasks, including SuperResolution, Random Inpainting, Gaussian Deblur, Temporal Average, and Spatio-Temporal Average. The results show that AVIS achieves the best restoration performance overall. Specifically:

- Computational Efficiency:

AVIS reduces initial latency from 114s to 4s and increases throughput from 0.71 to 1.18 FPS compared to the leading non-autoregressive solver LVTINO (0.71 FPS). AVIS Flash substantially boosts throughput to 5.91 FPS on a single RTX 4090 GPU, making it over 8× faster than VISION-XL and LVTINO in some tasks.

- Restoration Performance:

AVIS consistently attains the best or near-best PSNR, SSIM, LPIPS, and FVD across most tasks. AVIS Flash maintains competitive performance with baselines while reducing initial latency by over 100 seconds and boosting throughput by over 12× and 8×.

Ablation Study Insights

The ablation studies on AVIS Flash provide further insight into the framework's components:

- Autoregressive Propagation:

Removing the KV cache for subsequent chunks degrades all metrics, highlighting the benefit of autoregressive propagation, which provides performance gains complementary to the initialization. The authors show that autoregressive propagation alone preserves context but gradually drifts from the desired restoration, a drift mitigated by their initialization.

- Start Time (t0):

A smaller start time t0 consistently improves most metrics, with t0 = 0.1 yielding the best overall performance.

Improvements for AI systems

Here are specific improvements to AI systems based on the proposed Autoregressive Video Inverse problem Solver (AVIS) and AVIS Flash framework:


The proposed framework, AVIS/AVIS Flash, offers significant advancements in the deployment of zero-shot video inverse problem solvers by addressing the critical bottlenecks of high initial latency and low throughput. The improvements can be applied across several domains: Real-time Video Restoration, High-Throughput Generative AI Inference, and Novel View Synthesis.

Here are the specific improvements and capabilities:

  1. Real-Time, Low-Latency Video Restoration

The primary improvement is the transition from holistic restoration (high latency) to a streaming (autoregressive) approach.

The improved system can perform video restoration in near real-time by:

Achieving initial latency of only 4 seconds for high-resolution super-resolution tasks (compared to 167s for VISION-XL), and achieving throughput up to 5.91 FPS (AVIS Flash) or 10.2 FPS on an NVIDIA H100 GPU.

  1. High-Throughput Video Inpainting

The AVIS Flash variant drastically increases the speed of inpainting by eliminating iterative measurement consistency passes for subsequent chunks, making it suitable for high-demand applications like live video editing or real-time content moderation.

The improved system can perform:

Video inpainting with a throughput of 5.91 FPS (AVIS Flash), significantly faster than existing methods, while maintaining competitive perceptual quality metrics (e.g., LPIPS).

  1. Robust and Scalable Long Video Reconstruction

The framework includes a mechanism for periodic re-injection of measurement consistency guidance every 7 chunks, which prevents the temporal drift inherent in purely autoregressive propagation over long sequences.

The improved system can handle:

Long video restoration (e.g., one-minute videos consisting of 960 frames at 16 FPS) while successfully retaining ground truth context and preventing noticeable temporal drift in later frames.

  1. Inference-Time Novel View Synthesis

AVIS Flash can be adapted to serve as an inference operator for novel view synthesis, effectively acting as a high-quality inpainting mechanism to fill occluded regions during dynamic camera movement.

The improved system can perform:

Inference-time novel view synthesis by using AVIS Flash as an inpainting operator to plausibly synthesize disoccluded content under target camera trajectories.

  1. Improved Generative Priors via Autoregressive Modeling

By leveraging autoregressive diffusion models, the system establishes a strong temporal prior that is conditioned on previously generated frames, leading to superior temporal coherence and fidelity compared to non-autoregressive solvers.

The improved AI model can:

Generate perceptually realistic video sequences with superior Subject Consistency (Sub. C.) and Background Consistency (Bg. C.) metrics by utilizing the sequential conditioning of AR diffusion models.

  1. Enhanced Initialization for Zero-Shot Restoration

The use of a measurement-consistent initial estimate, derived via a multi-step optimization (CG), ensures that the reverse diffusion process starts from a better point than random noise, enhancing fidelity in zero-shot tasks.

The improved system can:

Solve complex video inverse problems in a zero-shot manner with superior restoration quality by initializing the reverse diffusion process with a measurement-consistent estimate derived from pixel-space optimization.

Abstract

Diffusion models provide powerful priors for zero-shot video inverse problems, but their real-time deployment is hindered by two inefficiencies: high initial latency caused by holistic video restoration, and low throughput resulting from multiple VAE passes to enforce measurement consistency in pixel space. To overcome these limitations, we propose Autoregressive Video Inverse problem Solver (AVIS). The AVIS framework leverages autoregressive video diffusion models to restore videos in a streaming manner, naturally eliminating latency bottlenecks. Specifically, AVIS initializes reverse diffusion with a measurement-consistent estimate, reducing the required sampling steps. Compared to leading non-autoregressive solvers, AVIS drastically reduces initial latency from 114s to 4s and increases throughput from 0.71 to 1.18 FPS while achieving superior restoration quality. We further introduce a highly accelerated variant, dubbed AVIS Flash, that enforces measurement consistency solely on the first chunk. AVIS Flash substantially boosts throughput to 5.91 FPS on a single RTX 4090 GPU while maintaining competitive performance and achieving a favorable efficiency-performance trade-off, paving the way toward real-time deployment.

Sources

Related papers