OmniVR: Audio-Video Conditional Generation for Archival Footage Restoration

arXiv:2608.04224 · cs.CV · Submitted 2026-08-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "OmniVR: Audio-Video Conditional Generation for Archival Footage Restoration".

Jane: Historical films suffer from co-occurring visual and audio degradations—blur, noise, flicker, hiss, clipping, and dropout—yet existing methods restore each modality independently, leaving quality gaps and cross-modal inconsistency.

Tom: First, who's behind it and why it matters.

Title and authors: Jane: So, looking at the title, "OmniVR: Audio-Video Conditional Generation for Archival Footage Restoration," it really tells us exactly what the paper is focused on—restoring old film footage by conditioning a generation process on both audio and video information.

Tom: It’s quite precise; they aren't just doing separate things anymore, they are explicitly using both streams as conditions to drive the restoration. That suggests a much deeper level of integration than previous work that might have only tried to improve the image or the audio independently.

Lu: The implications of that dual conditioning are huge because it forces the model to learn how visual details like blur and flicker relate directly to acoustic artifacts like hiss and clipping in a realistic way, which is something separate models miss.

Meng: I wonder if that joint approach means they have to manage two very different data distributions at once; ensuring the latent conditions for video and audio are effectively aligned during the denoising process must be quite delicate.

Lalam: For me, this points toward a future where media restoration isn't about patching up individual errors but about reconstructing the authentic sensory experience of what it was like to watch and hear that original historical clip.

Tom: That’s the essence, Jane; they are aiming for that holistic reconstruction where the visual fidelity and the acoustic detail support each other perfectly, which is a big step forward from previous work.

The paper's summary: Jane: The paper summarizes OmniVR as proposing a joint audio-video generative restoration model that uses a multimodal DiT to recover visual structure, temporal motion, and acoustic detail all at once by jointly denoising the low-quality video and audio using their degraded states as latent conditions.

Tom: Essentially, they’ve taken the concept of conditional generation and applied it across two modalities simultaneously, using a twenty-two billion parameter backbone to guide the recovery of both visual elements like sharpness and motion, and acoustic elements like hiss or dropout in one go.

Lu: What really stands out in their summary is how they are motivating this by pointing out empirical observations that show these defects co-occur across modalities in real old films, which justifies why a joint model is necessary instead of just independent fixes.

Meng: So, the core mechanism involves using those low-quality inputs as conditions combined with a fixed restoration prompt to guide the AI to jointly refine both streams under one objective, which sounds like a heavy lifting task for any generative backbone.

Lalam: The summary makes it clear that the goal isn't just making things look cleaner or sound better individually; it’s about achieving consistency across the entire sensory output of the historical footage.

Tom: Right, and they emphasize that this joint generation enables what they call "mutual promotion between modalities for stronger AV consistency and sound quality," which is a key mechanism they are highlighting.

The paper's improvements: Jane: Regarding the specific improvements mentioned in the paper, it highlights three main pillars: first, creating a joint audio-video degradation pipeline to synthesize realistic defects; second, an architecture-preserving transition from text-to-audio-video to audio-video-to-audio-video generation with prompt annealing; and third, using first-frame image anchors with loss reweighting and waveform supervision.

Tom: That second point about prompt annealing is particularly clever because it lets them start training from a pre-existing, functional T2AV model while smoothly shifting the task to an AV2AV restoration where you don't need captions at inference time.

Lu: The loss reweighting mechanism sounds smart for handling the varying severity of defects; by using composite severity scalars, they can dynamically adjust which details—visual or audio—need more recovery emphasis during training.

Meng: From a practical standpoint, the first-frame anchoring addresses temporal drift in long sequences, which is a major issue when working with extended historical footage; that makes it much more robust for recovering content over longer videos.

Lalam: I think the loss reweighting and waveform supervision are crucial because they ensure that even if one modality is weaker, the other isn't completely neglected during the refinement process, which leads to a more balanced outcome.

Tom: So, they’ve got this multi-layered approach: simulating realistic degradation to train on, using prompt annealing for a smooth transition in training setup, and anchoring with loss reweighting for stable long-video recovery.

Conclusion: Jane: To wrap up the discussion on OmniVR: it seems the paper concludes that this joint generation approach significantly outperforms prior methods across visual quality, audio quality, temporal consistency, and audio-visual synchrony when tested on their benchmark.

Tom: It’s clear that they achieved a strong result here; they surpassed existing baselines on all six visual metrics and managed to achieve superior audio quality compared to the clean reference in certain tests.

Lu: The implication is that we can move towards restoration systems that inherently understand the interdependency of audio and video data, which is a significant conceptual step for multi-modal AI applications in general.

Meng: For implementation, the fact that it handles both colorization and sharpness while keeping the audio fidelity high simultaneously suggests a more robust pipeline for archival work than anything we've seen before.

Lalam: I think this work shows us that when we tackle complex, real-world data like historical film, integrating modalities deeply doesn't just produce better individual outputs; it produces a result that is perceived as far more coherent to the human eye and ear.

Tom: Exactly; OmniVR demonstrates how focusing on jointly addressing all three aspects—colorization, sharpness, and audio fidelity—simultaneously leads to an overall performance boost that’s hard to ignore.

Jane: So, in summary, OmniVR provides a practical framework for joint audio-video restoration by using a unified DiT and sophisticated training techniques like prompt annealing and loss reweighting.

Lu: It opens up questions about how much more complex these conditional generation backbones can become before they start needing entirely new architectural ideas.

Meng: I'm just thinking about the future applications; if we can reliably restore these old films, it could unlock huge amounts of historical and cultural data that are currently inaccessible due to poor quality.

Lalam: That’s the real promise here—making those records accessible and understandable again through high-fidelity reconstruction.

Tom: Alright, folks, that covers OmniVR; a really compelling piece of work on how we can handle coupled data streams in restoration tasks. We'll take a quick break and then move on to the next paper.

Xin Lu, Zihao Fan, Mingchen Zhong, Jie Huang, Xueyang FuB, Zheng-Jun Zha

University of Science and Technology of China

cs.CV

Submitted: 2026-08-04

Updated: 2026-09-28

Project page: https://xin1u.github.io/OminiVR_PAGE/1

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: Historical films suffer from co-occurring visual and audio degradations—blur, noise, flicker, hiss, clipping, and dropout—yet existing methods restore each modality independently, leaving quality

Key concepts

Joint Audio-Video Generative Restoration
This is the core idea of OmniVR: instead of fixing video and audio separately, the model generates both at once. It uses a single large model to handle visual structure, motion, and acoustic details together. This addresses the problem where fixing one part leaves quality gaps in the other.
Unified Multimodal DiT
The model is built on a 22B-parameter backbone that functions as a unified Diffusion Transformer (DiT). It takes both low-quality video and audio as 'latent conditions' to guide the generation process. This allows the model to coordinate denoising across different data types in one integrated framework.
Prompt Annealing
This technique manages how the model receives instructions during training and inference. It gradually shifts from using descriptive text captions to a fixed restoration prompt, maximizing the use of a strong pre-trained generative prior while ensuring the model can perform clean restoration tasks without needing specific captions at runtime.

Terminology

Summary

Historical films suffer from co-occurring visual and audio degradations—blur, noise, flicker, hiss, clipping, and dropout—yet existing methods restore each modality independently, leaving quality gaps and cross-modal inconsistency. OmniVR is presented as the first joint audio-video generative restoration model that addresses this fundamental problem by formulating restoration as conditional generation within a unified multimodal DiT.

How it works

OmniVR is built upon a 22B-parameter audio-video generation backbone and formulates restoration as conditional generation within a unified multimodal DiT, where the low-quality video and audio are encoded as latent conditions combined with a fixed restoration prompt, jointly denoised to recover visual structure, temporal motion, and acoustic detail under one coordinated objective. This approach is motivated by empirical observations showing that real old films exhibit co-occurring defects across modalities and that joint generation enables mutual promotion between modalities for stronger AV consistency and sound quality.

Key Design Pillars

The model incorporates three key designs to adapt the large-scale AV generative foundation model for faithful restoration:

  1. A joint audio-video degradation pipeline that "synthesizes realistic audio-video degradation from high-quality clips collected from the Internet, simulating the visual characteristics (blur, noise, flicker, compression, low exposure) and audio characteristics (hiss, clipping, bandwidth loss, dropout) of real historical films."

  2. An architecture-preserving text-to-audio-video (T2AV) to audio-video-to-audio-video (AV2AV) transition with prompt annealing that maximally preserves the generative prior of the pretrained T2AV model while smoothly transitioning to an LQ video conditioned restoration task requiring no per-clip captioning at inference.

  3. First-frame image-to-video (I2V) anchoring with loss reweighting and waveform supervision for long-video extrapolation and audio fidelity.

Training and Optimization

The training process involves several sophisticated mechanisms to ensure fidelity:

Latent Encoding and Flow Matching:

Frozen VAE encoders map both modalities into latent spaces, which are reshaped into tokens. Training follows rectified flow Liu et al. (2023), where the degraded condition is injected via channel concatenation, and the model jointly denoises both modalities in a shared token sequence using bidirectional cross-modal attention and a shared timestep.

Prompt Annealing:

The text condition transitions from paired descriptive captions to a fixed restoration prompt via linear annealing, gradually decreasing caption probability while increasing the fixed prompt weight to maximize prior preservation. Classifier-free guidance is enabled by dropping the prompt to a null embedding with probability 0.1.

Loss Reweighting and Supervision:

The training objective combines modality-weighted velocity losses, where per-sample weights are determined by composite severity scalars (e.g., wv = NormB(1 + αvdv + βvda)), and an optional multi-resolution STFT loss provides direct waveform domain feedback to stabilize spectral detail.

Evaluation and Benchmarking

OmniVR introduces OmniVRBench, the first benchmark for joint audio-video restoration, which comprehensively assesses restoration quality across visual quality (MUSIQ), audio quality (DNSMOS), temporal consistency, and audio-visual synchrony on 200 real historical clips. The model surpasses all prior methods on all six visual metrics and achieves the best audio quality. Evaluation also includes a full-reference protocol for the Controlled track to assess fidelity against clean references, as well as human preference studies confirming that jointly restored outputs are consistently perceived as more coherent than separately processed streams.

Performance Summary

On the Real historical track, OmniVR leads all visual, audio, and sync metrics. Specifically, on OmniVRBench's Controlled track (talking-face only), it outperforms all baselines on every metric and matches or exceeds the clean reference on most visual axes while surpassing the clean reference in audio quality (DNSMOS 2.70 vs. 2.47) and lip sync (LSE-C 3.52 vs. 4.00). Human evaluation confirms a clear preference for jointly restored results, with OmniVR achieving an overall human win rate of 80% across visual, audio, and sync dimensions compared to the strongest baseline by over 54 percentage points. The model's success is attributed to its ability to jointly address all three aspects (colorization, sharpness, and audio fidelity) simultaneously.

Limitations and Future Work

Current limitations include the reliance on proxy-based evaluation due to the lack of clean references for real archival footage and potential temporal drift in long sequences due to first-frame anchoring. Future work plans include incorporating larger-scale paired real degraded/restored data to reduce the domain gap, and adopting streaming (causal) generation for unbounded length restoration with lower latency.

Improvements for AI systems

Based on the scientific paper OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films, here are specific improvements that can be made to existing AI systems, and a description of what these improved systems can achieve:


)

  1. Improvements to Existing Restoration Architectures (e.g., independent video SR or audio enhancement models):

  2. Improvements in Cross-Modal Consistency and Synchronization:

  3. Advancement in Generative Prior Preservation:

  4. Creation of a Unified, Joint Audio-Video Generative Restoration System (OmniVR Paradigm):

  5. Existing restoration architectures (like DeepRemaster or RealBasicVSR) can be fundamentally improved by integrating the audio stream directly into the visual denoising process. Instead of treating video and audio as separate problems, these systems should adopt a joint conditional generation framework where both modalities are encoded as latent conditions for a single multimodal diffusion model.

  6. Existing systems will gain the capability to perform true restoration—meaning they can simultaneously recover visual structure (denoising, sharpening) and acoustic detail (removing hiss, clipping) without the quality gaps or temporal misalignment inherent in independent processing. This leads to outputs that are perceptually superior across all modalities.

  7. Existing AI systems will achieve near-perfect audio-visual synchronization. By utilizing cross-modal routing within the backbone (as demonstrated by OmniVR's mutual attention), the system can ensure that visual motion, lip movements, and acoustic energy fluctuations are perfectly aligned in time, which is critical for tasks like historical film analysis or character performance verification.

  8. Existing generative models will be able to leverage learned generative priors more effectively during restoration. This means the restored output will not just look clean but will maintain the original scene's identity and structure while hallucinating high-fidelity details (like texture or color) that are missing from the degraded input, guided by a fixed restoration prompt.

  9. The new system can achieve realistic colorization. Unlike current baselines that introduce biased or temporally flickering colors, this improved system will jointly restore natural color fidelity across all frames while ensuring temporal consistency in the hue and saturation of restored objects.

  10. The system will be able to handle long-form extrapolation for video restoration. Using first-frame chaining (I2V anchoring), the AI can coherently recover and generate content over long sequences, overcoming the temporal drift issues that plague current windowed models, resulting in seamless, continuous video restoration from degraded source material.

  11. The system will be able to handle degradation-aware quality weighting. By dynamically adjusting loss coefficients based on the severity of degradation in each modality (joint audio-video degradation pipeline), the model can prioritize recovering details where they are most severely missing, leading to a more robust and artifact-free restoration process regardless of whether a clip is visually or acoustically dominant.

  12. The system will be capable of being evaluated using standardized, multi-dimensional benchmarks (like OmniVRBench). This allows researchers to move beyond single metrics (e.g., just visual quality) and assess the holistic performance across visual fidelity, audio clarity, temporal consistency, and audio-visual synchrony simultaneously.

Abstract

Archival footage often suffers from coupled visual and acoustic degradations, yet most restoration systems process the two modalities separately. To address this problem, we present OmniVR, the first systematic framework for joint audio-video restoration, covering data construction, model adaptation, efficient inference, and evaluation. We construct a high-quality audio-video corpus with detailed captions and use a joint degradation pipeline to produce aligned clean and degraded pairs. Using these pairs, we adapt a pretrained text-to-audio-video model (T2AV) by introducing degraded audio-video conditions (TAV2AV), then progressively replace sample captions with a fixed restoration prompt while retaining caption/null rehearsal. The resulting AV2AV model requires no user-provided text. Under a compatible residual-learning model, we prove that this condition-annealing schedule reduces gradient variance and expected restoration risk relative to direct fixed-prompt adaptation at the same training budget. For efficient deployment, OmniVR-Flash combines reduced-resolution video conditioning, MeanFlow-based one-step distillation, and Turbo VAE, achieving approximately 38 fps at 1K and 18 fps at 2K on a single B200 GPU. We further introduce OmniVRBench to evaluate four complementary dimensions: visual quality, audio quality, temporal consistency, and audio-visual synchrony. OmniVR achieves state-of-the-art results on public benchmarks and OmniVRBench. Data, code, and model weights will be released. Project Page: https://xin1u.github.io/OminiVR PAGE/

Sources

Related papers