VoRTeC: Taming Foundation Flow for One-step Real time Video Compression
cs.CV, cs.AI
Submitted: 2026-09-02
Updated: 2026-09-03
Code: https://github.com/microsoft/DCVChttps:
Project page: https://darc8-sun.github.io/VoRTec_compress
License: http://creativecommons.org/licenses/by/4.0/
The gist: Ultra-low bitrate video compression still faces critical challenges: traditional neural video compression inevitably introduces blurring artifacts, while diffusion-based generative video compression
Terminology
Abstract
Ultra-low bitrate video compression still faces critical challenges: traditional neural video compression inevitably introduces blurring artifacts, while diffusion-based generative video compression suffers from excessive decoding latency and poor temporal consistency. To address these issues, we propose, a Video Compression framework built upon a foundational flow model (Wan2.1). By compactly encoding latent video representations, predicting the positions of compressed representations along flow trajectories, and integrating multi-scale priors, enables the compressor to harness generative video flow priors effectively. Without accessing the parameters or gradients of flow matching networks, our framework achieves one-step decoding and reconstructions with high perceptual fidelity. Meanwhile, we maintain consistency across frame groups via tail-frame reuse and prior caching. Extensive experiments demonstrate that our method reduces bit consumption by 58% compared to prior diffusion-based approaches, with decoding speed boosted by 3 to 197 times: achieves a decoding speed of 13 FPS at 720p and 32 FPS at 480p.
Sources
- Wan: Open and Advanced Large-Scale Video Generative Models
- CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
- Generative Neural Video Compression via Video Diffusion Prior
- Free-GVC: Towards Training-Free Extreme Generative Video Compression with Temporal Coherence
- Improved Distribution Matching Distillation for Fast Image Synthesis
- Single-step Diffusion-based Video Coding with Semantic-Temporal Guidance
- Towards Efficient Low-rate Image Compression with Frequency-aware Diffusion Prior Refinement
- CoD-Lite: Real-Time Diffusion-Based Generative Image Compression
- YODA: Yet Another One-step Diffusion-based Video Compressor
- FloLPIPS: A Bespoke Video Quality Metric for Frame Interpoation
- End-to-end Optimized Image Compression
- Decoupled Weight Decay Regularization
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models