Neural Video Compression Based on Deformable Temporal Alignment and Difference-aware Fusion

arXiv:2609.03520 · cs.CV, cs.AI · Submitted 2026-09-03 · Read on arXiv

cs.CV, cs.AI

Submitted: 2026-09-03

Updated: 2026-09-03

License: http://creativecommons.org/licenses/by/4.0/

The gist: In conditional coding-based neural video compression, the quality of temporal context directly affects compression per- formance.

Terminology

Abstract

In conditional coding-based neural video compression, the quality of temporal context directly affects compression per- formance. Existing methods mostly construct context from prop- agated reference features, but they are vulnerable to motion esti- mation and local alignment errors in regions with complex mo- tion, occlusion, and high-frequency textures, resulting in inaccu- rate temporal information. To address this issue, this paper pro- poses a method combining deformable temporal alignment and difference-aware spatial selective fusion. A Context-aware Tem- poral Alignment Module is used to generate complementary tem- poral context, while a Difference-aware Spatial Selective Fusion module adaptively selects reliable temporal information and sup- presses misalignment. Experiments show that the proposed method achieves certain rate-distortion performance improve- ment over DCVC-DC.

Related papers