QuantWM: Temporally Consistent 2-Bit KV Cache Quantization for Video World Models
cs.CV, cs.AI
Submitted: 2026-09-22
Updated: 2026-09-28
Project page: https://quantwm-project.github.io/QuantWM
License: http://creativecommons.org/licenses/by/4.0/
The gist: KV cache memory has become a major deployment bottleneck for video generation and world models, which motivates low-bit quantization study for efficiency.
Terminology
Abstract
KV cache memory has become a major deployment bottleneck for video generation and world models, which motivates low-bit quantization study for efficiency. Existing 2-bit KV cache quantization methods can achieve nearly lossless performance on video benchmarks such as VBench, however, we find that they still cause severe temporal flickering and visual degradation. Meanwhile, deeper investigates show that Key quantization produces smaller reconstruction errors than Value, but surprisingly leads to much larger output degradation. We trace this discrepancy to attention: small Key perturbations can change the attention logits, i.e., QK, and shift the temporal-spatial tokens selected by Queries. These observations motivate us to explicitly preserve attention logits and temporal-spatial token selection during KV cache quantization to alleviate the visual degradation problem. To address this issue, we present QuantWM, a training-free and strictly causal 2-bit KV cache quantization framework. QuantWM introduces two complementary techniques to mitigate the attention shifts. Firstly, quantization-sensitivity-aware clustering (QSAC) jointly considers historical Query sensitivity and residual ranges to select INT2-friendly Key centroids, which reduces quantization errors in channels that are more critical to attention. In addition, principal-subspace attention compensation (PSAC) restores the remaining Key errors along the dominant Query subspace using low-rank projections, which provides a direct and efficient correction to stabilize attention logits. Extensive experiments on Causal-Forcing, LingBot-World-v2, HY-World 1.5, Matrix-Game-2 and Longcat-Video demonstrate that QuantWM significantly improves visual quality and temporal consistency, while outperforming existing methods across image and video quality metrics with up to 6.20x KV cache memory compression and limited additional overhead.
Sources
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- Infinite Worlds with Versatile Interactions
- Matrix-game 2.0: An open-source, real-time, and streaming interactive world model
- SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models
- CommVQ: Commutative Vector Quantization for KV Cache Compression
- KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
- Movie Gen: A Cast of Media Foundation Models
- Make-A-Video: Text-to-Video Generation without Text-Video Data
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
- LongCat-Flash-Omni Technical Report
- Quantized Keys Steal Attention: Bias Correction for KV-Cache Compression in Video Diffusion
- Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
- Scaling Autoregressive Video Models
- Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization
- LongLive: Real-time Interactive Long Video Generation
- Boost Post-Training Quantization via Null Space Optimization for Large Language Models
- Open-Sora: Democratizing Efficient Video Production for All
- Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models