Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation".
Tom: Real-time interactive video generation requires low-latency, streaming, and controllable rollout, and this paper addresses the limitations of existing autoregressive diffusion distillation methods by proposing Causal Forcing++,
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So, we’ve covered that Causal Forcing++ is a principled pipeline designed to solve the initialization bottleneck when generating videos with only one or two sampling steps, by using causal consistency distillation. Jane That means the paper argues that existing strategies are either misaligned with the causal rollout process, or they just aren't scalable enough for this aggressive regime.
Lu: The central claim of Causal Forcing++ is that it proposes a way to initialize a few-step AR student using causal consistency distillation, which aims to learn the same AR-conditional flow map as causal ODE distillation <ref:2605.15141#pg1>.
Meng: That sounds like it’s trying to bridge the gap between two different initialization techniques by focusing on what they both want to achieve: understanding the underlying flow map of the teacher model.
Lalam: It seems like a clever way to get supervision that is inherently local and real-time, which is exactly what we need when dealing with interactive video where we can't afford massive offline processing.
Tom: Exactly, and the authors highlight that causal ODE distillation requires generating full multi-step PF-ODE trajectories for every training sample, which they found to be very costly to scale <ref:2605.15141#pg1>.
Jane: And Causal Forcing++ avoids that cost by obtaining supervision from a single teacher ODE step between adjacent timesteps on real videos, avoiding the need to precompute and store those full trajectories <ref:2605.15141#pg2>.
Lu: The paper sets up the pipeline by inheriting Stage one and Stage three from previous work, but fundamentally changes Stage two by substituting initialization with Causal Consistency Distillation for few-step AR student initialization <ref:2605.15141#pg2>.
Meng: So the pipeline structure itself is mostly borrowed, but the crucial innovation lies in replacing that specific step with this more efficient causal CD method for getting the student started <ref:2605.15141#pg2>.
Lalam: This points to a design philosophy where we can retain the main training stages but inject a more intelligent, locally consistent way of setting up the student model for deployment <ref:2605.15141#pg2>.
Tom: The authors emphasize that Causal CD is a principled substitute for causal ODE initialization because it targets the same learning target—the AR-conditional flow map—but obtains supervision in a different, more manageable way <ref:2605.15141#pg1>.
Jane: And they quantify the efficiency gains by showing that this method reduces training cost by four times and requires no extra data curation or auxiliary storage, which is a significant practical improvement <ref:2605.15141#pg2>.
Lu: The comparison in Figure one shows that Causal Forcing++ outperforms Self Forcing and the Costly ODE initialization across several metrics like initialization time, cost, and latency <ref:2605.15141#pg2>.
Meng: That quantitative evidence is what we need; seeing a four times reduction in cost on top of achieving better VBench scores is concrete proof that this isn't just theoretical work <ref:2605.15141#pg2>.
Lalam: For me, the biggest implication here is that we can design AI systems that are inherently more efficient from the ground up, focusing on local consistency rather than relying on massive, expensive global trajectory calculations during training.
Conclusion: Tom: So, looking at the full picture, Causal Forcing++ isn't just another tweak; it’s a principled way to handle initialization that is both correct—it targets the right AR flow map—and more efficient because causal CD entirely sidesteps the need for offline trajectory generation <ref:2605.15141#pg2>.
Jane: I think the authors have successfully shown that we can achieve performance comparable to, and in some cases better than, previous state-of-the-art methods when running in the demanding frame-wise two-step setting <ref:2605.15141#pg2>.
Lu: The authors distinguished their work from prior research by being theoretically sound about targeting the correct AR flow map while avoiding that offline trajectory generation bottleneck, which is a key technical hurdle <ref:2605.15141#pg2>.
Meng: From an engineering standpoint, this suggests that future work in generative AI should prioritize objectives like Causal CD when designing models intended for real-time applications where latency and resource usage are primary concerns <ref:2605.15141#pg2>.
Lalam: The broader impact is that it lowers the barrier to entry for building sophisticated, high-quality, real-time interactive video systems by making the initialization process significantly more efficient and scalable <ref:2605.15141#pg2>.
Tom: So in short, Causal Forcing++ delivers strong results under aggressive low-step conditions while simultaneously slashing latency by fifty percent compared to earlier approaches <ref:2605.15141#pg2>.
Jane: It’s a significant step forward because it proves that we can get high fidelity in a very constrained, real-time environment without incurring the massive computational penalty of offline trajectory precomputation <ref:2605.15141#pg2>.
Min Zhao, Hongzhou Zhu, Kaiwen Zheng, Zihan Zhou, Bokai Yan, Xinyuan Li, Xiao Yang, Chongxuan Li, Jun Zhu
Tsinghua University · ShengShu · Renmin University of China
cs.CV
Submitted: 2026-05-14
Updated: 2026-10-04
Comments: fixed typos
Code: https://github.com/thu-ml/Causal-Forcing
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 82/100
The gist: Causal Forcing++ replaces the expensive causal ODE initialization of autoregressive diffusion distillation with causal consistency distillation, achieving frame-wise 2-step generation that surpasses
Key concepts
- Causal Consistency Distillation (Causal CD)
- This is the core innovation. Instead of requiring expensive precomputed full trajectories, Causal CD learns the same AR-conditional flow map as causal ODE distillation but gets supervision from just one teacher ODE step between adjacent video frames. This makes initialization both more efficient and easier to optimize.
- Few-step AR Students
- These are the small autoregressive models that are initialized using Causal CD. They are used in real-time generation where only a few steps (like 1 or 2) of the diffusion process are needed to produce an output, making them crucial for low-latency video tasks.
- AR Diffusion Distillation
- This is the overall training strategy. It involves using a larger, more powerful teacher model to guide and train a smaller student model. The goal is for the student to learn the same complex patterns as the teacher in an autoregressive diffusion setting, which is vital for high-quality video generation.
- Asymmetric DMD
- This refers to Stage 3 of the pipeline, where asymmetric Discrete Moving Difference (DMD) is used with self-rollout. In this stage, both the teacher and critic models remain bidirectional, allowing for effective training while maintaining temporal consistency in the final generation phase.
Terminology
Summary
Causal Forcing++ replaces the expensive causal ODE initialization of autoregressive diffusion distillation with causal consistency distillation, achieving frame-wise 2-step generation that surpasses prior 4-step chunk-wise methods while halving first-frame latency and cutting Stage 2 training cost roughly fourfold. Real-time interactive video generation requires low-latency, streaming, and controllable rollout, yet existing autoregressive (AR) diffusion distillation methods remain limited by coarse response granularity and non-negligible sampling latency <ref:2605.15141#pg1,requires low-latency, streaming, and controllable rollout>. The paper therefore studies a more aggressive setting: frame-wise autoregression with only 1–2 sampling steps, where the initialization of a few-step AR student is identified as the key bottleneck <ref:2605.15141#pg1,frame-wise autoregression with only 1–2 sampling steps>. The proposed pipeline, Causal Forcing++, surpasses the SOTA 4-step chunk-wise Causal Forcing under the frame-wise 2-step setting by 0.1 in VBench Total, 0.3 in VBench Quality, and 0.335 in VisionReward, while reducing first-frame latency by 50% and Stage 2 training cost by ∼4× <ref:2605.15141#pg1,reducing first-frame latency by 50% and Stage 2 training cost>. The work also extends the pipeline to action-conditioned world model generation in the spirit of Genie3 <ref:2605.15141#pg1,action-conditioned world model generation in the spirit of Genie3>.
The initialization bottleneck
The paper identifies the initialization of a few-step AR student before asymmetric DMD as the key bottleneck in this aggressive regime, where existing strategies fail in complementary ways <ref:2605.15141#pg1,the initialization of a few-step AR student before asymmetric DMD>. Three candidate strategies are examined and each falls short:
-
ODE initialization with a bidirectional teacher, as used in CausVid and Self Forcing, is architecturally misaligned with causal rollout, since the teacher trajectory depends on future frames that are unavailable to an AR student <ref:2605.15141#pg1,architecturally misaligned with causal rollout>.
-
Directly using a multi-step AR diffusion model, as in LiveAvatar and WorldPlay, avoids this mismatch but lacks few-step generation capability, and its per-frame approximation error is severely amplified during self-rollout <ref:2605.15141#pg2,its per-frame approximation error is severely amplified during self-rollout>.
-
Causal ODE initialization, as in Causal Forcing, corrects the learning target by distilling from an AR teacher, but requires generating full multi-step PF-ODE trajectories for every training sample, making it costly to scale <ref:2605.15141#pg2,requires generating full multi-step PF-ODE trajectories for every training sample>.
A satisfactory initialization must therefore be simultaneously AR, few-step, and scalable <ref:2605.15141#pg2,simultaneously AR, few-step, and scalable>.
How it works
The key observation is that causal ODE distillation and causal consistency distillation (causal CD) aim to learn the same object: the AR-conditional flow map, or consistency function, of the teacher <ref:2605.15141#pg2,the AR-conditional flow map (or namely the consistency function>; they differ only in how supervision is obtained <ref:2605.15141#pg2,they differ, however, in how the supervision is obtained>. Causal ODE distillation requires the AR teacher to generate an entire multi-step PF-ODE trajectory for each training sample, which must be precomputed and stored offline, whereas causal CD obtains supervision from a single online teacher ODE step between adjacent timesteps on real videos <ref:2605.15141#pg2,a single online teacher ODE step between adjacent timesteps on real videos>. Beyond efficiency, adjacent-timestep consistency yields a smaller per-step optimization gap than causal ODE distillation, which regresses noisy intermediate states directly to clean endpoints <ref:2605.15141#pg2,which regresses noisy intermediate states directly to clean endpoints>. The pipeline inherits Stage 1 (teacher forcing AR diffusion training) and Stage 3 (asymmetric DMD with self-rollout) from Causal Forcing, and replaces its Stage 2 with the causal CD objective <ref:2605.15141#pg7,replaces its Stage 2 with the causal CD objective>.
Results and efficiency
Under frame-wise 2-step generation, Causal Forcing++ achieves the best overall performance among existing AR diffusion distillation methods, improving VBench Total, VBench Quality, and VisionReward over prior methods while reducing first-frame latency by 50% <ref:2605.15141#pg2,improving VBench Total, VBench Quality, and VisionReward over prior methods>. Its Total score (84.14) and Quality score (84.89) surpass all previous methods, while its Semantic score remains comparable to Causal Forcing <ref:2605.15141#pg11>. Causal ODE requires around 11,600 A800 GPU· hours and 1,900 GiB extra storage, whereas causal CD requires only around 2,900 A800 GPU· hours and no extra storage, reducing Stage 2 time cost by roughly 4× <ref:2605.15141#pg12,reduces Stage 2 time cost by roughly 4×>.
Why not causal DMD
Causal DMD yields lower overall quality than causal CD, and although its first few frames appear sharper, later frames rapidly drift with severe camera shifts <ref:2605.15141#pg8,as autoregressive generation proceeds, the later frames from causal DMD rapidly drift>. Because DMD optimizes a reverse KL objective while CD optimizes a forward KL objective, mode-seeking DMD is more sensitive to accumulated history errors, causing exposure bias to rapidly amplify in later frames <ref:2605.15141#pg9,mode-seeking DMD is more sensitive to accumulated history errors>.
Improvements for AI systems
-
Replace causal ODE Stage 2 with online causal consistency distillation. Drop the offline PF-ODE trajectory generation entirely and instead supervise with a single online teacher ODE step between adjacent timesteps on ground-truth videos, per Eq. (3): pair
x t iwithx̂ t-Δt iconditioned on the ground-truth prefixx gt<i. This cuts Stage 2 from 11,600 to 2,900 A800-GPU hours (4×) and eliminates the 1,900 GiB of paired-trajectory storage, while matching or beating causal ODE on VBench. -
Adopt frame-wise autoregression with 1–2 sampling steps instead of chunk-wise 4-step. Generate with frame-wise AR and only 1–2 steps (keeping the first latent frame at 4 steps per the ASD trick), which yields 20.7 FPS at 1-step and 14.1 FPS at 2-step versus 10.4 FPS for CausVid/Self Forcing/Causal Forcing, and halves first-frame latency from 0.60 s to 0.27 s on an A800.
-
Enforce frame-level injectivity by initializing only from an AR teacher. Never initialize the few-step student from a bidirectional teacher's ODE trajectories, since that
violates the frame-level injectivity required by an AR student
and collapses the target towardE[x 0 i x t i, x t<i, t], producing blurred, misaligned initialization that catastrophically breaks down at frame-wise 1-step. -
Always insert an explicit few-step adaptation before asymmetric DMD. Do not feed the multi-step AR diffusion model directly into DMD; the ablation shows it collapses in the 1-step setting to 0 Dynamic Degree, 1.101 VisionReward, and −14 Instruction Following, because
reducing the chunk size increases the number of AR calls... while reducing the sampling step count increases the approximation error within each call,
and these compound during self-rollout. -
Use causal CD rather than causal DMD for few-step initialization. Although DMD produces sharper early frames, its reverse-KL mode-seeking behavior concentrates probability mass and is highly sensitive to accumulated history drift, so later frames
rapidly drift, accompanied by severe camera shifts
; causal CD's forward-KL mode-covering behavior keeps mass in good-quality regions under the same shift, giving higher Stage 2 and Stage 3 VBench and 0.5 higher VisionReward across all step settings. -
Keep Stage 3 as asymmetric DMD with student self-rollout and a bidirectional teacher/critic. Retain the Self Forcing correction—train the student on its own generated prefixes rather than ground-truth context—so that concatenated frames form a valid generated video, while the teacher and critic stay bidirectional (Wan2.1-14B real-score model) to transfer a stronger generative distribution.
-
Extend the same pipeline to action-conditioned world models via camera-pose conditioning. Finetune Wan2.1-1.3B into a bidirectional pose-conditioned diffusion model using PRoPE, then distill with Causal Forcing++ into a chunk-wise 4-step interactive AR world model, enabling Genie3-style control such as
moving forward continuously
ormove forward first, then tilt the camera downward.
-
Scale the pipeline to harder regimes without re-curating data. Because causal CD supervises on real videos with one online teacher step, teacher/data/chunk-size changes no longer force regeneration of stored trajectories, removing the structural bottleneck that previously made systematic exploration of aggressive low-step regimes costly.
Abstract
Real-time interactive video generation requires low-latency, streaming, and controllable rollout. Existing autoregressive (AR) diffusion distillation methods have achieved strong results in the chunk-wise 4-step regime by distilling bidirectional base models into few-step AR students, but they remain limited by coarse response granularity and non-negligible sampling latency. In this paper, we study a more aggressive setting: frame-wise autoregression with only 1--2 sampling steps. In this regime, we identify the initialization of a few-step AR student as the key bottleneck: existing strategies are either target-misaligned, incapable of few-step generation, or too costly to scale. We propose Causal Forcing++, a principled and scalable pipeline that uses causal consistency distillation (causal CD) for few-step AR initialization. The core idea is that causal CD learns the same AR-conditional flow map as causal ODE distillation, but obtains supervision from a single online teacher ODE step between adjacent timesteps, avoiding the need to precompute and store full PF-ODE trajectories. This makes the initialization both more efficient and easier to optimize. The resulting pipeline,, surpasses the SOTA 4-step chunk-wise Causal Forcing under the frame-wise 2-step setting by 0.1 in VBench Total, 0.3 in VBench Quality, and 0.335 in VisionReward, while reducing first-frame latency by 50% and Stage 2 training cost by about 4 times. We further extend the pipeline to action-conditioned world model generation in the spirit of Genie3. Project Page: https://github.com/thu-ml/Causal-Forcing and https://github.com/shengshu-ai/minWM.
Sources
- Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models
- Wan: Open and Advanced Large-Scale Video Generative Models
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
- Open-Sora Plan: Open-Source Large Video Generation Model
- Open-Sora: Democratizing Efficient Video Production for All
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
- Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length
- Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation
- StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human Avatars
- Vidarc: Embodied Video Diffusion Model for Closed-loop Control
- Pyramidal Flow Matching for Efficient Video Generative Modeling
- MAGI-1: Autoregressive Video Generation at Scale
- SkyReels-V2: Infinite-length Film Generative Model
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
- Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation
- Phased Consistency Models
- Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models
- Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency
- Score-Based Generative Modeling through Stochastic Differential Equations
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models