Mode Seeking meets Mean Seeking for Fast Long Video Generation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Mode Seeking meets Mean Seeking for Fast Long Video Generation".
Jane: The paper was written by the authors from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Summary: Tom: In the summary of Mode Seeking meets Mean Seeking for Fast Long Video Generation, they clearly articulate the fundamental problem of scaling from seconds to minutes.
Jane: The authors are saying that while short-clip data is abundant, high-quality long-form content is scarce and usually exists in narrow domains.
Tom: This creates a critical gap between local quality and long-range coherence, which I think most existing systems struggle to bridge effectively.
Lu: They argue that just trying to interpolate across different temporal horizons—like an image model does—doesn't work for videos because the time dimension is an extrapolation, not just a multi-resolution interpolation.
Meng: That’s a huge conceptual difference; they aren't just making the video longer, they are adding entirely new causal structure and events that were never in the training data.
Lalam: It’s about teaching something entirely new to the model while maintaining what it already knows about realism, which is a massive challenge.
Tom: The paper says their solution is a unified representation where global Flow Matching captures narrative structure and local Distribution Matching ensures local realism by aligning every sliding window segment to this frozen short-video teacher.
Jane: So, instead of trying to force the model to learn long-range coherence from limited data, they are giving it a "teacher" that already knows how to generate realistic short clips.
Tom: It’s a fantastic way to frame the problem so we can move onto how this decoupling actually works in practice.
Improvements/Methodology: Tom: Now, let's talk about the methodology in Mode Seeking meets Mean Seeking for Fast Long Video Generation, because it’s where the actual magic happens.
Jane: The core idea is that they use a Decoupled Diffusion Transformer, or DDT, which separates the two training goals into two distinct heads.
Tom: One head is responsible for the global structure using supervised flow matching on real long videos—what they call Mean Seeking.
Lu: And the other part of the process, Mode Seeking, is where they use a local Distribution Matching head to align sliding windows to that expert short-video teacher.
Meng: This alignment uses this specific technique called a mode-seeking reverse-KL divergence, which is brilliant because it means the student doesn's trying to average out possibilities but commit strongly to the teacher’s high-fidelity modes.
Lalam: It’s a way of saying that we let the long videos teach us *what* to do globally, and we let the short video teacher remind us *how* to do it locally.
Tom: The paper shows that by sharing a unified context encoder, they manage these two separate objectives without them interfering with each other, which is a big win.
Jane: This is important because typically forcing two different training goals into one single velocity predictor would cause gradient interference and ruin the results.
Tom: It’s really elegant how they utilize this dual-head architecture to achieve both long-range coherence and local quality simultaneously, setting up the next natural question.
Results & Ablation: Tom: The results presented in Mode Seeking meets Mean Seeking for Fast Long Video Generation show that this dual approach works incredibly well across various scenarios.
Jane: In the qualitative comparison, we see that while SFT baselines might achieve some long context, their outputs are often blurry and lack fine detail.
Tom: And those teacher-only methods, like CausVid or Self-Forcing, struggle to maintain realistic evolution over time because they can't model long-range structure.
Lu: The data suggests that without this decoupling, the model is simply trying to solve an unsolvable problem for both local realism and global narrative.
Meng: The ablation study really confirms that all three pieces—the decoupled DDT heads, the sliding window DMD technique, and the SFT on real long clips—are essential.
Lalam: It’s fascinating how the model manages to achieve both sharp local motion and consistent scene continuity across frames.
Tom: The paper demonstrates that by achieving this balance, they have effectively closed that fidelity-horizon gap.
Conclusion: Tom: So, as we wrap up our discussion on Mode Seeking meets Mean Seeking for Fast Long Video Generation, it’s clear this is a major step forward in video generation.
Jane: We've seen how the dual system successfully leverages both the abundance of short-clip data and the scarcity of long-form data to create high-fidelity, minute-scale videos.
Tom: It moves us beyond just trying to make things longer and into finding a way to make them *coherent* while maintaining sharpness.
Lu: I think this work opens up so many possibilities for complex storytelling that was previously impossible due to data constraints.
Meng: From an engineering standpoint, the fact that the Distribution Matching head can serve as a fast, few-step sampler is extremely promising for real-time applications.
Lalam: The ability to improve local realism while extending temporal coherence has profound implications for how we will interact with dynamic content in the future.
Tom: I think this paper truly manages to meet the Mean Seeking and Mode Seeking objectives perfectly.
Jane: It’s a relief to see a method that doesn't just compromise between sharpness and duration, but delivers both.
Lu: I can't wait to see what more creative applications emerge now.
Meng: I hope the next steps in scaling this work are focused on making it even faster for everyone.
Lalam: We will be watching how this changes the way we consume and create stories online.
cs.CV, cs.LG
Submitted: 2026-02-27
Updated: 2026-09-13
Code: https://github.com/NVlabs/FastGen
Project page: https://primecai.github.io/mmm
Importance score: 89/100
The gist: The paper details a method titled "Mode Seeking meets Mean Seeking for Fast Long Video Generation." Implementation Details: The model is trained on "A100 and GB200 GPUs." The training process
Key concepts
- Decoupled Diffusion Transformer (DDT)
- This is the core architecture that separates two distinct training goals into separate heads. By using this dual approach, the system avoids gradient interference and allows it to achieve both long-range coherence and local quality simultaneously.
- Mean Seeking
- This process utilizes supervised flow matching on real long videos. Its function is to capture the global narrative structure of the video, ensuring that the resulting output maintains consistent continuity over extended periods of time.
- Mode Seeking
- This technique employs a local Distribution Matching head. It aligns sliding window segments to a frozen short-video teacher, compelling the model to commit strongly to high-fidelity modes rather than averaging possibilities.
Terminology
Summary
The paper details a method titled Mode Seeking meets Mean Seeking for Fast Long Video Generation.
Implementation Details:
The model is trained on A100 and GB200 GPUs.
The training process utilizes dynamic batching, where training videos [are] treated as variable-length sequences and pre-processing them into length-based buckets for sampling.
Furthermore, Variable-length training is used across training, reducing padding waste and IDLE time.
To achieve long contexts, the authors employ the DeepSpeed Ulysses (Jacobs et al., 2023) sequence-parallelism strategy,
using a sequence-parallelism group size of 4 for A100 GPUs, and 2 for GB200 GPUs.
All implementation was conducted on top of the FastGen (Nie et al., 2026) repository.
Data:
The dataset collection is drawn from multiple sources. Specifically, the authors use all videos available from the Sekai dataset (Li et al., 2025c), and a subset from MiraData (Ju et al., 2024).
Additionally, they filter out single-shot videos from randomly collected internet videos.
Collectively, these datasets comprise more than 100k videos ranging from 10 seconds to minutes, with an average of 31 seconds.
The temporal scope is constrained by setting the temporal upper bound to 61 seconds and subsample when a video exceeds it.
Sliding Window DMD Implementation:
A critical technical challenge addressed is the semantic mismatch that occurs when applying DMD with sliding windows on long latent sequences
in large-scale video latent diffusion models, where the latent space includes both image and video frame latents. The issue arises because a window cropped from the middle begins with a video latent, while the teacher expects the first latent of a clip to be an image latent.
To resolve this, they adopt a strategy following LongLive (Yang et al., 2026): "for any window starting at offset p > 0, we decode the latent prefix [0,..., p - 1] with the frozen VAE, take the last decoded RGB frame, and re-encode it into an image latent; we then prepend this reconstructed image latent to the student’s windowed video latents before computing the DMD loss, masking out the reconstructed latent to avoid backpropagating through the VAE. The authors found this strategy
to be effective not only for causal AR models (Yang et al., 2026), but also for non-causal bidirectional models like ours."
Evaluation Criteria:
The evaluation task requires scoring the semantic consistency
of a video on a 0-100 scale. Semantic consistency is defined as when objects, identities, attributes, and the overall scene remain coherent over time; penalize sudden unrealistic changes, object identity swaps, content drift, or contradictions between frames.
Crucially, the evaluation must account for static content: "If the video is essentially a still image / frozen frame(s) with little-to-no motion or temporal change, do NOT give a high score. In that case, assign a low score because it does not demonstrate temporal consistency under motion."
Limitations and Future Work:
The method is noted to be orthogonal to causal autoregressive methods (Yin et al., 2025; Huang et al., 2025b).
Potential follow-ups include using the model as a base model for causal AR training, or, conversely, to distill our long-context bidirectional model into a causal sampler.
Due to the student being trained with substantially longer native temporal context than typical short-clip teachers,
the authors expect it to extrapolate well to minute-scale (and longer) generation,
particularly when combined with techniques such as Rolling Forcing (Liu et al., 2025), LongLive (Yang et al., 2026), or longer-context positional embedding schemes.
More broadly, the native long-context encoder provides a persistent history representation in the spirit of Genie-style models,
making adding interaction/action conditioning on top of this representation an especially promising direction.
Improvements for AI systems
The core improvement is a Decoupled Diffusion Transformer (DDT) training paradigm that resolves the fundamental conflict between global narrative coherence and local high-fidelity realism, enabling stable, fast long-form video synthesis.
1. Implementation of Decoupled Dual-Head Architecture:
-
Unified Backbone: Utilize a shared, full-range temporal attention encoder (E phi) to process the noisy latent representation (x long t), conditioning on text and time step t. This ensures the global context is represented consistently across all learning objectives.
-
Dual Decoders: Attach two distinct, lightweight velocity prediction heads atop E phi:
-
Flow Matching (FM) Head (D theta FM): Responsible for capturing long-range dynamics. This head is trained using a supervised flow-matching objective on the full range of real, long-form videos (x long about pi). This anchors the student model to real, complex temporal trajectories and narrative structure.
-
Distribution Matching (DM) Head (D psi): Responsible for local realism. This head is trained via a mode-seeking reverse-KL divergence. It operates on sliding windows of the student's generated rollouts, aligning them to a frozen, expert short-video teacher distribution (p teacher).
2. Refined Training Methodology (Mode Seeking Meets Mean Seeking):
-
The system utilizes a joint loss function: L total = L SFT(phi, theta) + lambda seg L seg(phi, psi).
-
Ablation of Gradient Interference: The decoupling ensures that the mean-seeking SFT objective (which encourages averaging under ambiguity when training on scarce long data) does not interfere with the mode-seeking teacher alignment (which forces fidelity to a high-density local mode). This allows both objectives to be optimized simultaneously via the shared encoder.
-
Inference Optimization: The final production model is trained such that only the DM head (D psi) is required for inference, eliminating the need for complex, multi-stage training or full diffusion sampling.
The resulting AI system possesses capabilities far superior to standard SFT or teacher-only methods:
1. Seamless Long-Form Generation:
- It can synthesize coherent video sequences extending from seconds to minutes (up to 60+ seconds), maintaining consistent narrative structure and causal flow over the entire duration.
2. Guaranteed Local High Fidelity:
- The system guarantees that every local segment (sliding window) maintains the sharp textures, fine details, and realistic motion patterns of a high-fidelity short-video expert teacher, preventing the
softening
orblurring
artifacts common in long-form SFT models.
3. Fast Inference and Low Computational Overhead:
- By utilizing the DM head as a few-step sampler, the system achieves rapid video generation at inference time without requiring complex autoregressive rollouts or extensive multi-stage distillation processes.
4. Robust Domain Generalization:
- It effectively leverages the
long-context anchor
from scarce real long videos while inheriting local realism from a frozen teacher, allowing it to generalize across diverse domains without needing massive amounts of expensive, high-resolution long-form training data.
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models