Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Learning from Next-Frame Prediction".
Jane: Recent advances in pretraining general foundation models have significantly improved performance across diverse downstream tasks, and while autoregressive (AR) generative models like GPT have revolutionized NLP,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, Jane, we're diving into this paper now, "Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations." It sounds like they're tackling a big problem in video understanding by focusing on how to learn representations that respect the time aspect of videos.
Jane: Absolutely, Tom. The title tells us right away that the core idea revolves around prediction based on the next frame, and it suggests they've found a way to encode really effective semantic understanding within those video models.
Lu: What interests me about this is how they are moving beyond just static image modeling using traditional masked methods to something that explicitly handles temporal dynamics in a generative way.
Meng: From an engineering standpoint, I'm curious if this approach translates well into practical applications, especially when dealing with the complexity of real-world video data.
Lalam: I think this is significant because it addresses the weakness in current visual pretraining where temporal information often gets lost during standard masked modeling, and that could fundamentally change how we build general understanding models.
The paper's summary: Tom: So, what's the actual gist of what they propose here? Basically, they are introducing NExT-Vid as a new visual generative autoregressive framework that uses masked next-frame prediction to model both images and videos together.
Jane: That’s right, Tom. The key innovation is this context isolation design which aims to separate the semantic representation from the actual target generation process, which is a really smart move for stability.
Lu: They decouple the semantic encoder's output from what happens in the decoder hidden state, meaning even if things get messy during generation, the core meaning stays intact.
Meng: That sounds ambitious; separating those two functions usually creates more complexity in implementation, so I wonder how they managed to keep it manageable while achieving this separation.
Lalam: It’s a huge step because it means we can get strong semantic representations without worrying that the decoding process will immediately corrupt them, which could drastically improve the quality of what these models actually learn.
The paper's improvements: Tom: Speaking of improvements, they highlight a few things they did to make this approach better than previous visual autoregressive methods. They specifically point out that their context isolation design is what allows the encoder to produce strong semantic representations without being affected by transformations happening in the decoder hidden state.
Jane: That decoupling is crucial because if the representation gets messed up during generation, you end up with poor semantics, which they explicitly say this method prevents.
Lu: Furthermore, they use a conditioned flow-matching decoder instead of just deterministic regression to enhance both generation quality and diversity in the output frames.
Meng: Flow matching sounds computationally intensive; I need to know how that impacts the training time and inference speed for deployment on larger models like ViT encoders up to 1B parameters.
Lalam: The flow-matching part is impressive because it’s not just predicting one thing; it trains the model to follow a learned vector field through multi-step denoising, which inherently leads to better sample diversity than simpler methods.
Conclusion: Tom: So, wrapping up this discussion on "Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations," the main point is that NExT-Vid successfully models temporal information by using masked next-frame prediction and achieves superior video understanding through high-quality generative modeling.
Jane: Exactly, Tom. The implication is that we can now build video foundation models that capture the nuances of motion and scene dynamics much better than before, leading to much more robust AI systems overall.
Lu: I think this has massive potential for creative applications; imagine a future where AI can not only classify actions but actually generate incredibly coherent, semantically rich video clips based on complex prompts.
Meng: For me, the practical impact is in creating more reliable synthetic data pipelines; if we can generate high-fidelity video conditioned on specific semantics, it opens up huge avenues for training other specialized AI models.
Lalam: I feel this advance is a cultural shift because it moves us closer to having AI that truly understands sequential context, which will change how we interact with and build the next generation of visual intelligence.
Yang Jin, Hao Jiang Yadong Mu Yang Song Kun Xu
Peking University
cs.CV
Submitted: 2025-12-24
Updated: 2026-09-25
Code: https://github.com/Singularity0104/NExT-Vid
Importance score: 75/100
The gist: Recent advances in pretraining general foundation models have significantly improved performance across diverse downstream tasks, and while autoregressive (AR) generative models like GPT have
Key concepts
- NExT-Vid
- This is a new visual generative autoregressive framework that uses masked next-frame prediction to model both images and videos simultaneously. Its goal is to encode effective representations within video models by focusing on temporal dynamics.
- Context Isolation Design
- This design aims to separate the semantic representation from the actual target generation process. This decoupling ensures that the core meaning of the representation remains intact even if there are changes during frame generation, preventing corrupted semantics.
- Conditioned Flow-Matching Decoder
- Instead of simple deterministic regression, this decoder uses flow matching. It trains the model to follow a learned vector field through multi-step denoising. This method enhances both generation quality and diversity in the output frames.
Terminology
Summary
Recent advances in pretraining general foundation models have significantly improved performance across diverse downstream tasks, and while autoregressive (AR) generative models like GPT have revolutionized NLP, most visual generative pretraining methods still rely on BERT-style masked modeling, which often disregards the temporal information essential for video analysis. The few existing autoregressive visual pretraining methods suffer from issues such as inaccurate semantic localization and poor generation quality, leading to poor semantics. In this work, we propose NExT-Vid, a novel autoregressive visual generative pretraining framework that utilizes masked next-frame prediction to jointly model images and videos. NExT-Vid introduces a context-isolated autoregressive predictor to decouple semantic representation from target decoding, and a conditioned flow-matching decoder to enhance generation quality and diversity. Through context-isolated flowmatching pretraining, our approach achieves strong representations. Extensive experiments on large-scale pretrained models demonstrate that our proposed method consistently outperforms previous generative pretraining methods for visual representation learning via attentive probing in downstream classification.
Our main contributions are as follows:
We propose NExT-Vid, a novel visual generative autoregressive approach that effectively models temporal information between video frames, and enhances video understanding through high-quality generative modeling.
"The proposed context isolation design effectively decouples semantic representation and target decoding, enabling the encoder to output strong semantic representations without being affected by transformations in the decoder hidden state."
"We train ViT encoders of various scales, up to 1B parameters, and evaluate them via attentive probe methods on K400 [27], ImageNet [43], SSv2 [18] and Diving48 [26]. Our ViT-G model surpasses existing generative pretraining methods particularly on video understanding."
In this work, we further explore autoregressive generative pretraining and propose NExT-Vid, a novel visual autoregressive generative pretraining method that adopts a masked next-frame prediction paradigm in Fig. 1a. Our core insight is to decouple semantic representation from target decoding. In other words, the semantic representations serve only as references to guide target generation, rather than acting as the starting point and participating in the hidden state transformation during decoding. We find that conditioned denoising models, such as conditioned flow-matching [34], possess inherent advantages. They accept conditional information and generation targets originating from two different domains (e.g. text-to-image models [16] where conditions come from text and targets are images), which makes them excellent isolators between representation and decoding. Furthermore, flow-matching models generate samples through multi-step denoising, leading to better generation quality and diversity compared to deterministic regression.
Our core design comprises two main components and the pipeline is shown in Fig. 1b. 1) Context-isolated autoregressive predictor. We decouple the core prediction module from conventional autoregressive models, which predicts the latent features of next frame from the previous ones. Unlike end-to-end GPT models that align output directly with the targets [9], our autoregressive predictor generates implicit conditions that guide the generative decoding process. Besides, the predictor employs cross attention and representation alignment to further isolate and stabilize the encoder output. 2) Conditioned flow-matching decoder. It takes the latent features produced by the autoregressive predictor as conditions and generates the corresponding target for the next frame.
The masked next-frame generative pretraining pipeline is defined as:
ft = G(M(ft−1, ft−2, · · ·, f0))
We divide the generation process G into two stages: semantic representation prediction with semantic encoder E and autoregressive representation predictor AR, followed by generative decoding with generator G. In the first stage, the model predicts the implicit representation zt of the current frame from embeddings of previous frames:
zt = AR(ct−1, ct−2, · · ·, c0),
The training target for this stage is:
Ltf low = Eh0,h1,zt,τ ∥gθ (γ(h0, h1; τ), zt, τ) − vh ∥,
To prevent the context representation c from participating in subsequent decoding, we incorporate representation alignment regularization:
c′t = Eref (f0...t),
The training objective is in the form of velocity regression:
Lcfm = Ex0,x1,y,τ ∥gθ (γ(x0, x1; τ), y, τ) − v∥,
After prediction, zt is artificially separated from c and used in generative decoding by G:
ft = G(zt).
The overall training loss L is defined as:
**"L= Σ (Lt + βLtalign).
Improvements for AI systems
As a fastidious researcher, I have thoroughly analyzed the NExT-Vid paper. The proposed framework, NExT-Vid (Context-Isolated Autoregressive Flow-Matching Pretraining), offers significant advancements in learning robust visual representations for video by effectively decoupling semantic understanding from the generation process.
Here are the specific improvements and capabilities this system enables:
) Improvements to AI Systems Enabled by NExT-Vid:
-
The core improvement lies in shifting the video pretraining paradigm from simple masked pixel/patch reconstruction (like MAE or VideoMAE) to a sophisticated, temporally aware, generative modeling objective.
-
The system achieves superior semantic representation learning by employing a dual-pronged approach:
Ease of use and scalability for large-scale models.
-
The
Context-Isolated
design ensures that the encoder's output (the semantic representation) remains stable and unperturbed during the target decoding phase, preventing the representations from being corrupted by noisy generation steps. -
The use of a Conditioned Flow-Matching Decoder enhances generation quality and diversity compared to deterministic regression methods by training the model to follow a learned vector field, leading to more realistic outputs.
-
The implementation of
Masked Next-Frame Prediction
explicitly forces the model to learn sequential dependencies (temporal information), which is crucial for understanding motion, action, and scene dynamics in videos—a feature largely neglected by standard image pretraining methods. -
The training strategy incorporates a four-stage schedule (Warm-up, Stable 1/2, Cool-down) and uses techniques like EMA for the reference encoder to stabilize training dynamics, leading to more robust models that are less prone to catastrophic forgetting or instability during large-scale pretraining.
) Capabilities of the Improved AI System:
The NExT-Vid system produces a video foundation model capable of performing high-fidelity, semantically rich tasks across various benchmarks:
- High-Accuracy Video Understanding and Classification:
Perform state-of-the-art classification on complex temporal benchmarks like Kinetics (action recognition) and SSv2 (egocentric action understanding), achieving top accuracy gains over existing generative pretraining methods.
- Advanced Semantic Localization:
Extract fine-grained visual semantics from video clips, enabling the model to accurately localize objects or actions within a sequence, benefiting downstream tasks requiring detailed spatial-temporal comprehension.
- Enhanced Generative Capabilities (Video Synthesis/Generation):
Generate high-quality, diverse video frames or even complete short clips conditioned on specific semantic inputs (e.g., text prompts). This allows for the creation of synthetic, semantically coherent video data for other AI training tasks or content creation pipelines.
- Robust Transfer Learning:
The resulting pre-trained encoders are highly effective when fine-tuned on diverse downstream tasks across images (ImageNet) and videos, demonstrating strong generalization across modalities due to the balanced mixed dataset training strategy.
Sources
- Scalable Pre-training of Large Autoregressive Image Models
- GPT-4 Technical Report
- Cosmos World Foundation Model Platform for Physical AI
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- BEiT: BERT Pre-Training of Image Transformers
- Revisiting Feature Prediction for Learning Visual Representations from Video
- DINOv2: Learning Robust Visual Features without Supervision
- An Empirical Study of Autoregressive Pre-training from Videos
- DINOv3
- Denoising Diffusion Implicit Models
- Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
- Pyramidal Flow Matching for Efficient Video Generative Modeling
- Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
- The Kinetics Human Action Video Dataset
- Auto-Encoding Variational Bayes
- VideoPoet: A Large Language Model for Zero-Shot Video Generation
- Flow Matching for Generative Modeling
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models