Future Video Generation Better Aligns with the Human Visual Cortex than Observed Video
cs.CV, cs.AI, q-bio.NC
Submitted: 2026-09-30
Updated: 2026-09-30
Terminology
Sources
- Self-Rectifying Diffusion Sampling with Perturbed-Attention Guidance
- The Invisible Hand of Physics: When Video Diffusion Models Know More Than They Show
- TokenFlow: Consistent Diffusion Features for Consistent Video Editing
- Diffusion-Based Action Recognition Generalizes to Untrained Domains
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
- Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling
- Gen4U: Unifying Video Generation and Understanding via Diffusion
- Flow Matching for Generative Modeling
- Emergent Temporal Correspondences from Video Diffusion Transformers
- Scalable Diffusion Models with Transformers
- Probing and Leveraging Video Diffusion Transformer Features for Robust Point Tracking
- Wan: Open and Advanced Large-Scale Video Generative Models
- Diffusion Models Generate Images Like Painters: an Analytical Theory of Outline First, Details Later
- From Slow Bidirectional to Fast Autoregressive Video Diffusion Models
- Video-Mirai: Autoregressive Video Diffusion Models Need Foresight
- Helios: Real Real-Time Long Video Generation Model
- Denoise to Track: Harnessing Video Diffusion Priors for Robust Correspondence
- VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models