How Do Video Foundation Models Encode Intuitive Physics? Probing Across Pretraining Paradigms
cs.CV, cs.AI, cs.LG
Submitted: 2026-06-08
Updated: 2026-09-13
License: http://creativecommons.org/licenses/by/4.0/
The gist: We study whether pretrained video foundation models encode intuitive-physics information in their frozen representations, and how this information varies across model families, layers, and probe
Terminology
Abstract
We study whether pretrained video foundation models encode intuitive-physics information in their frozen representations, and how this information varies across model families, layers, and probe types. Using frozen-feature probing on IntPhys2 and Minimal Video Pairs (MVP), we compare predictive joint-embedding models (V-JEPA), masked reconstruction models (VideoMAE), and a diffusion-based video generator (LTX-Video). V-JEPA achieves the strongest overall results across benchmarks, especially with probes that model temporal dynamics, while VideoMAE remains competitive and LTX-Video recovers weaker but non-trivial signal. Layerwise analyses show that physics-relevant information is weakest in early layers and becomes most accessible at intermediate-to-late depth, and temporal controls show that disrupting frame order substantially reduces performance, especially on MVP. Together, these results suggest that intuitive-physics knowledge emerges reliably in pretrained video representations, but its accessibility depends strongly on pretraining paradigm, representational depth, and readout mechanism.
Sources
- Understanding intermediate layers using linear classifier probes
- Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- LTX-Video: Realtime Video Latent Diffusion
- Revisiting Feature Prediction for Learning Visual Representations from Video
- Interpreting Physics in Video World Models
- Better plain ViT baselines for ImageNet-1k
- Linguistic Knowledge and Transferability of Contextual Representations
- Decoupled Weight Decay Regularization
- Attention, Please! Revisiting Attentive Probing Through the Lens of Efficiency
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
- VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking
- Video Diffusion Models are Training-free Motion Interpreter and Controller
- Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models