StrucPhysVideo: Learning Physical Dynamics from Structured Captions and Robot Actions
cs.CV
Submitted: 2026-09-16
Updated: 2026-09-16
Code: https://github.com/westlakedi-awomo/StrucPhysVideo
Project page: https://westlakedi-awomo.github.io/StrucPhysVideo-Page
Terminology
Sources
- Physics-IQ Verified
- Qwen3-VL Technical Report
- Wan: Open and Advanced Large-Scale Video Generative Models
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
- Classifier-Free Diffusion Guidance
- Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation
- Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation
- One-step Diffusion with Distribution Matching Distillation
- TransNet V2: An effective deep network architecture for fast shot transition detection
- Cosmos 3: Omnimodal World Models for Physical AI
- AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
- DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
- Learning Transferable Dynamics Priors from Action to World Modeling
- WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models