4DGS-JEPA: Temporally Compositional Joint-Embedding Prediction for Dynamic Gaussian Splatting
cs.AI
Submitted: 2026-08-24
Updated: 2026-08-24
Comments: 44 pages
License: http://creativecommons.org/licenses/by/4.0/
The gist: Dynamic Gaussian Splatting provides an explicit representation of evolving 3D scenes, but existing approaches are primarily optimized for reconstruction, future-state generation, or rendering rather
Terminology
Abstract
Dynamic Gaussian Splatting provides an explicit representation of evolving 3D scenes, but existing approaches are primarily optimized for reconstruction, future-state generation, or rendering rather than for learning reusable predictive dynamics. We propose 4DGS-JEPA, a Gaussian-native joint-embedding predictive architecture for causal multi-horizon prediction over dynamic Gaussian scenes. The model uses a hierarchical scene-, motion-group-, and Gaussian-level representation together with a horizon-conditioned transition operator that supports both direct prediction and recursive rollout. Its central principle is temporal composition: different chronological transition paths reaching the same future endpoint should produce compatible predictive states. Endpoint and multi-horizon path supervision anchor these predictions to future target embeddings, while a selective geometry decoder and geometry-level composition ground the learned dynamics in consistent group motion and Gaussian geometry without requiring complete future appearance reconstruction. We further introduce a hybrid correspondence mechanism that combines persistent canonical identity with residual optimal-transport matching under reordering and topology change. We characterize zero-loss path agreement and finite-error rollout accumulation theoretically. Three controlled experiments provide mechanism-level evidence that temporal composition reduces latent path dependence while retaining predictive accuracy, geometry-level composition improves consistency of decoded motion, and hybrid correspondence preserves reliable identity while remaining robust when correspondence becomes ambiguous. Together, 4DGS-JEPA provides a predictive, temporally compositional formulation of dynamic Gaussian worlds.
Sources
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation
- BiJEPA: Bi-directional Joint Embedding Predictive Architecture for Symmetric Representation Learning
- Space-Time Forecasting of Dynamic Scenes with Motion-aware Gaussian Grouping
- GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
- V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
- D-NeRF: Neural Radiance Fields for Dynamic Scenes
- GaussianDream: A Feed-Forward 3D Gaussian World Model for Robotic Manipulation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection