LaGSplat: Inferring Physics-Governed Interactive Simulation from Monocular Video Using Latent Lagrangian Gaussian Splatting
cs.CV, cs.LG
Submitted: 2026-08-17
Updated: 2026-09-15
Comments: 25 pages, 11 figures, 4 tables. Project page with interactive demo: https://louenpottier.github.io/lagsplat.html
Code: https://github.com/katfriedl/reduced_hamiltonians
Project page: https://louenpottier.github.io/lagsplat.html
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: We present LaGSplat (Latent Lagrangian Gaussian Splatting), a framework that infers interactive, physics-governed dynamics from one or a few monocular videos.
Terminology
Abstract
We present LaGSplat (Latent Lagrangian Gaussian Splatting), a framework that infers interactive, physics-governed dynamics from one or a few monocular videos. At inference it lets a user push on the filmed object, rigid or deformable, with an external force that was never measured, annotated, or seen during training. This is possible because a low-dimensional latent state q in R d plays two roles at once: it is the generalised coordinate of a learned dissipative Lagrangian and the conditioning variable of a Gaussian Splatting decoder. The inductive bias of this decoder, whose primitives are explicit points μ i(q) that move with the object, is what lets a force f applied in the image pull back into a latent generalised force J(q) f and enter the equations of motion, which pixel-space (CNN) or neural-field (NeRF) decoders cannot do. We validate LaGSplat on test cases of increasing difficulty, from rigid to deformable and from autonomous to forced real systems, combining monocular video and sensor measurements. We further demonstrate interactive use: forces of arbitrary magnitude and direction can be applied to the reconstructed object at any time, its response rendered in real time, in 2D or 3D. Assuming a dissipative Euler-Lagrange equation over a few generalised coordinates trades generality for a bounded, plausible response to unseen forces, where an unconstrained predictor diverges.
Sources
- 4D Gaussian Splatting as a Learned Dynamical System
- PHAST: Port-Hamiltonian Architecture for Structured Temporal Dynamics Forecasting
- Learning Physics From Video: Unsupervised Physical Parameter Estimation for Continuous Dynamical Systems
- Discovering State Variables Hidden in Experimental Data
- Learning Relativistic Geodesics and Chaotic Dynamics via Stabilized Lagrangian Neural Networks
- Learning mechanical systems from real-world data using discrete forced Lagrangian dynamics
- Category-Agnostic Neural Object Rigging
- A Riemannian Framework for Learning Reduced-order Lagrangian Dynamics
- Latent Neural ODEs with Sparse Bayesian Multiple Shooting
- Learning Hamiltonian Dynamics at Scale: A Differential-Geometric Approach
- NeuROK: Generative 4D Neural Object Kinematics
- Differentiable neural network representation of multi-well, locally-convex potentials
- Investigating Lagrangian Neural Networks for Infinite Horizon Planning in Quadrupedal Locomotion
- Learning Visually Interpretable Oscillator Networks for Soft Continuum Robots from Video
- Learning Physically Consistent Lagrangian Control Models Without Acceleration Measurements
- Pixie: Fast and Generalizable Supervised Learning of 3D Physics from Pixels
- Dynamic Appearance Particle Neural Radiance Field
- Sharp Monocular View Synthesis in Less Than a Second
- PokeFlex: A Real-World Dataset of Volumetric Deformable Objects for Robotics
- ParticleGS: Learning Neural Gaussian Particle Dynamics from Videos for Prior-free Physical Motion Extrapolation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models