Algebraic Consistency Alone Does Not Certify Temporal Structure in Latent Action Models
cs.CV, cs.RO
Submitted: 2026-09-20
Updated: 2026-09-20
Terminology
Sources
- From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data
- ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models
- Learning Additively Compositional Latent Actions for Embodied AI
- RotVLA: Rotational Latent Action for Vision-Language-Action Model
- DLAM: Distributional Latent Actions with Temporal Constraints
- ABot-M0.5: Unified Mobility-and-Manipulation World Action Model
- What Matters for Latent Actions in Robot Learning
- Latent Action Learning Requires Supervision in the Presence of Distractors
- Why Latent Actions Fail, and How to Prevent It
- MVP-LAM: Learning Action-Centric Latent Action via Cross-Viewpoint Reconstruction
- Causally Debiased Latent Action Model for Embodied Action-Conditioned World Models
- ATM: Why Latent World Models Can Fail to Plan
- ConLA: Contrastive Latent Action Learning from Human Videos for Robotic Manipulation
- CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- DINOv2: Learning Robust Visual Features without Supervision
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models