J-LAW: Joint Localization and Action-Conditioned World Modeling via Coupled Latent Factor Graphs

arXiv:2606.28712 · cs.RO, cs.LG · Submitted 2026-06-27 · Read on arXiv

cs.RO, cs.LG

Submitted: 2026-06-27

Updated: 2026-08-30

Comments: Under review

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

The gist: Classical simultaneous localization and mapping (SLAM) estimates metric poses and a geometric map but does not provide an action-conditioned predictive state.

Terminology

Abstract

Classical simultaneous localization and mapping (SLAM) estimates metric poses and a geometric map but does not provide an action-conditioned predictive state. Action-conditioned world models learn compact latent dynamics but ignore global metric consistency and accumulate drift under open-loop rollout. We introduce J-LAW (Joint Localization and Action-Conditioned World Modeling), a unified factor-graph formulation that connects metric pose variables, predictive latent states, and persistent latent landmarks in this letter.J-LAW represents each image as a compact predictive state and combines it with pose or motion measurements through a separately learned mapping. Its maximum a posteriori (MAP) factor graph enforces consistency between these complementary sources of information over time. Experiments on PushT and WildGS show that J-LAW's factor-graph representation can improve long-horizon latent consistency and recover more reliable predictive states under partial observations, forming a foundation for future integrated localization and planning systems.

Sources

Related papers