ACID: Action Consistency via Inverse Dynamics for Planning with World Models
summary
The gist
The gist The proposed ACID framework introduces cycle action consistency into decision-time planning for action-conditioned world models to ensure predicted trajectories are realizable, which
In short
ACID is a decision-time planning framework that improves action-conditioned world models by adding cycle action consistency to the planning cost. It ensures predicted trajectories are physically realizable by checking if an inverse dynamics model's inferred action matches the conditioning action. This consistently enhances planning performance across various tasks and world models.
Key concepts
- Cycle Action Consistency
- This is a check that verifies if the action used to predict a transition is consistent with what an inverse dynamics model would infer from that same transition. It detects when a predicted trajectory drifts away from the actions it was conditioned on, ensuring physical plausibility step-by-step.
- Inverse Dynamics Model (IDM) Gϕ
- The IDM is a model used backward in time to infer what action must have caused a specific state transition. By using this model, ACID can calculate the 'expected' action for every step of a trajectory, allowing it to measure the difference against the actual planned actions.
- Augmented Planning Cost
- This is the total cost used by the planner to select good action sequences. It combines two parts: a goal cost (how close you are to your desired outcome) and an action consistency cost (how much each step deviates from what physics suggests should happen). This forces the planner to choose paths that are both efficient and physically possible.
Terminology used across episodes
This episode discusses
- ACID: Action Consistency via Inverse Dynamics for Planning with World Models · Paper Radio
- World Models
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- GAIA-1: A Generative World Model for Autonomous Driving
- World Action Models are Zero-shot Policies
- FreeAction: Training-Free Techniques for Enhanced Fidelity of Trajectory-to-Video Generation
- GigaWorld-Policy: An Efficient Action-Centered World--Action Model
- DreamGen: Unlocking Generalization in Robot Learning through Video World Models
- Rectified Flow: A Marginal Preserving Approach to Optimal Transport
- pi 0: A Vision-Language-Action Flow Model for General Robot Control
- LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
- DeepMind Control Suite
- AdaptiGraph: Material-Adaptive Graph-Based Neural Dynamics for Robotic Manipulation
The paper
ACID: Action Consistency via Inverse Dynamics for Planning with World Models · Read on arXiv
POSTECH · KAIST
Decision-time planning with action-conditioned world models has become a popular paradigm for embodied control. However, the standard planning cost judges a candidate solely by how close its predicted terminal state lies to the goal, leaving the realizability of the intermediate transitions unchecked--a predicted trajectory can look convincing while the environment rollout drifts away from it. In this paper, we propose ACID, a decision-time planning framework that introduces cycle action consistency: the action inferred backward from a predicted transition by an inverse dynamics model should recover the one that was conditioned on. We fold this per-step residual into the planning cost via a scale-invariant adaptive weight. Across four action-conditioned world models and eight tasks encompassing object manipulation and articulated control in simulation, visual navigation, and real-robot manipulation, ACID consistently improves planning and matches the baseline's accuracy with substantially less planning compute.
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "ACID: Action Consistency via Inverse Dynamics for Planning with World Models".
Rosa: The gist The proposed ACID framework introduces cycle action consistency into decision-time planning for action-conditioned world models to ensure predicted trajectories are realizable,
Dev: First, who's behind it and why it matters.
Paper summary: Rosa: So, to summarize "ACID: Action Consistency via Inverse Dynamics for Planning with World Models," the authors are pointing out that standard planning cost functions only judge a candidate by how close its predicted terminal state is to the goal.
Dev: That leaves the realizability of the intermediate steps unchecked, meaning a trajectory can look convincing but physically impossible to execute in reality.
Rosa: The authors propose ACID to fix this by introducing cycle action consistency, which checks if the action inferred backward from a predicted transition by an inverse dynamics model matches the original conditioning action.
Dev: They fold this per-step residual—that difference—into the planning cost using a scale-invariant adaptive weight, w a = lambda times sigma g / sigma a.
Rosa: The central claim is that by costing the whole trajectory, not just the final state, they can target this blind spot and consistently improve planning performance across four action-conditioned world models and six tasks.
Dev: It matters because it provides a complementary decision-time mechanism that verifies action fidelity directly within the planning cost, leaving the world model itself untouched so it stays composable with other improvements.
Taro: So, they’re not retraining the entire world model; they are using an IDM as a verifier to check if the trajectory is physically consistent during planning.
Rosa: Right. They also use this verifier as an auxiliary task during training—like a pseudo-labeler—but the crucial part for decision-time planning is casting it as this cost signal that shapes which action sequence the planner commits to.
Dev: It shifts the focus from just predicting a good endpoint to ensuring you are actually following a physically executable path all along.
Conclusion: Rosa: Looking at "ACID: Action Consistency via Inverse Dynamics for Planning with World Models," the authors, Gawon Seo, Dongwon Kim, and Suha Kwak, are proposing a method to verify action fidelity during decision-time planning.
Dev: In simple terms, they take a trajectory candidate and ask an inverse dynamics model to see if the actions used to build that trajectory are actually consistent with the resulting physical state change at every single step.
Rosa: This consistency check is then added as a cost component, weighted adaptively so it balances against the goal cost, making sure you only discard paths that aren't physically realizable.
Dev: The implication for us is that we can start using decision-time planning in action-conditioned world models with a built-in mechanism to ensure the planned actions are actually executable by the underlying system.
Taro: It suggests that for autonomous systems, especially those dealing with complex physical interactions, we can build in a layer of physics-based verification right into the planning loop without needing massive retraining efforts.
Rosa: Exactly. They show this works across quite a few different environments and tasks, spanning everything from manipulating objects to visual navigation.
Dev: The paper suggests that by focusing on this action consistency, we get better planning results and significantly less total computation compared to just relying on the terminal state proximity alone.
More episodes
- 2610.11768-Narrow and Deep: An Ontology Tower as the Knowledge of an LLM Agent for an Industrial Equipment System
- 2610.11904-Large-Scale Partition-Based RIS Beamforming For Uplink RIS-Equipped Multi-User Systems: Asymptotic Analysis
- 2610.11885-Redefining fuel poverty: Introducing the temporal equity framework (TEF)
- 2610.11900-Reach-Stabilize Control of Control-Affine Systems with Unknown Affine Parameters
- 2610.11964-From Asymptotic to Designer-Assigned-Time Control: A Review of Stability Notions, Design Mechanisms, and Controller Architectures
- 2610.12226-Stabilization of Unidirectional First-Order PDE-ODE Coupled Systems with Boundary and Distributed Input Delays
- 2610.12028-Policy Synthesis for Finite Populations of MDP Agents under Aggregate Reach-Avoid Chance Constraints
- 2610.12103-Predefined-Time Integral Reinforcement Learning for Unknown Nonlinear Systems via Inverse-Optimal Design
- 2610.12110-Adaptive dynamic programming using Lyapunov function constraints
- 2610.12324-Convex Safety Filtering via Spectral Selection for Nonconvex Safe Sets