ACID: Action Consistency via Inverse Dynamics for Planning with World Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "ACID: Action Consistency via Inverse Dynamics for Planning with World Models".
Rosa: The gist The proposed ACID framework introduces cycle action consistency into decision-time planning for action-conditioned world models to ensure predicted trajectories are realizable,
Dev: First, who's behind it and why it matters.
Paper summary: Rosa: So, to summarize "ACID: Action Consistency via Inverse Dynamics for Planning with World Models," the authors are pointing out that standard planning cost functions only judge a candidate by how close its predicted terminal state is to the goal.
Dev: That leaves the realizability of the intermediate steps unchecked, meaning a trajectory can look convincing but physically impossible to execute in reality.
Rosa: The authors propose ACID to fix this by introducing cycle action consistency, which checks if the action inferred backward from a predicted transition by an inverse dynamics model matches the original conditioning action.
Dev: They fold this per-step residual—that difference—into the planning cost using a scale-invariant adaptive weight, w a = lambda times sigma g / sigma a.
Rosa: The central claim is that by costing the whole trajectory, not just the final state, they can target this blind spot and consistently improve planning performance across four action-conditioned world models and six tasks.
Dev: It matters because it provides a complementary decision-time mechanism that verifies action fidelity directly within the planning cost, leaving the world model itself untouched so it stays composable with other improvements.
Taro: So, they’re not retraining the entire world model; they are using an IDM as a verifier to check if the trajectory is physically consistent during planning.
Rosa: Right. They also use this verifier as an auxiliary task during training—like a pseudo-labeler—but the crucial part for decision-time planning is casting it as this cost signal that shapes which action sequence the planner commits to.
Dev: It shifts the focus from just predicting a good endpoint to ensuring you are actually following a physically executable path all along.
Conclusion: Rosa: Looking at "ACID: Action Consistency via Inverse Dynamics for Planning with World Models," the authors, Gawon Seo, Dongwon Kim, and Suha Kwak, are proposing a method to verify action fidelity during decision-time planning.
Dev: In simple terms, they take a trajectory candidate and ask an inverse dynamics model to see if the actions used to build that trajectory are actually consistent with the resulting physical state change at every single step.
Rosa: This consistency check is then added as a cost component, weighted adaptively so it balances against the goal cost, making sure you only discard paths that aren't physically realizable.
Dev: The implication for us is that we can start using decision-time planning in action-conditioned world models with a built-in mechanism to ensure the planned actions are actually executable by the underlying system.
Taro: It suggests that for autonomous systems, especially those dealing with complex physical interactions, we can build in a layer of physics-based verification right into the planning loop without needing massive retraining efforts.
Rosa: Exactly. They show this works across quite a few different environments and tasks, spanning everything from manipulating objects to visual navigation.
Dev: The paper suggests that by focusing on this action consistency, we get better planning results and significantly less total computation compared to just relying on the terminal state proximity alone.
POSTECH · KAIST
cs.RO, cs.AI, cs.CV
Submitted: 2026-07-02
Updated: 2026-10-08
Comments: Project page: https://gawon1224.github.io/ACID/
Project page: https://gawon1224.github.io/ACID
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 82/100
The gist: The gist The proposed ACID framework introduces cycle action consistency into decision-time planning for action-conditioned world models to ensure predicted trajectories are realizable, which
Key concepts
- Cycle Action Consistency
- This is a check that verifies if the action used to predict a transition is consistent with what an inverse dynamics model would infer from that same transition. It detects when a predicted trajectory drifts away from the actions it was conditioned on, ensuring physical plausibility step-by-step.
- Inverse Dynamics Model (IDM) Gϕ
- The IDM is a model used backward in time to infer what action must have caused a specific state transition. By using this model, ACID can calculate the 'expected' action for every step of a trajectory, allowing it to measure the difference against the actual planned actions.
- Augmented Planning Cost
- This is the total cost used by the planner to select good action sequences. It combines two parts: a goal cost (how close you are to your desired outcome) and an action consistency cost (how much each step deviates from what physics suggests should happen). This forces the planner to choose paths that are both efficient and physically possible.
Terminology
Summary
The gist The proposed ACID framework introduces cycle action consistency into decision-time planning for action-conditioned world models to ensure predicted trajectories are realizable, which consistently improves planning performance across diverse world models and tasks.
ACID Framework Overview
ACID is a decision-time planning framework that introduces cycle action consistency: the action inferred backward from a predicted transition by an inverse dynamics model should recover the one that was conditioned on. This per-step residual is folded into the planning cost via a scale-invariant adaptive weight. The framework targets this blind spot by costing the whole trajectory, not only its terminal state.
Cycle Action Consistency Mechanism
Cycle action consistency is introduced as a verifier-based cost that detects when a predicted trajectory drifts from its conditioning actions. This verifier performs backward inference by introducing an IDM Gϕ that infers the action responsible for a latent transition. The per-step consistency residual is measured between the conditioning action and the action recovered by the IDM. This cost aggregates this residual over a planning horizon to create a per-step realizability check on the trajectory.
Augmented Planning Cost and Weighting
The augmented planning cost is defined as c(a0:H−1) = cg(a0:H−1) + wa · ca(a0:H−1) where cg is the goal cost and ca is the action consistency cost. The scale-invariant adaptive weight wa is set adaptively as wa = λ · σg/σa, which equalizes the spread of the consistency cost and the goal cost up to a factor λ. This adaptive weight is set at every CEM iteration to ensure comparable influence on elite selection regardless of variations.
Experimental Results and Efficiency
ACID consistently improves planning and matches the baseline’s accuracy with substantially less planning compute across four action-conditioned world models and six tasks. The method is robust to hyperparameter choices because the adaptive weight wa = λ · σg/σa removes dependence on fixed values. The per-step overhead of the inverse dynamics verifier is bounded, as it reuses offline trajectories with no additional environment interaction. Despite this overhead, ACID reaches target quality with less total planning compute because it suppresses non-realizable candidates by checking each against the actions an inverse dynamics model infers from it.
Conclusion
In this paper, we introduce ACID, a decision-time planning framework that augments planning cost with cycle action consistency for action-conditioned world models. The framework achieves consistent improvements in decision-time planning across diverse action-conditioned world models and tasks spanning rigid and deformable object manipulation, articulated control, and visual navigation.
Limitation
ACID depends on the property that a pair of consecutive observations identifies the action between them. Under partial observability, or when exogenous interventions perturb the transition so that it is not explained by the conditioning action alone, this property weakens. However, since our method intervenes only on the planning cost, it is orthogonal to advances in action-conditioned world models.
How it works
-
Encode states: z0 ← Eθ(o0), zg ← Eθ(og)
-
Sample action sequences N (µj−1, Σj−1) from the distribution.
-
For each candidate sequence a(n)0:H−1 in N, unroll the trajectory zˆ(n)t+1 ← Fθ(ˆz(n)t, a(n)t), where zˆ0 = z0.
-
For each step t, infer the action â(n)t ← Gϕ(ˆz(n)t, zˆ(n)t+1), where â is the IDM inferred action.
-
Calculate the goal cost cg = (zˆH − zg) squared / 2.
-
Calculate the action consistency cost ca = (1/H) Σ a(n)t - â(n)t) squared / 2.
-
Compute the augmented planning cost c(n) ← cg + wa · ca for n = 1,..., N.
-
Select the K sequences with lowest c(n) as elites E: µj ← 1/K P n∈E a(n)0:H−1.
-
Update the sampling distribution µj ← 1/K Σ varn∈E a(n)0:H−1 <ref:2607.
Improvements for AI systems
- Bold header: Cycle Action Consistency in Decision-Time Planning
This introduces cycle action consistency,
where the action inferred from a predicted transition by an inverse dynamics model should recover the one that was conditioned on,
which is folded into the planning cost via a scale-invariant adaptive weight.
This allows the planner to prioritize sequences whose predicted trajectory is both goal-reaching and step-by-step realizable.
- Bold header: Robustness Across World Models and Tasks
The framework achieves consistent gains across four action-conditioned world models and six tasks, spanning rigid and deformable manipulation, articulated control, and visual navigation,
demonstrating that the method's effectiveness is robust to hyperparameter choices
by using an adaptive weight where wa = λ · σg/σa.
- Bold header: Enhanced Planning Efficiency
ACID delivers results with substantially less planning compute than the baseline,
reaching target quality with a net compute of approximately 0.7× compared to the original method, despite adding a per-step overhead of the inverse dynamics verifier.
Abstract
Decision-time planning with action-conditioned world models has become a popular paradigm for embodied control. However, the standard planning cost judges a candidate solely by how close its predicted terminal state lies to the goal, leaving the realizability of the intermediate transitions unchecked--a predicted trajectory can look convincing while the environment rollout drifts away from it. In this paper, we propose ACID, a decision-time planning framework that introduces cycle action consistency: the action inferred backward from a predicted transition by an inverse dynamics model should recover the one that was conditioned on. We fold this per-step residual into the planning cost via a scale-invariant adaptive weight. Across four action-conditioned world models and eight tasks encompassing object manipulation and articulated control in simulation, visual navigation, and real-robot manipulation, ACID consistently improves planning and matches the baseline's accuracy with substantially less planning compute.
Sources
- World Models
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- GAIA-1: A Generative World Model for Autonomous Driving
- World Action Models are Zero-shot Policies
- FreeAction: Training-Free Techniques for Enhanced Fidelity of Trajectory-to-Video Generation
- GigaWorld-Policy: An Efficient Action-Centered World--Action Model
- DreamGen: Unlocking Generalization in Robot Learning through Video World Models
- Rectified Flow: A Marginal Preserving Approach to Optimal Transport
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
- DeepMind Control Suite
- AdaptiGraph: Material-Adaptive Graph-Based Neural Dynamics for Robotic Manipulation
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving