IADD-TR: Intervention-Aware Dynamics Decoupling with Targeted Regularization for Model-Based Reinforcement Learning
Zefeng Liang, Jie Qiao, Ruichu Cai, Weilin Chen, Zhifeng Hao
Guangdong University of Technology · Shantou University
cs.LG
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: This paper addresses confounding bias in model-based reinforcement learning (MBRL).
Terminology
Summary
This paper addresses confounding bias in model-based reinforcement learning (MBRL). The authors identify two interrelated challenges: (1) learning an accurate dynamics model from limited agent-collected data, and (2) learning an effective policy from imperfect model-generated data. They argue that standard approaches treat the transition model and critic as monolithic predictors,
which overlooks policy-induced data bias.
This causes action can become entangled with environmental evolution, while uneven action coverage may distort the counterfactual value estimates used for policy improvement.
The paper illustrates the problem with a car-driving scenario: "the policy accelerates primarily near the starting region and brakes as the car approaches the destination. As a result, these policy-generated data would create spurious correlations such that a larger acceleration is correlated with smaller subsequent positions and a smaller acceleration is correlated with larger subsequent positions." The position acts as a confounder, influencing both action and resulting state.
The paper proposes a unified framework combining two components:
IADD factorizes transitions into an action-intervention stage and an action-free natural evolution stage.
The key insight is that an agent's action is fundamentally a causal intervention in many environments.
The framework introduces an intermediate representation s̃t = (s̃t,o, s̃t,l) where:
-
s̃t,o is the observable post-intervention state with the same dimensionality as the original state
-
s̃t,l is an auxiliary latent component
The action-intervention stage is modeled as: s̃t,o = pact,o(st, at), s̃t,l = pact,l(st, at, εt). The natural evolution stage is: st+1 = penv(pact,o(st, at), pact,l(st, at, εt)).
Zero-action anchor: To resolve non-uniqueness of the two-stage factorization, the authors introduce a reference zero action a0, defined such that pact,o(st, a0) = st.
This condition is enforced by construction rather than through an auxiliary loss
— when the input action equals a0, the action-intervention stage directly copies st into the observable coordinates.
Identifiability results: The paper provides two theorems. Theorem 1 establishes pointwise identifiability
of the observable post-intervention state s̃t,o under injectivity and support inverse conditions. Theorem 2 establishes component-wise identifiability
of the latent component s̃t,l up to permutation and one-dimensional invertible transformations,
using conditions from auxiliary-variable nonlinear ICA.
TR addresses policy learning bias. The paper derives the efficient influence function (EIF) of the replay-state policy-gradient functional, revealing that the inverse-probability term 1/πβ(ãs) in Eq. (16) amplifies the critic's residual error for actions rarely observed in the replay buffer.
The method constructs an adjusted critic: Qadj ω,ξ(s, a, eπβ) = Qω(s, a) + ϵξ(a)/eπβ, where eπβ is the logged behavior-policy density and ϵξ(a) is a learnable residual correction. The overall critic objective combines standard TD loss with targeted regularization: LIADD−TR(ω, ξ) = LTD(ω) + β RTR(ξ).
Theorem 4 establishes double robustness
: the estimator is consistent if either the targeted critic is consistent for Qπθ or the replay action-density estimate is consistent for the true replay density.
The paper makes four contributions:
-
An intervention-aware dynamics decoupling model separating action-induced effects from natural evolution
-
A zero-action-based identification strategy for two-stage decoupling
-
Targeted regularization for Q-function learning to reduce confounding bias in policy-gradient estimation
-
Theoretical and empirical validation
Experiments on five MuJoCo tasks (HalfCheetah, Hopper, Walker2d, Ant, Humanoid) show IADD-TR attains the highest or competitive returns across all five tasks while improving more rapidly than SAC, PPO, and SLBO.
The advantage is particularly pronounced on Ant and Humanoid, where the higher-dimensional dynamics make policy learning more sensitive to errors.
Adding either IADD or TR individually improves over MBPO. IADD provides the clearer individual gain on HalfCheetah, whereas MBPO+TR is particularly strong on Hopper. On Walker2d and Ant, combining the components produces the clearest improvement.
Using a controlled synthetic environment, increasing zero-action coverage from λ0=0 to λ0=0.40 reduced post-intervention-state MSE from 4.043 to 0.0109, a 99.7% reduction
and increased action-effect correlation from 0.346 to 0.973.
Held-out next-state MSE remained low, showing recovery doesn't compromise prediction.
On Hopper, TR achieved higher mean cosine similarity than No-TR at every actor checkpoint,
with gains ranging from 0.009 to 0.044
across six checkpoints.
The paper concludes that separating intervention effects from natural evolution and improving the alignment between critic learning and policy optimization can effectively mitigate learning bias in model-based reinforcement learning.
The applicability is bounded by the identifiability assumptions and the availability of a meaningful zero-action anchor.
Improvements for AI systems
Improvements to AI systems:
-
Causal dynamics decoupling for model-based RL: AI systems can learn two separate transition components—an action-intervention stage (how actions causally alter the state) and an action-free natural evolution stage (how the environment changes on its own). This prevents spurious correlations (e.g., a policy that brakes near a goal creating a false link between braking and position) from contaminating the learned world model. The improved system can make more accurate predictions under counterfactual actions, especially in high-dimensional control tasks like robotics or autonomous driving.
-
Zero-action anchored identification: By enforcing that a reference
no-op
action leaves the observable state unchanged by construction, the AI system can uniquely decompose transitions without requiring auxiliary losses or hand-crafted supervision. This enables the model to disentangle action effects from environmental drift even when data is collected under a biased policy, improving generalization to unseen action distributions. -
Targeted regularization for policy-gradient critics: The AI system can adjust its Q-function learning by adding a correction term proportional to the inverse of the logged behavior-policy density, specifically targeting rare actions in the replay buffer. This reduces the amplification of critic errors for under-sampled actions, leading to more stable and accurate policy-gradient estimates. The improved system achieves double robustness—it remains consistent if either the corrected critic or the replay density estimate is accurate, making it more reliable in off-policy settings.
-
Intervention-aware representation learning: The system learns an intermediate latent representation that separates observable post-intervention states from auxiliary latent components, with theoretical guarantees on identifiability (up to permutation and invertible transformations). This allows the AI to build a more structured world model that can answer
what would happen if I took action X here?
versuswhat would happen anyway?
—enabling better planning, counterfactual reasoning, and safe exploration in model-based RL. -
Improved sample efficiency and final performance: By combining both components, the AI system can achieve higher returns faster across diverse continuous-control benchmarks (e.g., Ant, Humanoid), with particularly strong gains in high-dimensional tasks where dynamics errors compound. The system can also validate its own decomposition quality (e.g., via zero-action coverage checks) to detect when its causal assumptions hold, enabling adaptive trust in its predictions.
Sources
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Proximal Policy Optimization Algorithms
- VCNet and Functional Targeted Regularization For Learning Causal Effects of Continuous Treatments
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks