TDC: Sim-to-Real Transferable Directional Compliance for Contact-Rich Manipulation

arXiv:2602.14174 · cs.RO · Submitted 2026-02-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "TDC: Sim-to-Real Transferable Directional Compliance for Contact-Rich Manipulation".

Rosa: Sim-to-real transfer for contact-rich manipulation remains challenging due to inherent discrepancies in contact dynamics,

Dev: First, who's behind it and why it matters.

Title and authors: Rosa: So, let's start with the specifics of who wrote this paper and what they are trying to achieve with "TDC: Sim-to-Real Transferable Directional Compliance for Contact-Rich Manipulation." The authors include a team from Zhejiang University.

Dev: I see the list of contributors, and it looks like a solid group tackling the robotics side, which is always good to see when you're dealing with these kinds of complex interaction dynamics.

Taro: The research itself is centered on using expert-designed controller logic to bridge that gap between simulation and physical reality for contact tasks.

Rosa: That’s right, Taro; they are moving away from relying solely on data collected in the real world, which is often slow and risky, by incorporating that expert knowledge directly into the learning process.

Dev: It sounds like a clever way to handle the discrepancy between simulated and real contact dynamics without having to perfectly model every tiny friction coefficient or material property.

The paper's summary: Rosa: The main summary of "TDC: Sim-to-Real Transferable Directional Compliance for Contact-Rich Manipulation" boils down to their method of predicting the end-effector pose, contact state, and most importantly, the desired contact force direction alongside those poses.

Dev: So, they aren't just learning where the robot should be; they’re also learning how it should feel like it's being pushed or pulled in terms of direction during contact.

Taro: That force direction prediction is what makes them robust because that vector is determined by the task geometry, which stays the same whether you're in simulation or on a real workbench.

Rosa: Precisely; they found that predicting this direction allows the policy to focus on "where to go" geometrically, letting a separate controller handle "how much force," which is a really nice separation of concerns.

Dev: And they use this prediction to configure a force-aware admittance controller during deployment, which lets them blend the learned intelligence with some manually tuned parameters for real-world adaptation.

The paper's improvements: Rosa: Thinking about what makes this approach an improvement, the authors emphasize that they bypass the high costs and safety risks associated with collecting real-world data by using their simulation environment extensively.

Dev: That’s a huge practical benefit; if we can get good performance from pure simulation data, it significantly cuts down on our need for expensive real-world trials.

Taro: They also suggest a hierarchy where the policy handles the geometric "where to go" part, and then the controller takes over to manage the dynamics of "how much force" is needed during contact interaction.

Rosa: And they introduce a finite state machine within the policy itself, meaning it dynamically switches its behavior depending on whether it's in free motion or actively interacting with an object.

Dev: That switching between position control and hybrid position/force control based on the contact state is where I see the most immediate benefit for loop rate management; it allows for tailored control strategies.

Conclusion: Rosa: So, to wrap up "TDC: Sim-to-Real Transferable Directional Compliance for Contact-Rich Manipulation," they successfully showed that predicting force direction as a transferable signal from simulation is key to robust contact manipulation.

Dev: I agree, the stability analysis they provided on the force-aware admittance controller, showing it’s input-to-state stable even when there's a disturbance, gives me confidence in its real-world application.

Taro: From an autonomy perspective, this means we have a method where the system can handle unexpected misbehavior during contact by having that state machine switch control modes appropriately while maintaining stability.

Rosa: It really suggests that we can get away from needing perfect force magnitude models and instead leverage structural geometric properties for reliable transfer.

Dev: I think the combination of leveraging pure simulation data and having lightweight manual tuning makes this a very scalable approach for deploying these kinds of policies.

Taro: Overall, it gives us a solid framework for tackling complex contact interactions by focusing on directionality rather than trying to learn every dynamic detail from scratch.

Zhejiang University

cs.RO

Submitted: 2026-02-15

Updated: 2026-10-05

Comments: Accepted by CoRL 2026

Code: https://github.com/isaac-sim/IsaacSim

Project page: https://yifei-y.github.io/project-pages/DirectionMatters

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 87/100

The gist: Sim-to-real transfer for contact-rich manipulation remains challenging due to inherent discrepancies in contact dynamics, and this work proposes a framework that leverages expert-designed controller

Key concepts

Dynamics Invariance
This concept refers to identifying parts of interaction forces that remain consistent regardless of simulation inaccuracies or real-world dynamics. The paper decomposes interaction forces into tangent and normal components; the direction of these components—the tangent 't' and normal 'n'—is determined by the task geometry, making them robust signals for learning.
Force-Aware Admittance Control
This is a real-world control mechanism that uses the policy's output to manage robot compliance. It combines a predicted force direction (the normal vector) with a manually specified target force magnitude. This allows the robot to establish and maintain contact stably while adapting its stiffness based on task requirements, such as during resistance overcoming tasks.
Tangent Force Direction (t)
The tangent force direction is one component of the interaction force decomposition that aligns with the feasible motion direction of the robot. In this framework, it is crucial because it relates to task geometry and provides a reliable signal for policy training, as its direction is independent of simulation dynamics.
Privileged State-Based Expert FSM
This is a control logic used in simulation to generate diverse training demonstrations. It follows human-designed kinematic rules tailored to specific tasks during contact phases. This expert logic provides the necessary ground truth for training the neural network policy, ensuring it learns task-specific interaction behaviors.

Terminology

Summary

Sim-to-real transfer for contact-rich manipulation remains challenging due to inherent discrepancies in contact dynamics, and this work proposes a framework that leverages expert-designed controller logic to enable robust policy transfer from simulation to the real world. The core finding is that predicting the contact force direction, which encodes high-level task geometry and is invariant to simulation inaccuracies, allows for training policies exclusively on simulation data, which are then deployed with a lightweight, manually tuned force-aware admittance controller for adaptive compliance in reality.

The Core Insight: Dynamics Invariance

The central motivation stems from the analysis of interaction forces to identify components robust to the dynamics gap. The paper focuses on translational interaction forces and decomposes them into a Tangent Space (Ft) and a Normal Space (Fn). While these subspaces are correlated with dynamics, their directions—the tangent force direction 't' and normal force direction 'n'—are determined by task geometry, which is unrelated to the dynamics. Specifically, the tangent force direction 't' aligns with the feasible motion direction, and the normal force direction 'n' aligns with the surface normal vector at the contact point. The ground-truth force directions in simulation provide robust, dynamics-invariant supervision signals for policy learning.

Policy Formulation and Training

The proposed framework formulates a policy to predict not only reference poses but also the contact state and force direction alongside end-effector poses, acting as a Finite State Machine (FSM) to modulate control modes. The action space consists of three components: the Reference Pose and Gripper Command (Xd), the Normal Force Direction (nt), and the Contact State (ct). To train this policy, a Privileged State-Based Expert FSM is used in simulation to generate diverse demonstrations. In contact interaction phases, the expert FSM follows human-designed kinematic rules tailored to each task. The network architecture utilizes E2VLA, modified by extending the input and output projection layers to accommodate joint prediction of reference poses, normal force directions, and contact states within a unified generative framework. The training loss function balances objectives: L = λ1∥Xd − Xˆ d∥1 + λ2∥n − nˆ∥1 + λ3∥c − cˆ∥1.

Force-Aware Admittance Control

For real-world execution, the policy outputs configure a force-aware admittance controller that combines the predicted force direction with a manually specified target force magnitude. The controller uses translational dynamics: Mx¨r + Dx˙ r + K(xr − xcmd) = Fext − Fcmd. When contact is active (contact state c=1), the commanded force is set along the predicted normal direction 'n' as Fcmd = fn, where f = fH + n · K(xcmd − xr) + n · Dx˙ r. This drives the robot to steadily establish and maintain contact with the target force magnitude fH, which is a task-specific manually tuned scalar. For resistance overcoming tasks, tangent stiffening is enabled by scaling the translational stiffness along the tangent direction 't' by a factor (e.g., 4×).

Stability and Validation

Theoretical analysis confirms the robustness of the controller across different contact conditions. The paper presents three propositions characterizing stability:

  1. Under disturbance-free conditions, the closed-loop normal-direction dynamics is asymptotically stable and converges to the equilibrium point xn = xe − fH/ke such that fext,n = fH.

  2. When contact is lost due to an external disturbance, the closed-loop normal-direction dynamics are asymptotically stable in velocity, driving the robot toward the environment for re-establishment.

  3. When a disturbance occurs during contact, the closed-loop normal-direction dynamics is input-to-state stable (ISS) with respect to the disturbance.

Experiments on four tasks—Microwave Opening, Peg-in-Hole, Whiteboard Wiping, and Door Opening—demonstrate superior performance compared to baselines. The method achieves a 91% overall success rate, significantly outperforming strong baselines in both success rate and robustness. Ablation studies confirm that the synergy of policy and controller is critical; removing either normal force regulation or tangent stiffening leads to substantial drops in success rates, validating the approach's effectiveness in managing task-relevant forces.

Conclusion

The framework successfully identifies force direction as a transferable signal that specifies task-relevant interaction intent and can be reliably learned in simulation, enabling adaptive compliance in the real world with minimal manual tuning. This approach overcomes the limitations of existing paradigms by leveraging pure simulation data for policy training while ensuring stable, task-aligned compliance through a carefully designed force-aware admittance controller. The results validate that predicting dynamics-invariant quantities allows for scalable sim-to-real transfer without relying on unreliable simulated force magnitudes.

Improvements for AI systems

Based on the provided scientific paper, here are specific improvements to existing AI systems and what those improved systems could achieve:


) Enhanced Sim-to-Real Transfer for Contact Dynamics:

By training policies exclusively in simulation using dynamics-invariant signals (force direction and contact state) instead of fragile force magnitudes, the system can achieve high performance on real-world contact tasks with minimal or no real-world data collection.

— Improved AI System Capability: Enables deployment of highly generalizable manipulation policies that are robust to the inherent discrepancies in physical contact dynamics between simulation and reality, drastically reducing the cost and safety risks associated with real-world data acquisition.

) Task-Aware Adaptive Compliance Control:

The system generates a force-aware admittance controller where the policy predicts both the necessary contact state and force direction, allowing for selective stiffness tuning (tangent stiffening vs. normal force regulation) based on task intent, rather than using blind, isotropic compliance.

— Improved AI System Capability: Enables robots to exhibit intelligent compliance—e.g., providing high resistance along the direction of a required pull (tangent stiffening during door opening) while maintaining necessary pressure for contact maintenance (normal force regulation during whiteboard wiping). This leads to superior task success rates and execution quality compared to baseline controllers that use fixed, isotropic stiffness settings.

) Robust Contact Maintenance Under Disturbance:

The controller is proven to be Input-to-State Stable (ISS), ensuring that the robot can maintain a target contact force magnitude during interaction even when subjected to external disturbances (e.g., surface displacement or external forces).

— Improved AI System Capability: Enables robots to reliably complete tasks in dynamic, uncertain real-world environments without triggering safety stops due to minor environmental perturbations, significantly enhancing operational robustness and reliability across all four tested contact-rich tasks.

) Fine-Grained Task Execution Quality Enhancement:

The combination of force direction prediction and adaptive compliance (normal force regulation + tangent stiffening) results in measurable improvements in execution quality metrics, such as deeper peg insertion and more thorough ink removal.

— Improved AI System Capability: Enables high-precision manipulation for tasks like peg-in-hole (deeper insertions) and whiteboard wiping (more complete ink removal), moving beyond simple geometric success to achieve higher fidelity interaction outcomes.

) Scalable Policy Training with Low Supervision Overhead:

The framework utilizes privileged supervision via a human-designed Finite State Machine (FSM) in simulation to generate diverse demonstrations, allowing the policy to learn complex state-switching behaviors without requiring explicit force profiles or stiffness mappings in the training data.

— Improved AI System Capability: Allows for rapid skill acquisition and generalization across different contact scenarios by leveraging structured expert knowledge (the FSM), enabling large-scale policy training using only simulation data, which is significantly more scalable and cost-effective than methods relying on expensive real-world force supervision.

Sources

Related papers