A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI

arXiv:2505.01458 · cs.RO, cs.AI · Submitted 2025-05-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI".

Dev: Merging these two sources requires careful synthesis to create a comprehensive, detailed overview that captures both the foundational concepts and specific technical contributions mentioned in each text.

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: Now that we've talked about the setup, let’s go into what this survey actually covers in detail regarding navigation and manipulation tasks. It lays out a taxonomy of the different simulators used for these two core embodied AI capabilities.

Dev: I think it systematically reviews classical engines like MuJoCo for precision and Isaac Sim for its GPU acceleration, but it also stresses that the choice depends entirely on how complex the physical interaction needs to be for the specific task at hand.

Taro: That’s interesting because navigation might just need rigid-body collision modeling, whereas manipulation definitely demands things like accurate contact and friction modeling.

Rosa: Precisely, and they detail the benchmarks available for each area, including those for rigid object manipulation versus deformable object manipulation, which is a big focus.

Dev: They also cover methods used for both navigation—like goal-driven vs. task-driven approaches—and manipulation techniques, such as dexterous grasping and whole-body control strategies.

Taro: I want to hear more about how they treat the specific challenges of deformable objects, because that seems like a major sticking point in real-world tasks today.

Rosa: They review specific benchmarks like SoftGym and PlasticineLab for deformable object manipulation, showing the evolution of methods from just rigid object handling to modeling soft bodies.

Dev: The paper emphasizes that accurate contact dynamics and force interactions are critical for manipulation, pointing to datasets like Grip that combines deformable-rigid coupling.

Taro: It seems like the authors are mapping out exactly where the current research strengths and weaknesses lie when it comes to simulating these complex physical realities for embodied AI.

Rosa: They also cover how different simulation settings, such as actuation timing and sensor models, constrain what policies can be learned during training, which is a very practical limitation.

Dev: It’s about understanding that even if the policy looks good in simulation, its performance on hardware will be capped by the quality of those underlying physical parameters.

Taro: So, the implication is that we aren't just training policies; we are constraining them based on what the simulator can accurately model about physics, which is a tight feedback loop.

Rosa: Right, and they also look at how different simulation types impact the required hardware constraints for running these complex models.

The paper's summary: Dev: The survey itself suggests some clear ways researchers can improve their work by focusing on simulator properties that haven't been studied as much before. It’s not just about using a better simulator, but understanding how to tune the existing ones better.

Rosa: One major suggestion they make is incorporating an adaptive calibration layer, where the system automatically adjusts low-level physics parameters like friction or sensor noise based on real-world interaction data or uncertainty metrics.

Taro: That sounds like a proactive approach to handling those discrepancies we discussed earlier; instead of training once for one specific environment, you adapt as you encounter different conditions.

Dev: From an engineering standpoint, that would mean the agent can achieve higher correlation coefficients across very different physical domains without needing a complete overhaul of the high-level policy.

Rosa: It also points toward using multimodal fusion architectures for manipulation, specifically integrating tactile or force feedback to correct visual policies during contact phases.

Taro: That addresses my concern about ambiguity; if the agent can "see to touch," it gains that high-frequency signal needed for stable grasping when visual cues are unclear.

Dev: That would be a significant step up in capability for dexterous manipulation, allowing for near-perfect stability during complex insertion or handling tasks.

Rosa: Then there’s the idea of using hierarchical memory systems to bridge spatial reasoning and high-level task execution, combining graph methods with foundation models like VLMs.

Taro: I think that hybrid approach makes sense for navigation; using a VLM for a semantic plan and then a graph structure to execute the path gives us both big picture context and local steering.

Dev: And they also suggest moving toward differentiable physics frameworks specifically for manipulation tasks, allowing gradients to flow directly through contact models.

Rosa: That would significantly increase sample efficiency for learning force-sensitive skills, like knowing exactly how much grip force is needed to prevent slip on a specific material.

Taro: If we can learn those physical parameters directly from the simulation gradients, it means the agent learns the optimal physics behavior much faster than through traditional reinforcement learning methods.

The paper's improvements: Rosa: So, as we wrap up this discussion of "A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI," it really boils down to how essential these simulators are for making embodied AI work at all. They are the primary tool we have to manage that difficult sim-to-real gap.

Dev: I agree, this paper reinforces that we can't just rely on training in simulation and hoping for the best; we need a deeper look at the underlying physics models before deploying these agents to hardware.

Taro: My final thought is that the future involves integrating these rigorous physical constraints with sophisticated memory structures so agents can handle unexpected world misbehaves gracefully.

Rosa: That’s right, and by using this survey, we get a roadmap for selecting the right simulators and designing evaluation metrics that actually measure deployment quality.

Dev: We need to keep pushing the loop rate and latency models as we move toward those adaptive calibration layers suggested in this work because real-world performance is all about responsive control.

Taro: I think focusing on those procedural quality metrics, like energy efficiency, will be vital for ensuring that these agents operate safely and naturally when they are actually out in the field for extended periods.

Rosa: That’s our summary of this paper today: it shows us exactly how to use physics simulators to build more reliable navigation and manipulation systems for embodied AI. Thanks to Rosa, Dev, and Taro for joining us on this deep dive into the survey.

Conclusion: Rosa: So, we’ve gone through this comprehensive overview of "A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI," which really maps out how physics simulators are the backbone for training embodied AI agents.

Dev: It certainly shows us that the choice of simulator isn't arbitrary; it directly dictates what kind of errors we’re going to have when moving from simulation to reality, which is critical for our loop rate and failure modes.

Taro: I think the focus on both perception and action-dynamics gaps really highlights that the sim-to-real challenge isn't just about rendering things looking pretty; it’s about correctly modeling how a robot actually touches and moves in a physical space.

Rosa: Exactly, and when you look at the application areas, it becomes clear that for navigation, we need accurate terrain response, while manipulation requires precision in contact dynamics and deformable object modeling.

Dev: And I see the implication for our control systems: if we can get better models of friction and actuation timing in simulation through differentiability, those learned policies will be much more robust when deployed on real hardware.

Taro: That robustness is what matters most; it means when the AI encounters something unexpected in the physical world, its underlying physics understanding gives it a better chance to recover instead of just failing.

Rosa: It’s inspiring to see how they're categorizing memory into explicit and implicit structures, which suggests that future embodied agents will need hybrid systems combining semantic maps with learned world models.

Dev: That hybrid approach is exactly what we need for complex tasks; you can use the learned model for long-term prediction while using a metric map for immediate local path planning.

Taro: And thinking about those safety metrics they mentioned, I believe moving toward procedural quality checks rather than just success rates will be key to ensuring these agents are trustworthy when they interact with humans or complex environments.

Rosa: It’s exciting to think about the practical impact on field robotics; if we can solve this sim-to-real issue effectively, we open up a whole new class of reliable robotic systems that can operate in diverse, real-world settings for longer periods.

Dev: I agree; improving the accuracy of contact and deformation models directly translates to more predictable control inputs, which lowers the risk of catastrophic failures during deployment.

Taro: It’s about enabling agents to execute complex instructions reliably across different physical domains, which is a big step toward true autonomy in unstructured environments.

Rosa: So that's our look at "A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI." We'll be looking closely at how these simulation techniques shape the next generation of embodied AI systems.

City University of Hong Kong · Nanyang Technological University · Universität Hamburg

cs.RO, cs.AI

Submitted: 2025-05-01

Updated: 2026-10-03

Comments: 35 pages, 9 figures. Accepted to ACM Computing Surveys (CSUR)

DOI: 10.1145/3856802

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 89/100

The gist: Merging these two sources requires careful synthesis to create a comprehensive, detailed overview that captures both the foundational concepts and specific technical contributions mentioned in each

Key concepts

Sim-to-Real Gap
This is the performance difference between an AI agent in a simulation and its actual performance on real hardware. It occurs because the simulation often fails to perfectly match reality regarding how sensors perceive the world or how physics (like friction or collisions) behave, making it hard to trust simulated training for real robots.
Action-Dynamics Gap
This gap specifically relates to the robot's physical interaction with its environment. It involves inaccuracies in modeling crucial physical aspects like how surfaces react when a robot touches them, the precise timing of motor movements versus simulation time, and accurately simulating complex objects that deform during interaction.
Differentiable Simulators
These are advanced simulators that allow researchers to calculate gradients directly from the physics simulation results. This capability is vital because it enables neural networks to learn better controllers by optimizing their performance based on real physical interactions within the simulation, significantly improving transferability to the real world.

Terminology

Summary

Merging these two sources requires careful synthesis to create a comprehensive, detailed overview that captures both the foundational concepts and specific technical contributions mentioned in each text.

Here is the detailed, synthesized summary of the paper A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI, integrating information from both provided excerpts:


This survey provides a deep examination of how physics simulators are being utilized to address the critical sim-to-real gap inherent in training embodied AI agents for complex robotic tasks, specifically focusing on navigation and manipulation. The paper frames physics simulation as a central enabling tool because it allows for scalable data generation, repeatable evaluation, controller prototyping, debugging, and safer testing before costly hardware deployment.

The fundamental challenge addressed is the sim-to-real transfer, where policies trained in simulation often degrade when deployed on physical hardware due to discrepancies between simulated and real environments. This discrepancy is formally defined as the **sim-to-real gap, G(pi) = psi sim(pi) - psi real(pi) **, where psi measures performance under a given task metric.

The survey systematically categorizes these sim-to-real discrepancies into two primary domains:

  • Perception Gap: Differences arising from sensory observations, including variations in RGB-D images, LiDAR data, lighting conditions, texture fidelity, sensor calibration issues, and inherent noise.

  • Action-Dynamics Gap: Discrepancies related to the robot–environment interaction itself. This encompasses challenges such as inaccurate collision response modeling, friction dynamics (especially for contact), precise actuation timing mismatches between simulation and reality, modeling of deformable object dynamics (soft bodies), terrain response variations, and general sensor noise.

The survey reviews the landscape of simulators used to bridge this gap, distinguishing between classical engines and more advanced methods:

  • Classical Physics Engines: These established tools are categorized by their strengths in different aspects:

  • MuJoCo [175]: Noted for prioritizing high precision in contact dynamics.

  • Isaac Sim [133]: Leverages GPU acceleration, often used for photorealistic rendering.

  • PyBullet [40]: Emphasizes speed and computational efficiency.

  • Other notable engines include Gazebo [94] and CoppeliaSim [150].

  • Differentiable Simulators: A growing area of focus is differentiable simulators (e.g., Genesis [8]), which are crucial because they provide gradients with respect to physical states, directly enabling the training of neural network controllers with improved sim-to-real transferability via gradient-based optimization.

The paper explores how agents utilize memory within simulated environments, categorizing representations into two main types:

  • Explicit Memory: Structured representations designed for planning, such as Metric Map-Based approaches (e.g., occupancy grid maps) and Graph-Based approaches (e.g., topological graphs).

  • Implicit Memory: Learned representations that capture complex environmental knowledge, including Latent Representation-Based methods (LSTMs/Transformers), Foundation Model-Based methods (LLMs/VLMs), and sophisticated World Model-Based methods that learn predictive dynamics models.

The survey details how these simulators are applied across different robotic capabilities:

Navigation tasks involve safely moving an agent through an environment, relying on the simulator to accurately model terrain response and collision dynamics to ensure safe locomotion, especially when dealing with actuation or terrain mismatches that cause failure cases.

Successful manipulation requires simulators to accurately model geometric and contact details, object motion physics (including stability), and force interactions. The survey highlights specific challenges in this domain:

  • Deformable Object Manipulation: Simulators are being developed to handle soft objects, with methods like Phystwin [85] focusing on physics-informed reconstruction of deformable objects from video, and datasets like SoftGym and PlasticineLab.

  • Contact Dynamics: Simulators must accurately model contact stability and force interactions. This is addressed by engines prioritizing precision (like MuJoCo) or specialized datasets like Grip [119], which combines deformable-rigid coupled grasping simulations.

  • Advanced Manipulation Techniques: Research is moving toward more complex control paradigms, such as Whole-body Loco-manipulation Control (e.g.

Improvements for AI systems

As a fastidious and diligent AI researcher, I have analyzed this survey to identify critical areas where current research is lacking or where simulator properties need more rigorous investigation to bridge the sim-to-real gap.

Here are specific, actionable improvements for AI systems based on this paper:


)1. System Improvement: Task-Specific Simulator Calibration and Domain Randomization Strategy

The paper emphasizes that simulator properties (contact solvers, friction models, actuation timing, sensor noise) critically affect policy transfer. Instead of using a single best simulator, the system should incorporate an adaptive calibration layer.

  • Improvement: Implement a meta-learning or self-calibration module that automatically adjusts low-level physics parameters (e.g., friction coefficients for different surfaces, contact stiffness, and sensor noise profiles) based on initial real-world interaction data or high uncertainty metrics (like high SRCC).

  • Improved AI System Capability: The agent can achieve significantly higher Sim-to-Real Correlation Coefficients (SRCC) across vastly different physical domains (e.g., switching from indoor navigation to outdoor driving) without requiring a complete re-training of the high-level policy, as the system adapts its perception and action dynamics models on the fly.

)2. System Improvement: Multimodal Fusion for Robust Contact-Rich Manipulation

The survey highlights that vision-only manipulation fails during contact phases because it lacks local, high-frequency interaction signals (slip, force).

  • Improvement: Integrate a Visuo-Tactile fusion architecture (as proposed in Section 3.4.1) where tactile/force feedback is used to correct the trajectory generated by a slow visual diffusion policy or VLA module. This should be implemented via See to Touch architectures that use tactile tokens for closed-loop, high-frequency force regulation during grasping and insertion.

  • Improved AI System Capability: The agent can perform complex, dexterous manipulation tasks (like unscrewing a bottle cap or handling fragile objects) with near-perfect stability and zero drop/slip failures in the real world, even when visual cues are occluded or ambiguous.

)3. System Improvement: Hierarchical Memory for Long-Horizon Task Execution

The paper distinguishes between Explicit and Implicit memory structures (metric maps vs. latent vectors vs. World Models). For complex, task-driven navigation (like embodied Q&A), a hybrid approach is necessary to handle both spatial reasoning and linguistic context.

  • Improvement: Develop a hierarchical memory system that uses Graph-Based methods for long-range planning in large environments (e.g., using Knowledge Graphs) combined with Foundation Models (VLMs) for high-level goal decomposition and symbolic reasoning. This should be augmented with World Model predictive capabilities for look-ahead trajectory ranking.

  • Improved AI System Capability: The agent can successfully execute Task-Driven Navigation by interpreting complex, multi-step natural language instructions (Find the blue book on the second shelf) by first using a VLM to generate a semantic plan, then using a Graph/Metric Map to navigate the path, and finally relying on its World Model to predict dynamic obstacles during execution.

)4. System Improvement: Safety-Aware Evaluation Metrics

The survey critiques current success rates (SR) for hiding unsafe behavior or poor temporal safety during deployment.

  • Improvement: Replace simple Success Rate metrics with procedural quality metrics inspired by human execution, such as Energy Efficiency and Smoothness, and introduce explicit Temporal Safety Violation counts into the evaluation pipeline. This requires moving beyond single-metric leaderboards to property-level runtime monitors.

  • Improved AI System Capability: The agent will not only achieve high task completion but will also demonstrate natural or safe motion (low oscillation, smooth force application) while operating under real-world constraints like latency and dynamic terrain changes, ensuring deployment is trustworthy.

)5. System Improvement: Differentiable Physics for Sample Efficiency in Manipulation

For tasks involving complex contact dynamics (deformable objects), classical simulators are computationally expensive and prone to physics inaccuracies.

  • Improvement: Transition the core manipulation policy training from model-free RL/Behavior Cloning to a differentiable simulation framework (like Genesis or DiffTaichi) where gradients flow directly through the contact and deformation models.

  • Improved AI System Capability: The agent can learn highly optimized, force-sensitive grasping and manipulation skills much faster than traditional RL methods, as it learns the optimal physical parameters (e.g., required grip force to prevent slip on a specific material) directly from the physics simulation gradients, leading to superior performance in deformable object handling.

Sources

Related papers