Robotics papers — 2026-09-25

The focus today was on improving robot movement when the exact shape or structure is unknown, which is important because real-world robots are rarely perfectly modeled. MorphIK attempts to use the robot's shape to condition neural inverse kinematics for unknown robots. This means the system tries to figure out the physical structure from what it sees and then uses that information to calculate how its joints should move for a desired action.

A related piece explored World Action Agent, which uses large vision-language models for robot manipulation by having them rehearse actions in a simulated world before actually performing them. This rehearsal approach aims to improve the agent's ability to handle complex tasks through experience.

RAPID focuses on robot agentic programming directly from demonstrations, which means learning how to program robots just by watching someone do the task. This is important because it bypasses the need for explicit low-level control code.

The work on Rolling-WAM deals with world action models that incorporate rolling imagination, suggesting a way for these models to explore potential future actions dynamically during operation. This feeds into how we might build more robust systems like those being developed in Coding Agents for Generalized Task and Motion Planning Problems, which aim to solve general planning issues.

RAPID's approach contrasts with uncertainty-gated exploration noise suppression in online reinforcement learning fine-tuning of a flow-matching vision-language-action policy, which tackles task collapse by managing exploration noise during real-time policy updates. This seems to be the cutting edge for making these complex systems reliable in uncertain environments.

The most significant development today involves the work on RotVLA, which tackles controlling vision language action models by introducing a rotational latent action. This is important because it directly addresses how these models can better plan and execute physical movements in real-world scenarios. The abstract describes learning this rotational latent action to improve performance in vision-language-action modeling.

This is supported by the work on Learning to Navigate with Minimal Parameters, which decomposes visual navigation into closed-form geometric interfaces. This method aims to reduce the parameter count needed for visual navigation tasks by leveraging these geometric structures rather than relying solely on large models. This approach builds upon the idea of creating structured representations that make complex tasks more manageable.

Another piece of work contributing to this direction is Representation World Model, which focuses on learning states, transitions, and executable plans within a representation framework. This research seeks to create a coherent internal model that allows an agent to reason about its environment and plan actions sequentially. This modeling capability is crucial for any system attempting autonomous navigation or complex task execution.

The abstract also touches upon the application of physics-informed solvers through RAPTOR, which is designed as a random-projection physics-informed transient solver. This tool is relevant because it allows for solving physical problems with constraints derived from known laws of physics, offering more accurate simulations than purely data-driven methods.

GridSFM presents a foundation model for solving AC optimal power flow problems. This work addresses the need for efficient solutions in electrical engineering by using a foundation model approach to tackle complex optimization challenges in power systems. This provides a different kind of structured problem-solving that complements the agent planning and physical simulation efforts seen elsewhere.

The work on Physics Guided Residual Reinforcement Learning for Humanoid Narrow Path Traversal is particularly important because it directly addresses the challenge of robots navigating complex, confined spaces safely. This approach involves using physics to guide reinforcement learning policies so that humanoid robots can move through tight corridors without colliding.

This method tries to learn how to traverse narrow paths by incorporating physical constraints into the reinforcement learning process. The results show that this guided learning leads to more stable and successful traversal compared to standard methods, suggesting a tangible improvement in real-world robot mobility. This finding builds upon the work of EgoSpeedUp, which focused on transferring human manipulation tempo into robot policies, showing how mimicking human movement can improve robotic control.

Another significant piece of research is BeyondRetarget, which aims to learn executable humanoid motions directly from monocular video inputs. This means the system learns how to perform specific movements just by watching a person do them in a video, bypassing traditional modeling steps. This capability complements the efforts in creating novel view synthesis, such as M3GD, which uses multi-modal data to generate new geometric views for cameras and LiDAR systems.

The work on Trajectory Induced Self Calibration for Hidden Target Localization through an Unknown Pose Range Bearing Relay is also crucial because it allows systems to accurately locate targets even when the robot's pose is unknown. This technique uses the trajectory itself to calibrate the system, which is a clever way to overcome sensor uncertainty. This calibration method connects conceptually with Free-Init, which deals with initialization for Doppler LiDAR-Inertial Systems by removing scan and motion dependencies.

Finally, Synthetic Enclosed Echoes introduces a new dataset designed to bridge the gap between simulated sonar data and real-world sonar data. This dataset is important because it helps train systems like Self Adaptive VLA for Robust Robot Deployment in more realistic scenarios. This entire collection of research shows a trend toward making robotic systems more robust by either improving motion planning, learning from human demonstrations, or better handling sensor uncertainty.

The most significant piece of work today involved StageCraft, which addresses the problem of failures caused by distractions and obstructions in virtual laboratory environments for visual learning agents. This method attempts to improve execution awareness in models by mitigating these failures, which is crucial because accurate execution awareness helps robots learn robust control policies.

This stems from the preceding work on coordinate-independent robot model identification, which seeks to create a general representation of a robot's dynamics regardless of its specific setup. This general model information is then fed into StageCraft to enhance the agent's ability to handle real-world execution issues.

Another important contribution is GenPHRI, which focuses on agentic generative simulation for physical human-robot interaction. This work explores how agents can generate realistic simulations for interacting with humans, which opens up new avenues for safe and intuitive robot collaboration.

We also saw progress on sampling-based Model Predictive Control for Double-Pendulum Sway Suppression on a Shipboard Crane, which uses MuJoCo to stabilize a crane system against swaying motion. This is important because stabilizing dynamic systems like this is fundamental for practical mobile manipulator operation.

Finally, there was research into FingerViP, which focuses on learning dexterous manipulation skills by incorporating fingertip visual perception to understand real-world contact. This sensory input feeds into the broader goal of object reconstruction awareness, suggesting a pathway toward more capable manipulation systems.

The most significant development today involves MPC-Injection, which attempts to bias off-policy locomotion reinforcement learning toward behaviors that align with what a controller would induce. This is important because it aims to bridge the gap between learned policies and physically executable control strategies.

This work builds upon earlier efforts by using memory-guided agents to steer frozen visual latent agents into reliable manipulation primitives. These agents are essentially visual representations that are guided by stored experiences, which helps them perform complex actions like grasping or moving objects reliably.

Another area of focus is modeling robot velocity fields as probability velocity fields for flow-based object manipulation, which seeks to make the motion planning process more probabilistic and robust. This connects to the work on ContactWorld, which investigates what kinds of representations matter for vision-tactile latent world models in contact-rich manipulation, suggesting that how we represent physical interactions is crucial for these flow models to succeed.

Finally, there is research into enabling robust cloth manipulation through inference-time simulator-in-the-loop refinement. This technique involves using a simulator during the actual operation of the agent to refine its actions on the fly, which helps overcome inaccuracies in purely learned models. This refinement process is also related to implicit behavior coordination from sub-task demonstrations, as both methods explore ways to improve complex movements by incorporating external guidance or simulation feedback.

The work on human-in-the-loop geospatial annotation for rapid dataset construction is crucial because it directly impacts how quickly we can build robust training data for field deployed UAV systems. This approach involves having people annotate images in the field, which speeds up the creation of large, diverse datasets needed for training.

This moves into the realm of vision and control with OCC4M, which aims to give spacecraft long-horizon manipulation capabilities by incorporating object-centric four dimensional memory to handle complex spatial reasoning. This is significant because it suggests a way for robots to remember things over long periods in space.

Then there is the work on Tendon-Driven Continuum Robots with modular stiffness and in situ self pose estimation, which deals with making soft robots more adaptable through stiffness control and figuring out their own position without external sensors. This builds on the idea of complex physical interaction.

We also have OA-MPPI, which focuses on occlusion aware model predictive path integral control for UAV flight, trying to make drones navigate better when parts of the view are blocked during flight. This is important for reliable autonomous aerial navigation.

This connects to Excitation-Supervised Closed-Loop Self-Calibration and Target Seeking for an Unknown-Pose Range-Bearing Relay, which tackles how a relay can find its target even when it doesn't know its exact position beforehand by using self calibration guided by excitation. This shows progress in autonomous positioning under uncertainty.

Finally, SCoCaT addresses spacecraft docking using success conditioned constrained reinforcement learning, which is a key step toward reliably executing precise docking maneuvers in space environments.

The most significant work this morning involved streaming deep reinforcement learning applied to adaptive continual learning within robotics, because this directly addresses the challenge of robots needing to learn new tasks while operating under communication constraints. A study on streaming deep reinforcement learning explored how a robot could adapt its policies continuously as it encounters novel situations, suggesting a framework for real-world deployment where retraining is impossible.

This concept builds upon the work of Streaming-WAM, which developed an action-conditioned world-action model designed for asynchronous robot manipulation, offering a way to handle the continuous nature of learning. Another piece focused on Koopman-accelerated model-based diffusion for real-time robot control, attempting to speed up how robots plan actions by using these mathematical models.

Then there is the work on FlyCNS, which focuses on connectome-grounded information organization for communication-constrained embodied control; this means structuring the robot's knowledge based on its physical connections to manage limited data flow effectively. TactileStep looked at sole tactile learning for regulating foot-terrain interaction in humanoid locomotion, which is crucial for stable movement on uneven surfaces.

Finally, RoboRecover benchmarked robot policy recovery under execution deviations, which tests how well a learned policy can recover when the real world doesn't perfectly match the simulation or training environment. This work connects to ActGaze, which learns action-grounded gaze through counterfactual visual interventions for high-precision manipulation by focusing on what the robot should look at during complex tasks.

The most significant development concerns the online adaptation of simulation models to real-world conditions through closed-loop systems, which is crucial for making autonomous systems reliable outside of controlled environments. This work involved testing a system that uses current aligned link manipulation techniques to adapt how a single arm lifts oversized objects, and the results showed promising alignment in handling these complex physical interactions.

This adaptation process builds on prior efforts in modular reconfigurable aerial-ground platforms designed for field operations, which explored how different components can be rearranged for varied tasks. Furthermore, research into outcome-sensitive motion search for impact-aware dexterity in catching objects suggests a path toward more robust real-world interaction planning.

A simpler torque observation alignment method was also investigated to achieve zero shot sim to real grasping with a direct drive gripper, which is foundational for making these adaptations work seamlessly. This is complemented by work on support-enhanced granular jamming grippers that improve reinforcement learning based grasping when using continuum manipulators. Finally, interactive bi-directional tracing of monochrome cables amidst clutter provides a method for navigating complex physical environments during operation.

Today's papers

The papers

Important terms

MorphIK
This technique uses a robot's shape to adjust neural inverse kinematics, allowing it to calculate joint movements even when the exact physical structure of an unknown robot is not perfectly known.
World Action Agent
This system uses large vision-language models to rehearse actions in a simulated world before performing them in reality, improving its ability to handle complex tasks through experience.
RAPID
This focuses on agentic programming directly from demonstrations, letting robots learn how to perform tasks just by watching someone do them, bypassing the need for explicit low-level control code.
RotVLA
This development introduces a rotational latent action to vision-language action models, helping them better plan and execute physical movements in real-world scenarios.
Physics Guided Residual Reinforcement Learning
This method uses physics constraints to guide reinforcement learning policies, enabling humanoid robots to safely navigate complex and confined spaces without collisions.