Robotics papers — 2026-09-23

The focus today is on making robots better at understanding and interacting with complex, messy real worlds, as this presents the biggest challenges right now. CODA is a generative model that tries to reconstruct a whole scene geometry from just one picture containing both color and depth information. The main problem it tackles is that when robots look at cluttered scenes, simply detecting objects separately and then putting them back together often leads to errors where things drift or overlap incorrectly.

To keep the reconstructed shape accurate against what is actually seen, CODA employs two specific 3D grounding mechanisms that act like checks to ensure the geometry stays consistent with observed surfaces while filling in the parts we cannot see. This approach has shown better accuracy and a higher success rate for keeping objects in place compared to other methods.

This scene reconstruction work connects to how robots plan their movements around people, which is covered by Destination Support Restoration. DSR is a method that repairs the set of possible future destinations a robot considers when planning its path around pedestrians. It does this by intelligently reallocating hypotheses, ensuring that even if some options are less likely, the crucial ones remain represented in the robot's decision-making set.

Furthermore, we are also looking at how to make vision-language models more reliable when they control physical actions through VLAQuantBench. This evaluation method tests different ways of reducing model size by changing numerical precision, showing that careful selection of these settings can dramatically boost performance on certain tasks. This connects to the broader goal of building robust systems, as we also have GINIO which provides a geometric interface for neural inertial odometry, ensuring that robot motion predictions respect the physical laws governing sensor mounting.

The work on MAVP is particularly important because it tackles the fundamental problem of reliable execution in mobile manipulation, where simply having a good plan is not enough; the robot needs to accurately move its base while performing complex arm movements. MAVP addresses this by reconstructing a static map from demonstrations and then using that map to predict explicit base-pose targets, which are then tracked with localization feedback during operation. This means the policy is constantly checking if its base movement matches what it learned from the demonstrations, allowing for corrections when deviations occur.

This framework builds upon earlier work that focused on improving execution reliability through pose-noise augmentation during training, which helps make the system robust to errors in its pose input. The overall approach involves jointly predicting target base poses alongside arm and gripper actions, and a low-level controller uses feedforward motion combined with pose error feedback to correct any deviations from those predicted targets. This entire system is tested across six real-world manipulation tasks and three different policy families, showing that MAVP achieves higher task success rates than unanchored velocity control in every test.

In the realm of human-robot collaboration, the PROACT framework is significant because it moves beyond simply making a robot responsive to user input or efficient alone by incorporating predictions of human collaborative behavior into its control loop. By training on a large dataset of dyadic transport demonstrations, PROACT uses a transformer architecture to distill this complex behavior into predictions about future object motion. This anticipation allows the robot to adjust its compliant whole-body control proactively, leading to substantial reductions in interaction work and completion time compared to prior methods like compliance-only or MPC baselines.

This predictive modeling of human intent is complemented by the focus on geometric supervision in vision-language-action models, where Geometry-Change VLA learns to predict future geometry changes from current observations. This helps ground the high-level planning in actual physical changes, and when combined with a residual flow recovery policy, it achieves very high success rates on benchmarks.

Meanwhile, there is a separate line of research focusing on autonomous microrobot navigation, which is crucial for minimally invasive procedures because it separates long-range geometric planning from short-range reactive control. The analytic geometry planner generates collision-free global routes quickly, while rule-based or reinforcement learning local controllers handle immediate obstacle avoidance before handing control back to the global path. This modular design allows the system to operate within tight video-rate control budgets and has been demonstrated in both static and dynamic microfluidic settings.

Another area where predictive modeling is key is in monocular drone navigation, where the Skytopia framework uses an action-conditioned latent world model to predict observation changes based on intended motion. Instead of relying solely on a prediction that feeds into action generation, Skytopia focuses on the representation needed to produce the next observation, which allows it to perform well across various navigation goals in simulation and even on physical drones without fine-tuning.

Finally, concerning vision-language-action models, research is exploring what kind of action representations actually matter for closed-loop control rather than just reconstruction fidelity. Studies show that while certain representations might have lower nominal reconstruction error, they can lead to less predictable token sequences and lower success rates when evaluated on policy performance across different training seeds.

The most significant finding from today’s work concerns how an attacker can plant a hidden backdoor directly into the instructions of an LLM controlling a robot. This method matters because it bypasses existing defenses that only look for external triggers, showing a new way to compromise autonomous systems internally.

We demonstrated this by manipulating the robot controller's instructions to embed a backdoor that activates based on a specific, rare sequence of the robot's own past actions. This means the malicious behavior, like causing a collision or stopping entirely, only happens when the robot has performed that exact sequence of movements previously.

This history-based attack proved highly effective in our simulations, achieving nearly perfect success rates while remaining very hard to spot during normal operation. This is a critical vulnerability because it exploits the agent's internal state rather than relying on easily detectable external cues.

Today's papers

The papers

Important terms

CODA
A generative model that reconstructs a complete 3D scene geometry from a single image containing both color and depth information, helping robots understand cluttered environments.
Destination Support Restoration (DSR)
A method that repairs the set of possible future destinations a robot considers when planning paths around people, keeping crucial options in its decision-making set.
MAVP
A framework for mobile manipulation that reconstructs a static map from demonstrations to predict explicit base-pose targets, allowing the robot to correct its base movement during operation.
PROACT
A human-robot collaboration framework that uses predictions of human collaborative behavior, distilled via a transformer, to proactively adjust the robot's whole-body control.
History-based Attack
A new method for compromising robot instructions by embedding a backdoor that activates based on a specific sequence of the robot's own past actions.