Robotics papers — 2026-09-23
The focus today is on making robots better at understanding and interacting with complex, messy real worlds, as this presents the biggest challenges right now. CODA is a generative model that tries to reconstruct a whole scene geometry from just one picture containing both color and depth information. The main problem it tackles is that when robots look at cluttered scenes, simply detecting objects separately and then putting them back together often leads to errors where things drift or overlap incorrectly.
To keep the reconstructed shape accurate against what is actually seen, CODA employs two specific 3D grounding mechanisms that act like checks to ensure the geometry stays consistent with observed surfaces while filling in the parts we cannot see. This approach has shown better accuracy and a higher success rate for keeping objects in place compared to other methods.
This scene reconstruction work connects to how robots plan their movements around people, which is covered by Destination Support Restoration. DSR is a method that repairs the set of possible future destinations a robot considers when planning its path around pedestrians. It does this by intelligently reallocating hypotheses, ensuring that even if some options are less likely, the crucial ones remain represented in the robot's decision-making set.
Furthermore, we are also looking at how to make vision-language models more reliable when they control physical actions through VLAQuantBench. This evaluation method tests different ways of reducing model size by changing numerical precision, showing that careful selection of these settings can dramatically boost performance on certain tasks. This connects to the broader goal of building robust systems, as we also have GINIO which provides a geometric interface for neural inertial odometry, ensuring that robot motion predictions respect the physical laws governing sensor mounting.
The work on MAVP is particularly important because it tackles the fundamental problem of reliable execution in mobile manipulation, where simply having a good plan is not enough; the robot needs to accurately move its base while performing complex arm movements. MAVP addresses this by reconstructing a static map from demonstrations and then using that map to predict explicit base-pose targets, which are then tracked with localization feedback during operation. This means the policy is constantly checking if its base movement matches what it learned from the demonstrations, allowing for corrections when deviations occur.
This framework builds upon earlier work that focused on improving execution reliability through pose-noise augmentation during training, which helps make the system robust to errors in its pose input. The overall approach involves jointly predicting target base poses alongside arm and gripper actions, and a low-level controller uses feedforward motion combined with pose error feedback to correct any deviations from those predicted targets. This entire system is tested across six real-world manipulation tasks and three different policy families, showing that MAVP achieves higher task success rates than unanchored velocity control in every test.
In the realm of human-robot collaboration, the PROACT framework is significant because it moves beyond simply making a robot responsive to user input or efficient alone by incorporating predictions of human collaborative behavior into its control loop. By training on a large dataset of dyadic transport demonstrations, PROACT uses a transformer architecture to distill this complex behavior into predictions about future object motion. This anticipation allows the robot to adjust its compliant whole-body control proactively, leading to substantial reductions in interaction work and completion time compared to prior methods like compliance-only or MPC baselines.
This predictive modeling of human intent is complemented by the focus on geometric supervision in vision-language-action models, where Geometry-Change VLA learns to predict future geometry changes from current observations. This helps ground the high-level planning in actual physical changes, and when combined with a residual flow recovery policy, it achieves very high success rates on benchmarks.
Meanwhile, there is a separate line of research focusing on autonomous microrobot navigation, which is crucial for minimally invasive procedures because it separates long-range geometric planning from short-range reactive control. The analytic geometry planner generates collision-free global routes quickly, while rule-based or reinforcement learning local controllers handle immediate obstacle avoidance before handing control back to the global path. This modular design allows the system to operate within tight video-rate control budgets and has been demonstrated in both static and dynamic microfluidic settings.
Another area where predictive modeling is key is in monocular drone navigation, where the Skytopia framework uses an action-conditioned latent world model to predict observation changes based on intended motion. Instead of relying solely on a prediction that feeds into action generation, Skytopia focuses on the representation needed to produce the next observation, which allows it to perform well across various navigation goals in simulation and even on physical drones without fine-tuning.
Finally, concerning vision-language-action models, research is exploring what kind of action representations actually matter for closed-loop control rather than just reconstruction fidelity. Studies show that while certain representations might have lower nominal reconstruction error, they can lead to less predictable token sequences and lower success rates when evaluated on policy performance across different training seeds.
The most significant finding from today’s work concerns how an attacker can plant a hidden backdoor directly into the instructions of an LLM controlling a robot. This method matters because it bypasses existing defenses that only look for external triggers, showing a new way to compromise autonomous systems internally.
We demonstrated this by manipulating the robot controller's instructions to embed a backdoor that activates based on a specific, rare sequence of the robot's own past actions. This means the malicious behavior, like causing a collision or stopping entirely, only happens when the robot has performed that exact sequence of movements previously.
This history-based attack proved highly effective in our simulations, achieving nearly perfect success rates while remaining very hard to spot during normal operation. This is a critical vulnerability because it exploits the agent's internal state rather than relying on easily detectable external cues.
Today's papers
- CODA: Depth-Aligned Scene Completion and Object Decomposition from a Single RGB-D Image We introduce a generative model that reconstructs the complete scene geometry from one image and then separates surfaces into environment and objects. [paper]
- Destination Support Restoration for Finite-Set Multimodal Trajectory Prediction This method repairs limited destination predictions by reallocating redundant hypotheses to deficient modes without retraining the host predictor. [paper]
- VLAQuantBench Closed-Loop Evaluation of Post-Training Quantization for Vision-Language-Action Models This benchmark evaluates how different quantization methods affect the performance of vision language action models. [paper]
- GINIO A Geometric SO(3)-Equivariant Interface for Neural Inertial Odometry We provide a geometric interface that ensures neural inertial odometry measurements transform correctly under any rotation of the sensor frame. [paper]
- Teaching Reinforcement Learning and Humanoid Robotics to High-School Students An integrated course framework organizes robotics research workflows into a structured curriculum for precollege students.
- TriWorldBench A Tri-View Consistency Perspective on Embodied World Models This benchmark evaluates how well different camera views consistently describe the same action and object state in embodied world models. [paper]
- IndustrialVLA-Bench A Traceable Multi-Axis Evaluation of Open Robot Policy Models This paper provides a unified evaluation schema to compare the capabilities of vision language action and world action models. [paper]
- Provably Safe Neural Network Controllers via Differential Dynamic Logic We introduce methods to verify the infinite-time safety of neural network control systems by combining control theory with differential dynamic logic. [paper]
- MAVP Map-Aware Visuomotor Policies for Mobile Manipulation This framework improves robot manipulation reliability by predicting base pose targets and tracking them using localization feedback. [paper]
- Learning from Humans for Proactive Assistance in Human-Robot Collaborative Transport PROACT introduces a framework that uses human behavior models to enable compliant whole-body control during collaborative transport. [paper]
- Real-time autonomous magnetic microrobot navigation across dynamic and biologically relevant environments This closed-loop framework separates long-range geometric planning from short-range reactive control for microrobot navigation in complex biological settings. [paper]
- Beyond Reconstruction Error: Analytical and Data-Driven Action Tokenization for Autoregressive Vision-Language-Action Models This work investigates which action representation properties are important for closed-loop control beyond simple reconstruction error. [paper]
- Skytopia Monocular Drone Navigation with Action-Conditioned Latent World Models We propose a policy built on an action conditioned latent world model that uses prediction to guide navigation in unseen environments. [paper]
- HABILIS Brain 0 Geometry-Change Supervision for Vision-Language-Action and Residual Flow Recovery This method learns geometry change tokens to provide geometric supervision for vision language action policies during manipulation. [paper]
- AgenticDiffusion Multi-View Reasoning with View-Conditioned Diffusion Planning for Vision-Based UAV Navigation This framework semantically coordinates different camera views to achieve mission goals in complex UAV navigation. [paper]
- StepTrigger Contact-State-Triggered Backdoor Attacks on VLM-Powered Legged Robots We present an attack that exploits foot contact patterns as a backdoor trigger for legged robot vision language models. [paper]
- Silent Sabotage Internal State Triggered Backdoor Attacks on LLM-Powered Robotic Systems This paper demonstrates how an attacker can embed a stealthy backdoor into LLM controllers triggered by rare sequences of the robot's own past actions. [paper]
The papers
- Provably Safe Neural Network Controllers via Differential Dynamic Logic —
- Real-time autonomous magnetic microrobot navigation across dynamic and biologically relevant environments —
- AgenticDiffusion: Multi-View Reasoning with View-Conditioned Diffusion Planning for Vision-Based UAV Navigation —
- GINIO: A Geometric SO(3)-Equivariant Interface for Neural Inertial Odometry —
- Learning from Humans for Proactive Assistance in Human-Robot Collaborative Transport —
- VLAQuantBench: Closed-Loop Evaluation of Post-Training Quantization for Vision-Language-Action Models —
- HABILIS Brain 0: Geometry-Change Supervision for Vision-Language-Action and Residual Flow Recovery —
- IndustrialVLA-Bench: A Traceable Multi-Axis Evaluation of Open Robot Policy Models —
- CODA: Depth-Aligned Scene Completion and Object Decomposition from a Single RGB-D Image —
- Teaching Reinforcement Learning and Humanoid Robotics to High-School Students: An Expert-Validated Curriculum Design on a Low-Cost Open Platform —
- Beyond Reconstruction Error: Analytical and Data-Driven Action Tokenization for Autoregressive Vision-Language-Action Models —
- Destination Support Restoration for Finite-Set Multimodal Trajectory Prediction —
- Skytopia: Monocular Drone Navigation with Action-Conditioned Latent World Models —
- StepTrigger: Contact-State-Triggered Backdoor Attacks on VLM-Powered Legged Robots —
- Silent Sabotage: Internal State Triggered Backdoor Attacks on LLM-Powered Robotic Systems —
- TriWorldBench: A Tri-View Consistency Perspective on Embodied World Models —
- MAVP: Map-Aware Visuomotor Policies for Mobile Manipulation —
Important terms
- CODA
- A generative model that reconstructs a complete 3D scene geometry from a single image containing both color and depth information, helping robots understand cluttered environments.
- Destination Support Restoration (DSR)
- A method that repairs the set of possible future destinations a robot considers when planning paths around people, keeping crucial options in its decision-making set.
- MAVP
- A framework for mobile manipulation that reconstructs a static map from demonstrations to predict explicit base-pose targets, allowing the robot to correct its base movement during operation.
- PROACT
- A human-robot collaboration framework that uses predictions of human collaborative behavior, distilled via a transformer, to proactively adjust the robot's whole-body control.
- History-based Attack
- A new method for compromising robot instructions by embedding a backdoor that activates based on a specific sequence of the robot's own past actions.