Robotics papers — 2026-09-29

Today's work centers on understanding the gap between what we intend for a large language model to do and what it actually executes when controlling physical robots, which is crucial because this gap directly impacts how reliably we can deploy autonomous systems in real-world settings. We explored how to make these LLM-based robots more robust by focusing on better planning methods for physical tasks.

One key area involved developing dynamic buffers for cost-efficient planning when rearranging objects on a tabletop using stacking techniques, which is about figuring out the cheapest way to move things around. This work connects to learning whole-body control methods like FastGrasp, which focuses on teaching mobile manipulators how to perform fast and dexterous grasping.

Another piece of the puzzle is RoboAlign-R1, which uses distilled multimodal reward alignment to improve robot video world models. This helps the robot better understand its visual environment when it needs to make decisions about action execution.

We also looked at adaptive action execution for world action models, which addresses when a system should trust its own imagination versus when it needs external guidance. This contrasts with work on recursive self-improvement in robotics, where we examined what stops agents from endlessly discovering new skills through agentic skill discovery rounds.

Finally, we touched upon robot manipulation using GPT-6-Astra to see how body knowledge and experience reuse translate into emergent skills during simtoreal transfer. This all feeds into the broader challenge of timed rule-based supervision for end-to-end autonomous parking policies, which shows how we can impose structure on complex behaviors.

The most crucial development this morning involves PHIRL, which tackles aligning learned rewards with task progress in inverse reinforcement learning. This problem is vital because it allows agents to learn optimal behaviors from demonstrations rather than just trial and error. This work attempts to bridge the gap between what an agent learns through experience and what the desired outcome actually requires.

A related effort is DS-VLA, which introduces a dendritic-inspired vision-language-action model designed for robust action control. This aims to make these complex models more reliable when interacting with the real world. This is significant because it moves beyond simple imitation by incorporating visual and linguistic understanding directly into the action planning loop.

Then there is RECAST, which focuses on recasting vision-language semantics into an actionable cost map specifically for robot navigation. This helps robots understand what costs are associated with different areas of a scene. This builds upon the need for better semantic grounding in embodied systems.

We also see work on Affordance-Conditioned Decision Making, which bridges the semantic-spatial gap in zero-shot cross-floor vision and language navigation by focusing on how physical affordances guide decisions across different environments. This is important for making navigation strategies more generalizable.

DRAM focuses on delta-rule recurrent associative memory to improve robot manipulation policies. This suggests a way to store and recall relevant past experiences efficiently during complex tasks. This method seeks to enhance the memory capabilities of embodied agents.

The most critical piece of work today involves figuring out how to distill the complex behavior of foundation models into policies that robots can actually use in the real world because this unlocks deployable intelligence. We saw some promising initial attempts with VPTwin, which focuses on real-sim-real video prediction for robotic manipulation planning. This suggests a way to bridge the gap between simulation and physical action.

Then there is ProcVLM, which tackles learning procedure-grounded progress rewards for robotic manipulation. This means it teaches robots what to do by rewarding them based on how well they follow a specific procedure during the task. This builds upon that by GT-VLA, which introduces target-conditioned trace guidance for generalizable robotic manipulation, aiming to make the learned skills work across different scenarios.

A more recent direction explored federated subspace guided vision-language-action policy distillation for non independent and identically distributed multi robot manipulation. This is important because it addresses the challenge of making policies work when robots have different experiences. This connects to how we are trying to stabilize online learning with neural ordinary differential equations, a method that provides Lyapunov guarantees for stable learning as the system evolves.

Finally, there was work on think fast plan selectively through adaptive deliberation for efficient data-driven model predictive control. This is about making robots decide what to do quickly and intelligently based on the data they have. This seems like it could feed into methods like Schur-Neural Kalman filters that learn consistent corrections to traditional extended Kalman filters.

The most critical development this morning is the work on what matters for latent actions in robot learning because understanding these underlying representations is key to creating policies that generalize across different tasks. The research explored what truly drives successful manipulation policies, and one line of inquiry focused on identifying these crucial latent actions.

A study investigated what matters for viewpoint-generalizable policies in visual imitation learning, which suggests there are specific aspects of the visual input that are most important for a robot to learn from when it needs to perform a task regardless of its exact viewing angle. This finding is significant because it moves beyond simple pixel matching toward understanding the semantic information that truly guides action selection.

Another piece of work tackled the representation for robust robot manipulation, focusing on copper-policy, which seems to be trying to distill the essential visual features needed for reliable physical interaction with objects. This is important because if you can capture the right representation, you can build a system that doesn't fail when things move slightly out of position.

Then there was collisiongat, which introduced a controller-agnostic one-step collision screening method for multi-agent motion. This aims to prevent robots from bumping into each other while moving around. This is a practical step toward safer deployment in shared environments.

SPIDER addressed scalable physics-informed dexterous retargeting by incorporating physical constraints directly into the learning process. This helps the robot understand how forces and gravity affect its movements during manipulation. This connects to the viewpoint work because understanding physics helps ground those visual representations in real-world dynamics.

Stereopolicy improved robotic manipulation policies by integrating stereo perception, meaning it used depth information from multiple views to make better decisions about how to grasp or move an object. This is a direct application of leveraging richer sensory input for better control.

Finally, tac2pix introduced image-space visuo-tactile fusion for dexterous manipulation. This combines visual and tactile data to give the robot a richer sense of touch while it's working on something intricate. This builds upon the representation work by adding a crucial sense of physical contact.

The most significant development today involves the AquaBEV-Nav system, which tackles the challenge of learning occupancy in underwater environments for navigation and exploration. This work is crucial because accurately mapping an unknown underwater space is fundamental for any autonomous vehicle operating in such conditions.

This effort built upon prior work on World SLAM Model, which aimed to achieve joint world modeling for simultaneous localization and mapping. Building on that foundation, the AquaBEV-Nav approach seems to have made a specific leap by learning BEV occupancy directly within the underwater context. This means it's not just mapping the environment but understanding where things are likely to be in a bird's eye view underwater.

Another key piece of research is FINE, which focuses on future-informed navigation encoding for vision-language navigation tasks. This work attempts to make navigation more data efficient by incorporating knowledge about what might happen next into the visual and language processing pipeline. This connects this forward-looking capability to the broader goal of robust autonomous movement.

Then there is TriDrive, which integrates driver, vehicle, and road modeling for better forecasting and driver monitoring in terrestrial driving scenarios. This provides a different kind of context for modeling dynamic interactions compared to the underwater focus of AquaBEV-Nav.

Moving into manipulation, Dynamic Manipulation with World-Action Models via Counterfactual Planning was explored to enable planning actions in complex dynamic scenes by using counterfactual reasoning. This is distinct from the purely navigational focus of the other papers, but both contribute to a larger toolkit for autonomous systems.

Finally, there is DeltaWAM, which introduces change-centric visual foresight using delta tokens for more efficient world-action models. This suggests a method for updating world understanding incrementally rather than re-processing everything every time something changes. This incremental learning idea complements the persistent memory goals suggested by PORTER, which aims to create edge-cloud residency for persistent 3D scene graph memory.

The most significant piece of work today involved SocialHumanoid, which attempts to create expressive humanoid behavior through one-step co-speech motion generation. This matters because it moves beyond simple task execution to imbue robots with a more natural, communicative presence. The core idea is that by generating speech and motion simultaneously from a single input signal, we can achieve a level of embodied interaction previously unattainable.

This approach builds upon earlier explorations into recursive harness distillation across agents for robot manipulation, which focused on improving how different robotic systems coordinate their physical actions. That work provided the necessary framework for understanding agent-to-agent communication in complex settings. Furthermore, the development of a multi-modal tactile fingertip design for robotic hands aimed to enhance dexterous manipulation by giving the robots better sensory feedback during physical interaction.

The integration challenges are becoming clearer when looking at AI-driven collaborative assembly line inspection, which deals with the practical hurdles of deploying these sophisticated systems in real industrial environments. This system aims to solve problems related to system integration and deployment, showing how theoretical models translate into tangible operational constraints. This operational reality is informed by human-guided planning for complex manipulation tasks using the screw geometry of motion, which provides a geometric understanding necessary for precise physical execution.

Finally, the underlying cognitive architecture is being refined through unifying deep predicate invention with pre-trained foundation models. This seeks to give AI a more robust way to reason about actions and concepts. This cognitive advancement is complemented by LogicEnvGen, which generates diverse simulated environments driven by task logic, creating the necessary training ground for embodied AI agents to practice these complex behaviors.

The most critical piece of work today involves learning geometrically grounded amodal three dimensional representations for view generalizable robotic manipulation because it directly addresses the core challenge of making robots understand and interact with the physical world in a way that generalizes across different viewpoints. This research explored methods to create these representations, specifically focusing on how they can be used for manipulation tasks.

A significant development was the work on self-evolutionary replanning for failure aware motion planning. This attempts to make robot movement smarter when things go wrong by allowing the plan to change dynamically based on observed failures. This is important because it moves beyond static plans that fail easily in real-world scenarios.

Another area of focus was dynamic model identification and gravity compensation for the dVRK-Si patient side manipulator. This deals with making a specific robotic arm function better by figuring out how its physical properties change over time and compensating for gravity effects. This helps ensure precise control during delicate procedures.

The work on task driven co design of heterogeneous multi robot systems is also relevant as it tackles how different robots can work together effectively toward a common goal. This connects to the effort in learning control policies to provably satisfy hard affine constraints for black box hybrid dynamical systems, as both aim to ensure reliable system behavior under tight physical rules.

Finally, there was some exploration into premoe which focuses on robust preference modeling with mixture of experts reward learning. This is a technique used to train agents based on human preferences in complex decision-making environments. This work complements the broader goal of creating more capable and adaptable robotic systems.

The most significant development today involves DriveAnchor, which tackles the core challenge of planning for autonomous driving by using progressive anchor-based flow learning. This approach aims to build robust plans by learning how to transition between different states, which is crucial because it addresses the fundamental difficulty of long-horizon decision-making in complex driving scenarios.

This method builds upon work like PACE, which focuses on phase-aware chunk execution for robot policies through action chunking. PACE attempts to break down complex tasks into smaller, manageable chunks based on the current phase of operation. This helps manage computational load during execution.

Efficient-WAM presents a 1 billion parameter world-action model designed for low-cost future imagination, offering a way to simulate what might happen next without requiring massive computational resources for every possible outcome. This efficiency contrasts with Assistron, which explores Bayesian shared autonomy by integrating off-the-shelf vision language action models into the system.

Assistron leverages pre-trained vision language action models within a Bayesian framework to enable shared autonomy, meaning it lets the robot and human collaborate more effectively in real-time situations. Similarly, SurgVIL scales surgical robot imitation learning by utilizing open-source surgical videos to improve performance in complex medical tasks.

RoboEdit focuses on turning human manipulation videos into scalable robot experience. This is a method for generating diverse training data for robotic systems from existing demonstrations. This contrasts with reduced Cartesian kinetostatics, which deals with residual stabilization during full-shape propagation in tendon-driven continuum robots.

Finally, Path Planning with Motion Primitives in Dynamic Environments using SIPP on Lattices addresses path planning in dynamic settings by employing motion primitives within lattice structures to handle environmental changes effectively.

Today's papers

The papers

Important terms

PHIRL
This method aligns learned rewards with actual task progress in inverse reinforcement learning. It helps agents learn optimal behaviors directly from demonstrations instead of relying solely on trial and error.
DS-VLA
A dendritic-inspired model that combines vision, language, and action for robust control. It aims to make complex models more reliable when interacting with the physical world.
VPTwin
This focuses on real-sim-real video prediction specifically for robotic manipulation planning. It seeks to bridge the gap between what happens in simulation and what actually occurs physically.
AquaBEV-Nav
A system designed to learn occupancy in underwater environments for navigation. This is crucial for mapping unknown spaces and allowing autonomous vehicles to explore underwater conditions.