Robotics papers — 2026-09-30

Gondola tries to create grounded vision language planning for robotic manipulation, aiming to connect high-level language understanding with low-level physical actions. This work is important because it addresses the difficulty of making a robot understand what a command means in a real workspace.

AlignDrive explores aligned lateral and longitudinal planning for end-to-end autonomous driving, suggesting ways to make driving decisions more consistent across different spatial dimensions. This idea is built upon by EgoPriMo, which generates egocentric motion for interactive humanoid control, focusing on how a robot should move when interacting with a person.

Don't Drop the BATON uses agentic subtask exploration and transition-aware memory to achieve long-horizon robot manipulation. This is important because it tackles the challenge of planning complex tasks over many steps by allowing the agent to explore its options and remember past events.

Hydra presents a navigation world action model that combines discrete latent planning with continuous flow-matching execution, suggesting a way to handle both abstract planning and smooth physical movement at the same time. This contrasts with FineART, which focuses on creating a fine-grained annotated robotic trajectory dataset along with a vision language action model specifically for bimanual manipulation.

Losing the name before the box measures the cost of narrow fine-tuning when deploying a detector outside its initial training vocabulary. This is a foundational piece that examines how robust models are when they encounter novel situations not explicitly covered in their original training.

The most significant work today involved distilling privileged control barrier functions into RGB-only safety filters because it directly addresses the need for robust, real-time safety in dynamic visual navigation systems. This approach seeks to create a lightweight filtering mechanism that relies only on visual input, which is crucial when high-fidelity sensor data might be unavailable or too computationally expensive during operation.

iTeach explored interactive teaching for failure-driven adaptation of robot perception, suggesting a method where the robot learns by interacting with human feedback when it encounters unexpected situations. This contrasts with the filtering work by focusing on learning and adaptation rather than just pre-defined safety constraints.

Soft yet Effective Robots via Holistic Co-Design looked at designing robots where the physical structure and control systems are co-designed from the start, suggesting a synergy between mechanical design and control strategy for achieving soft yet effective movement. This idea of holistic design connects to how trajectory parametrization in learning on the job work might influence system behavior under uncertainty.

Learning On The Job tackled zero-shot task execution under parametric uncertainty using trajectory-parametrized dual control, meaning the robot learns to perform new tasks even when it does not know all its exact physical parameters beforehand. This learning capability is further supported by Stein-based optimization of sampling distributions in Model Predictive Path Integral Control, which refines how the system samples possible paths based on model uncertainties.

Temporal Cascading of Planning and Control for Quadrotor MPC deals with sequencing planning and control actions over time for quadrotors, which is a crucial step in ensuring smooth, temporally consistent movement. This temporal sequencing builds upon the foundational path integral control methods that handle sampling distributions.

The most pressing work involves developing methods to ensure AI systems can handle unexpected situations reliably because current models often fail outside their training domains. One line of research focused on governing capability evolution by implementing lifecycle-time compatibility checking and rollback mechanisms for AI-component based systems, which included a proof of concept evaluation on embodied agents to see if this control could actually work in practice.

This relates to the exploration aspect, where ContactExplorer was developed to guide general purpose dexterous manipulation through contact coverage guided exploration, meaning the robot learns how to touch things effectively by focusing its movements on areas it has not explored yet. Moving down slightly in importance is the work on Manifold-Constrained MPPI, which provides real time sampling based control for nonlinear equality constrained robotic systems, a technique that helps robots move smoothly even when dealing with complex physical constraints.

Another piece of research tackled reasoning chain as a control surface for a vision language action policy, exploring how altering thoughts can lead to altered actions in AI systems. This builds on the idea of structured decision making, which is somewhat connected to the ADMM based continuous trajectory optimization in graphs of convex sets, which deals with optimizing continuous paths within defined geometric shapes.

Finally, there is the work on RobotValues, which attempts to evaluate household robots when human values conflict, suggesting a deeper dive into aligning robot behavior with complex human ethical frameworks.

The most significant development today involves the work on Elastic ODYN, which tackles the problem of learning control when desired actions are physically impossible. This method uses differentiable optimization to find a path through infeasible control spaces, essentially teaching a robot how to attempt movements that violate physical constraints in a learnable way. This is crucial because real-world robotics often encounters limits that standard reinforcement learning struggles with.

We also saw progress on IR-SIM, which introduces a lightweight declarative simulator designed for navigation learning and benchmarking. This simulator allows researchers to test and compare different control strategies in a controlled environment without needing massive computational resources for full physical simulations. This provides a scalable testing ground for the more complex control algorithms being developed elsewhere.

Temporal Self-Imitation Learning showed promise in modeling dynamic systems by having an agent learn to imitate its own past behavior over time. This technique is important because it helps agents develop temporal reasoning skills, which are necessary for tasks requiring sequential decision-making.

The work on Monocular 3D Occupancy Perception for Robots on Sidewalks via Hybrid 2D-3D Learning addresses how robots can perceive their surroundings effectively when only using a single camera. This hybrid approach combines 2D and 3D learning to build a robust understanding of the environment, which is vital for safe navigation.

GPU-Accelerated Polygonal Signed Distance Functions for Real-Time Collision Avoidance provides a fast way for robots to detect potential collisions by using signed distance functions on polygonal meshes, running quickly enough for real-time applications. This speed is essential when dealing with dynamic obstacles in motion.

RynnWorld-Teleop introduces an action-conditioned world model specifically designed to facilitate digital teleoperation, allowing human operators to interact with simulated environments more intuitively. This model helps bridge the gap between high-level commands and low-level robotic execution.

Finally, RynnWorld-4D presents 4D embodied world models for robotic manipulation, which are a step toward creating comprehensive models that capture both spatial and temporal dynamics relevant to complex physical tasks.

The most important development today concerns the framework for indoor UAV swarms, which is crucial because it moves us closer to reliable autonomous navigation inside complex structures. We explored a mission-oriented coordinated navigation framework that attempts to guide multiple aerial robots together. This work involves developing a system where the swarm can coordinate its flight paths based on shared goals.

A significant piece of this effort was SAKI, which focuses on skill assembly and kinematic imitation from human videos for long-horizon mobile manipulation tasks. This means the robot learns complex actions by watching people perform them in video, allowing it to plan multi-step movements over a long distance. This builds upon the idea of learning skills directly from demonstrations.

Another area of progress involves S2A2, which uses audio-visual imitation learning for manipulation tasks by incorporating acoustic spatial information. This suggests that robots can learn how to interact with objects not just by seeing them, but also by understanding the sound cues associated with those interactions. This auditory input adds a new layer to visual learning.

We also looked at PAC-MAN, which is a perception-aware collision avoidance framework using CBF reinforcement learning for whole-body safety in humanoid dodgeball scenarios. This research tackles the problem of ensuring physical safety during dynamic human interaction by using learned policies that consider perception and collision risk across the entire body.

For ground robots facing challenging environments, TASG-Explore was introduced, which is a traversability-aware sector-guided exploration method for uneven terrain. This system helps robots decide where to go next by considering how easy or difficult the ground is to traverse in a specific direction.

The study on passive-dynamic walking inspired dynamics guidance aims at creating energy-efficient locomotion for humanoids. This involves guiding the robot's movement using principles inspired by how humans walk passively, which should lead to more efficient power usage during movement.

Finally, MagNav presents a dual-core magnetic track guidance framework designed for lighting-invariant navigation in two-wheeled robots. This framework provides robust navigation even when visual cues are poor or changing due to lighting conditions.

The most important work today involves how we can make robot policies adapt quickly when the physical hardware starts to fail because this is crucial for real-world deployment. We looked at test-time adaptation of manipulation policies under actuator degradation, where researchers found that by using specific feedback signals from the actions and outcomes, they could modify the policy in real time to maintain performance even as parts wore out. This means robots can keep doing delicate tasks without needing a full retraining cycle every time a motor degrades.

Another significant piece of research focused on creating scalable data for these systems through skillweaver, which is an agentic exploration method designed to generate robot data efficiently by focusing on neural interaction skills rather than brute-force exploration. This helps build the necessary datasets for learning robust behaviors. Following that, there was work on design and validation of an antagonistic tendon-driven dexterous robotic hand with bidirectional operation, which deals with the physical construction of complex grippers capable of both grasping and releasing objects in a controlled manner.

The concept of bilinear world models also surfaced as something important because it aims to learn representations using structured dynamics, which should lead to more efficient control methods. This relates closely to atlas, which focuses on aligned transport of latent structure for reliable world model planning, suggesting that understanding the underlying structure of the environment is key for good planning. Finally, dora addresses divergence-oriented data-relay algorithms for partially connected robot teams, which tackles how different robot groups can share and coordinate information effectively when they are not fully integrated.

The most significant piece of work today involved the attempt to actualize futures from pretrained world models into robot actions because this moves beyond mere prediction into actionable intelligence for autonomous systems. This effort was explored through the work on One from Infinity, which investigates how to translate these large-scale world models directly into executable robot policies.

A related but more specific piece of research focused on outcome-grounded world modeling for autonomous driving, specifically World4Scorer. This work tried to create a system that scores outcomes based on the world model's understanding of the environment, meaning it is trying to make decisions based on predicted results rather than just raw perception.

Then there was DQ-MPCC, which tackled dual-quaternion MPCC for quadrotor racing; this is about developing a better way for quadrotors to navigate complex maneuvers by using dual quaternions for motion control. This builds upon the foundational modeling work seen in other areas of robotics.

We also saw research into closed-form Cartesian forward kinetostatics for spatial multi-segment tendon-driven continuum robots, which is a mathematical approach to precisely controlling the movement of flexible, soft robots. This provides the low-level physical control necessary for complex manipulation tasks.

The question of policy adaptation was addressed by LIBERO-MAX, which examines whether robot policies can successfully adapt when the world itself changes unexpectedly. This is important because it tests the robustness of learned behaviors in dynamic settings.

Finally, there was work on trajectory-level mode guidance for controllable diffusion-based multi-robot motion planning, which deals with guiding multiple robots through complex paths using diffusion models to ensure coordinated movement. This connects the high-level planning concepts to practical multi-agent execution.

The work on planning oriented three dimensional scene completion using coupled TUDF occupancy representation learning is the most significant because it directly addresses how robots can build a complete understanding of an environment when they only have partial observations. This method attempts to learn a dense representation of the scene by coupling two different types of occupancy grids, which helps in inferring missing parts of the world.

This approach builds upon earlier efforts in learning to explore hidden kinematics for articulated object manipulation, suggesting that understanding how joints move can help fill in gaps in visual data. Furthermore, the work on equipDP3 presents a SIM(3)-invariant point-cloud encoder specifically designed for data-efficient humanoid locomotion manipulation, which is crucial for making robots move realistically.

Then there is the development of a robust single sensing element tactile sensor that can detect both pressure and tackiness simultaneously while decoupling the signals in real time. This sensor information feeds into inferring soil friction angle from robot foot-ground force histories, which uses a Bayesian inverse approach to figure out how slippery the ground is based on what the robot feels when it steps.

Finally, there is the cooperative multi-agent vision language action model that uses reinforced fine tuning to allow agents to work together in complex tasks. This work complements the simple agentic memory for generalist robot policies, which aims to give robots a basic form of long-term memory so they can generalize their actions better across different situations.

Today's papers

The papers

Important terms

Grounded Vision Language Planning
This focuses on helping robots connect high-level language commands with actual physical actions in a real workspace, solving how robots interpret what a command means physically.
Control Barrier Functions (CBF) Distillation
This is about creating lightweight safety filters for visual navigation by taking complex control functions and simplifying them to only use visual input, ensuring fast, real-time safety.
Elastic ODYN
This method teaches robots how to learn control when the desired actions are physically impossible. It uses optimization to find a path through spaces that violate physical constraints in a learnable way.
Trajectory-Parametrized Dual Control
This involves learning new tasks without knowing all the robot's exact physical parameters beforehand. It refines how the system samples possible paths based on model uncertainties.