Robotics papers — 2026-09-14

Building a robust whole-body control system for humanoid robots is key to making them useful tools for interacting with the real world reliably. This work involves GigaBrain-WBC-0.5, which uses a causal Transformer to jointly predict the next action, next state, and the distribution of future latent behaviors. This unified policy aims to handle real-time commands while staying robust against implausible instructions and physical disturbances.

This model is built upon an automatic terrain-annotation pipeline that recovers full three dimensional contact geometry from motion data. This allows researchers to annotate terrain at a scale comparable to existing motion datasets, and this geometric understanding feeds into the prediction process. During deployment, the distribution of next behaviors is reused to flag implausible commands and retract them onto learned behaviors. The result is a policy that attempts tasks in a best effort manner while remaining robust to falls and disturbances.

Another area of work focuses on improving how robots perceive their surroundings when things are blocked, which is crucial for human-centered operations like search and triage. The OA-NBV pipeline autonomously selects the next traversable viewpoint by scoring candidate views using a target-centric visibility model that specifically accounts for occlusion, target scale, and completeness. This approach has shown success in both simulation and real world trials with over ninety percent success rates, significantly improving observation quality compared to baseline methods.

Furthermore, researchers are exploring how to generate accurate physical simulations from natural language descriptions of scenes using PhysCodeBench. This benchmark helps translate text into executable simulation code by measuring physical correctness through conservation-law residuals and expert assertions rather than just checking if the code runs. A self-corrective multi agent refinement framework is being used as a reference method because it shows that targeted correction drives physical accuracy better than generic iterative refinement, nearly tripling the pass rate of proprietary baselines.

Finally, work on car-following modeling introduces the Markov Chain Car-Following model to improve how robots follow other vehicles. This approach represents state transitions as a Markov process and predicts behavior by sampling accelerations from empirical distributions within discretized state bins. This method outperforms several physics based baselines on datasets like WOMD, providing a robust foundation for simulating population level stochastic traffic behavior without needing manual parameter calibration.

The work that matters most right now is Pelican-Sim one point zero because it creates a general world model simulator for embodied intelligence. This means the simulator can predict what will happen next based on what the robot sees and does to help it learn and make decisions. Its unified action representation across many different robot types keeps the model valid everywhere, and its action visual injection gives it much better control over how the robot behaves in different scenes.

The sparse mixture of experts layer helps with handling different dynamics while reducing conflicts between modalities, which is a key part of making the simulation more robust. This improved simulation capability is then used to train downstream applications that have shown huge gains in success rates, such as raising policy success from seventy percent to ninety-three percent when using fifty generated trajectories added to fifty demonstrations per task.

Language guided terrain adaptive neural mp control for autonomous traversal addresses the difficult problem of how articulated tracked robots should move reliably through complex, contact-rich environments like stairwells. This framework combines a learned kinematics model that predicts short-horizon movements with neural kinematically model predictive control that plans with multi-objective costs, and a large language model to propose updates to the weights.

VertexCBF improves safety in autonomous systems by learning neural control barrier functions in a scalable way, which helps ensure systems stay within safe bounds without being overly conservative. This method uses GPU parallel vertex restricted tree search to generate supervision points efficiently, allowing it to recover large safe sets where other methods might fail or be too cautious.

The comfort by construction work tackles the issue of safety metrics being inflated by abrupt maneuvers in driving simulators. It proposes an adaptive action parameterization that adjusts the control grid at every step to match the actual feasible control set. This approach keeps comfort violations below one percent while maintaining navigability on challenging routes.

The most important takeaway is the development of the FLOAT Drone, which tackles the fundamental problem of enabling aerial robots to operate up close by solving the dual challenge of generating manipulation forces while fighting gravity. This is crucial because it unlocks new possibilities for tasks requiring precise interaction with nearby objects.

The core difficulty lies in how propulsion systems must simultaneously create pushing or pulling forces and keep the drone airborne, which causes dynamic coupling effects during physical contact. Existing fully-actuated unmanned aerial vehicles manage these coupling issues through six-degree-of-freedom force-torque decoupling, but these current designs are often too large, which hurts their ability to maneuver in real situations.

FLOAT Drone addresses this by introducing two main structural changes: it integrates control surfaces into fully-actuated systems for the first time, which significantly reduces disruptive lateral airflow disturbances during operation. This is paired with a coaxial dual-rotor setup that keeps the robot compact while still being very efficient at hovering.

To handle these complexities, researchers developed hierarchical position and attitude controllers capable of switching between fully-actuated and underactuated modes depending on the task. Experimental testing in real-world scenarios confirmed that this system successfully performs its intended close-proximity operations as designed.

Today's papers

The papers

Important terms

Causal Transformer
A type of transformer model used to jointly predict the next action, next state, and future latent behaviors in humanoid robots. It helps create a unified policy that is robust against unpredictable real-world commands.
Automatic Terrain-Annotation Pipeline
A system that recovers full three-dimensional contact geometry from robot motion data. This allows researchers to accurately map terrain at a scale comparable to existing datasets for better simulation and prediction.
Target-Centric Visibility Model
A model used by robots to autonomously select the next traversable viewpoint. It scores candidate views based on occlusion, target scale, and completeness, significantly improving how robots perceive blocked environments.
PhysCodeBench
A benchmark that translates natural language descriptions of scenes into executable simulation code. It measures physical correctness using conservation-law residuals to ensure the generated simulations are accurate.
Markov Chain Car-Following Model
A model for car-following that treats state transitions as a Markov process. It predicts vehicle behavior by sampling acceleration from empirical distributions, offering a robust way to simulate traffic without manual tuning.