Robotics papers — 2026-09-14
Building a robust whole-body control system for humanoid robots is key to making them useful tools for interacting with the real world reliably. This work involves GigaBrain-WBC-0.5, which uses a causal Transformer to jointly predict the next action, next state, and the distribution of future latent behaviors. This unified policy aims to handle real-time commands while staying robust against implausible instructions and physical disturbances.
This model is built upon an automatic terrain-annotation pipeline that recovers full three dimensional contact geometry from motion data. This allows researchers to annotate terrain at a scale comparable to existing motion datasets, and this geometric understanding feeds into the prediction process. During deployment, the distribution of next behaviors is reused to flag implausible commands and retract them onto learned behaviors. The result is a policy that attempts tasks in a best effort manner while remaining robust to falls and disturbances.
Another area of work focuses on improving how robots perceive their surroundings when things are blocked, which is crucial for human-centered operations like search and triage. The OA-NBV pipeline autonomously selects the next traversable viewpoint by scoring candidate views using a target-centric visibility model that specifically accounts for occlusion, target scale, and completeness. This approach has shown success in both simulation and real world trials with over ninety percent success rates, significantly improving observation quality compared to baseline methods.
Furthermore, researchers are exploring how to generate accurate physical simulations from natural language descriptions of scenes using PhysCodeBench. This benchmark helps translate text into executable simulation code by measuring physical correctness through conservation-law residuals and expert assertions rather than just checking if the code runs. A self-corrective multi agent refinement framework is being used as a reference method because it shows that targeted correction drives physical accuracy better than generic iterative refinement, nearly tripling the pass rate of proprietary baselines.
Finally, work on car-following modeling introduces the Markov Chain Car-Following model to improve how robots follow other vehicles. This approach represents state transitions as a Markov process and predicts behavior by sampling accelerations from empirical distributions within discretized state bins. This method outperforms several physics based baselines on datasets like WOMD, providing a robust foundation for simulating population level stochastic traffic behavior without needing manual parameter calibration.
The work that matters most right now is Pelican-Sim one point zero because it creates a general world model simulator for embodied intelligence. This means the simulator can predict what will happen next based on what the robot sees and does to help it learn and make decisions. Its unified action representation across many different robot types keeps the model valid everywhere, and its action visual injection gives it much better control over how the robot behaves in different scenes.
The sparse mixture of experts layer helps with handling different dynamics while reducing conflicts between modalities, which is a key part of making the simulation more robust. This improved simulation capability is then used to train downstream applications that have shown huge gains in success rates, such as raising policy success from seventy percent to ninety-three percent when using fifty generated trajectories added to fifty demonstrations per task.
Language guided terrain adaptive neural mp control for autonomous traversal addresses the difficult problem of how articulated tracked robots should move reliably through complex, contact-rich environments like stairwells. This framework combines a learned kinematics model that predicts short-horizon movements with neural kinematically model predictive control that plans with multi-objective costs, and a large language model to propose updates to the weights.
VertexCBF improves safety in autonomous systems by learning neural control barrier functions in a scalable way, which helps ensure systems stay within safe bounds without being overly conservative. This method uses GPU parallel vertex restricted tree search to generate supervision points efficiently, allowing it to recover large safe sets where other methods might fail or be too cautious.
The comfort by construction work tackles the issue of safety metrics being inflated by abrupt maneuvers in driving simulators. It proposes an adaptive action parameterization that adjusts the control grid at every step to match the actual feasible control set. This approach keeps comfort violations below one percent while maintaining navigability on challenging routes.
The most important takeaway is the development of the FLOAT Drone, which tackles the fundamental problem of enabling aerial robots to operate up close by solving the dual challenge of generating manipulation forces while fighting gravity. This is crucial because it unlocks new possibilities for tasks requiring precise interaction with nearby objects.
The core difficulty lies in how propulsion systems must simultaneously create pushing or pulling forces and keep the drone airborne, which causes dynamic coupling effects during physical contact. Existing fully-actuated unmanned aerial vehicles manage these coupling issues through six-degree-of-freedom force-torque decoupling, but these current designs are often too large, which hurts their ability to maneuver in real situations.
FLOAT Drone addresses this by introducing two main structural changes: it integrates control surfaces into fully-actuated systems for the first time, which significantly reduces disruptive lateral airflow disturbances during operation. This is paired with a coaxial dual-rotor setup that keeps the robot compact while still being very efficient at hovering.
To handle these complexities, researchers developed hierarchical position and attitude controllers capable of switching between fully-actuated and underactuated modes depending on the task. Experimental testing in real-world scenarios confirmed that this system successfully performs its intended close-proximity operations as designed.
Today's papers
- GigaBrain-WBC-0.5: A Behavior World Model for Robust Humanoid Whole-Body Tracking with Environment Interaction Whole-body motion tracking policies can be made robust by modeling how the environment changes what the robot can do next. [paper] [episode]
- PhysCodeBench: Benchmarking Physics-Aware Symbolic Simulation of 3D Scenes via Self-Corrective Multi-Agent Refinement This paper introduces a benchmark to test if language models can generate simulation code that correctly follows physical laws. [paper] [episode]
- OA-NBV: Occlusion-Aware Next-Best-View Planning for Human-Centered Active Perception on Mobile Robots This method helps robots choose the best viewpoint to see occluded people by considering occlusion and target completeness. [paper] [episode]
- Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model This model uses a shared trajectory model to predict actions, observations, and goals from language, vision, and dynamics simultaneously. [paper]
- IMPLY: Physically Anchored Consistency for World-Model Rollouts This framework checks the consistency of world models by anchoring them to physical evidence rather than just internal predictions. [paper]
- Scenario-Independent Criticality Assessment and Prediction for Vulnerable Road Users in Autonomous Driving This research proposes a metric to assess safety risks from different object types, especially vulnerable road users, regardless of the specific scenario. [paper]
- Very Exciting: Zero-Shot Model Predictive Control of Buildings via Excitation-Based Generalized Transfer Learning Models This technique allows model predictive control to work on new buildings without specific training data by using generalized models pre-trained on diverse source buildings. [paper]
- An Empirical Markov Chain Car-Following (MC-CF) Model This paper proposes a probabilistic model for car following that uses empirical distributions to predict traffic behavior without needing manual behavioral parameter calibration. [paper]
- Pelican-Sim 1.0: A General World Model Simulator for Embodied Intelligence This simulator is designed to predict future observations from robot actions and visual context across different robot types and scenes. [paper]
- Language-Guided Terrain-Adaptive Neural MPC for Autonomous Traversal of Articulated Tracked Robots This framework uses language to guide a neural model predictive control system to help articulated robots navigate complex, contact-rich environments. [paper]
- VertexCBF: Improving Neural Control Barrier Functions via Vertex-Restricted Control Search This method learns safe control functions for robots by efficiently searching for critical points in the control space. [paper]
- Adaptive Agent Design We study how agents can design their own transition rules and policies when interacting with uncertain or non-Markovian environments. [paper]
- Comfort by Construction: Adaptive, Comfort-Bounded Action Spaces for Learned Driving Policies This approach adjusts action spaces in real-time to ensure that learned driving policies produce smooth, human-like motions. [paper]
- A Multi-Vehicle Dataset with Camera, LiDAR, and Radar Sensors and Scanned 3D Models for Custom Auto-Annotation using RTK-GNSS This paper describes a dataset that combines sensor data with 3D scans to provide highly accurate reference information for autonomous driving perception. [paper]
- Computing at Sea: Floating and Offshore Data Centres as a Pathway to Sustainable AI Infrastructure This article explores using offshore data centers to sustainably power AI by utilizing ocean resources and renewable energy. [paper]
- ARC: Autonomous Robotics Compliance A Three-Layer Governance Architecture for Deployed Autonomous Systems This paper proposes a three-layer framework for governing the safety and compliance of deployed autonomous robots. [paper]
- FLOAT Drone: A Fully-actuated Coaxial Aerial Robot for Close-Proximity Operations This robot is designed to operate closely with humans by using a coaxial dual-rotor configuration to manage airflow disturbances. [paper]
The papers
- PhysCodeBench: Benchmarking Physics-Aware Symbolic Simulation of 3D Scenes via Self-Corrective Multi-Agent Refinement — The paper introduces PhysCodeBench, a novel and comprehensive benchmark designed to evaluate an AI model's ability to perform physics-aware symbolic simulation within complex 3D environments. [episode]
- GigaBrain-WBC-0.5: A Behavior World Model for Robust Humanoid Whole-Body Tracking with Environment Interaction — I apologize, but the text for "GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction" was not included in your request. [episode]
- OA-NBV: Occlusion-Aware Next-Best-View Planning for Human-Centered Active Perception on Mobile Robots — This paper introduces OA-NBV, an occlusion-aware Next-Best-View planning pipeline designed for human-centered active perception on mobile robots. [episode]
- Pelican-Sim 1.0: A General World Model Simulator for Embodied Intelligence —
- IMPLY: Physically Anchored Consistency for World-Model Rollouts —
- Adaptive Agent Design —
- Computing at Sea: Floating and Offshore Data Centres as a Pathway to Sustainable AI Infrastructure —
- VertexCBF: Improving Neural Control Barrier Functions via Vertex-Restricted Control Search —
- Very Exciting: Zero-Shot Model Predictive Control of Buildings via Excitation-Based Generalized Transfer Learning Models —
- A Multi-Vehicle Dataset with Camera, LiDAR, and Radar Sensors and Scanned 3D Models for Custom Auto-Annotation using RTK-GNSS —
- ARC: Autonomous Robotics Compliance A Three-Layer Governance Architecture for Deployed Autonomous Systems —
- Comfort by Construction: Adaptive, Comfort-Bounded Action Spaces for Learned Driving Policies —
- Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model —
- FLOAT Drone: A Fully-actuated Coaxial Aerial Robot for Close-Proximity Operations —
- Language-Guided Terrain-Adaptive Neural MPC for Autonomous Traversal of Articulated Tracked Robots —
- An Empirical Markov Chain Car-Following (MC-CF) Model —
- Scenario-Independent Criticality Assessment and Prediction for Vulnerable Road Users in Autonomous Driving —
Important terms
- Causal Transformer
- A type of transformer model used to jointly predict the next action, next state, and future latent behaviors in humanoid robots. It helps create a unified policy that is robust against unpredictable real-world commands.
- Automatic Terrain-Annotation Pipeline
- A system that recovers full three-dimensional contact geometry from robot motion data. This allows researchers to accurately map terrain at a scale comparable to existing datasets for better simulation and prediction.
- Target-Centric Visibility Model
- A model used by robots to autonomously select the next traversable viewpoint. It scores candidate views based on occlusion, target scale, and completeness, significantly improving how robots perceive blocked environments.
- PhysCodeBench
- A benchmark that translates natural language descriptions of scenes into executable simulation code. It measures physical correctness using conservation-law residuals to ensure the generated simulations are accurate.
- Markov Chain Car-Following Model
- A model for car-following that treats state transitions as a Markov process. It predicts vehicle behavior by sampling acceleration from empirical distributions, offering a robust way to simulate traffic without manual tuning.