Robotics papers — 2026-09-16
Today we are focusing on how robots can handle dynamic tasks where things move. This is crucial because current vision language action models struggle with this because they only look at one moment in time and cannot predict what will happen next. We are looking at motion ambiguity, where a single view doesn't show the movement of objects, and state aliasing, where similar views require different actions.
The TEMPO approach tackles this by adding two temporal inputs. It uses a motion summary from a video foundation model to fix the motion ambiguity. It also uses a compact history of proprioceptive data to resolve state aliasing. This method improved bottle handover success from forty-four percent up to seventy-four percent across four dynamic tasks, and it is unique in solving state aliasing.
This idea of using temporal context is also important when we consider how world models improve embodied intelligence. These models aim to connect perception and decision-making by anticipating consequences. The key question is which predictions actually help behavior, whether they are plausible, controllable, or actionable.
The work suggests that moving toward actionable models means ensuring predictions capture task-relevant state and reflect how interventions change that state. Another area of focus involves continual learning for robot navigation in uncertain terrains. This framework adapts to new surfaces without forgetting old ones by using a generative model to recall past experiences while incorporating uncertainty from those generated samples.
This is vital because robots need to adapt when they encounter unexpected ground conditions, like traction loss, which can cause instability. Finally, we are seeing how different approaches build upon existing vision language action systems. For instance, intrinsic robot rewarding proposes reusing the visual representations from a VLA system to evaluate the robot's own outcomes and guide policy improvement without needing a separate evaluator.
This reuse is promising because it lowers integration effort while connecting internal outcome evaluation directly to physical policy changes. The work that matters most is ProxiDex because it tackles the fundamental problem of unstable hand-object interactions. It treats proximity as a learned interaction state, which is crucial for dexterous manipulation in practice.
This framework reconstructs interaction point clouds and converts geometric distances into proximity cues. This creates a hardware-agnostic contact representation that gives immersive feedback during virtual teleoperation. This representation allows ProxiDex to learn action-conditioned proximity dynamics using a coupled forward-inverse design where future observations are predicted from actions and proximity variations are decoded from latent changes.
This dynamic understanding is then used to adaptively reweight proximity tokens across manipulation phases. Dynamics-consistency supervision guides policy inference to stabilize action generation even when visual feedback is unreliable. This method shows improved success rates and robustness compared to standard baselines in both simulation and real-world tests involving unseen objects or perturbation scenarios.
Weave builds upon this by creating a unified framework for learning whole-body dexterous humanoid-object interaction from human demonstrations. This is significant because it coordinates locomotion, balance, and hand contact. Weave first converts captured human interactions into executable robot-object references using contact-aware retargeting and approach motion completion.
At its core is a policy that jointly commands twenty-nine body joints and twelve finger joints across multiple objects. This system achieved a ninety-two point five percent success rate on trained interactions, and without further training, it still managed sixty-five point zero percent on sequences never seen during training.
TARC addresses the efficiency trade-off in robotic control by introducing a reinforcement learning framework. The policy predicts both a control action and its duration of application. TARC learns temporally extended actions by optimizing task performance under constraints on the number of control switches, allowing for adaptive modulation of control rates.
This approach matches the performance of high-frequency discrete-time controllers while operating at less than half their control frequency across different hardware platforms like an RC car and a quadruped robot. FluxVLA Engine aims to solve the engineering bottlenecks that separate promising embodied-learning algorithms from reliable robot systems.
It provides an open, configuration-driven platform for deployment by standardizing interfaces for datasets, visual-language models, and action heads. It integrates features like compositional dual-arm simulation and model-decoupled human-in-the-loop rollout. This engine connects offline learning to real robot execution through shared contracts and accelerated inference backends.
SafeFlow tackles the issue of physical hallucinations in text-driven humanoid motion generation by combining physics-guided motion generation with a three stage safety gate driven by explicit risk indicators. It uses Physics Guided Rectified Flow Matching in a VAE latent space to improve executability. The safety gate detects semantic out-of-distribution prompts and enforces hard kinematic constraints before the trajectory reaches the low level controller.
HumanEgo bridges the embodiment gap between human demonstrations and robot policies by lifting each human demonstration to an entity-level representation of hand-object interaction. This framework trains a flow matching policy with dense auxiliary objectives that amplify supervision from every trajectory. It achieves ninety-two point five percent average success across four real world tasks using only thirty minutes of human videos per task.
The most important thing we need to look at is how large language models can safely be put into control systems without breaking stability guarantees. The slow, unpredictable nature of these models clashes directly with the strict safety needs of networked control or cyber-physical systems. This work suggests that the LLM should only act as a slow supervisor setting high-level goals and constraints.
It should only act as a slow supervisor setting high-level goals and constraints, while a fast, certified inner loop handles the actual physical stability. This idea is important because it frames how we can manage risks when integrating these powerful models into critical infrastructure.
The integration maps directly to classical control problems where things like inference delay are treated like network delays. Model hallucinations are treated as bounded disturbances in the system. A related piece of work shows how uncertainty drives better path planning for unmanned ground vehicles using vision-language models.
This planner, UDAV, takes multiple uncertain route predictions from a vision-language model and uses the median prediction as its main route. It checks the spread of those predictions; if the uncertainty gets too high in any part of the path, it triggers a re-evaluation.
This uncertainty handling is really useful because when the model's confidence drops significantly, it gives an actionable signal to adjust its plan rather than blindly following a potentially flawed prediction. This approach successfully reduced the mean average displacement error from 147.4 pixels down to 115.9 pixels compared to a purely deterministic plan.
Another area of research focuses on improving how we test robot locomotion controllers in real-world settings by creating a new evaluation suite called The Neverwhere Benchmark Suite. This suite consists of over sixty three dimensional Gaussian Splatting reconstructions of various indoor and outdoor scenes. It aims to make it easier to create reproducible tests for visual locomotion policies.
This work is significant because it addresses the gap between training performance and real-world deployment by forcing policies to be tested in diverse, hyper-realistic environments. It also points out that relying only on data generated from Gaussian splats might be risky, emphasizing the need for varied data sources to ensure robust performance across different scenes.
Today's papers
- TEMPO: Learning Temporal Context for Dynamic Robot Manipulation TEMPO adds temporal inputs to vision-language-action models to fix issues related to motion ambiguity and state aliasing in dynamic manipulation tasks. [paper]
- World Models for Embodied Intelligence: From Plausible to Controllable to Actionable World models connect perception and decision-making by predicting consequences of interventions at different levels of capability. [paper]
- Continual Learning for Traversability Prediction with Uncertainty-Aware Adaptation This framework uses a generative experience recall model to allow traversability prediction methods to adapt incrementally without forgetting past experiences while accounting for uncertainty. [paper]
- Unified Heterogeneous Graph Neural Network solver for Power Flow, Optimal Power Flow and State Estimation This paper proposes a single shared graph neural network backbone that can solve power flow, optimal power flow, and state estimation problems simultaneously. [paper]
- AssemblyGrid v1: A Benchmark for Multi-Robot Production with Temporary Coalitions, Local Information, and Geometric Constraints AssemblyGrid v1 is a benchmark for testing cooperative multi-robot production under decentralized control with various workload families. [paper]
- QDTraj: Exploration of Diverse Trajectory Primitives for Articulated Objects Robotic Manipulation QDTraj uses quality-diversity algorithms to generate diverse low-level trajectory primitives for manipulating articulated objects to improve real-world performance. [paper]
- Intrinsic Robot Rewarding: Reusing VLA Representations for Autonomous Evaluation and Policy Improvement IRR proposes reusing visual representations from vision-language action models to evaluate robot outcomes and improve policies efficiently. [paper]
- Visual Cue Guided Video Planning for Generalizable Robot Navigation CueNav uses visual cues from a bird's-eye view map combined with an inverse dynamics model to enable longer-horizon, embodiment-aware video planning for robot navigation. [paper]
- ProxiDex: Learning Dynamics-Guided Proximity Policy for Dexterous Manipulation ProxiDex learns action-conditioned proximity dynamics by treating hand-object proximity as an interaction state to improve dexterous manipulation robustness. [paper]
- TARC: Time-Adaptive Robotic Control TARC is a reinforcement learning framework that jointly predicts a control action and its duration to adapt the control frequency of robotic systems online. [paper]
- Weave: Learning Whole-Body Dexterous Loco-Manipulation from Human-Object Interactions Weave learns whole-body humanoid interaction by converting human demonstrations into contact-aware references to command coordinated body and finger movements. [paper]
- FluxVLA Engine: A One-Stop VLA Engineering Platform for Embodied Intelligence FluxVLA Engine is an open platform that standardizes interfaces to connect various embodied policy components into a reproducible data-to-deployment workflow. [paper]
- Auto-HSI: Personalized human control of a robot swarm on demand by using LLMs for online automatic code generation Auto-HSI uses natural language and gestures to automatically generate personalized state machines for controlling robot swarms. [paper]
- HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos HumanEgo transfers skills from short human egocentric videos to robots by lifting demonstrations to entity-level representations and training a flow matching policy. [paper]
- SafeFlow: Real-Time Text-Driven Humanoid Whole-Body Control via Physics-Guided Rectified Flow and Selective Safety Gating SafeFlow generates physically feasible motion trajectories for humanoid control while using a safety gate to filter out unsafe text prompts. [paper]
- The Latent That Never Was: A Forensic Re-run of the CVAE Ablation in Action Chunking Transformers This paper investigates whether the encoder in action chunking transformers provides meaningful information for policy reconstruction during training. [paper]
- Large Language Models in the Loop: A Stability- and Network-Aware Survey in Networked Control, Cyber-Physical, and Multi-Agent Systems This survey analyzes how large language models can safely supervise physical control systems by operating as a slow supervisor adjusting high-level goals. [paper]
- Kernel-Based Metrics Learning for Uncertain Opponent Vehicle Trajectory Prediction in Autonomous Racing This study proposes heterogeneous kernel metrics to robustly capture diverse opponent vehicle policies for precise trajectory prediction and uncertainty estimation in racing. [paper]
- The Neverwhere Visual Parkour Benchmark Suite This work develops a suite of hyper-photo-realistic evaluation environments to improve the reproducibility and large-scale testing of visual locomotion controllers. [paper]
- UDAV: Uncertainty-Driven Adaptive VLM Waypoint Planner UDAV uses multiple stochastic trajectory predictions from vision-language models to select a nominal route and estimate uncertainty for adaptive navigation. [paper]
The papers
- TARC: Time-Adaptive Robotic Control —
- SafeFlow: Real-Time Text-Driven Humanoid Whole-Body Control via Physics-Guided Rectified Flow and Selective Safety Gating —
- QDTraj: Exploration of Diverse Trajectory Primitives for Articulated Objects Robotic Manipulation —
- HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos —
- AssemblyGrid v1: A Benchmark for Multi-Robot Production with Temporary Coalitions, Local Information, and Geometric Constraints —
- Auto-HSI: Personalized human control of a robot swarm on demand by using LLMs for online automatic code generation —
- UDAV: Uncertainty-Driven Adaptive VLM Waypoint Planner —
- The Neverwhere Visual Parkour Benchmark Suite —
- ProxiDex: Learning Dynamics-Guided Proximity Policy for Dexterous Manipulation —
- Large Language Models in the Loop: A Stability- and Network-Aware Survey in Networked Control, Cyber-Physical, and Multi-Agent Systems —
- Weave: Learning Whole-Body Dexterous Loco-Manipulation from Human-Object Interactions —
- World Models for Embodied Intelligence: From Plausible to Controllable to Actionable —
- Visual Cue Guided Video Planning for Generalizable Robot Navigation —
- Unified Heterogeneous Graph Neural Network solver for Power Flow, Optimal Power Flow and State Estimation —
- The Latent That Never Was: A Forensic Re-run of the CVAE Ablation in Action Chunking Transformers —
- TEMPO: Learning Temporal Context for Dynamic Robot Manipulation —
- Intrinsic Robot Rewarding: Reusing VLA Representations for Autonomous Evaluation and Policy Improvement —
- Continual Learning for Traversability Prediction with Uncertainty-Aware Adaptation —
- Kernel-Based Metrics Learning for Uncertain Opponent Vehicle Trajectory Prediction in Autonomous Racing —
- FluxVLA Engine: A One-Stop VLA Engineering Platform for Embodied Intelligence —
Important terms
- Motion Ambiguity
- This occurs when a single view of a scene doesn't clearly show how objects are moving over time, making it hard for robots to understand dynamic tasks.
- State Aliasing
- This problem arises when different visual views that look similar require the robot to take completely different actions, confusing its decision-making process.
- Temporal Context
- Using past and future data helps robots make better decisions by providing a sense of time. This temporal context is key for improving world models and embodied intelligence.
- ProxiDex
- This framework treats proximity as a learned interaction state, helping robots handle unstable hand-object interactions during dexterous manipulation.