Robotics papers — 2026-09-16

Today we are focusing on how robots can handle dynamic tasks where things move. This is crucial because current vision language action models struggle with this because they only look at one moment in time and cannot predict what will happen next. We are looking at motion ambiguity, where a single view doesn't show the movement of objects, and state aliasing, where similar views require different actions.

The TEMPO approach tackles this by adding two temporal inputs. It uses a motion summary from a video foundation model to fix the motion ambiguity. It also uses a compact history of proprioceptive data to resolve state aliasing. This method improved bottle handover success from forty-four percent up to seventy-four percent across four dynamic tasks, and it is unique in solving state aliasing.

This idea of using temporal context is also important when we consider how world models improve embodied intelligence. These models aim to connect perception and decision-making by anticipating consequences. The key question is which predictions actually help behavior, whether they are plausible, controllable, or actionable.

The work suggests that moving toward actionable models means ensuring predictions capture task-relevant state and reflect how interventions change that state. Another area of focus involves continual learning for robot navigation in uncertain terrains. This framework adapts to new surfaces without forgetting old ones by using a generative model to recall past experiences while incorporating uncertainty from those generated samples.

This is vital because robots need to adapt when they encounter unexpected ground conditions, like traction loss, which can cause instability. Finally, we are seeing how different approaches build upon existing vision language action systems. For instance, intrinsic robot rewarding proposes reusing the visual representations from a VLA system to evaluate the robot's own outcomes and guide policy improvement without needing a separate evaluator.

This reuse is promising because it lowers integration effort while connecting internal outcome evaluation directly to physical policy changes. The work that matters most is ProxiDex because it tackles the fundamental problem of unstable hand-object interactions. It treats proximity as a learned interaction state, which is crucial for dexterous manipulation in practice.

This framework reconstructs interaction point clouds and converts geometric distances into proximity cues. This creates a hardware-agnostic contact representation that gives immersive feedback during virtual teleoperation. This representation allows ProxiDex to learn action-conditioned proximity dynamics using a coupled forward-inverse design where future observations are predicted from actions and proximity variations are decoded from latent changes.

This dynamic understanding is then used to adaptively reweight proximity tokens across manipulation phases. Dynamics-consistency supervision guides policy inference to stabilize action generation even when visual feedback is unreliable. This method shows improved success rates and robustness compared to standard baselines in both simulation and real-world tests involving unseen objects or perturbation scenarios.

Weave builds upon this by creating a unified framework for learning whole-body dexterous humanoid-object interaction from human demonstrations. This is significant because it coordinates locomotion, balance, and hand contact. Weave first converts captured human interactions into executable robot-object references using contact-aware retargeting and approach motion completion.

At its core is a policy that jointly commands twenty-nine body joints and twelve finger joints across multiple objects. This system achieved a ninety-two point five percent success rate on trained interactions, and without further training, it still managed sixty-five point zero percent on sequences never seen during training.

TARC addresses the efficiency trade-off in robotic control by introducing a reinforcement learning framework. The policy predicts both a control action and its duration of application. TARC learns temporally extended actions by optimizing task performance under constraints on the number of control switches, allowing for adaptive modulation of control rates.

This approach matches the performance of high-frequency discrete-time controllers while operating at less than half their control frequency across different hardware platforms like an RC car and a quadruped robot. FluxVLA Engine aims to solve the engineering bottlenecks that separate promising embodied-learning algorithms from reliable robot systems.

It provides an open, configuration-driven platform for deployment by standardizing interfaces for datasets, visual-language models, and action heads. It integrates features like compositional dual-arm simulation and model-decoupled human-in-the-loop rollout. This engine connects offline learning to real robot execution through shared contracts and accelerated inference backends.

SafeFlow tackles the issue of physical hallucinations in text-driven humanoid motion generation by combining physics-guided motion generation with a three stage safety gate driven by explicit risk indicators. It uses Physics Guided Rectified Flow Matching in a VAE latent space to improve executability. The safety gate detects semantic out-of-distribution prompts and enforces hard kinematic constraints before the trajectory reaches the low level controller.

HumanEgo bridges the embodiment gap between human demonstrations and robot policies by lifting each human demonstration to an entity-level representation of hand-object interaction. This framework trains a flow matching policy with dense auxiliary objectives that amplify supervision from every trajectory. It achieves ninety-two point five percent average success across four real world tasks using only thirty minutes of human videos per task.

The most important thing we need to look at is how large language models can safely be put into control systems without breaking stability guarantees. The slow, unpredictable nature of these models clashes directly with the strict safety needs of networked control or cyber-physical systems. This work suggests that the LLM should only act as a slow supervisor setting high-level goals and constraints.

It should only act as a slow supervisor setting high-level goals and constraints, while a fast, certified inner loop handles the actual physical stability. This idea is important because it frames how we can manage risks when integrating these powerful models into critical infrastructure.

The integration maps directly to classical control problems where things like inference delay are treated like network delays. Model hallucinations are treated as bounded disturbances in the system. A related piece of work shows how uncertainty drives better path planning for unmanned ground vehicles using vision-language models.

This planner, UDAV, takes multiple uncertain route predictions from a vision-language model and uses the median prediction as its main route. It checks the spread of those predictions; if the uncertainty gets too high in any part of the path, it triggers a re-evaluation.

This uncertainty handling is really useful because when the model's confidence drops significantly, it gives an actionable signal to adjust its plan rather than blindly following a potentially flawed prediction. This approach successfully reduced the mean average displacement error from 147.4 pixels down to 115.9 pixels compared to a purely deterministic plan.

Another area of research focuses on improving how we test robot locomotion controllers in real-world settings by creating a new evaluation suite called The Neverwhere Benchmark Suite. This suite consists of over sixty three dimensional Gaussian Splatting reconstructions of various indoor and outdoor scenes. It aims to make it easier to create reproducible tests for visual locomotion policies.

This work is significant because it addresses the gap between training performance and real-world deployment by forcing policies to be tested in diverse, hyper-realistic environments. It also points out that relying only on data generated from Gaussian splats might be risky, emphasizing the need for varied data sources to ensure robust performance across different scenes.

Today's papers

The papers

Important terms

Motion Ambiguity
This occurs when a single view of a scene doesn't clearly show how objects are moving over time, making it hard for robots to understand dynamic tasks.
State Aliasing
This problem arises when different visual views that look similar require the robot to take completely different actions, confusing its decision-making process.
Temporal Context
Using past and future data helps robots make better decisions by providing a sense of time. This temporal context is key for improving world models and embodied intelligence.
ProxiDex
This framework treats proximity as a learned interaction state, helping robots handle unstable hand-object interactions during dexterous manipulation.