Robotics papers — 2026-09-18

MaskHarness-WAM links high-level planning with low-level movement policies using target masks that update as the scene changes. This harness continuously checks and verifies these masks at task boundaries, feeding the updated instance information to the low-level policy so it knows exactly which object it needs to focus on next. This approach substantially improves performance over simpler limited-horizon methods when performing sequential multi-object manipulation tasks on real robots.

We are also looking into how vision language models can be used for post-training robot policies without needing a physical robot constantly present, which is addressed by HIL-UMI. Another area of work involves making closed-loop robot software easier to learn and reuse. We found that using execution experience from one task to help acquire policies for new tasks significantly boosts success rates. This connects with how models can be used to check safety in complex systems, as LLM-Falsifier shows promise in finding counterexamples for formal specifications.

The most significant work here is Agile-WAM because it tackles the efficiency problem in tactile World Action Models by using a direct vision-tactile-to-action flow matching process to generate action chunks and future latents. This matters because it allows for precise, high-frequency control in contact-rich scenarios without relying on massive pretrained backbones that limit deployment flexibility.

This approach is built upon the observation that visual and tactile signals operate on different timescales. So, the system uses multi-horizon multimodal prediction to supervise visual latents over a longer time while predicting tactile latents for fine contact dynamics in the next frame. This is more specific than just looking at one modality; it is about intelligently combining what we see with what we feel.

The results show that this agile architecture performs strongly across nine simulated and five real-world tasks. It achieves a relative gain of twenty-nine point four percent in overall success rates while keeping inference latency low at eleven point nine milliseconds in five real-world experiments. This means the model is both accurate and fast enough for practical robot control.

This success contrasts with other methods; for instance, TacSushi showed that using future-consequence supervision on training data yields a thirty-seven point five percent out-of-distribution success rate compared to twenty-five point zero percent when only using direct tactile concatenation. This suggests that training the model to predict future consequences is a valuable way to improve its ability to handle novel situations.

The EmbodiedMind system’s work on efficient training paradigms matters because it tackles the fundamental bottlenecks in building large embodied foundation models. Specifically, it addresses how to use limited data effectively and how to handle the complex credit assignment problem in long-horizon planning. This is crucial for making these models practical rather than just impressive demonstrations.

The most significant contribution is Trie-GRPO, a novel reinforcement learning algorithm based on action prefix trees that solves the credit assignment issue by estimating step-level advantages. This allows for better exploration efficiency and depth than traditional search trees, resulting in EmbodiedMind achieving a state-of-the-art average performance of seventy point zero two percent across eighteen benchmarks. This significantly outperforms other embodied foundation models in long-horizon task planning accuracy.

This work is built upon the preceding stage where Rejection Sampling-based Fine-Tuning filters out low informative samples to establish robust behavioral priors while avoiding distributional collapse. Following this filtering, Iterative Rejection GRPO balances datasets across reinforcement learning iterations using task-specific queues stratified by difficulty, coupled with a hybrid reward mechanism for precise cross-task feedback. This balanced training feeds into the Trie-GRPO stage, which then refines the policy using action prefix trees to manage long sequences of actions. This entire pipeline demonstrates how strategic data selection and hierarchical policy optimization can lead to superior performance in complex embodied tasks.

The most significant work from yesterday is GAVEL because it tackles the fundamental problem of making long-horizon planning reliable when using large language models. This is crucial for any complex robot task. GAVEL introduces an explicit graph world model to verify and repair plans generated by LLMs, meaning it checks the consequences of actions before they happen and only lets the LLM replan when deep semantic reasoning is actually needed. This approach shows that harnessing an explicit graph world model can substantially improve the reliability and efficiency of long-horizon embodied planning across compact and frontier hosted LLM capabilities.

This idea builds upon earlier work in agent-centered architectures like Teach and Grow, which aimed to make robot learning more efficient by turning successful demonstrations into reusable skills without needing gradient updates. TGL achieved very high success rates on various LIBERO suites, suggesting that grounding learned behaviors in explicit skill blocks is a powerful way to build persistent knowledge. This contrasts with GAVEL's focus on planning verification during execution, showing two different paths to improving agent capability: one through learning reusable skills and the other through runtime plan correction.

Another area of progress involves making robot control more socially aware, as seen in Learn2Drive, where a neural network-based framework incorporates social value orientation to make autonomous vehicles adapt their driving profiles based on traffic flow efficiency rather than just individual vehicle performance. This contrasts with the planning focus of GAVEL by looking at real-time decision-making in dynamic environments.

Furthermore, research into manipulation has shown that combining simulation and human data, as in SimHum, offers a way to get data-efficient learning by extracting kinematic priors from simulation and visual priors from human observations. This complements the planning verification work of GAVEL by providing richer scene understanding for the LLM's world model.

Finally, coding agents are being made safer through SafeHarness, which equips language models with obstacle-aware harnesses to prioritize safety constraints during route planning and execution. This directly addresses a major flaw in current coding agent paradigms where safety is often neglected in favor of task completion, demonstrating how explicit constraint grounding can improve performance metrics significantly.

The work on action similarity supervision is particularly important because it addresses the core challenge of making latent action models usable across different robot bodies. This method trains the similarity between two latent actions to match the similarity of their corresponding ground-truth robot actions, meaning the latent space learns to respect how similar those actual movements are. This approach outperforms using an auxiliary loss that tries to predict the ground-truth action during training, suggesting a more robust way to align representations between different physical embodiments.

This technique is significant because it allows for better cross-embodiment transfer when policies are trained on demonstrations from one robot and then tested on another, as seen in the RoboTwin 2.0 evaluation where predicting latent actions more than doubled cross-embodiment success compared to predicting ground-truth actions. This success stems from computing similarities based on end-effector motion rather than joint space motion, which seems to be the best way to compare movements across different robots.

Another key finding is that reliable scenario generation in air-ground co-simulation requires verifying realized behavior, not just executable code, a concept highlighted by the AURORA framework. This framework uses an Air-Ground Scenario Graph to explicitly map out all dependencies between agents and events. This allows for runtime verification to expose silent failures that simple execution checks miss.

Post-training fine-tuning methods show significant gains in driving performance; OPTED decouples reinforcement learning from policy post-training by using a privileged teacher trained on vectorized inputs to supervise the student during closed-loop deployment. This resulted in driving scores increasing by factors of 1.6 times and 9.5 times for the TransFuser and VaVAM models, while requiring far fewer simulator interactions than direct reinforcement learning post-training.

Finally, when considering deployment constraints like quantization for world action models, the PreDE framework offers a policy-calibrated way to predict task degradation from offline action deviations before costly closed-loop evaluations. This system successfully issues decisions on new quantization configurations based on predicted performance, identifying candidates that require testing while avoiding configurations likely to fail under deployment conditions.

Today's papers

The papers

Important terms

Agile-WAM
This method improves tactile World Action Models by using a direct vision-tactile-to-action flow matching process to generate action chunks and future latents. It enables precise, high-frequency control in contact scenarios without needing massive pretraining.
Trie-GRPO
A novel reinforcement learning algorithm based on action prefix trees that solves the credit assignment problem by estimating step-level advantages. This leads to better exploration efficiency and depth for long-horizon planning.
GAVEL
Introduces an explicit graph world model to verify and repair plans generated by large language models. It checks action consequences before execution, allowing LLMs to replan only when deep reasoning is necessary.
Action Similarity Supervision
This technique trains the similarity between two latent actions to match the similarity of their real movements. This helps latent spaces respect how similar actual robot motions are across different physical bodies.