Robotics papers — 2026-09-18
MaskHarness-WAM links high-level planning with low-level movement policies using target masks that update as the scene changes. This harness continuously checks and verifies these masks at task boundaries, feeding the updated instance information to the low-level policy so it knows exactly which object it needs to focus on next. This approach substantially improves performance over simpler limited-horizon methods when performing sequential multi-object manipulation tasks on real robots.
We are also looking into how vision language models can be used for post-training robot policies without needing a physical robot constantly present, which is addressed by HIL-UMI. Another area of work involves making closed-loop robot software easier to learn and reuse. We found that using execution experience from one task to help acquire policies for new tasks significantly boosts success rates. This connects with how models can be used to check safety in complex systems, as LLM-Falsifier shows promise in finding counterexamples for formal specifications.
The most significant work here is Agile-WAM because it tackles the efficiency problem in tactile World Action Models by using a direct vision-tactile-to-action flow matching process to generate action chunks and future latents. This matters because it allows for precise, high-frequency control in contact-rich scenarios without relying on massive pretrained backbones that limit deployment flexibility.
This approach is built upon the observation that visual and tactile signals operate on different timescales. So, the system uses multi-horizon multimodal prediction to supervise visual latents over a longer time while predicting tactile latents for fine contact dynamics in the next frame. This is more specific than just looking at one modality; it is about intelligently combining what we see with what we feel.
The results show that this agile architecture performs strongly across nine simulated and five real-world tasks. It achieves a relative gain of twenty-nine point four percent in overall success rates while keeping inference latency low at eleven point nine milliseconds in five real-world experiments. This means the model is both accurate and fast enough for practical robot control.
This success contrasts with other methods; for instance, TacSushi showed that using future-consequence supervision on training data yields a thirty-seven point five percent out-of-distribution success rate compared to twenty-five point zero percent when only using direct tactile concatenation. This suggests that training the model to predict future consequences is a valuable way to improve its ability to handle novel situations.
The EmbodiedMind system’s work on efficient training paradigms matters because it tackles the fundamental bottlenecks in building large embodied foundation models. Specifically, it addresses how to use limited data effectively and how to handle the complex credit assignment problem in long-horizon planning. This is crucial for making these models practical rather than just impressive demonstrations.
The most significant contribution is Trie-GRPO, a novel reinforcement learning algorithm based on action prefix trees that solves the credit assignment issue by estimating step-level advantages. This allows for better exploration efficiency and depth than traditional search trees, resulting in EmbodiedMind achieving a state-of-the-art average performance of seventy point zero two percent across eighteen benchmarks. This significantly outperforms other embodied foundation models in long-horizon task planning accuracy.
This work is built upon the preceding stage where Rejection Sampling-based Fine-Tuning filters out low informative samples to establish robust behavioral priors while avoiding distributional collapse. Following this filtering, Iterative Rejection GRPO balances datasets across reinforcement learning iterations using task-specific queues stratified by difficulty, coupled with a hybrid reward mechanism for precise cross-task feedback. This balanced training feeds into the Trie-GRPO stage, which then refines the policy using action prefix trees to manage long sequences of actions. This entire pipeline demonstrates how strategic data selection and hierarchical policy optimization can lead to superior performance in complex embodied tasks.
The most significant work from yesterday is GAVEL because it tackles the fundamental problem of making long-horizon planning reliable when using large language models. This is crucial for any complex robot task. GAVEL introduces an explicit graph world model to verify and repair plans generated by LLMs, meaning it checks the consequences of actions before they happen and only lets the LLM replan when deep semantic reasoning is actually needed. This approach shows that harnessing an explicit graph world model can substantially improve the reliability and efficiency of long-horizon embodied planning across compact and frontier hosted LLM capabilities.
This idea builds upon earlier work in agent-centered architectures like Teach and Grow, which aimed to make robot learning more efficient by turning successful demonstrations into reusable skills without needing gradient updates. TGL achieved very high success rates on various LIBERO suites, suggesting that grounding learned behaviors in explicit skill blocks is a powerful way to build persistent knowledge. This contrasts with GAVEL's focus on planning verification during execution, showing two different paths to improving agent capability: one through learning reusable skills and the other through runtime plan correction.
Another area of progress involves making robot control more socially aware, as seen in Learn2Drive, where a neural network-based framework incorporates social value orientation to make autonomous vehicles adapt their driving profiles based on traffic flow efficiency rather than just individual vehicle performance. This contrasts with the planning focus of GAVEL by looking at real-time decision-making in dynamic environments.
Furthermore, research into manipulation has shown that combining simulation and human data, as in SimHum, offers a way to get data-efficient learning by extracting kinematic priors from simulation and visual priors from human observations. This complements the planning verification work of GAVEL by providing richer scene understanding for the LLM's world model.
Finally, coding agents are being made safer through SafeHarness, which equips language models with obstacle-aware harnesses to prioritize safety constraints during route planning and execution. This directly addresses a major flaw in current coding agent paradigms where safety is often neglected in favor of task completion, demonstrating how explicit constraint grounding can improve performance metrics significantly.
The work on action similarity supervision is particularly important because it addresses the core challenge of making latent action models usable across different robot bodies. This method trains the similarity between two latent actions to match the similarity of their corresponding ground-truth robot actions, meaning the latent space learns to respect how similar those actual movements are. This approach outperforms using an auxiliary loss that tries to predict the ground-truth action during training, suggesting a more robust way to align representations between different physical embodiments.
This technique is significant because it allows for better cross-embodiment transfer when policies are trained on demonstrations from one robot and then tested on another, as seen in the RoboTwin 2.0 evaluation where predicting latent actions more than doubled cross-embodiment success compared to predicting ground-truth actions. This success stems from computing similarities based on end-effector motion rather than joint space motion, which seems to be the best way to compare movements across different robots.
Another key finding is that reliable scenario generation in air-ground co-simulation requires verifying realized behavior, not just executable code, a concept highlighted by the AURORA framework. This framework uses an Air-Ground Scenario Graph to explicitly map out all dependencies between agents and events. This allows for runtime verification to expose silent failures that simple execution checks miss.
Post-training fine-tuning methods show significant gains in driving performance; OPTED decouples reinforcement learning from policy post-training by using a privileged teacher trained on vectorized inputs to supervise the student during closed-loop deployment. This resulted in driving scores increasing by factors of 1.6 times and 9.5 times for the TransFuser and VaVAM models, while requiring far fewer simulator interactions than direct reinforcement learning post-training.
Finally, when considering deployment constraints like quantization for world action models, the PreDE framework offers a policy-calibrated way to predict task degradation from offline action deviations before costly closed-loop evaluations. This system successfully issues decisions on new quantization configurations based on predicted performance, identifying candidates that require testing while avoiding configurations likely to fail under deployment conditions.
Today's papers
- MaskHarness-WAM: Instance-Grounded Harnessing for Long-Horizon Robot Manipulation Long-horizon robot manipulation requires connecting high-level planning with low-level control using visual feedback to schedule subtasks. [paper]
- Worst-Case Hidden-Vehicle Trajectory Search in Spatiotemporal Occlusion Regions This method searches for the worst possible hidden vehicle trajectory by combining temporal occlusion reasoning with response-aware search. [paper]
- HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface This framework allows human operators to post-train vision language action models without requiring repeated physical robot execution. [paper]
- Learning and Transferring Closed-Loop Robot Software This study shows that reusing code from source tasks can significantly improve the success rate when acquiring policies for new tasks. [paper]
- Large Language Models as Falsifiers for Cyber-Physical Systems This work introduces an LLM approach to falsify formal specifications by minimizing robustness using semantic information naturally present in language. [paper]
- REACT: A Fully Spiking State-Space Model for Real-Time Event-Driven Temporal Perception This model uses a spiking state-space approach to process raw event camera streams with microsecond temporal resolution for fast perception. [paper]
- VLN on the Fly: An Onboard Vision-Language Navigation Stack for Aerial Robots This stack fuses vision, language, planning, and control into one network for robust vision-language navigation on aerial robots. [paper]
- MILER: Semantic Mid-Level Representation for Sim-to-Real Reinforcement Learning in Unstructured Autonomous Driving This framework uses a semantic mid-level representation to enable zero-shot sim-to-real transfer for reinforcement learning in unstructured driving. [paper]
- Agile Tactile World Action Model for Contact-Rich Robot Control This model uses vision and tactile data fusion to predict future world states and robot actions with an agile architecture suitable for precise control.
- Drag-Aware Aerodynamic Manipulability for Torque-Limited Redundant Multirotors This paper develops a metric based on rotor acceleration capacity to quantify aerodynamic manipulability in redundant multirotor systems.
- GLAMDRING: Gait Learning And Morphology co-Design via Reinforcement Learning of CPGs This framework jointly learns the optimal robot body morphology and gait policy for locomotion tasks. [paper]
- Demystifying Linear Operator Learning for Control Systems This paper proposes a structured approach using group theory to learn linear operators for control systems from data. [paper]
- GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies This method adaptively adjusts the action horizon based on the geometric reliability of flow-based trajectories to improve VLA policy learning. [paper]
- A Simulation Platform for AUV Fault Recovery: Exploring LLM-Based Diagnostic Strategies This architecture uses a closed-loop simulation platform with an LLM to diagnose and recover from failures in autonomous underwater vehicles. [paper]
- TacSushi: Tactile-Grounded World-Action Modeling for Dexterous Sushi Manipulation This model learns world actions by predicting future visual observations and tactile summaries during training on real robot trials. [paper]
- From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation This harness automates scene resetting and scoring for long-horizon manipulation evaluation using a graph world model. [paper]
- EmbodiedMind: Adaptive Data Curation and Prefix-Tree Reinforcement Learning for Efficient Embodied Intelligence This approach uses data filtering and prefix trees to efficiently train embodied foundation models for long-horizon task planning. [paper]
- Kinematics-Grounded Agentic AI for Robotic Additive Manufacturing Process Planning This framework uses an agentic tool to plan robot additive manufacturing processes by integrating manufacturing constraints with kinematic realization. [paper]
- LLM-Guided Transformation of Non-Critical Driving Scenes into Safety-Critical Scenarios Using Augmented Reality This system uses LLMs and AR to transform safe driving scenes into realistic safety-critical scenarios for autonomous vehicle testing. [paper]
- UniExo: Unified Multi-Skill Policies for Musculoskeletal Locomotion and Co-Adaptive Exoskeleton Control This framework trains a single controller to perform diverse locomotor activities by co-adapting with a multi-skill human policy. [paper]
- PhysCoRe: Physics-Corrected Residual World Models for Material-Aware Deformable Dynamics This model uses a physics simulator coupled with neural networks to predict the evolution of deformable objects under manipulation. [paper]
- Workspace Models: Lightweight Robotic Memory via Saliency-Driven Supervision This method learns a lightweight latent memory representation by using VLM queries during training to store task-salient information. [paper]
- Accelerating Visual Policy Learning with Sampling-Based Model Predictive Control This technique couples first-order policy gradients with sampling-based model predictive control to improve visual policy learning efficiency. [paper]
- CoreSense: Traceable Failure Recall and Conflict-Aware Belief Gating for Auditable Robot Decisions This architecture uses belief gating to manage robot decision recall based on evidence scope, provenance, and conflict awareness. [paper]
- GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning This framework uses an explicit graph world model to verify and repair long-horizon plans generated by large language models. [paper]
- Teach and Grow: An Agent-Centered Architecture for General Robot Learning This training-free architecture turns successful demonstrations into reusable skills by using an agent to identify skill blocks. [paper]
- Learn2Drive: A neural network-based framework for socially compliant automated vehicle control This study proposes a control framework that uses social value orientation to make autonomous vehicles behave more efficiently in traffic. [paper]
- AntiGrounding: Executable Robot Trajectories as Visual Prompts for VLM-Guided Manipulation This method uses trajectory evaluation scores from a digital twin to guide vision language models in selecting executable manipulation trajectories. [paper]
- Sim-and-Human Co-training for Data-Efficient and Scene-Generalizable Bimanual Manipulation This recipe extracts kinematic priors from simulation and visual priors from human data to improve scene generalization in bimanual manipulation. [paper]
- Tackling Snow-Induced Challenges: Safe Autonomous Lane-Keeping with Robust Reinforcement Learning This paper proposes action-robust deep reinforcement learning algorithms to maintain lane keeping in snowy road conditions. [paper]
- Coding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation This framework adds obstacle-aware harnesses to coding agents to ensure safety constraints are prioritized during robot manipulation planning.
- StarVLA- alpha: Reducing Complexity in Vision-Language-Action Systems This paper introduces a simple baseline VLA model designed to study the impact of minimal architectural complexity on performance. [paper]
- Improving Cross-embodiment Transfer in Latent Action Models with Action-Similarity Supervision This method uses action similarity supervision to improve cross-embodiment transfer in latent action models by aligning latent actions across different robots. [paper]
- AURORA: A Natural Language-Driven Agentic Framework for Understanding, Reasoning, and Orchestrating Reliable Air-Ground Co-Simulation This framework treats air-ground scenario generation as a compilation process verified by an explicit scenario graph. [paper]
- OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher This method uses an on-policy fine-tuning approach with a privileged teacher to improve end-to-end driving policies with fewer simulator interactions. [paper]
- Predict Before You Deploy: Offline Prediction of Quantization-Induced Task Degradation for World Action Models This framework predicts task degradation caused by quantization configurations before closed-loop testing is performed. [paper]
The papers
- AntiGrounding: Executable Robot Trajectories as Visual Prompts for VLM-Guided Manipulation —
- Learn2Drive: A neural network-based framework for socially compliant automated vehicle control —
- Tackling Snow-Induced Challenges: Safe Autonomous Lane-Keeping with Robust Reinforcement Learning —
- Sim-and-Human Co-training for Data-Efficient and Scene-Generalizable Bimanual Manipulation —
- Drag-Aware Aerodynamic Manipulability for Torque-Limited Redundant Multirotors: Aerodynamic Promptness based on the Symmetric Acceleration Capacity —
- StarVLA- alpha: Reducing Complexity in Vision-Language-Action Systems —
- PhysCoRe: Physics-Corrected Residual World Models for Material-Aware Deformable Dynamics —
- Teach and Grow: An Agent-Centered Architecture for General Robot Learning —
- REACT: A Fully Spiking State-Space Model for Real-Time Event-Driven Temporal Perception —
- GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning —
- Kinematics-Grounded Agentic AI for Robotic Additive Manufacturing Process Planning —
- From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation —
- Demystifying Linear Operator Learning for Control Systems —
- Predict Before You Deploy: Offline Prediction of Quantization-Induced Task Degradation for World Action Models —
- GLAMDRING: Gait Learning And Morphology co-Design via Reinforcement LearnING of CPGs —
- CoreSense: Traceable Failure Recall and Conflict-Aware Belief Gating for Auditable Robot Decisions —
- AURORA: A Natural Language-Driven Agentic Framework for Understanding, Reasoning, and Orchestrating Reliable Air-Ground Co-Simulation —
- TacSushi: Tactile-Grounded World-Action Modeling for Dexterous Sushi Manipulation —
- EmbodiedMind: Adaptive Data Curation and Prefix-Tree Reinforcement Learning for Efficient Embodied Intelligence —
- UniExo: Unified Multi-Skill Policies for Musculoskeletal Locomotion and Co-Adaptive Exoskeleton Control —
- Improving Cross-embodiment Transfer in Latent Action Models with Action-Similarity Supervision —
- Learning and Transferring Closed-Loop Robot Software —
- MaskHarness-WAM: Instance-Grounded Harnessing for Long-Horizon Robot Manipulation —
- VLN on the Fly: An Onboard Vision-Language Navigation Stack for Aerial Robots —
- LLM-Guided Transformation of Non-Critical Driving Scenes into Safety-Critical Scenarios Using Augmented Reality —
- Worst-Case Hidden-Vehicle Trajectory Search in Spatiotemporal Occlusion Regions —
- Accelerating Visual Policy Learning with Sampling-Based Model Predictive Control —
- A Simulation Platform for AUV Fault Recovery: Exploring LLM-Based Diagnostic Strategies —
- HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface —
- MILER: Semantic Mid-Level Representation for Sim-to-Real Reinforcement Learning in Unstructured Autonomous Driving —
- Large Language Models as Falsifiers for Cyber-Physical Systems —
- OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher —
- Agile-WAM: An Agile Tactile World Action Model for Contact-Rich Robot Control —
- GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies —
- Workspace Models: Lightweight Robotic Memory via Saliency-Driven Supervision —
- Coding Agents with Harness for Safe Robot Control —
Important terms
- Agile-WAM
- This method improves tactile World Action Models by using a direct vision-tactile-to-action flow matching process to generate action chunks and future latents. It enables precise, high-frequency control in contact scenarios without needing massive pretraining.
- Trie-GRPO
- A novel reinforcement learning algorithm based on action prefix trees that solves the credit assignment problem by estimating step-level advantages. This leads to better exploration efficiency and depth for long-horizon planning.
- GAVEL
- Introduces an explicit graph world model to verify and repair plans generated by large language models. It checks action consequences before execution, allowing LLMs to replan only when deep reasoning is necessary.
- Action Similarity Supervision
- This technique trains the similarity between two latent actions to match the similarity of their real movements. This helps latent spaces respect how similar actual robot motions are across different physical bodies.