Robotics papers — 2026-09-11
Today we are diving into how we can make robots do things dynamically without needing massive amounts of real-world training data. This is crucial because a single error in a dynamic throw can ruin an entire operation. The Wiggle and Go! framework tackles this by observing a brief, safe wiggle action to predict rope parameters. It then uses those predictions to guide the trajectory optimizer for the actual goal-conditioned movement.
This identification module is designed to be task-agnostic, meaning it supports different manipulation policies without needing retraining. This parameter prediction is key because it allows us to transfer those predicted rope dynamics—with a Pearson correlation of 0.95 between simulation and reality—to unseen motions. This means the identification module generalizes well across different tasks.
This predictive capability then conditions the trajectory optimizer for zero-shot execution, allowing the system to perform well on multi-objective tasks like lobbing and draping with over fifty percent success. Show-Harness addresses a different challenge by showing how foundation vision-language models can directly control robots through a compact semantic interface. It works by exposing discrete semantic action units that the VLM can reason about.
These units are then grounded into specific robot actions by an embodiment-specific interpreter, keeping the VLM responsible for fine physical decisions. This setup allows for zero-shot control of closed-source frontier VLMs and even enables adapting smaller open-source models with minimal fine-tuning. Meanwhile, we are also looking at how to make these agents more robust when they interact with the physical world by introducing ReactHuman.
ReactHuman is a benchmark designed to test if multimodal large language models can turn physical understanding into immediate, safe action in response to sudden hazards. While these models show promise in general tasks, our results indicate that reactive safety is still far from solved. Many models mishandle hazards or trust appearance over actual motion.
Finally, we are exploring how to structure the context fed into large language models for engineering design by introducing a framework of formal operations for assembling modular context units like policy prompts and reference units. This systematic structuring helps us evaluate how well these LLMs support systems architecture modeling by assessing their compliance to the intended design intent. The work on formal verification for automated driving is most significant because it directly tackles the fundamental gap between simulation success and real-world failure, which is critical for safety.
We trained two end-to-end steering networks in CARLA, one under clear conditions and another under adverse weather like fog or night. Using bound propagation, a formal method that reads the trained weights, we found conditions that broke the clear model without needing further simulation testing. This calculation covered a massive scope on the arterial road, spanning 133 poses where ten intensities each would be 10 to 133 combinations in minutes on one GPU.
This finding suggests that formal verification is a viable partner to simulation for verifying automated driving systems. This complements the work on muscle-driven locomotion, which uses a reflex-informed framework to create physically plausible human movement by modulating reflex gains based on the current state. This approach improves kinematic accuracy and symmetry under nominal walking conditions while remaining robust to muscle weakness without retraining.
Similarly, HiRAD addresses the routing challenges for large fleets of autonomous guided vehicles by proposing a hierarchical reinforcement learning framework for continuous-space routing. This method uses a step-level spatiotemporal representation and an asynchronous event-driven pipeline to reduce inference complexity from O(n squared) down to O(n). This cuts per-step latency by as much as seventy one percent, which in turn reduces makespan by forty five percent on two warehouse maps.
Finally, the CT-SAFR framework offers a multi-layered verification method for autonomous robots that uses chain-of-thought prompting to detect unsafe reasoning outputs. This system achieved ninety four point two percent hallucination detection with sub five hundred milliseconds of latency. It demonstrated an eighty seven percent reduction in unsafe reasoning outputs in a warehouse robot case study.
Today's papers
- Wiggle and Go! System Identification for Zero-Shot Dynamic Rope Manipulation: This framework uses a brief wiggle to predict rope parameters to enable zero-shot manipulation without retraining. [paper] [episode]
- Show-Harness: Just a VLM Agent Can Play Robots: Show-Harness lets foundation vision-language models control robots by linking intent to action through a compact semantic interface. [paper]
- ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs: ReactHuman is a benchmark testing how multimodal models react safely and physically grounded to sudden hazards in simulated environments. [paper]
- Context operations to architecture modelling output from large language models and evaluation criteria for their use in systems engineering design: This paper introduces a framework for structuring context when using large language models for engineering design and evaluating their outputs. [paper]
- 2AM: Grounding Agent-Side Memory as Guidance for Steerable Action Models in Long-Horizon Manipulation: 2AM keeps task memory on the agent side to steer action models effectively during long-horizon manipulation tasks. [paper]
- ActSafeGuard: Differentiable and Training-Aligned Constraint Enforcement for Flow-Matching Policies: ActSafeGuard adds a safety layer to flow-matching policies that enforces physical constraints during training, ensuring safe robot actions. [paper]
- Compact Visuotactile World Models for Lifting: This study develops a world model that uses vision and touch to predict force constraints for accurate lifting in robotic manipulation.
- ObstaDiff: Generalizable Diffusion Policy Learning via Obstacle-aware Representations: ObstaDiff uses obstacle-aware representations to help diffusion policies generate successful trajectories in cluttered, real-world scenes. [paper]
- Testing Between the Test Cases: Proving End-to-End Steering in Conditions You Never Drove: This work uses formal verification to test how well automated vehicle steering policies generalize across different driving conditions beyond simulation. [paper]
- Reflex-Informed Neuromuscular Reinforcement Learning for Muscle-Driven Locomotion: This framework combines a fixed reflex controller with reinforcement learning to create physically plausible and robust muscle-driven locomotion. [paper]
- HiRAD: A Flexible Large-Scale AGV Routing System: HiRAD proposes a hierarchical reinforcement learning system to solve complex, real-time routing problems for large fleets of autonomous guided vehicles. [paper]
- CT-SAFR: Safe and Interpretable Chain-of-Thought Reasoning for Autonomous Robots: CT-SAFR is a verification framework that checks the safety and faithfulness of reasoning outputs from large language models in robotics.
- HuRo: Robotizing Human Videos for Scalable VLA Pretraining: HuRo creates a dataset by robotizing human videos to provide scalable supervision for training vision-language-action policies. [paper]
- Multi-Agent Reinforcement Learning for Autonomous UAV Exploration in Wildfire Response: This study uses deep reinforcement learning to train unmanned aerial vehicles to effectively navigate and monitor simulated wildfire environments. [paper]
The papers
- Wiggle and Go! System Identification for Zero-Shot Dynamic Rope Manipulation — This paper introduces a novel framework for performing system identification of dynamic rope manipulation in a zero-shot manner. [episode]
- Compact Visuotactile World Models for Lifting: Prediction, Reward Alignment, and Force Constraints —
- CT-SAFR: Safe and Interpretable Chain-of-Thought Reasoning for Autonomous Robots: A Multi-Layered Verification Framework for Trustworthy AI-Driven Robotic Decision Making —
- HiRAD: A Flexible Large-Scale AGV Routing System —
- Context operations to architecture modelling output from large language models and evaluation criteria for their use in systems engineering design —
- Multi-Agent Reinforcement Learning for Autonomous UAV Exploration in Wildfire Response —
- Show-Harness: Just a VLM Agent Can Play Robots —
- HuRo: Robotizing Human Videos for Scalable VLA Pretraining —
- ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs —
- ObstaDiff: Generalizable Diffusion Policy Learning via Obstacle-aware Representations —
- Testing Between the Test Cases: Proving End-to-End Steering in Conditions You Never Drove —
- 2AM: Grounding Agent-Side Memory as Guidance for Steerable Action Models in Long-Horizon Manipulation —
- ActSafeGuard: Differentiable and Training-Aligned Constraint Enforcement for Flow-Matching Policies —
- Reflex-Informed Neuromuscular Reinforcement Learning for Muscle-Driven Locomotion —
Important terms
- Wiggle and Go! framework
- A method that observes a brief, safe wiggle action to predict rope parameters. This prediction then guides the trajectory optimizer for actual goal-conditioned movement, allowing robots to move dynamically without extensive real-world training data.
- Show-Harness
- Uses foundation vision-language models to control robots via a compact semantic interface. It exposes discrete semantic action units that the VLM reasons about, which are then grounded into physical actions by an interpreter.
- ReactHuman
- A benchmark testing if multimodal large language models can turn physical understanding into immediate, safe action when facing sudden hazards. It highlights current limitations in reactive safety for these models.
- Formal verification
- A formal method that reads trained model weights to check conditions that break a system without further simulation. This is proposed as a viable partner to simulation for verifying automated driving systems.
- HiRAD
- A hierarchical reinforcement learning framework for routing autonomous guided vehicles. It uses an event-driven pipeline to reduce inference complexity significantly, cutting latency and improving routing efficiency.