Robotics papers — 2026-09-11

Today we are diving into how we can make robots do things dynamically without needing massive amounts of real-world training data. This is crucial because a single error in a dynamic throw can ruin an entire operation. The Wiggle and Go! framework tackles this by observing a brief, safe wiggle action to predict rope parameters. It then uses those predictions to guide the trajectory optimizer for the actual goal-conditioned movement.

This identification module is designed to be task-agnostic, meaning it supports different manipulation policies without needing retraining. This parameter prediction is key because it allows us to transfer those predicted rope dynamics—with a Pearson correlation of 0.95 between simulation and reality—to unseen motions. This means the identification module generalizes well across different tasks.

This predictive capability then conditions the trajectory optimizer for zero-shot execution, allowing the system to perform well on multi-objective tasks like lobbing and draping with over fifty percent success. Show-Harness addresses a different challenge by showing how foundation vision-language models can directly control robots through a compact semantic interface. It works by exposing discrete semantic action units that the VLM can reason about.

These units are then grounded into specific robot actions by an embodiment-specific interpreter, keeping the VLM responsible for fine physical decisions. This setup allows for zero-shot control of closed-source frontier VLMs and even enables adapting smaller open-source models with minimal fine-tuning. Meanwhile, we are also looking at how to make these agents more robust when they interact with the physical world by introducing ReactHuman.

ReactHuman is a benchmark designed to test if multimodal large language models can turn physical understanding into immediate, safe action in response to sudden hazards. While these models show promise in general tasks, our results indicate that reactive safety is still far from solved. Many models mishandle hazards or trust appearance over actual motion.

Finally, we are exploring how to structure the context fed into large language models for engineering design by introducing a framework of formal operations for assembling modular context units like policy prompts and reference units. This systematic structuring helps us evaluate how well these LLMs support systems architecture modeling by assessing their compliance to the intended design intent. The work on formal verification for automated driving is most significant because it directly tackles the fundamental gap between simulation success and real-world failure, which is critical for safety.

We trained two end-to-end steering networks in CARLA, one under clear conditions and another under adverse weather like fog or night. Using bound propagation, a formal method that reads the trained weights, we found conditions that broke the clear model without needing further simulation testing. This calculation covered a massive scope on the arterial road, spanning 133 poses where ten intensities each would be 10 to 133 combinations in minutes on one GPU.

This finding suggests that formal verification is a viable partner to simulation for verifying automated driving systems. This complements the work on muscle-driven locomotion, which uses a reflex-informed framework to create physically plausible human movement by modulating reflex gains based on the current state. This approach improves kinematic accuracy and symmetry under nominal walking conditions while remaining robust to muscle weakness without retraining.

Similarly, HiRAD addresses the routing challenges for large fleets of autonomous guided vehicles by proposing a hierarchical reinforcement learning framework for continuous-space routing. This method uses a step-level spatiotemporal representation and an asynchronous event-driven pipeline to reduce inference complexity from O(n squared) down to O(n). This cuts per-step latency by as much as seventy one percent, which in turn reduces makespan by forty five percent on two warehouse maps.

Finally, the CT-SAFR framework offers a multi-layered verification method for autonomous robots that uses chain-of-thought prompting to detect unsafe reasoning outputs. This system achieved ninety four point two percent hallucination detection with sub five hundred milliseconds of latency. It demonstrated an eighty seven percent reduction in unsafe reasoning outputs in a warehouse robot case study.

Today's papers

The papers

Important terms

Wiggle and Go! framework
A method that observes a brief, safe wiggle action to predict rope parameters. This prediction then guides the trajectory optimizer for actual goal-conditioned movement, allowing robots to move dynamically without extensive real-world training data.
Show-Harness
Uses foundation vision-language models to control robots via a compact semantic interface. It exposes discrete semantic action units that the VLM reasons about, which are then grounded into physical actions by an interpreter.
ReactHuman
A benchmark testing if multimodal large language models can turn physical understanding into immediate, safe action when facing sudden hazards. It highlights current limitations in reactive safety for these models.
Formal verification
A formal method that reads trained model weights to check conditions that break a system without further simulation. This is proposed as a viable partner to simulation for verifying automated driving systems.
HiRAD
A hierarchical reinforcement learning framework for routing autonomous guided vehicles. It uses an event-driven pipeline to reduce inference complexity significantly, cutting latency and improving routing efficiency.