Robotics papers — 2026-10-08
Today's work centers on figuring out how to make vision language action models better at understanding and responding to human instructions, which is crucial because if these models can truly interpret complex commands, they open the door for robots to perform more nuanced tasks in real environments. NovaPlan attempts zero-shot long-horizon manipulation using closed-loop video language planning, meaning the system tries to plan a sequence of actions based on a natural language command without needing specific training for every single scenario.
We also explored CIRRA, which focuses on dual-level continual instruction reconciliation with ongoing execution for embodied robot agents in interactive household tasks. This work tackles the problem of robots forgetting previous instructions while trying to complete multi-step chores like cleaning or cooking, and this is related to how we build robust systems. ED3R uses energy-aware distributed disaster detection via cooperative agents in robotic systems, which is about making sure a team of robots can detect emergencies efficiently while managing their power consumption.
UniCross addresses unified cross-skill dexterous manipulation synthesis, which is about combining different skills so a robot can perform complex actions that require multiple abilities at once. Furthermore, TAVIS provides a benchmark for egocentric active vision and anticipatory gaze in imitation learning, helping us measure how well robots look ahead and decide where to focus their attention before acting. Finally, Agentic Scene Policies suggests a framework for scene policies that allows agents to make decisions based on the context they perceive.
The most crucial development concerns the creation of agentic policies grounded in scene reconstruction, which addresses how robots can operate reliably when the environment changes. Agentic RSR attempts to bridge the gap between simulation and reality by using scene reconstruction to inform execution-grounded robot policies, suggesting a path toward more robust real-world deployment. This work is significant because it moves beyond purely reactive systems toward models that can reason about their surroundings dynamically.
This relates closely to efforts in small-object navigation within shifting layouts, where a new benchmark and method were introduced to handle the challenge of navigating changing object arrangements over time. Furthermore, research into visual representations for autonomous driving asks whether simply having better visual data always translates into superior end-to-end performance in self-driving systems.
A related area focuses on improving vision-language action models through automated video-language grounding, specifically with YUBI-STAG, which aims to achieve contact and semantic richness by aligning video and language data. This builds upon the foundational work of Juno, which tackles the problem of taming predictive latents within vision-language action models.
Finally, there is work on temporal visuo-tactile learning designed to enhance dexterous grasp stability by incorporating both visual and tactile information over time. This contrasts with the broader goal of creating multimodal aerial datasets like MultiFly, which focuses on annotation efficiency and cross-modal semantic consistency for aerial robotics.
The most significant development today concerns the framework for robotic failure analysis and correction, specifically RoboFAC. This work is crucial because it moves beyond simple task completion to actively understanding and fixing when robots go wrong in complex physical settings.
RoboFAC introduces a comprehensive system designed to diagnose failures and then propose corrective actions, which is a big step forward from just observing errors. It builds upon prior work in world models that try to predict outcomes, suggesting that by integrating failure detection directly into the planning loop, we can make these systems much more robust.
Then there is the effort on transition path sampling using Koopman operators and exit-time optimal control. This method is important because it allows us to find the best way for a robot to move from one state to another while minimizing time, which directly impacts how quickly a physical task can be performed.
Instrumentation for imitation learning also made progress today, focusing on enhancing training datasets for clothes hanger insertion. This work provides better sensory input specifically tailored to teaching robots delicate manipulation skills, which feeds into the broader goal of making generalist agents capable of handling varied physical interactions.
The unification of object-centric world models and diffusion policy represents another key direction; this hierarchical framework aims to give robots a unified way to plan multi-stage tasks, linking the high-level understanding of objects with the low-level control policies. This connects nicely to how SAPS is attempting to steer policies by blending teleoperation with a pretrained vision language agent.
The most significant development today involves the work on robotic ultra-long-horizon manipulation skills via human guided lifelong code generation, because it directly addresses the challenge of teaching robots complex, multi-step tasks that require continuous learning over extended periods. This approach attempts to build these skills by having humans guide the robot through a process of generating and refining code for its actions.
This method is being explored to achieve these long-horizon skills. The research focuses on using human guidance to create lifelong code generation for robotic manipulation tasks, which suggests a way for robots to acquire complex abilities incrementally rather than through pre-programmed scripts.
Another important area is the development of dynamic neural koopman distillation for fast robot control using diffusion models, as it promises faster and more robust control mechanisms by leveraging these generative models. This work aims to distill knowledge from large diffusion models into a model that can be deployed for real-time robotic control, which is crucial for dynamic interactions.
We also saw some progress in targeting world models to compromise robot learning pipelines, which means researchers are actively trying to find ways to intentionally introduce errors or constraints into the internal representations of robots so they become more robust when encountering novel situations. This is an attempt at adversarial training to improve safety and generalization.
Finally, there is the work on safe unified slip and fracture detection with low-cost acoustic sensing in robotic grasping, which matters because it directly enhances the physical interaction capabilities of robots by allowing them to detect slippage or breakage during grasping using simple sound data. This builds upon previous efforts by providing a tangible way for robots to assess contact quality.
The most significant development today concerns the work on MimicX, which refines policy-in-the-loop supervision for tracking humanoid motion driven by video. This is important because it directly addresses the need for more robust and adaptable control systems when dealing with complex visual inputs in real-world scenarios. The research involved refining how a policy supervises itself based on video data to improve tracking accuracy.
This refinement builds upon earlier efforts, such as those exploring the transfer of co-evolved communication from two dimensional to three dimensional simulations, which provided foundational understanding for how control signals propagate across different spatial dimensions. Furthermore, the work on PhysEvo shows an attempt to allow Astra robots to act autonomously based on its capabilities.
A related piece of research focused on ClimbLab, a MATLAB simulation platform designed specifically for legged climbing robotics, which provides a controlled environment for testing locomotion strategies. This simulation work feeds into the broader goal of creating responsive noise-relaying diffusion policies that offer efficient visuomotor control.
Finally, there is the RoboPilot project, which aims to achieve generalizable dynamic robotic manipulation through dual-thinking modes. This approach seeks to give robots flexible decision-making capabilities in manipulation tasks, connecting back to the autonomous navigation challenges posed by quadruped systems.
The most pressing work today involves the development of self mixing laser interferometry for robotic tactile sensing because it directly addresses the need for high fidelity in how robots perceive physical contact. Researchers explored a method where laser interferometry is used to create a self mixing system, which aims to improve the accuracy of force and motion sensing on robot hands by integrating multiple light paths. This work builds upon prior efforts that focused on improving the robustness of these sensing modalities.
A significant piece of progress was made in SurGE, which uses surrogate gradient guidance for co-designing legged robots with parallel elasticity. This approach seeks to optimize the physical structure and control laws simultaneously, meaning they are designed together rather than separately. This is important because it moves away from purely sequential design methods toward a more holistic system architecture.
Then there is FAR, which focuses on failure aware retry for test time recovery and continual policy improvement in robotic systems. This technique attempts to make robots more resilient when things go wrong during operation by intelligently retrying actions based on observed failures. This is connected to the work on adapting generalist vehicle models for high speed MPC across terrains, as both aim to improve real-time performance under challenging conditions.
Another area of exploration involved bridging reinforcement learning and optimal control through feasible action mapping. This research tries to connect the abstract decision-making of machine learning with the precise control required for physical movement. This is a step toward creating systems that can learn complex behaviors while still adhering to strict physical constraints.
Finally, there is work on trajectory planning without trajectory data using a manifold guided approach, which focuses on generating paths even when specific prior path data is unavailable. This complements the efforts in evidence driven human agent robot teaming for anomaly triage by providing better foundational motion planning capabilities for autonomous agents operating in unstructured environments.
The most pressing work involves understanding how robot world models fail when they encounter unexpected physical interactions, which is crucial because current systems often lack the necessary sensitivity to adapt. geodex attempts to build a library for motion planning on Riemannian manifolds, which means it's trying to create smarter ways for robots to navigate complex curved spaces. This foundational mapping work is significant because it provides the mathematical framework for movement that other planning systems will eventually use.
Building upon this, eGRAP tackles the problem of coordinated dual-arm robotic disassembly of electronic devices using graph-based adaptive planning. This means instead of following a fixed plan, the system dynamically adjusts its sequence of actions based on what it observes during the physical breakdown process. This is important because real-world electronics rarely follow textbook assembly procedures, so this adaptability is key to success.
Another area focuses on multisensory continual learning, which adapts pretrained visuomotor policies to handle force feedback. This research looks at how robots can learn to adjust their actions when they feel unexpected resistance during manipulation tasks. This builds on the idea of improving policy robustness by incorporating tactile information into the learning loop.
Then there is VIA, which develops a visual interface agent specifically for robot control. This agent aims to give human operators a better way to guide complex robotic movements through visual input. This is a direct attempt to bridge the gap between high-level human intent and low-level robotic execution.
ModPack explores an extensible teleoperation interface designed for bimanual mobile manipulation, focusing on how humans can control robots with two hands in a flexible way. This work addresses the practical challenge of giving humans intuitive, dexterous control over complex objects using multiple limbs.
Finally, the research on contact shifts and tactile representations delves into moving beyond wearable interfaces to create truly dexterous policies by focusing on how robots perceive physical contact changes. This is about getting the robot's sense of touch much more nuanced so it can react intelligently to subtle physical cues.
The most significant development was the work on HULK, which focuses on learning whole-body forceful locomotion manipulation for humanoids. This matters because it directly addresses how robots can move with the kind of dynamic, powerful movement we see in humans, moving beyond simple pre-programmed motions. The research attempted to learn these complex movements through some form of learning process that resulted in a new method for controlling humanoid bodies.
FlashNeRD introduced performance-first contact-rich neural robot dynamics, which is important because it seems to focus on how robots should react when they make physical contact during tasks. This approach builds upon the idea that a unified kinematic representation can be used to estimate joint moments in a reusable biological joint estimation framework. That estimation method is key because it allows for more flexible control strategies, and this connects directly to the work exploring what matters in action tokenization for robot policies.
The tokenization research looked at what truly matters when deciding on robot policies, suggesting a way to distill complex actions into meaningful units. This idea relates to embedded evaluation of task admission coalescing in decentralized multi-robot systems, which deals with how robots coordinate their decisions. Furthermore, adaptive risk-certified event-triggered replanning for dynamic navigation shows how robots can safely adjust their paths when unexpected situations arise.
RobotAPO focused on adversarial physics preference optimization for robotic manipulation video generation, which is important because it tries to make robot actions look more realistic by optimizing them against physical constraints. This contrasts with the context-aware adaptive pesticide spraying for agricultural robots under changing weather and terrain, which uses vision-language models to adapt spraying based on visual input and environmental changes.
Today's papers
- Rephrase Before You Act: Characterizing and Mitigating Language Sensitivity in Vision-Language-Action Models. [paper]
- NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning. [paper] [episode]
- CIRRA: Dual-Level Continual Instruction Reconciliation with Ongoing Execution for Embodied Robot Agents in Interactive Household Tasks. [paper]
- ED3R: Energy-Aware Distributed Disaster Detection via Cooperative Agents in Robotic Systems. [paper] [episode]
- UniCross: Unified Cross-Skill Dexterous Manipulation Synthesis. [paper] [episode]
- Agentic Scene Policies. [paper] [episode]
- TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning. [paper] [episode]
- Modeling Robotics Dataset Construction as an Artifact-Based Build Process. [paper] [episode]
- A Review Of Robotic World Models For Dynamic Environments Based On Factor And Scene Graphs. [paper]
- Juno: Taming Predictive Latents for Vision-Language-Action Models. [paper]
- Lifelong small-object navigation in changing object layouts: a benchmark and method. [paper]
- Do Better Visual Representations Always Lead to Better End-to-End Autonomous Driving?. [paper]
- YUBI-STAG: Contact and Semantic-Rich Alignment for VLAs via Automated Video-Language Grounding. [paper]
- Temporal Visuo-Tactile Learning for Dexterous Grasp Stability. [paper]
- Agentic RSR: Real-to-Sim-to-Real through Scene Reconstruction and Execution-Grounded Robot Policies. [paper]
- MultiFly: A Real-World Multimodal Aerial Dataset with Annotation-Efficient Label Transfer and Cross-Modal Semantic Consistency. [paper]
- RoboQuest: Generalist Physical Agents that Search, Inspect and Test. [paper]
- Long-WAM: Scaling the Context of World-Action Models. [paper]
- Transition Path Sampling Using Koopman Operators and Exit-Time Optimal Control. [paper]
- RoboFAC: A Comprehensive Framework for Robotic Failure Analysis and Correction. [paper] [episode]
- Instrumentation for Imitation Learning: Enhancing Training Datasets for Clothes Hanger Insertion. [paper] [episode]
- Resolving Conflicts Where and When They Arise: Reactive Composition of Multi-Goal Behavior. [paper] [episode]
- Unifying Object-Centric World Models and Diffusion Policy: A Hierarchical Framework for Multi-Stage Robotic Tasks. [paper] [episode]
- SAPS: Shared Autonomy for Policy Steering by Blending Teleoperation with a Pretrained VLA. [paper] [episode]
- SAIN: Structure-Aware Interactive Navigation with Active Dialogue Grounding for Mobile Robot. [paper] [episode]
- Where Success Breaks: Failure-Boundary Learning for Robust Vision-Language-Action Models. [paper] [episode]
- Robotic Ultra-Long-Horizon Manipulation Skills via Human-guided Lifelong Code Generation. [paper] [episode]
- Dynamic Neural Koopman Distillation for Fast Robot Control Using Diffusion Models. [paper] [episode]
- Targeting World Models to Compromise Robot Learning Pipelines. [paper] [episode]
- The Impact of Operational-Data Fidelity when Assessing Safety-Critical Autonomous-Vehicle Software. [paper] [episode]
- Visible Touch: Rendering Contact for Visuomotor Policies. [paper] [episode]
- SAFE: Unified Slip and Fracture Detection with Low-Cost Acoustic Sensing in Robotic Grasping. [paper]
- Taming an End-to-End Autonomous Driving Policy for Urban Navigation of Quadruped Robots. [paper]
- PhysEvo: Astra Can Act, Let It. [paper]
- MimicX: Policy-in-the-Loop Supervision Refinement for Video-Driven Humanoid Motion Tracking. [paper]
- Fast Planning for Multi-object Multi-target Throwing. [paper]
- Evaluating the Transfer of Co-Evolved Communication from 2D to 3D Simulation. [paper]
- ClimbLab: MATLAB Simulation Platform for Legged Climbing Robotics. [paper]
- Responsive Noise-Relaying Diffusion Policy: Responsive and Efficient Visuomotor Control. [paper] [episode]
- RoboPilot: Generalizable Dynamic Robotic Manipulation with Dual-thinking Modes. [paper] [episode]
- Self-Mixing Laser Interferometry for Robotic Tactile Sensing. [paper] [episode]
- SurGE: Surrogate Gradient-guided Evolution for Co-design of Legged Robots with Parallel Elasticity. [paper] [episode]
- FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement. [paper] [episode]
- Adapting Generalist Vehicle Models for High-Speed MPC Across Terrains. [paper] [episode]
- Bridging Reinforcement Learning and Optimal Control via Feasible Action Mapping. [paper] [episode]
- Trajectory Planning without Trajectory Data: A Manifold-Guided Approach. [paper]
- Toward Evidence-Driven Human-Agent-Robot Teaming for Earth-Independent Anomaly Triage. [paper]
- Workhorse: Learning Robust Whole-Body Humanoid Loco-Manipulation from Human Data. [paper]
- World Models Dream of Success: Diagnosing and Repairing Failure Insensitivity in Robot World Models. [paper]
- geodex: A Library for Motion Planning on Riemannian Manifolds. [paper]
- Graph-Based Adaptive Planning for Coordinated Dual-Arm Robotic Disassembly of Electronic Devices (eGRAP). [paper] [episode]
- TriDeliver: Cooperative Air-Ground Instant Delivery with UAVs, Couriers, and Crowdsourced Ground Vehicles. [paper] [episode]
- Multisensory Continual Learning: Adapting Pretrained Visuomotor Policies to Force. [paper] [episode]
- VIA: Visual Interface Agent for Robot Control. [paper] [episode]
- ModPack: Extensible Teleoperation Interface for Bimanual Mobile Manipulation. [paper] [episode]
- From Wearable Interfaces to Dexterous Policies: Contact Shifts and Tactile Representations. [paper]
- HULK: Learning Whole-Body Forceful Loco-Manipulation for Humanoids. [paper]
- FlashNeRD: Performance-First Contact-Rich Neural Robot Dynamics. [paper]
- A Unified Kinematic Representation Enables Reusable Biological Joint Moment Estimation. [paper]
- Beyond Reconstruction: What Matters in Action Tokenization for Robot Policies?. [paper]
The papers
- Visible Touch: Rendering Contact for Visuomotor Policies — Integrating contact information into visuomotor policies remains an open problem because most modern policies operate from vision and proprioception alone, yet touch is essential for robust manipulation. [episode]
- FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement — Failure-Aware Retry (FAR) is a framework designed to enable robot policies to learn from previous failures at test time, adapt their behavior accordingly, and eventually complete tasks autonomously. [episode]
- RoboFAC: A Comprehensive Framework for Robotic Failure Analysis and Correction — Vision-Language-Action (VLA) models are advanced in robotic manipulation but lack structured supervision for failure diagnosis and recovery, which limits their robustness in open-world scenarios. [episode]
- Robotic Ultra-Long-Horizon Manipulation Skills via Human-guided Lifelong Code Generation — Large language models (LLMs)-based code generation for robotic manipulation has recently shown promise by directly translating human instructions into executable code, but existing approaches are limited by language ambiguity, noisy outputs, and limited context windows, which mak [episode]
- RoboPilot: Generalizable Dynamic Robotic Manipulation with Dual-thinking Modes — RoboPilot introduces a dual-thinking closed-loop framework for dynamic robotic manipulation that enables adaptive reasoning by dynamically switching between fast and slow thinking modes to balance efficiency and accuracy in complex, real-world environments. [episode]
- ModPack: Extensible Teleoperation Interface for Bimanual Mobile Manipulation — ModPack introduces a modular and extensible teleoperation system designed to support diverse robot embodiments and task requirements within a unified framework, addressing the limitations of existing systems that are often tailored to specific hardware. [episode]
- Modeling Robotics Dataset Construction as an Artifact-Based Build Process — Modeling robotics dataset construction as an artifact-based build process substantially reduces dataset update latency while maintaining a deterministic build design that supports reproducibility. [episode]
- Resolving Conflicts Where and When They Arise: Reactive Composition of Multi-Goal Behavior — Reactive control is often considered insufficient for multi-objective tasks because conflicting objectives give rise to local minima. [episode]
- SAIN: Structure-Aware Interactive Navigation with Active Dialogue Grounding for Mobile Robot — Most existing vision-language navigation tasks assume that instructions are complete and unambiguous, but real-world robots often encounter natural human instructions that are ambiguous, underspecified, or incomplete, requiring them to resolve such uncertainties through active qu [episode]
- Unifying Object-Centric World Models and Diffusion Policy: A Hierarchical Framework for Multi-Stage Robotic Tasks — Visual world models have shown great potential in learning complex system dynamics, but they are currently limited to single-stage robotic tasks and struggle with multi-stage ones that demand complex sequential planning. [episode]
- VSC-HVDC setpoint adjustment for maximum grid utilisation under voltage constraints — Voltage-source converter high-voltage direct current (VSC-HVDC) links offer controllable active and reactive power output, making them a promising asset for emergency voltage support. [episode]
- Graph-Based Adaptive Planning for Coordinated Dual-Arm Robotic Disassembly of Electronic Devices (eGRAP) — Electronic-device Graph-based Adaptive Planning (eGRAP) is a perception-driven, graph-based planning framework that enables coordinated dual-arm robotic disassembly of electronic devices by converting live detections into a precedence graph and applying simple rules to coordinate [episode]
- Adapting Generalist Vehicle Models for High-Speed MPC Across Terrains — High-speed off-road autonomy requires precise closed-loop control for a target vehicle while remaining robust across changing terrains. [episode]
- SurGE: Surrogate Gradient-guided Evolution for Co-design of Legged Robots with Parallel Elasticity — Co-design of legged robots with elastic elements is challenging due to the non-differentiability of contact dynamics and mechanism engagement. [episode]
- Collaboration in Multi-Robot Systems: Taxonomy and Survey of Frameworks for Collaboration — Collaboration is a central theme in multi-robot systems as tasks and demands increasingly require capabilities that go beyond what any one individual robot possesses, yet terminology surrounding collective interaction remains inconsistent across research communities. [episode]
- SAPS: Shared Autonomy for Policy Steering by Blending Teleoperation with a Pretrained VLA — Recent advancements in Vision-Language-Action (VLA) models demonstrate impressive generalist capabilities in robot manipulation, yet these policies can be brittle under out-of-distribution spatial and semantic perturbations. [episode]
- ED3R: Energy-Aware Distributed Disaster Detection via Cooperative Agents in Robotic Systems — Robotics are expected to support environmental monitoring and natural disaster management, where decisions must be made under uncertainty, resource limitations, and strict operational constraints. [episode]
- Reinforcement Learning-Based Traffic Signal Control for IoT-Enabled Intersections — Urban traffic congestion remains a persistent challenge, and this research investigates reinforcement learning (RL) as an edge-intelligent approach for adaptive traffic signal operation at a signalized urban intersection in Kuwait. [episode]
- Multisensory Continual Learning: Adapting Pretrained Visuomotor Policies to Force — Robot manipulation often depends on sensory data beyond vision, especially in contact-rich tasks where force, tactile, or audio feedback reveals interaction states not directly visible from images. [episode]
- VIA: Visual Interface Agent for Robot Control — Robot manipulation requires complex skills like visual understanding, physical reasoning, and planning, which have been enhanced by foundation models (FMs). [episode]
- UniCross: Unified Cross-Skill Dexterous Manipulation Synthesis — A unified framework for cross-skill dexterous manipulation synthesis has been presented that models grasping, relocation, in-hand rotation, and in-hand translation from a shared hand-object relational perspective. [episode]
- Responsive Noise-Relaying Diffusion Policy: Responsive and Efficient Visuomotor Control — Responsive Noise-Relaying Diffusion Policy (RNR-DP) addresses a key limitation in Diffusion Policy, which suffers from poor responsiveness due to its reliance on a large action horizon. [episode]
- NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning — Given that solving long-horizon manipulation requires integrating high-level semantic reasoning with low-level physical interaction, NovaPlan introduces a hierarchical framework that unifies closed-loop VLM and video planning with geometrically grounded robot execution for zero-s [episode]
- TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning — Active vision—where a policy controls its own gaze during manipulation—has emerged as a key capability for imitation learning, with multiple independent systems demonstrating its benefits in the past year. [episode]
- Global Geolocated Realtime Data of Interfleet Urban Transit Bus Idling — Urban transit bus idling contributes significantly to ecological stress, economic inefficiency, and medically hazardous health outcomes due to emissions. [episode]
- Feasibility and Explicit Safety Filters for Control Barrier Functions in Linear Systems — Safety filters based on control barrier functions (CBFs) and high-order control barrier functions (HOCBFs) are often implemented through quadratic programs (QPs), but feasibility certification can be difficult, especially with multiple constraints, as feasibility may be lost as t [episode]
- Agentic Scene Policies — Executing open-ended natural language queries is a core problem in robotics, and this work presents Agentic Scene Policies (ASP), an agentic framework that leverages advanced semantic, spatial, and affordance-based querying capabilities of modern scene representations to implemen [episode]
- Stability Analysis in Multi-Constraint Safety Filters for Linear Systems — Multi-constraint safety filters based on control barrier functions for linear systems with affine state constraints yield continuous piecewise-affine closed-loop dynamics and may introduce boundary equilibria and unstable active-set modes, raising questions about whether such loc [episode]
- Targeting World Models to Compromise Robot Learning Pipelines — World models introduce a uniquely stealthy and effective data poisoning entry point into the robot learning supply chain that can result in the deployment of unsafe or otherwise compromised robotic policies despite training on seemingly safe ground truth training data. [episode]
- Privacy-Preserving Cram'er-Rao Lower Bound — The paper establishes a privacy-preserving Cramér-Rao lower bound (CRLB) theory to characterize the fundamental limit of identification accuracy under general stochastic obfuscation mechanisms, which is crucial for designing optimal system identification algorithms by quantifyin [episode]
- Bridging Reinforcement Learning and Optimal Control via Feasible Action Mapping — Operating constrained dynamical systems requires controllers to efficiently solve complex tasks while enforcing recursive feasibility and safety constraints, which this work addresses by presenting Feasible Action for Optimal Control (FAOC), a novel control framework integrating [episode]
- Where Success Breaks: Failure-Boundary Learning for Robust Vision-Language-Action Models — Where Success Breaks: Failure-Boundary Learning for Robust Vision-Language-Action Models proposes a novel training paradigm, DLS, to enhance the robustness of flow-based Vision-Language-Action (VLA) models by learning where closed-loop behavior transitions from recoverable deviat [episode]
- Instrumentation for Imitation Learning: Enhancing Training Datasets for Clothes Hanger Insertion — Instrumentation for imitation learning provides invaluable state information that enables efficient learning for robotic manipulation, as demonstrated in this study focusing on clothes hanger insertion. [episode]
- Circuit realization and hardware linearization of monotone operator equilibrium networks — It is shown that resistor–diode networks correspond to monotone operator equilibrium networks, providing a parsimonious construction for analog hardware and enabling direct hardware computation of gradients. [episode]
- On observer forms for hyperbolic PDEs with boundary dynamics — A hyperbolic observer canonical form (HOCF) for linear hyperbolic Partial Differential Equations (PDEs) with boundary dynamics is presented, offering a systematic method to transform system descriptions into an observer canonical form using observability coordinates. [episode]
- Learning Robust Control Lyapunov Functions through Lipschitz Neural Networks — This work presents a novel framework for learning robust control Lyapunov functions and stabilizing controllers for nonlinear dynamical systems subject to additive disturbances upper bounded by a state-dependent function. How it works 1. [episode]
- Generalized Model Predictive Path Integral Control as Expectation--Maximization — Model Predictive Path Integral (MPPI) control can be interpreted as an Expectation–Maximization (EM) algorithm applied to a probabilistic inference formulation of optimal control, providing a unified optimization-theoretic perspective that extends MPPI beyond Gaussian parameter [episode]
- Dynamic Neural Koopman Distillation for Fast Robot Control Using Diffusion Models — Dynamic Neural Koopman Distillation (DNK) is a framework that distills multistep diffusion inference into a single forward pass using state-dependent factorized latent dynamics, enabling millisecond-level inference latency suitable for high-frequency closed-loop robot control whi [episode]
- Confidence-Aware Safe and Stable Control of Control-Affine Systems — Designing control inputs that satisfy safety requirements is crucial in safety-critical nonlinear control, and this task becomes particularly challenging when full-state measurements are unavailable. [episode]
- TriDeliver: Cooperative Air-Ground Instant Delivery with UAVs, Couriers, and Crowdsourced Ground Vehicles — Instant delivery, shipping items before critical deadlines, is essential in daily life. [episode]
- Collision Avoidance for Convex Primitives via Differentiable Optimization Based High-Order Control Barrier Functions — Ensuring system safety through collision avoidance is a critical challenge in robotics and autonomous systems, and this work introduces a novel framework that addresses this by transforming nonconvex safety constraints into linear constraints using high-order control barrier func [episode]
- Self-Mixing Laser Interferometry for Robotic Tactile Sensing — Self-mixing interferometry (SMI) has been adapted for robotic fingertip sensing to detect object slip and extrinsic contact, offering a novel, non-contact tactile sensing modality. [episode]
- The Impact of Operational-Data Fidelity when Assessing Safety-Critical Autonomous-Vehicle Software — For safety-critical software, data from its operational past can provide statistical support for reliability claims, but this data might lack sufficient detail to capture important failure features. [episode]
- Real-time Coordination of Cascaded Hydropower under Decision-Dependent Uncertainty — Real-time coordination frameworks for cascaded hydropower systems are essential for balancing generation reliability and water management constraints amidst complex, coupled uncertainties. [episode]
- Immiscible Diffusion Policy: Preserving Multimodal Robot Actions through Label-Free Noise Assignment —
- Trajectory Sensitivity-Based Dynamic Assessment of an AI Data Center Under Contingencies —
- Distributionally Robust Frequency Response Estimation Using Non-Steady-State Frequency-Domain Data —
- IVG-UAV: An Intelligent Voice-Guided UAV System for Autonomous Ripe Fruit Harvesting with Vision-Based Classification and Adaptive Path Planning —
- Hierarchical Defense in Leader-Defender-Attacker Games: Nested Equilibria and Receding-Horizon Control —
- Restorative Control and the Limits of Safe Model Discrimination —
- TempoBridge: Language-Guided Tempo Control for Vision-Language-Action Policies —
- RobotAPO: Adversarial Physics Preference Optimization for Robotic Manipulation Video Generation —
- TMT: Runtime Backdoor Detection for Vision-Language-Action Policies on Unseen Tasks —
- Precise SE(3) End-Effector Tracking in Whole-Body Humanoid Control —
- SiGNgapore - An Interactive Dataset for Sign-based Visual Navigation —
- Sparse Feature Policy Unlearning Mitigates State Hallucination in Vision-Language-Action Models —
- Not All Uncertainty Matters: Simulation-in-the-Loop Fast-Slow Reasoning for Decision-Critical Autonomous Driving System —
- Contact-Aware Imitation Learning Through Contact Factorization —
- MagCilia: A Compact Magnetociliary Tactile Sensor with 3D Force Sensing for Robotic Contact Perception and Grasping Feedback —
- Targeted Modality Dropout for Real-Robot Manipulation Robust to Intermittent Vision Loss —
- Point It, Strike It: Direction-Conditioned Dynamic Manipulation of Deformable Linear Objects —
- Adaptive Code Generation for Controlling Robots —
- Fast and Robust Teach-and-Repeat Navigation Using MixVPR Visual Place Recognition* —
- Safe Control of Semi-Explicit Differential-Algebraic Systems: Control Barrier Functions with SOS Verification —
- Design optimization of tendon-driven robots considering tendon wrapping and shortcut —
- NAViLoss: An Underwater Navigation-Aware Dual-Residual Objective for Physics-Consistent Learning —
- Do Better Visual Representations Always Lead to Better End-to-End Autonomous Driving? —
- RoboPace: Contact-Aware Time-Optimal Retiming for Action-Chunk Policies —
- YUBI-STAG: Contact and Semantic-Rich Alignment for VLAs via Automated Video-Language Grounding —
- WAM: Distilling Action Tangent Fields into World Action Models —
- AeroEval: Staged Program and Execution Validation for AI-Generated Drone Missions —
- On-Demand Robotic Assembly via Differentiable Geometric Part Repair —
- End-to-End Autonomous Generation of Human Assembly Plans —
- Enhancing Robotic Perception and Adaptability through Sensor Fusion and Origami-Inspired Designs —
- Towards Accurate End-Effector Localization for UMI-Style Robotic Manipulation Teaching —
- Juno: Taming Predictive Latents for Vision-Language-Action Models —
- Many Ways to Succeed: Diversity-Driven RL Fine-Tuning for VLA Generalization —
- Borrowed Eyes: Markerless Nano-UAV Flight with an Active Quadruped Observer —
- Decoding Neural Population Dynamics through Robotic Analog —
- Continuous-Time Critic as a Lyapunov Function: Projected Adaptation and Cart-Pole Stabilization —
- Tactile Reconstruction of Contact Task Frames and Forces for Hybrid Force/Motion Control —
- Video-to-Model: Automatic Modeling of Deformable Linear Objects —
- Transition Path Sampling Using Koopman Operators and Exit-Time Optimal Control —
- Distributed Motion Planning for Multi-Robot Systems under Topological Constraints —
- RealtimeWAM: How Fast Can I Run My World Action Model? —
- Lifelong small-object navigation in changing object layouts: a benchmark and method —
- Do Vision-Language-Action Models Understand Instructions? A Mechanistic Interpretability Study on Language Grounding —
- Design of a Fully Actuated 4-DOF Robotic Finger With Joint-Specific Hybrid Remote Actuation —
- Benchmarking Behavioral Steerability in Behavior Foundation Models —
- Making Task Abstractions Executable: Control-Aware Layout Repair for a Fixed Controller —
- Temporal Visuo-Tactile Learning for Dexterous Grasp Stability —
- Energy-Efficient Gait Adaptation via Hierarchical Reinforcement Learning for Quadrupedal Locomotion Across Diverse Terrains —
- Hall Effect-Based Tactile Force Detection Sensor for Robot-Assisted Minimally Invasive Surgery —
- MultiFly: A Real-World Multimodal Aerial Dataset with Annotation-Efficient Label Transfer and Cross-Modal Semantic Consistency —
- Semantic-Aware Predictive Mapping for Exploration and Navigation —
- OpenViTac: Learning and Benchmarking Visuo-Tactile Policies in a Unified Sim-and-Real Framework —
- NeRFifyMesh: Optimizing Neural Radiance Fields from Textured Meshes for Robotics Scene Building —
- RoboQuest: Generalist Physical Agents that Search, Inspect and Test —
- RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments —
- AirGroundVLN: A Large-Scale Benchmark for Goal-Oriented Air-Ground Collaborative Vision-and-Language Navigation —
- RFPO: Rectified Flow Policy Optimization for Embodied Control —
- FoldBack: Self-Correcting Masked Generative Policy for Long-Horizon Garment Folding —
- LLA-MPPI: Rapidly Adaptive Whole-body Control of Legged Robots with GPU-Accelerated Parallel Simulations —
- CMP-IRRT*: A Perception-Assisted Height-Adaptive Planner for Quadruped Robots —
- Robotic Boomerang Throwing via Model-Based Release Design —
- Agentic RSR: Real-to-Sim-to-Real through Scene Reconstruction and Execution-Grounded Robot Policies —
- HuMBLE: Human Motion-Driven Behavior Learning for Embodied Locomotion —
- Factorized Tactile Representation and Control for Sim-to-Real Manipulation —
- Rephrase Before You Act: Characterizing and Mitigating Language Sensitivity in Vision-Language-Action Models —
- Long-WAM: Scaling the Context of World-Action Models —
- RoboPrompt: Intuitive Robot Policy Steering with Sparse Human Input —
- A Review Of Robotic World Models For Dynamic Environments Based On Factor And Scene Graphs —
- SAFE: Unified Slip and Fracture Detection with Low-Cost Acoustic Sensing in Robotic Grasping —
- Context-Aware Adaptive Pesticide Spraying for Agricultural Robots under Changing Weather and Terrain Using Vision-Language Models —
- Taming an End-to-End Autonomous Driving Policy for Urban Navigation of Quadruped Robots —
- Autonomous Droplet Navigation via Model-Based Reinforcement Learning: Zero-Shot Transfer and Emergent Dynamics —
- CIRRA: Dual-Level Continual Instruction Reconciliation with Ongoing Execution for Embodied Robot Agents in Interactive Household Tasks —
- Trajectory Planning without Trajectory Data: A Manifold-Guided Approach —
- From Wearable Interfaces to Dexterous Policies: Contact Shifts and Tactile Representations —
- Toward Evidence-Driven Human-Agent-Robot Teaming for Earth-Independent Anomaly Triage —
- Quantifying Grid-Forming Requirements for System Strength: How Much and Where? —
- HULK: Learning Whole-Body Forceful Loco-Manipulation for Humanoids —
- The internal model principle in reservoir computing —
- PhysEvo: Astra Can Act, Let It —
- Electric Racing Kart with BLDC Drive, LiFePO4 Battery System, Traction Control and Regenerative Braking —
- MimicX: Policy-in-the-Loop Supervision Refinement for Video-Driven Humanoid Motion Tracking —
- AI Data Centers Meet Electrical Grids: A Review of Power Challenges and Coordinated Solutions —
- Local Reproduction Numbers for Analysis and Optimal Mitigation of Epidemic Processes —
- Localized Thermal Management of Induction-Heated Microrobots for Hyperthermia Applications —
- Workhorse: Learning Robust Whole-Body Humanoid Loco-Manipulation from Human Data —
- FlashNeRD: Performance-First Contact-Rich Neural Robot Dynamics —
- World Models Dream of Success: Diagnosing and Repairing Failure Insensitivity in Robot World Models —
- A Unified Kinematic Representation Enables Reusable Biological Joint Moment Estimation —
- geodex: A Library for Motion Planning on Riemannian Manifolds —
- Beyond Reconstruction: What Matters in Action Tokenization for Robot Policies? —
- CAP: Codebook-Aligned Prediction for Tokenized Robot Policies —
- Embedded Evaluation of Task Admission Coalescing in Decentralized Multi-Robot Systems —
- Belief-Space Planning with Planner-Conditioned Estimator Error under Intermittent Observations —
- Mission-critical spectrum sharing with decentralized Multi-Agent Reinforcement Learning —
- Fast Planning for Multi-object Multi-target Throwing —
- Co-Evolving Robot Orchestrators and Policies through Deployment —
- RoboRender: Robot-Oriented Video Generation for Visual Sim-to-Real Transfer —
- Cooperative Dueling DQN SAC Learning for Energy Efficiency in Dynamic OWC Networks —
- A Single-Inductor N:1 Resonant Switched-Capacitor Ladder Converter with Efficient Continuous Voltage Conversion Capability —
- Evaluating the Transfer of Co-Evolved Communication from 2D to 3D Simulation —
- LACE-CRAFT: Robot Co-Design with Actor Inheritance and Blackboard Collaboration —
- Co squared Skill: Whole-Body Control via Skill Composition for Long-Horizon Human-Environment Interaction —
- Adaptive Risk-Certified Event-Triggered Replanning for Dynamic Navigation —
- Predicted Futures Are Not Enough: Learning Executable Goals for Robot Manipulation —
- ClimbLab: MATLAB Simulation Platform for Legged Climbing Robotics —
- MCFR: A Mask-Guided Coarse-to-Fine Regression Framework for Robust Multi-Variant Board-to-Board Connector Assembly —
- Learning Unknown Constraints without Unsafe Data via Optimality and Counterfactual Regularization —
- COOL: Curiosity-Driven Object Ownership Learning for Personalized Robotic Assistance —
- Learning Stability of Replay-Based Co-Optimization for Transmission Expansion under Strategic Bidding —
Important terms
- Vision Language Action Models
- These models are being improved to better understand and respond to complex human instructions, which is key for enabling robots to perform nuanced tasks in real environments.
- Zero-shot Long-horizon Manipulation
- This involves planning a sequence of actions based on natural language commands without needing specific training for every single scenario, allowing robots to handle complex, multi-step tasks.
- RoboFAC
- A crucial framework that diagnoses robot failures and proposes corrective actions. It moves beyond just observing errors to actively understanding and fixing mistakes during operation.
- Scene Reconstruction Grounded Policies
- This technique connects simulation with reality by using reconstructed scenes to inform robot policies, allowing robots to reason dynamically about changing environments.