Robotics papers — 2026-10-01
Today's focus is on figuring out the practical limits of latent world models to see when these abstract representations break down in real-world tasks. Researchers looked at MotionWeave, which tries to learn motion-centered future dynamics for vision-language-action policies. This means training a system to predict movement based on what it sees and understands.
There is also work on UniWAM, which focuses on unified mobile manipulation using mixed-stream world-action modeling and supervision of manipulation anchor poses. This aims to make mobile robots more capable in complex environments by building upon an understanding of dynamics.
Ego4WAM investigates what matters when scaling egocentric human data for robot learning. This suggests that the quality and relevance of that input data dictate how well a robot learns to navigate or act, which feeds into how world models are designed.
Research on Multi-Link Safety Filtering for VLA policies around moving hazards adds a layer of safety checks to vision-language-action policies when dealing with dynamic risks. This safety layer is what needs consideration when pushing these predictive models into operational settings.
The most significant piece of work from the day is DexHoldem, which sets a new benchmark for how agents can handle dexterous manipulation in complex scenarios like Texas Hold'em. This work matters because it pushes the boundaries of what is expected from robotic dexterity in real-world, dynamic environments.
DexHoldem introduced an agentic robotics benchmark specifically designed to test how well robots can perform intricate tasks requiring fine motor skills. The results showed that agents using this framework demonstrated superior performance compared to prior methods when executing these manipulative challenges. This finding provides a standardized way to measure the success of learning-based manipulation policies.
Following that, there was progress on OGPO, which focuses on real-time robot control using one-step generative policy optimization. This method attempts to generate control actions directly during execution, which is crucial for fast responses in dynamic situations. This approach builds upon the idea of generating policies quickly to address immediate control needs.
Another key development involves Learning-Based Progressive Barrier Control for robot manipulators that start with errors outside their expected tracking bounds. This work tackles the problem of robots failing when they encounter unexpected deviations during movement, showing how to gradually guide them back into safe operational limits. This is a necessary step toward making physical systems more robust.
This concept of dynamic correction connects well with DSDyn-VLA, which employs a dual-stream dynamic manipulation framework incorporating motion perception and future awareness for real-time correction. DSDyn-VLA seems to be taking the barrier control idea and adding predictive elements to handle the uncertainty inherent in movement. The research is still open regarding how effectively these predictive streams integrate with the low-level control loops.
The most pressing work this week centers on developing interactive human humanoid planning for long horizon surgical assistance. This matters because it directly addresses the safety and efficacy of future medical robotics. RoboAssist explored this by creating a framework for interactive human-humanoid planning, suggesting a way for humans to guide robotic assistants during complex procedures.
This builds upon foundational work in scaling humanoid dexterous manipulation through camera-space ego-centric pretraining, which IronMind achieved by training models on visual data to improve how robots interact with their environment. A related effort involved using a biophysically detailed C. elegans circuit as a task-agnostic dynamical core for visually robust robot manipulation, providing a stable internal model for movement regardless of the specific task at hand.
Further refinement came from Discrete Forcing, which infused discrete guidance into continuous denoising processes to create few-step action experts capable of handling complex tasks efficiently. This approach complements Sparse Planner, a hybrid planner that uses a conditional variational autoencoder for efficient sampling when dealing with sparse environments. Finally, MVP-SLAM focused on multi-camera visual-inertial floorplan prior SLAM, which is crucial for building accurate spatial maps in dynamic settings where robots operate.
The most significant development today involves the work on RealSimReal loops, which addresses the critical gap between simulated and real-world robot performance. This framework attempts to bridge this divide by creating a loop where policies trained in simulation are adapted to perform reliably in physical environments.
FlowDPG introduced a deterministic policy gradient method applied to flow matching policies for real-world manipulation tasks. This means they developed a way for robots to learn how to physically move objects by using the flow matching concept, which is essentially guiding the learned policy toward a desired outcome in the real world. This work builds upon prior efforts in learning control strategies.
Another important piece of research focused on scale and selection within automatic harness evolution for visual-interface robot agents. They investigated what specific characteristics make these agents effective when they are automatically evolving their interaction methods based on visual input. This helps determine which evolutionary paths lead to better performance in complex visual tasks, linking directly to how the policy transfer framework might select the most robust control strategies.
Then there was HiWE, which builds a hierarchical world knowledge model enhanced with visual keypoint information for zero-shot 3D path planning. This system allows agents to plan movement through unseen three-dimensional spaces by using learned knowledge about the world structure and specific visual markers. This capability is crucial because it provides the necessary spatial understanding for the policy transfer loop to operate effectively in novel physical settings.
Finally, there is ECHO-G, which deals with embodied co-speech humanoid motion generation. This research focuses on creating realistic human-like movements for robots that are also capable of interacting verbally with humans. This adds a layer of complex interaction capability to the control systems being developed alongside the policy transfer and planning methods.
The most critical piece of work this morning is the development of ChunkTrust, which addresses a major hurdle in making robot policies robust when they encounter situations outside their initial training scope. This method adapts execution horizons for vision-language-action models by incorporating action-expert evidence to help the model make better decisions during runtime. It means that instead of failing completely when things get unexpected, the system can use this expert guidance to recover gracefully.
This is supported by research into learning from runtime feedback through failure-bank self evolution for vision-language-action models. This shows how these models can improve their behavior by actively learning from their own mistakes during operation. This iterative improvement builds on the idea that when instructions retrieve trajectories, we can diagnose and mitigate generalization failures in VLA models.
Another important piece is Magic-W0, a structured world action foundation model designed to serve as a foundation for physical intelligence. This model aims to provide a more coherent understanding of how actions relate to the physical world, which is crucial for complex manipulation tasks. This structural approach contrasts with purely reactive learning methods by providing a more organized framework for physical reasoning.
We also see work on magnetic based in-situ self three dimensional pose estimation for a modular soft tendon-driven continuum robot using IMU fusion. This helps robots understand their own position in real time, and this is complemented by active mapping of underwater litter using camera sonar fusion, which allows systems to build environmental maps while operating in challenging aquatic conditions.
The most important thing from today was the work on passive stiffness shaping in cable-suspended aerial manipulation because it directly addresses how robots can interact safely and flexibly with their environment. Researchers explored using movable compliant anchors to control the passive stiffness of these systems. This is crucial for handling delicate objects without damaging them during aerial tasks.
This investigation involved designing a system where compliant anchors could be moved to alter the mechanical properties of the cable suspension, and they found that this manipulation significantly influenced the robot's ability to maintain stable contact. This finding is important because it shows a pathway toward more intuitive physical interaction for aerial robots.
Another piece of work focused on TCBiRRT, which is a rapid motion planning method for tightly coupled dual-arm space manipulators using task-space random expansion. They tested this planner and found that it could generate collision-free trajectories much faster than existing methods, suggesting a significant speedup in planning complex movements.
This planning speedup complements the work on teaching vision language action models with spatial supervision and demonstration conditioning, which is XS-VLA. This latter model focuses on how tiny vision language action models learn to perform tasks by observing demonstrations and receiving spatial guidance.
WorldToken, which is a time-first sequence modeling approach for robotic imitation learning, was also examined today. This method attempts to capture the temporal dependencies in sequences better than standard models, which is key for making robots learn complex behaviors over time.
Finally, FORTE provided forecasting occupancy for spatiotemporal risk-aware planning in dynamic environments. This work helps robots anticipate potential hazards in changing settings before they happen, which builds upon the foundational concepts of safe control explored in neuro-symbolic predicate learning for semantic safe robot control.
The most critical piece of work today involves TACTIC, which tackles the problem of roadside LiDAR attacks by using a temporal and context-aware large language model for tactical planning. This matters because it addresses real-time security challenges in autonomous systems. The research explored how this LLM can plan appropriate responses when facing these adversarial inputs.
This planning work builds upon foundational models like EWAM, which focuses on emergent depth-wise specialization within a unified embodied model, moving from semantic understanding to visual foresight and finally to action. EWAM’s approach allows the system to gain a deeper grasp of the environment before taking steps. This is then complemented by SplineWAM, which introduces adaptive action horizons for world action models using B-spline representations.
SplineWAM refines how the model decides on actions by adapting its horizon based on these spline representations. This refinement is connected to RoboCoach, which uses world models as active coaches to improve compositional robot skills. RoboCoach aims to teach robots complex movements by letting them practice and receive guidance from simulated or real demonstrations within a world model framework.
Another significant piece of work is Experience-Driven Continual Learning of Terrain Traversability for Quadruped Robots. This focuses on improving how quadruped robots navigate uneven ground over time. This learning process allows the robots to adapt their movement strategies based on accumulated experience, which is a key step toward robust locomotion.
Finally, there is research into Identifiable Decomposition of Submovements in Human Hand Trajectories. This work seeks to break down complex human hand movements into smaller, identifiable components, which could inform better control policies for robotic manipulation.
The most significant development is the work on closing the planning and learning loop for robot control with learned world models. This addresses how robots can better navigate and interact in unpredictable settings over time rather than relying solely on pre-programmed instructions.
This concept connects to DiffWAM, which introduces a fast and efficient navigation world action model designed for this purpose. It suggests that by having this learned model, the robot can make better decisions about its immediate movements in real-time.
Another piece of work focuses on RL-guided PAC-NMPC for probabilistically safe perception-based navigation in unknown environments. This tackles the challenge of robots needing to perceive and move safely when they don't fully understand their surroundings. This is complemented by PhasePlan, which deals with ordered future-phase planning for robot brain models, suggesting a structured way for the robot to think about its long-term actions.
These navigation efforts are supported by research into rethinking legibility in social robot hallway navigation. This looks at how representing intent can affect human distraction during movement. This is contrasted by work on making waves with a membrane-coupled delta array for manipulating objects below the actuator spacing, which focuses on precise physical interaction capabilities.
Finally, tool-policy co-design for powder weighing in laboratory automation provides a practical example of applying these control strategies to specific tasks. This shows how learned models and planning can be tailored for real-world applications.
Today's papers
- The Planning Limits of Latent World Models This paper explores how far an AI can plan using a learned representation of the world. [paper]
- MotionWeave Learning Motion-Centered Future Dynamics for Vision-Language-Action Policies This work focuses on learning future movement patterns to guide robot actions based on vision and language. [paper]
- Beyond the Current Scene Event-Referential Grasping with Active View Selection This research shows how robots can select the best camera views to identify objects of interest in a scene. [paper] [episode]
- UniWAM Technical Report Unified Mobile Manipulation via Mixed-Stream World-Action Modeling and Manipulation Anchor Pose Supervision This report details a unified model for mobile manipulation that combines different streams of information. [paper]
- Multi-Link Safety Filtering for VLA Policies Around Moving Hazards This paper introduces a method to keep robot policies safe when dealing with moving obstacles in vision-language-action tasks. [paper]
- Ego4WAM What Matters When Scaling Egocentric Human Data for Robot Learning? This study investigates what aspects of egocentric human data are most important when training robots. [paper] [episode]
- Autonomous Human-Robot Interaction via Operator Imitation This work focuses on how robots can learn to interact with humans by imitating the actions of an operator. [paper] [episode]
- BIM Informed Visual SLAM for Construction Environments This paper uses building information models to improve visual simultaneous localization and mapping in construction settings. [paper] [episode]
- Learning to Build Autonomous Robotic Assembly of Stable Structures Without Predefined Plans This research shows how robots can build stable structures by learning assembly skills without a fixed plan. [paper] [episode]
- Grounded World Model Latent Planning with Language Goals This paper discusses using a world model that understands language goals to guide the robot's planning process. [paper] [episode]
- DexHoldem An Agentic Robotics Benchmark for Dexterous Manipulation in Texas Hold'em This is a benchmark designed to test how well agents can perform dexterous manipulation in a game like Texas Hold'em. [paper] [episode]
- Learning-Based Progressive Barrier Control for Robot Manipulators with Initial Errors Outside Prescribed Tracking Bounds This paper develops a method to control robot arms even when they start with errors outside the expected bounds. [paper] [episode]
- OGPO One-Step Generative Policy Optimization for Real-Time Robot Control This technique allows robots to generate and optimize policies quickly for real-time control tasks. [paper] [episode]
- SafeVLA-Bench A Benchmark for the Success-Safety Gap in Vision-Language-Action Models This paper creates a benchmark to measure the safety issues that arise when training vision-language models. [paper] [episode]
- DSDyn-VLA A Dual-Stream Dynamic Manipulation Framework with Motion Perception, Future Awareness, and Realtime Correction This framework uses two streams of data to improve dynamic manipulation by predicting motion and correcting errors in real time. [paper]
- ASENA Self-evolving Agents for Embodied Navigation This work presents self-improving agents that learn how to navigate environments through embodied experience. [paper]
- ReWAM Reciprocal World Action Models for Interactive Autonomous Driving This paper proposes models where vehicles can predict and react to each other's actions in interactive driving scenarios. [paper] [episode]
- A Biophysically Detailed C. elegans Circuit as a Task-Agnostic Dynamical Core for Visually Robust Robot Manipulation This research uses a biological circuit from worms to create a robust core for robot manipulation tasks. [paper] [episode]
- RoboAssist Interactive Human-Humanoid Planning for Long-Horizon Surgical Assistance This paper focuses on planning long sequences of actions for robots assisting in complex surgical procedures with humans. [paper] [episode]
- IronMind Scaling Humanoid Dexterous Manipulation via Camera-Space Ego-Centric Pretraining This work shows how pretraining humanoid manipulation skills using camera views centered on the robot improves performance. [paper]
- Discrete Forcing Infusing Discrete Guidance into Continuous Denoising for Few-Step Action Experts This method uses discrete guidance to help experts perform better in tasks requiring only a few steps. [paper]
- Sparse Planner A Hybrid Planner for Efficient Sampling via a Conditional Variational Autoencoder This paper introduces a hybrid planning approach that uses an autoencoder to sample efficiently. [paper]
- MVP-SLAM Multi-Camera Visual-Inertial Floorplan-Prior SLAM This system combines multiple cameras and inertial measurements to create accurate floor plans in SLAM. [paper]
- Towards Agile Vision-Based Multi-UAV Flight Revisiting State Estimation This work looks at improving state estimation methods for agile flight using multiple unmanned aerial vehicles. [paper] [episode]
- An Real-Sim-Real RSR Loop Framework for Generalizable Robotic Policy Transfer This framework creates a loop between real and simulated environments to make robot policies more generalizable. [paper] [episode]
- FlowDPG Deterministic Policy Gradient on Flow Matching Policies for Real-World Manipulation This method uses flow matching to create deterministic policies that work well in the real world. [paper] [episode]
- Scale and Selection What Makes Automatic Harness Evolution Work for Visual-Interface Robot Agents This paper examines how agents can automatically evolve their visual interfaces based on scale and selection criteria. [paper]
- HiWE Hierarchical World Knowledge Model with Visual Keypoint Enhancement for Zero-Shot 3D Path Planning This model uses a hierarchical knowledge structure to plan paths in 3D space without prior training data. [paper]
- Communication-Free Distributed Multi-Robot Task Allocation under Partial Observations Using Labeled Multi-Bernoulli Filtering This method allows robots to coordinate tasks without communication by using probabilistic filtering. [paper]
- General Performance Guarantee for Human Torque Estimation-Based Task-Agnostic Assistive Exoskeleton Control This work provides a mathematical guarantee for controlling exoskeletons based on estimated human torque. [paper] [episode]
- ECHO-G Embodied Co-speech Humanoid mOtion Generation This system generates realistic human motion for humanoid robots that can follow spoken commands. [paper] [episode]
- Text-to-3D Policy Fine-Grained Language-Behavior Alignment for Unseen Specification Generalization This paper shows how to align text descriptions with 3D robot policies for generalizing to new tasks. [paper]
- TO-mdiSPAs Topology Optimization of multi-directional Soft Pneumatic Actuators This research optimizes the shape of soft pneumatic actuators for various manipulation needs. [paper]
- From Local Whole-Body VLA Behaviors to Scene-Scale Aerial Manipulation This work bridges the gap between small, local robot behaviors and large, scene-scale aerial movements. [paper]
- ChunkTrust Adapting Execution Horizons for Robot Policies with Action-Expert Evidence This method adjusts how far a robot looks ahead based on evidence from expert actions during execution. [paper]
- Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models This system improves vision models by learning from failures encountered during runtime. [paper]
- Magic-W0 A Structured World-Action Foundation Model for Physical Intelligence This paper introduces a foundation model designed to understand and act in the physical world. [paper]
- Active Mapping of Underwater Litter Using Camera-Sonar Fusion This research uses camera and sonar data combined to map underwater litter efficiently. [paper]
- Magnetic based In-situ Self 3D Pose Estimation for a Modular Soft Tendon-Driven Continuum Robot via IMU-Fusion This system estimates the 3D pose of a soft robot using magnetic fields and inertial measurements. [paper] [episode]
- When Instructions Retrieve Trajectories Diagnosing and Mitigating Generalization Failures in VLA Models This paper analyzes why vision-language models fail when instructions lead to unexpected trajectories. [paper] [episode]
- Passive Stiffness Shaping in Cable-Suspended Aerial Manipulation via Movable Compliant Anchors This technique uses compliant anchors to control the stiffness of cables during aerial manipulation. [paper]
- TCBiRRT Rapid Motion Planning for Tightly Coupled Dual-arm Space Manipulator Using Task-space Random Expansion This method plans fast motions for dual robotic arms by randomly expanding the task space. [paper] [episode]
- XS-VLA Teaching Tiny Vision-Language-Action Models with Spatial Supervision and Demonstration Conditioning This work shows how to train small vision models using spatial guidance and demonstration data. [paper] [episode]
- WorldToken Time-First Sequence Modeling for Robotic Imitation Learning This model uses time as a key input to improve imitation learning for robots. [paper] [episode]
- Benchmarking EMlog Calibration for Autonomous Surface Vehicles This paper establishes a benchmark for calibrating the Extended Kalman Filter in autonomous surface vehicles. [paper]
- FORTE Forecasting Occupancy for Spatiotemporal Risk-Aware Planning in Dynamic Environments This system forecasts where obstacles will be and plans safely in dynamic environments. [paper]
- LIBERO-Agent Evaluating General-Purpose Agents for Direct Embodied Manipulation This is an evaluation framework to test the capabilities of general agents for direct physical manipulation. [paper] [episode]
- Neuro-Symbolic Predicate Learning for Semantic Safe Robot Control This method combines neural networks with symbolic logic to learn safe robot control predicates. [paper] [episode]
- Prediction is Better than Detection Traffic Congestion Control using Drones This paper suggests that predicting traffic flow is a better way to control drones than just detecting congestion. [paper] [episode]
- RoboCoach World Models as Active Coaches for Compositional Robot Skills This work uses world models to actively teach robots how to combine simple skills into complex tasks. [paper] [episode]
- Experience-Driven Continual Learning of Terrain Traversability for Quadruped Robots This system continuously learns how good a robot is at walking on different terrains based on experience. [paper]
- Toward Real-Time VLAs Stage-Aware Two-Step Flow Denoising and System-Level Evaluation This work focuses on making vision-language models fast enough for real time by using two denoising steps. [paper]
- SplineWAM Adaptive Action Horizons for World Action Models via B-Spline Representations This method uses spline representations to adapt the planning horizon of world action models. [paper] [episode]
- TACTIC Temporal and Context-Aware LLM Tactical Planning for Roadside LiDAR Attacks This paper uses large language models to plan tactical responses against attacks detected by roadside LiDAR. [paper]
- EWAM Emergent Depth-Wise Specialization in a Unified Embodied Model From Semantic Understanding through Visual Foresight to Action This model shows how one unified system can specialize in depth understanding and action. [paper] [episode]
- Identifiable Decomposition of Submovements in Human Hand Trajectories This research breaks down complex human hand movements into simpler, identifiable submovements. [paper] [episode]
- Tactile Curiosity Drives Robot Interaction This system uses tactile feedback to drive a robot's curiosity and improve its interaction with the world. [paper]
- Rethinking Legibility in Social Robot Hallway Navigation Impact of Intent Representation and Human Distraction This paper explores how robots should represent intent to navigate social spaces better. [paper] [episode]
- Making Waves A Membrane-Coupled Delta Array for Manipulating Objects Below the Actuator Spacing This research uses a membrane array to manipulate objects very close to the actuator. [paper]
- Beyond Policy Alignment Closing the Planning-Learning Loop for Robot Control with Learned World Models This work focuses on connecting learned world models back into the planning process for better control. [paper]
The papers
- ECHO-G: Embodied Co-speech Humanoid mOtion Generation — Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. [episode]
- ReWAM: Reciprocal World Action Models for Interactive Autonomous Driving — In interactive autonomous driving scenarios, existing World Action Models (WAMs) are limited because they typically model other agents as components of the world model rather than as decision-makers that fundamentally shape the ego agent's action. [episode]
- A Reachability-based Safety Certificate for Dynamical System Motion Policies — Dynamical systems (DS) are first-order autonomous systems used to define motion policies in robotics, but their local safety modifications often fail in complex environments. [episode]
- Yggdrasil: a Layer-First 3D Scene Graph for Real-Time Querying — YGGDRASIL introduces a novel 3D scene graph architecture specifically designed to be efficient for both generation and consumption, addressing the latency costs incurred by existing perception pipelines that must work around their scene graphs. [episode]
- LIBERO-Agent: Evaluating General-Purpose Agents for Direct Embodied Manipulation — General-purpose agents can plan, use tools, and revise their behavior from feedback, but it remains unclear whether these capabilities transfer from digital environments to embodied manipulation. [episode]
- Neuro-Symbolic Predicate Learning for Semantic Safe Robot Control — Neuro-Symbolic Predicate Learning for Semantic Safe Robot Control (NEUPRO) proposes a neuro-symbolic framework that represents safety specifications as interpretable first-order logic rules, allowing for flexible, interpretable safety reasoning directly from visual inputs. [episode]
- Collision-Aware Humanoid Whole-Body Control under Imperfect Tracking Targets — Humanoid robots often execute motion commands through whole-body controllers (WBCs) that track targets while maintaining balance and stability, but these controllers are typically blind to scene geometry, which can lead to collisions from imperfect target motions. [episode]
- SafeVLA-Bench: A Benchmark for the Success-Safety Gap in Vision-Language-Action Models — SafeVLA-Bench is a post-hoc safety-evaluation framework designed to measure the critical success–safety gap in Vision-Language-Action (VLA) models, addressing the limitation that binary task success metrics often hide dangerous or unsafe trajectory behaviors. [episode]
- RoboCoach: World Models as Active Coaches for Compositional Robot Skills — Long-horizon robot manipulation reuses skills across many task compositions, but improving these compositions with additional end-to-end demonstrations is costly. [episode]
- OGPO: One-Step Generative Policy Optimization for Real-Time Robot Control — This paper introduces Dispersive MeanFlow Policy Optimization (DMPO), a unified framework designed to enable true one-step generation for real-time robotic control, which is crucial for time-critical applications. [episode]
- TCBiRRT: Rapid Motion Planning for Tightly Coupled Dual-arm Space Manipulator Using Task-space Random Expansion — This paper introduces TCBiRRT, a novel Task-space Constrained Bidirectional Rapidly-exploring Random Tree algorithm designed to rapidly plan motion paths for tightly coupled dual-arm space manipulators under closed-chain constraints. [episode]
- FAST-Sync: Fast Group Synchronization for any Matrix Lie Group — Group synchronization (GS) is a fundamental problem in robotics and computer vision that involves estimating unknown group elements from noisy relative measurements, and this paper introduces Fast-Sync, a fast linear approximation method for GS suitable for initializing local man [episode]
- XS-VLA: Teaching Tiny Vision-Language-Action Models with Spatial Supervision and Demonstration Conditioning — Tiny Vision-Language-Action (VLA) models are crucial for real-time robotic control, but scaling them down often compromises essential capabilities like task-conditioned spatial grounding and coherent action generation. [episode]
- Learning to Fix: Optimisation-Aware Machine Learning for Accelerated Unit Commitment — Unit Commitment is a computationally demanding mixed-integer linear optimisation problem that requires many binary commitment decisions across a scheduling horizon, and this paper introduces an optimisation-aware framework to translate probabilistic predictions from machine learn [episode]
- Bridge-WA: Learning Action-Relevant World Dynamics for Robotic Manipulation — BRIDGE-WA is a lightweight world-action framework designed to predict where and how scenes will change during robotic action, thereby enabling more robust manipulation policies that are less sensitive to nuisance appearance factors. [episode]
- Data-Efficient Adaptation of a Driving VLA to Class 8 Trucks — Class 8 trucks present unique geometric and dynamic challenges that require specialized adaptation for vision-language-action (VLA) models trained on passenger vehicles, particularly in complex scenarios like construction zones and accident scenes. [episode]
- DiffWAM: A Fast and Efficient Navigation World Action Model — Pretrained video foundation models encode rich semantic and spatiotemporal priors for embodied navigation, yet converting these priors into UAV motion typically requires expensive future-video synthesis and geometric reconstruction. [episode]
- Tri-Info: Generalizable, Interpretable Failure Prediction for VLA Models via Information Theory — Vision-Language-Action (VLA) models are increasingly deployed in robotics, yet they remain black boxes whose physical interactions can cause irreversible harm, necessitating generalizable and interpretable failure detection. [episode]
- BatSLAM 2.0: Sequence-Verified Sonar Place Recognition in a Robust Pose Graph — Echolocating bats navigate dark and cluttered spaces using echolocation, and this research introduces BatSLAM 2.0, a novel sonar-only SLAM system that achieves robust topological map creation by addressing the ambiguity of sonar place recognition through sequence verification and [episode]
- Globally Certified Invariant-Ellipsoid Control from Data — This letter develops a data-based method for designing state feedback for discrete-time linear systems under bounded disturbances by optimizing an invariant ellipsoid to minimize a trace-based measure of its output enclosure. [episode]
- Enforcing Human-like Kinematics in Dexterous Piano Playing via Adversarial Posture Regularization — This paper introduces Adversarial Posture Regularization (APR), a novel bimanual reinforcement learning framework designed to enforce human-like kinematics in high-degree-of-freedom dexterous piano playing. [episode]
- SplineWAM: Adaptive Action Horizons for World Action Models via B-Spline Representations — World action models (WAMs) are large embodied policies that jointly predict future video and actions, and SplineWAM introduces an adaptive action representation using cubic B-splines to allow these models to handle varying motion complexities efficiently. [episode]
- Bimanual Robot Manipulation via Multi-Agent In-Context Learning — Language Models (LLMs) are emerging as powerful reasoning engines for embodied control, and this paper introduces BiCICLe, the first framework enabling standard LLMs to perform few-shot bimanual manipulation without fine-tuning. [episode]
- Beyond the Current Scene: Event-Referential Grasping with Active View Selection — A robot must be able to carry out later requests that refer back to past interactions, even when those objects are no longer visible, and this paper presents BeyondCSe, a zero-shot grasping system that achieves this by combining event history with active view selection. [episode]
- RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning — RoboHarness is a unified framework designed for long-horizon robotic tasks that require diverse capabilities, addressing the limitations of existing planning methods which assume homogeneous skills and fixed applicability. [episode]
- Rethinking Legibility in Social Robot Hallway Navigation: Impact of Intent Representation and Human Distraction — Legibility in social robot navigation is crucial for ensuring human safety and smooth coordination in dynamic, constrained environments where human attention can be divided. [episode]
- On the Input Sensitivity of METANET Models and the Robustness of Dynamic Calibration — A calibrated METANET model can amplify small additive perturbations to boundary conditions along the corridor, causing the simulated state to diverge from the nominal baseline, which compromises its utility for counterfactual analysis. [episode]
- A Numerical Investigation of Indirect Adaptive Predictive Control with Structure-Informed Nonlinear Regressors — This research introduces Adaptive Behavioral Predictive Control (ABPC), a novel, kernel-based indirect adaptive controller designed for online operation on streaming data. [episode]
- H-WM: Robotic Task and Motion Planning Guided by Hierarchical World Model — The Hierarchical World Model (H-WM) is a novel framework designed to provide informative, grounded, and long-horizon–robust guidance for Vision–Language–Action (VLA) models in robotic task and motion planning. [episode]
- EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action — Based on the provided excerpts, I have meticulously synthesized a detailed summary of the paper "EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model." Here is the comprehensive analysis: * The paper introduces EWAM, an action-centric, unified embodied model desig [episode]
- Dense Temporal Motion Retargeting for Legged Robots — Legged robots can learn expressive whole-body skills from human and animal motions, but adapting these motions to a robot's specific dynamics requires careful adjustment of timing and control, especially for dynamic movements like jumps. [episode]
- Arm2Air: Cross-Embodiment Skeleton Transfer for 3D Relay Formation — Arm2Air addresses the complex challenge of 3D UAV relay network formation by transferring obstacle-avoidance skeletons from robot arms to UAV placement through cross-embodiment transfer. [episode]
- FlowDPG: Deterministic Policy Gradient on Flow Matching Policies for Real-World Manipulation — FlowDPG is a novel DDPG-style method specifically designed for flow matching policies, addressing the computational and numerical fragility associated with backpropagating through time (BPTT) along multi-step ODEs. [episode]
- Identifiable Decomposition of Submovements in Human Hand Trajectories — Submovement-Identifiable Decomposition (Sub-ID) proposes a novel method to decompose human hand trajectories into discrete submovements by using spatiotemporal kernel correlation as an identifiability criterion, allowing it to recover ground-truth parameter distributions in compl [episode]
- ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors — EXPERTGEN is a framework designed to automate expert policy learning in simulation to enable scalable sim-to-real transfer by learning generalizable and robust behavior cloning policies from imperfect demonstrations. [episode]
- RoboAssist: Interactive Human-Humanoid Planning for Long-Horizon Surgical Assistance — RoboAssist presents an agent-based framework for interactive human–humanoid planning designed to enable long-horizon surgical assistance by coordinating robot actions with evolving human activities while ensuring safety across planning and execution. [episode]
- Timing Sensitivity in Actuated Traffic Signal Control: A simulation study on one urban network — A comparison between adaptive and fixed-time traffic signals can change when only their timing settings change, and this study examines that sensitivity in 240 simulation runs on an OpenStreetMap-derived urban network to determine if performance differences can be attributed to a [episode]
- Condition-Based Maintenance of Degrading Assets underIntermittent Accessibility — Many maintenance models implicitly assume that maintenance can be performed whenever intervention is warranted, but in practice, environmental uncertainty can make maintenance opportunities intermittent and dynamically evolving. [episode]
- A Biophysically Detailed C. elegans Circuit as a Task-Agnostic Dynamical Core for Visually Robust Robot Manipulation — A biophysically detailed Caenorhabditis elegans sensorimotor circuit is embedded as a task-agnostic dynamical core for robot manipulation, suggesting that visual robustness can be inherited from evolved circuit dynamics rather than learned by task-specific controllers. [episode]
- Learning to Build: Autonomous Robotic Assembly of Stable Structures Without Predefined Plans — This paper presents a novel autonomous robotic assembly framework designed to construct stable structures without relying on predefined architectural blueprints, addressing the limitations of rigid planning in dynamic construction environments. [episode]
- Grounded World Model: Latent Planning with Language Goals — This work proposes a Grounded World Model (GWM) to enhance planning in Model Predictive Control (MPC) by leveraging a vision-language-aligned latent space, offering a novel approach to semantic generalization for visuomotor agents. [episode]
- Eigenspace-Based Clustering for Personalized System Identification — This paper proposes a novel, one-shot, training-free clustering method for personalized federated system identification that leverages the structural information within locally observed data to identify systems with shared underlying dynamics. [episode]
- When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models — Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action. [episode]
- Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling — Video generation models (VGMs) offer strong spatiotemporal priors for embodied observation–action modeling, but existing joint-space action vectors lack explicit image-space structure and vary across embodiments, making it challenging to directly leverage the rich spatiotempora [episode]
- Optimize, Learn, Refine: Whole-Body Grasping and Pick-and-Throw with a Spiral Soft Robot — Soft continuum robots can exploit distributed compliance for whole-body manipulation, but synthesizing behavior through changing contacts remains difficult. [episode]
- Computing Scaled Relative Graphs of Discrete-Time LTI Systems: A Frequency-Domain Approach — The scaled relative graph (SRG) analysis provides a powerful geometric tool for characterizing operators, and this paper characterizes the SRG closure of causal, stable square discrete-time linear time-invariant (LTI) systems by relating it to the convex hull of transformed frequ [episode]
- Autonomous Human-Robot Interaction via Operator Imitation — This paper proposes a novel framework for creating autonomous human-robot interactions by training a model to imitate expert operator data, aiming to enable robots to perform expressive, mood-varying behaviors comparable to those of human operators. [episode]
- WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning — As a fastidious and diligent researcher, I have meticulously analyzed both provided texts. [episode]
- Skill-Based AI Agents for Power-System Studies — This paper describes a skill-based agentic framework for power-system studies using Model Context Protocol (MCP)-connected engineering tools, demonstrating that agentic systems can greatly accelerate power system dynamic simulation processes for transmission planning studies. [episode]
- Memorize, Adapt, Ignore: Diagnosing Robot Learning Mechanisms under Training Data Variation — Training data variation, whether through designing a domain randomization (DR) scheme in simulation or curating demonstrations for imitation learning, is a primary lever for improving the robustness of robotic manipulation policies. [episode]
- Social-WM: Safety-Aware Latent World Models for Robot Social Navigation — Safe social navigation requires a robot to anticipate not only the future consequences of its actions, but also whether a nominal action can actually be executed under surrounding physical and social constraints. [episode]
- A Frequency Domain Approach to Bounding Riccati Equation Perturbations — The discrete algebraic Riccati equation (DARE) is used to solve for optimal feedback gain in linear-quadratic regulator (LQR) control, and this paper presents a new frequency domain approach to derive an explicit bound on the difference between DARE solutions when the LQR data is [episode]
- Observability Analysis and Online Calibration of Visual-Inertial-Wheel Odometry for 4WIS4WID Mobile Robots — In this work, a visual-inertial-wheel odometry (VIWO) framework with online calibration is developed for four-wheel independently steered and driven (4WIS4WID) mobile robots to address the increased kinematic complexity and calibration challenges introduced by this platform. [episode]
- General Performance Guarantee for Human Torque Estimation-Based Task-Agnostic Assistive Exoskeleton Control — Accurate human torque estimation is crucial for enabling task-agnostic control in robotic exoskeleton systems because estimation errors can cause mismatches between robot assistance and human intention, degrading controllability and task performance. [episode]
- RL-Guided PAC-NMPC for Probabilistically-Safe Perception-Based Navigation in Unknown Environments — In this paper, an approach combining stochastic nonlinear model predictive control (SNMPC) and reinforcement learning (RL) enables probabilistically safe perception-based navigation in unknown environments. [episode]
- Onboard Vision and MPC Navigation for Underwater Robots: An Open BlueROV2 Platform for Multi-Robot Experiments & Docking — Autonomous underwater robots require robust perception, estimation and control to operate in confined environments. This paper presents an open-source BlueROV2 platform combining onboard vision with nonlinear Model Predictive Control (NMPC) for autonomous navigation and docking. [episode]
- Forward-Invariant Policy Classes for Safe Reinforcement Learning in Multicopter Control — Safety is embedded directly into the policy class by constructing a finite library of feedback controllers sharing a common Lyapunov certificate that establishes forward invariance of a prescribed admissible set under arbitrary switching. [episode]
- Towards Agile Vision-Based Multi-UAV Flight: Revisiting State Estimation — Agile multi-UAV flight requires accurate and low-latency onboard estimation of neighboring UAVs' kinematic states for critical tasks like collision avoidance and motion coordination. [episode]
- L1-MPPI: L1 Adaptive Model Predictive Path Integral for Agile UAV Control — L1-MPPI proposes an L1 Adaptive Model Predictive Path Integral framework that cascades L1 adaptive control with MPPI to enhance trajectory tracking for high-speed UAVs under model uncertainties and external disturbances. [episode]
- Ego4WAM: What Matters When Scaling Egocentric Human Data for Robot Learning? — Egocentric human data provides a scalable source of experience for robot learning, but varies substantially in human-robot alignment, behavioral coverage, and available supervision. [episode]
- Demonstrating a Robust Walking Algorithm for Underactuated Bipedal Robots in Non-flat, Non-stationary Environments — This work presents an innovative control algorithm designed to significantly enhance the mobility of underactuated bipedal robots, specifically addressing challenges in navigating non-flat, non-stationary environments where foot support opportunities are constrained. [episode]
- Parameter-Robust Sensorless Control of IPMSM Drives With Adaptive Flux Observer — To address parameter sensitivity commonly found in interior permanent magnet synchronous motor (IPMSM) sensorless control, this paper proposes a parameter-robust control framework by extending an adaptive flux observer from surface-mounted PMSM to salient-pole machines. [episode]
- An Real-Sim-Real (RSR) Loop Framework for Generalizable Robotic Policy Transfer — This paper introduces a novel Real-Sim-Real (RSR) loop framework designed to bridge the critical sim-to-real gap in robotics by iteratively refining simulation parameters using real-world data and simultaneously training policies. [episode]
- Systematically Exploring the Capabilities of GPT-6 Astra as Embodied Policies — As a diligent researcher, I have meticulously analyzed both provided texts concerning GPT-6 Astra's capabilities as an embodied policy. [episode]
- Draft: A Parametric Tool for Robot Design Exploration — Robot performance is often limited by the cost of iterating on morphology and control together, since every computer-aided design (CAD) change has to be carried into a simulation-ready model before control work begins. [episode]
- Duration-Aware Ramp Adequacy Screening — Ramp products are widely used in regional electricity markets to procure intertemporal flexibility in anticipation of net demand changes, but their design often lacks clear specification regarding ramping duration, potentially leading to infeasible dispatch solutions. [episode]
- Low-Rank and Lifted Semidefinite Programming for Mixed-Integer Polynomial Power Grid Optimization — Can a local solver return the guaranteed globally optimal solution to a nonconvex mixed-integer polynomial power grid optimization problem? By mixing low-rank semidefinite programming (SDP) and moment-based lifting in the Lasserre hierarchy, this paper provides anecdotal evidence [episode]
- BIM Informed Visual SLAM for Construction Environments — This research introduces ivS-Graphs, a novel visual Simultaneous Localization and Mapping (SLAM) system designed to monitor building construction sites by integrating structural priors derived from Building Information Models (BIM). [episode]
- Drone Soccer: Learning to Manipulate with Multicopter Downwash — Aerial manipulation performance can be impacted by “downwash,” the airflow produced by propellers, and this work explores using downwash actively as a tool during manipulation. [episode]
- Surrogate Modeling of Interconnector Flows: A Machine Learning Alternative to Full-Scale Power System Simulations with Application to Cross-Border Electricity Exchange — This paper proposes a machine-learning (ML) surrogate framework designed to generate synthetic, interconnector-level flow time series from readily available nodal data (demand and renewable generation). [episode]
- CADeT: Causal-Aware Deformation Transmission for Indirect Robotic Manipulation of Soft Tissue — Indirect manipulation of deep-seated deformable anatomy inaccessible to robots is challenging because intervening tissues spatially filter deformation transmission, leading to observational ambiguity between modes. [episode]
- A Goal-Oriented Approach for Active Object Detection with Exploration-Exploitation Balance — Active object detection is a critical capability for autonomous robots tasked with executing operations in unknown environments, such as manufacturing tasks, where identifying objects of interest through controlled camera movements is essential. [episode]
- PneuTac: Tactile Manipulation with Soft Pneumatic Robots via Unified MPM-Gaussian Splatting Simulation — Soft robots and tactile sensors have demonstrated great potential in delicate manipulation tasks, but learning tactile manipulation with compliant robots has been challenging due to the lack of efficient simulation tools. [episode]
- Prediction is Better than Detection: Traffic Congestion Control using Drones — A central question in deploying teams of mobile robots for persistent monitoring is how task performance scales with fleet size, and whether this scaling holds once sensing drives downstream action rather than mere observation. [episode]
- Combinatorial Optimization of Robotic Hand Kinematic Structures Using a Potential-Dexterity-Based QUBO Formulation — This research presents a quadratic unconstrained binary optimization (QUBO)-based formulation framework for robot design optimization, specifically applied to kinematic structures like robotic hands. [episode]
- Diffusion-2BC: Hybrid Diffusion and Regression Training for Offline Behavior Cloning in Autonomous Driving — Diffusion-2BC presents a hybrid training architecture that combines a diffusion denoising objective with an auxiliary deterministic behavior-cloning loss to improve closed-loop reliability in offline behavior cloning for autonomous driving. [episode]
- GestAdapt: Workspace-Conditioned Co-Speech Gesture Generation for Humanoid Robots — Co-speech gestures for robots must adapt not only to speech and embodiment, but also to the workspace available for performing the motion. [episode]
- ART-TEB: Adaptive Trajectory Planning for Mobile Robots in Cluttered Environments — This paper introduces an adaptive trajectory refinement algorithm designed to enhance the reliability and efficiency of Timed Elastic Band (TEB) planning, particularly for mobile robots navigating challenging, narrow passages in cluttered environments. [episode]
- Learning-Based Progressive Barrier Control for Robot Manipulators with Initial Errors Outside Prescribed Tracking Bounds — This paper presents a reinforcement learning-based neuroadaptive control framework designed for robotic manipulators operating under deferred constraints, addressing the limitations of traditional methods that struggle with initial constraint violations and system uncertainties. [episode]
- VHF Reconfigurable Intelligent Surfaces for Meteor Burst Communication — Meteor Burst Communication (MBC) utilizes transient ionized trails left by meteors to reflect Very High Frequency (VHF) signals, enabling long-range, beyond-line-of-sight communication without reliance on terrestrial or satellite infrastructure. [episode]
- Enhanced SIRRT*: A Structure-Aware RRT* for 2D Path Planning with Hybrid Smoothing and Bidirectional Rewiring — Enhanced SIRRT (E-SIRRT) is an advanced structure-aware motion planner that builds upon the Skeletonization-Informed RRT (SIRRT) framework by introducing hybrid path smoothing and bidirectional rewiring to improve initial solution quality and tree connectivity. [episode]
- Sharing the Gains of Aggregation: Cooperative Imbalance Cost Allocation — Pooling imperfectly correlated residuals nets consumers’ imbalances and reduces the portfolio’s total imbalance cost, but raises an allocation question: how should these savings be divided among heterogeneous consumers? This paper formulates this as a cooperative game, the im [episode]
- Non-Invasive Inspection of Water Canals Using Dronar — Open concrete canals play a vital role in water transportation, serving as primary water infrastructure for millions of people across the Phoenix, Arizona metro area. [episode]
- AURA: Asymptotically Optimal Uncertainty-Robust Replanning Algorithm for Kinodynamic Systems — AURA is presented as an asymptotically optimal meta-planner framework designed to enhance both path quality and tracking performance for kinodynamic systems operating under motion uncertainty. [episode]
- Behavioral Persistence and Incomplete Functional Transfer of Co-evolved Communication in Evolutionary Robotics — This work evaluates whether a co-evolved communication protocol can be directly transferred from a 2D simulation to a 3D physical environment without retraining the network weights, revealing that while some behavioral elements persist, full functional transfer is contingent upon [episode]
- DexHoldem: An Agentic Robotics Benchmark for Dexterous Manipulation in Texas Hold'em — DexHoldem introduces a novel system-level benchmark for evaluating embodied agents that couples dexterous manipulation skills with agentic perception within a real-world Texas Hold'em tabletop setting. [episode]
- Probability-Based Collision Risk Evaluation of Trajectories for Optimal Control Problems with Moving Obstacles — This paper presents a method for approximating occupancy distributions using smooth B-spline surfaces to enable time-dependent quantification of collision risk within optimal control problems. [episode]
- Multidisciplinary Design Optimization for Wave-Driven Desalination Systems — This scientific paper presents a holistic, multidisciplinary design optimization (MDO) framework for wave-driven desalination systems, addressing the high costs that currently hinder widespread adoption of this innovative technology. [episode]
- Data-Driven Communication Topology and Distributed Controller Synthesis: Control-Aware and Co-Design — In distributed control schemes, designing an optimal communication topology that guarantees controller existence while balancing communication costs and control performance is crucial for practical implementation. [episode]
- Assessing Modeling Fidelity for Long-Term Battery Energy Storage Planning: Operation, Degradation, and Temporal Representation — Long-term battery energy storage system (BESS) planning often relies on simplified representations that may obscure critical long-term effects, making it essential to assess which modeling fidelity—such as degradation mechanisms, health updates, or temporal resolutions—is nec [episode]
- Magnetic based In-situ Self 3D Pose Estimation for a Modular Soft Tendon-Driven Continuum Robot via IMU-Fusion — Continuum robots are well suited for gentle manipulation because of their inherent compliance and ability to adapt to complex environments, but their continuously deformable structure makes accurate configuration estimation challenging, particularly when external vision systems a [episode]
- Centralized Multi-UAV Exploration and 3D Reconstruction Using Single-UAV Planners —
- STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction —
- GPU-Accelerated Path-Dependent Marginal Information Gain for Autonomous Exploration —
- DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents —
- TACTIC: Temporal and Context-Aware LLM Tactical Planning for Roadside LiDAR Attacks —
- Fiatlux: A Long-Horizon Benchmark for Humanoid Ladder Climbing and Light-Bulb Replacement —
- SynIL: Leveraging Synergy for Offline Imitation Learning from Imperfect Demonstration Datasets —
- A Two-Echelon Covering Tour Vehicle Routing Problem with Drones for Post-Disaster Relief —
- TALK-Dem: Benchmarking Embodied Task Planning under Dementia-Associated Communication Patterns —
- Embodiment-aware control by inference over the operator: a simulation study —
- BIND: Binding 3D Robot Actions to 2D Image Features —
- TrafficSignBench: Rule-Centric Closed-Loop Evaluation of Traffic-Sign Compliance in Autonomous Driving —
- What to Attend, What to Keep: Skill-Conditioned Visuotactile Representation with Progress-Guided Event Memory —
- Correcting WHERE, Preserving HOW: Compositional Generalization for Vision-Language-Action Models via Referential Guidance —
- TERRA: Terrain-Aware Reconstruction, Retargeting and Control for Musculoskeletal Locomotion —
- Residual Wrench Certification and Margin-Aware Control Synthesis for Aerial Physical Interaction —
- CEER2: Directional and Tunable End-Effector and Root Compliance for Humanoid Loco-Manipulation —
- Path-Following Control and Terramechanics Analysis for Planetary Rovers Under Wheel-to-Wheel Traction Asymmetry —
- FlapKAD: A Simulation Dataset of Coupled Wing Kinematics and Aerodynamic Dynamics for Flapping-Wing Aerial Vehicles —
- On the NP-Hardness of Unconstrained Static Output Feedback Stabilization —
- Bounded Channel-Adaptive Spectral Learning for Forward-Consistent Inverse Flapping-Wing Aerodynamics —
- AIfred: Augmented Learning through Functional Robotic Embodiment at the Desk —
- HALO: Heterogeneous Allocation Via Localized Observations for the Vehicle Routing Problem —
- PhaseSync-Exo: Human Clock Anchored Reference Adaptation for Dynamic Gait Tracking —
- Coral Grow-out Robotic Assessment System (CGRAS): Scaling Coral Recruit Monitoring Through Robotics and Computer Vision —
- Locomotion-Grounded Humanoid Soccer: Task-Gated Reinforcement Learning of a Multi-Directional Kicking Library —
- Online Evolution Strategy for Flow-Matching VLA Policies via Self-Supervised Trajectory Distribution Optimization —
- Plan-Conditioned Imitation for Robust Object Retrieval under Self-Occlusion in Dense Clutter —
- Efficient Multi-Modal Planning with Reward-Guided Preference Optimization for Autonomous Driving —
- DODGER: Safety-Guided Reinforcement Learning for Robot Navigation Among Dynamic Obstacles —
- Benchmarking and Enhancing Skill-Level Memory for Partially Observable Robotic Manipulation —
- PRICE the Action Chunks: Physical Relational Credit Assignment for Embodied Reinforcement Learning —
- EmbodiRSI: Recursive Self-Improvement for Data-Efficient Robot Adaptation —
- Local-Minimum Escaper: Programmatic Subgoal Generation for Robust Navigation in Unknown Environments —
- DrivingBench: Can Vision-Language Models Drive a Toyota Corolla? —
- Video2SwimFish: An Automated Pipeline for Reconstructing Controllable Fish Models and Biological Locomotion from Real Fish Videos —
- SimEX: Simulation-Integrated Robotics AutoResearch —
- Sparse-WAM: Accelerating World Action Models via Action-Guided Sparse Imagination —
- Cue the Flow: Steering Flow-Matching Policies for Open-World Delivery Manipulation —
- NEXUS: Perceptive Whole-Body Control for Terrain-Adaptive Teleoperation —
- Function beyond Form: Functional Correspondence for Cross-Embodiment Dexterous Grasp Generation —
- OccluDex: Hierarchical 3D Visuo-Tactile Representation Learning for Egocentric Dexterous Manipulation under Self-Occlusion —
- Make Code as Policy Great Again: Frontier Agents Write, Call, and Evolve Robot Tools —
- Looking Back to Move Forward: Temporal Verification for Generative Robot Policies —
- CLIPPER Beyond Shortlisting: Auditable Decision Support for Changing Municipal Micromobility Policies —
- SteerQuant: Steering Quantization Error with Action-Guided Scaling in World-Action Models —
- DiFF: Doppler-informed Flow Matching for Human Motion Flow —
- Occlusion-Aware, Quasi-Static, Stability-Oriented Trajectory Planning on Uneven Terrain —
- LBDU-VIO: Learned Bias Dynamics and Uncertainty for Visual-Inertial Odometry with Unreliable Vision —
- Drape-Compatible Tool-Tip Localization for Hand-Held Laparoscopic Instruments via UWB Carrier-Phase Ranging and Trocar-Constrained Geometry —
- Blackout vs. Freeze: Analyzing Physical Failure Modes of VLAs under Camera Faults —
- Linear Recurrent Memory Suffices to Distil a World-Model Policy for Robot Air Hockey —
- Concurrent Semantic Search and Mission Execution for LTL Missions in Unknown Environments —
- Exploiting Vulnerabilities: Universal Adversarial Attacks on Vision-Language-Action Models in Robotics —
- LocoWM: High-Precision Locomotion through World-Model-Guided Residual Adaptation —
- DSDyn-VLA: A Dual-Stream Dynamic Manipulation Framework with Motion Perception, Future Awareness, and Realtime Correction —
- Benchmarking EMlog Calibration for Autonomous Surface Vehicles —
- ASENA: Self-evolving Agents for Embodied Navigation —
- Fast and Sample Efficient Safety Verification via Extreme Learning Machine —
- The Planning Limits of Latent World Models —
- Scale and Selection: What Makes Automatic Harness Evolution Work for Visual-Interface Robot Agents —
- FORTE: Forecasting Occupancy for Spatiotemporal Risk-Aware Planning in Dynamic Environments —
- HiWE: Hierarchical World Knowledge Model with Visual Keypoint Enhancement for Zero-Shot 3D Path Planning —
- MotionWeave: Learning Motion-Centered Future Dynamics for Vision-Language-Action Policies —
- UniWAM Technical Report: Unified Mobile Manipulation via Mixed-Stream World-Action Modeling and Manipulation Anchor Pose Supervision —
- IronMind: Scaling Humanoid Dexterous Manipulation via Camera-Space Ego-Centric Pretraining —
- Communication-Free Distributed Multi-Robot Task Allocation under Partial Observations Using Labeled Multi-Bernoulli Filtering —
- Discrete Forcing: Infusing Discrete Guidance into Continuous Denoising for Few-Step Action Experts —
- Sparse Planner: A Hybrid Planner for Efficient Sampling via a Conditional Variational Autoencoder —
- MVP-SLAM: Multi-Camera Visual-Inertial Floorplan-Prior SLAM —
- Text-to-3D Policy: Fine-Grained Language-Behavior Alignment for Unseen Specification Generalization —
- Circuit-Based Dispersion Analysis of Periodic Cross-Shaped Unit Cells with Dirac Characteristics —
- TO-mdiSPAs: Topology Optimization of multi-directional Soft Pneumatic Actuators —
- Making Waves: A Membrane-Coupled Delta Array for Manipulating Objects Below the Actuator Spacing —
- From Local Whole-Body VLA Behaviors to Scene-Scale Aerial Manipulation —
- STELLA: A 16nm Spatio-Temporal Elastic Low-Latency CGRA for Multi-Stage Pipelined Applications —
- Transient Triggering Grid-Forming Synchronization Control Under Voltage and Frequency Dips —
- Beyond Policy Alignment: Closing the Planning-Learning Loop for Robot Control with Learned World Models —
- ChunkTrust: Adapting Execution Horizons for Robot Policies with Action-Expert Evidence —
- Experience-Driven Continual Learning of Terrain Traversability for Quadruped Robots —
- Tool-Policy Co-Design for Powder Weighing in Laboratory Automation —
- Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models —
- Toward Real-Time VLAs: Stage-Aware Two-Step Flow Denoising and System-Level Evaluation —
- Flatness-based Neural Network Control of DC-DC Converters —
- Concept and Rationale for Stratospheric Balloon-Based Laser Debris Removal —
- Magic-W0: A Structured World-Action Foundation Model for Physical Intelligence —
- Exact Noise Limits for Bounded Scalar Stabilization Experiments —
- Active Mapping of Underwater Litter Using Camera-Sonar Fusion —
- PhasePlan: Ordered Future-Phase Planning for Robot Brain Models —
- Multi-Link Safety Filtering for VLA Policies Around Moving Hazards —
- Learning Higher Order DC-DC Converter Control from Inversion —
- Passive Stiffness Shaping in Cable-Suspended Aerial Manipulation via Movable Compliant Anchors —
- Tactile Curiosity Drives Robot Interaction —
- A Modular State-Machine Based Event PID Controller —
- PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors —
Important terms
- Latent World Models
- These are abstract representations that robots use to understand their environment and predict future dynamics. Researchers are testing their practical limits to see when these models fail in real-world tasks.
- DexHoldem
- This is a new benchmark for testing how well robots can perform intricate, dexterous manipulation in complex scenarios like Texas Hold'em. It sets a high bar for robotic dexterity.
- OGPO
- This method uses one-step generative policy optimization to generate control actions directly during robot execution. This is key for fast responses needed in dynamic situations.
- ChunkTrust
- This technique helps robot policies recover gracefully when they encounter unexpected situations outside their training scope by using expert guidance during runtime.