Robotics papers — 2026-09-29
Today's work centers on understanding the gap between what we intend for a large language model to do and what it actually executes when controlling physical robots, which is crucial because this gap directly impacts how reliably we can deploy autonomous systems in real-world settings. We explored how to make these LLM-based robots more robust by focusing on better planning methods for physical tasks.
One key area involved developing dynamic buffers for cost-efficient planning when rearranging objects on a tabletop using stacking techniques, which is about figuring out the cheapest way to move things around. This work connects to learning whole-body control methods like FastGrasp, which focuses on teaching mobile manipulators how to perform fast and dexterous grasping.
Another piece of the puzzle is RoboAlign-R1, which uses distilled multimodal reward alignment to improve robot video world models. This helps the robot better understand its visual environment when it needs to make decisions about action execution.
We also looked at adaptive action execution for world action models, which addresses when a system should trust its own imagination versus when it needs external guidance. This contrasts with work on recursive self-improvement in robotics, where we examined what stops agents from endlessly discovering new skills through agentic skill discovery rounds.
Finally, we touched upon robot manipulation using GPT-6-Astra to see how body knowledge and experience reuse translate into emergent skills during simtoreal transfer. This all feeds into the broader challenge of timed rule-based supervision for end-to-end autonomous parking policies, which shows how we can impose structure on complex behaviors.
The most crucial development this morning involves PHIRL, which tackles aligning learned rewards with task progress in inverse reinforcement learning. This problem is vital because it allows agents to learn optimal behaviors from demonstrations rather than just trial and error. This work attempts to bridge the gap between what an agent learns through experience and what the desired outcome actually requires.
A related effort is DS-VLA, which introduces a dendritic-inspired vision-language-action model designed for robust action control. This aims to make these complex models more reliable when interacting with the real world. This is significant because it moves beyond simple imitation by incorporating visual and linguistic understanding directly into the action planning loop.
Then there is RECAST, which focuses on recasting vision-language semantics into an actionable cost map specifically for robot navigation. This helps robots understand what costs are associated with different areas of a scene. This builds upon the need for better semantic grounding in embodied systems.
We also see work on Affordance-Conditioned Decision Making, which bridges the semantic-spatial gap in zero-shot cross-floor vision and language navigation by focusing on how physical affordances guide decisions across different environments. This is important for making navigation strategies more generalizable.
DRAM focuses on delta-rule recurrent associative memory to improve robot manipulation policies. This suggests a way to store and recall relevant past experiences efficiently during complex tasks. This method seeks to enhance the memory capabilities of embodied agents.
The most critical piece of work today involves figuring out how to distill the complex behavior of foundation models into policies that robots can actually use in the real world because this unlocks deployable intelligence. We saw some promising initial attempts with VPTwin, which focuses on real-sim-real video prediction for robotic manipulation planning. This suggests a way to bridge the gap between simulation and physical action.
Then there is ProcVLM, which tackles learning procedure-grounded progress rewards for robotic manipulation. This means it teaches robots what to do by rewarding them based on how well they follow a specific procedure during the task. This builds upon that by GT-VLA, which introduces target-conditioned trace guidance for generalizable robotic manipulation, aiming to make the learned skills work across different scenarios.
A more recent direction explored federated subspace guided vision-language-action policy distillation for non independent and identically distributed multi robot manipulation. This is important because it addresses the challenge of making policies work when robots have different experiences. This connects to how we are trying to stabilize online learning with neural ordinary differential equations, a method that provides Lyapunov guarantees for stable learning as the system evolves.
Finally, there was work on think fast plan selectively through adaptive deliberation for efficient data-driven model predictive control. This is about making robots decide what to do quickly and intelligently based on the data they have. This seems like it could feed into methods like Schur-Neural Kalman filters that learn consistent corrections to traditional extended Kalman filters.
The most critical development this morning is the work on what matters for latent actions in robot learning because understanding these underlying representations is key to creating policies that generalize across different tasks. The research explored what truly drives successful manipulation policies, and one line of inquiry focused on identifying these crucial latent actions.
A study investigated what matters for viewpoint-generalizable policies in visual imitation learning, which suggests there are specific aspects of the visual input that are most important for a robot to learn from when it needs to perform a task regardless of its exact viewing angle. This finding is significant because it moves beyond simple pixel matching toward understanding the semantic information that truly guides action selection.
Another piece of work tackled the representation for robust robot manipulation, focusing on copper-policy, which seems to be trying to distill the essential visual features needed for reliable physical interaction with objects. This is important because if you can capture the right representation, you can build a system that doesn't fail when things move slightly out of position.
Then there was collisiongat, which introduced a controller-agnostic one-step collision screening method for multi-agent motion. This aims to prevent robots from bumping into each other while moving around. This is a practical step toward safer deployment in shared environments.
SPIDER addressed scalable physics-informed dexterous retargeting by incorporating physical constraints directly into the learning process. This helps the robot understand how forces and gravity affect its movements during manipulation. This connects to the viewpoint work because understanding physics helps ground those visual representations in real-world dynamics.
Stereopolicy improved robotic manipulation policies by integrating stereo perception, meaning it used depth information from multiple views to make better decisions about how to grasp or move an object. This is a direct application of leveraging richer sensory input for better control.
Finally, tac2pix introduced image-space visuo-tactile fusion for dexterous manipulation. This combines visual and tactile data to give the robot a richer sense of touch while it's working on something intricate. This builds upon the representation work by adding a crucial sense of physical contact.
The most significant development today involves the AquaBEV-Nav system, which tackles the challenge of learning occupancy in underwater environments for navigation and exploration. This work is crucial because accurately mapping an unknown underwater space is fundamental for any autonomous vehicle operating in such conditions.
This effort built upon prior work on World SLAM Model, which aimed to achieve joint world modeling for simultaneous localization and mapping. Building on that foundation, the AquaBEV-Nav approach seems to have made a specific leap by learning BEV occupancy directly within the underwater context. This means it's not just mapping the environment but understanding where things are likely to be in a bird's eye view underwater.
Another key piece of research is FINE, which focuses on future-informed navigation encoding for vision-language navigation tasks. This work attempts to make navigation more data efficient by incorporating knowledge about what might happen next into the visual and language processing pipeline. This connects this forward-looking capability to the broader goal of robust autonomous movement.
Then there is TriDrive, which integrates driver, vehicle, and road modeling for better forecasting and driver monitoring in terrestrial driving scenarios. This provides a different kind of context for modeling dynamic interactions compared to the underwater focus of AquaBEV-Nav.
Moving into manipulation, Dynamic Manipulation with World-Action Models via Counterfactual Planning was explored to enable planning actions in complex dynamic scenes by using counterfactual reasoning. This is distinct from the purely navigational focus of the other papers, but both contribute to a larger toolkit for autonomous systems.
Finally, there is DeltaWAM, which introduces change-centric visual foresight using delta tokens for more efficient world-action models. This suggests a method for updating world understanding incrementally rather than re-processing everything every time something changes. This incremental learning idea complements the persistent memory goals suggested by PORTER, which aims to create edge-cloud residency for persistent 3D scene graph memory.
The most significant piece of work today involved SocialHumanoid, which attempts to create expressive humanoid behavior through one-step co-speech motion generation. This matters because it moves beyond simple task execution to imbue robots with a more natural, communicative presence. The core idea is that by generating speech and motion simultaneously from a single input signal, we can achieve a level of embodied interaction previously unattainable.
This approach builds upon earlier explorations into recursive harness distillation across agents for robot manipulation, which focused on improving how different robotic systems coordinate their physical actions. That work provided the necessary framework for understanding agent-to-agent communication in complex settings. Furthermore, the development of a multi-modal tactile fingertip design for robotic hands aimed to enhance dexterous manipulation by giving the robots better sensory feedback during physical interaction.
The integration challenges are becoming clearer when looking at AI-driven collaborative assembly line inspection, which deals with the practical hurdles of deploying these sophisticated systems in real industrial environments. This system aims to solve problems related to system integration and deployment, showing how theoretical models translate into tangible operational constraints. This operational reality is informed by human-guided planning for complex manipulation tasks using the screw geometry of motion, which provides a geometric understanding necessary for precise physical execution.
Finally, the underlying cognitive architecture is being refined through unifying deep predicate invention with pre-trained foundation models. This seeks to give AI a more robust way to reason about actions and concepts. This cognitive advancement is complemented by LogicEnvGen, which generates diverse simulated environments driven by task logic, creating the necessary training ground for embodied AI agents to practice these complex behaviors.
The most critical piece of work today involves learning geometrically grounded amodal three dimensional representations for view generalizable robotic manipulation because it directly addresses the core challenge of making robots understand and interact with the physical world in a way that generalizes across different viewpoints. This research explored methods to create these representations, specifically focusing on how they can be used for manipulation tasks.
A significant development was the work on self-evolutionary replanning for failure aware motion planning. This attempts to make robot movement smarter when things go wrong by allowing the plan to change dynamically based on observed failures. This is important because it moves beyond static plans that fail easily in real-world scenarios.
Another area of focus was dynamic model identification and gravity compensation for the dVRK-Si patient side manipulator. This deals with making a specific robotic arm function better by figuring out how its physical properties change over time and compensating for gravity effects. This helps ensure precise control during delicate procedures.
The work on task driven co design of heterogeneous multi robot systems is also relevant as it tackles how different robots can work together effectively toward a common goal. This connects to the effort in learning control policies to provably satisfy hard affine constraints for black box hybrid dynamical systems, as both aim to ensure reliable system behavior under tight physical rules.
Finally, there was some exploration into premoe which focuses on robust preference modeling with mixture of experts reward learning. This is a technique used to train agents based on human preferences in complex decision-making environments. This work complements the broader goal of creating more capable and adaptable robotic systems.
The most significant development today involves DriveAnchor, which tackles the core challenge of planning for autonomous driving by using progressive anchor-based flow learning. This approach aims to build robust plans by learning how to transition between different states, which is crucial because it addresses the fundamental difficulty of long-horizon decision-making in complex driving scenarios.
This method builds upon work like PACE, which focuses on phase-aware chunk execution for robot policies through action chunking. PACE attempts to break down complex tasks into smaller, manageable chunks based on the current phase of operation. This helps manage computational load during execution.
Efficient-WAM presents a 1 billion parameter world-action model designed for low-cost future imagination, offering a way to simulate what might happen next without requiring massive computational resources for every possible outcome. This efficiency contrasts with Assistron, which explores Bayesian shared autonomy by integrating off-the-shelf vision language action models into the system.
Assistron leverages pre-trained vision language action models within a Bayesian framework to enable shared autonomy, meaning it lets the robot and human collaborate more effectively in real-time situations. Similarly, SurgVIL scales surgical robot imitation learning by utilizing open-source surgical videos to improve performance in complex medical tasks.
RoboEdit focuses on turning human manipulation videos into scalable robot experience. This is a method for generating diverse training data for robotic systems from existing demonstrations. This contrasts with reduced Cartesian kinetostatics, which deals with residual stabilization during full-shape propagation in tendon-driven continuum robots.
Finally, Path Planning with Motion Primitives in Dynamic Environments using SIPP on Lattices addresses path planning in dynamic settings by employing motion primitives within lattice structures to handle environmental changes effectively.
Today's papers
- Easier Said Than Done Unpacking Intent-Behavior Gap in Jailbreaking LLM-based Robots This paper explores how to understand and fix the gap between what an LLM intends to do and what it actually does when controlling robots. Dynamic Buffers Cost-Efficient Planning for Tabletop Rearrangement with Stacking This work focuses on planning efficient ways for robots to stack and rearrange objects on a table while managing costs. FastGrasp Learning-based Whole-Body Control Method for Fast Dexterous Grasping with Mobile Manipulators This paper presents a learning method to help mobile robots quickly grasp objects using whole-body control. RoboAlign-R1 Distilled Multimodal Reward Alignment for Robot Video World Models This research uses distilled multimodal rewards to align the behavior of robot world models seen in videos. When to Trust Imagination Adaptive Action Execution for World Action Models This paper develops adaptive action execution strategies for world action models to decide when it is safe to act based on imagination. What Stops Recursive Self-Improvement in Robotics Lessons from 123 Rounds of Agentic Skill Discovery This study investigates what limits the self-improvement capabilities of agents through a series of skill discovery rounds. Robot Manipulation with GPT-6-Astra Body Knowledge Experience Reuse Emergent Skills and Sim2Real Transfer This paper shows how robots can use large language models to learn and transfer skills for manipulation from simulation to the real world. Timed Rule-Based Supervision of an End-to-End Autonomous Parking Policy This work describes a method for supervising autonomous parking policies by using rules that are timed based on the situation. PHIRL Aligning Learned Rewards with Task Progress for Inverse Reinforcement Learning This paper introduces a technique to align learned rewards with the actual progress of a task when learning through inverse reinforcement learning. DS-VLA A Dendritic-inspired Vision-Language-Action Model for Robust Action Control This model uses a dendritic structure to create robust vision language action models for controlling robots. Affordance-Conditioned Decision Making Bridging the Semantic-Spatial Gap in Zero-Shot Cross-Floor Vision and Language Navigation This work shows how robots can navigate between floors by using affordances to bridge the gap between what they see and what they need to do. RE-0 Verified Recursive Improvement of Embodied Code as Policy Agents through Local On-Policy Distillation This method allows embodied agents to improve their code policies recursively by distilling knowledge locally during on-policy training. DRAM Delta-rule Recurrent Associative Memory for Robot Manipulation Policies This paper proposes a recurrent memory system using delta rules to help robots learn and maintain manipulation policies. RECAST Recasting Vision-Language Semantics into an Actionable Cost Map for Robot Navigation This work converts vision and language information into a cost map that robots can use to navigate effectively. Scanning While Imagining A Scene-Graph World Model for Robotic Ultrasound Navigation This model uses a scene graph world model to help robots navigate and scan environments using ultrasound data. RoboFoundry System-as-Policy Evolution for Self-Learning Embodied Agents This paper describes how embodied agents can evolve their entire system policy through self-learning processes. CAPEX Efficiently Distilling Foundation Model Behavior into Deployable Robot Policies through Experience-Adaptive Reasoning This work focuses on distilling large foundation model behavior into robot policies using experience that adapts to the reasoning process. VPTwin Real-Sim-Real Video Prediction for Robotic Manipulation Planning This method uses video prediction to improve planning by bridging the gap between simulation and reality in robotic manipulation. ProcVLM Learning Procedure-Grounded Progress Rewards for Robotic Manipulation This research learns progress rewards grounded in the specific procedures a robot follows during manipulation tasks. GT-VLA Target-Conditioned Trace Guidance for Generalizable Robotic Manipulation This paper uses target guidance to create generalizable robotic manipulation policies by conditioning them on a desired outcome. Federated Subspace Guided Vision-Language-Action Policy Distillation for Non-IID Multi-Robot Manipulation This method distills vision language action policies across multiple robots with non-independent and identically distributed data. Neural ODEs Meet Concurrent Learning Stable Online Learning with Lyapunov Guarantees This paper combines neural ordinary differential equations with concurrent learning to achieve stable online learning with mathematical guarantees. Think Fast Plan Selectively Adaptive Deliberation for Efficient Data-Driven MPC This work develops an adaptive deliberation system for model predictive control that plans efficiently by selecting actions based on the data available. Schur-Neural KF Learned Schur-Consistent Corrections to the Extended Kalman Filter This paper introduces learned corrections to the extended Kalman filter using a Schur-Neural framework. An Empirical Study on What Matters for Viewpoint-Generalizable Policies in Visual Imitation Learning This study investigates which factors are most important for creating robotic policies that work well from different viewpoints during visual imitation learning. Copper-Policy Focus on the Representation for Robust Robot Manipulation This paper focuses on designing better representations to make robot manipulation policies more robust. CollisionGAT Controller-Agnostic One-Step Collision Screening for Multi-Agent Motion This method provides a fast way to screen for collisions in multi-agent motion without needing a specific controller. SPIDER Scalable Physics-Informed Dexterous Retargeting This paper presents a scalable method for retargeting dexterous movements while incorporating physical constraints. From Instruction to Event Sound-Triggered Mobile Manipulation This work enables mobile manipulation tasks to be triggered by sound events instead of explicit instructions. StereoPolicy Improving Robotic Manipulation Policies via Stereo Perception This research improves robotic manipulation policies by incorporating stereo vision data into the planning process. Tac2Pix Image-Space Visuo-Tactile Fusion for Dexterous Manipulation This method fuses image and tactile information in image space to enhance dexterous manipulation capabilities. What Matters for Latent Actions in Robot Learning This paper explores the importance of latent actions when learning policies for robots. AquaBEV-Nav Learned BEV Occupancy for Underwater Navigation and Exploration This work learns bird's eye view occupancy maps to enable navigation and exploration in underwater environments. World SLAM Model Joint World Modeling for SLAM and Navigation This paper proposes a joint world modeling approach that combines simultaneous localization and mapping with navigation. FINE Future-Informed Navigation Encoding for Data-Efficient Vision-Language Navigation This method uses future information to create more data-efficient encoding for vision language navigation tasks. TriDrive Joint Driver Vehicle and Road Modeling for Forecasting and Driver Monitoring This work models the driver, vehicle, and road together to improve forecasting and driver monitoring in autonomous driving. Dynamic Manipulation with World-Action Models via Counterfactual Planning This paper uses counterfactual planning with world action models to perform dynamic manipulation tasks. DeltaWAM Change-Centric Visual Foresight via Delta Tokens for an Efficient World-Action Model This work introduces delta tokens to create an efficient world action model that focuses on changes in the environment. PORTER Edge-Cloud Residency for Persistent 3D Scene Graph Memory This paper discusses storing persistent 3D scene graph memory across edge and cloud devices. Q-WAM 4-Bit Quantization of World Action Models with Action-Subspace Protection This method quantizes world action models to use less memory while protecting the action subspace. SocialHumanoid Towards Expressive Humanoid Behavior via One-Step Co-Speech Motion Generation This work generates expressive humanoid behavior by using one step of co-speech motion generation. Recursive Harness Distillation across Agents for Robot Manipulation This paper uses recursive harness distillation to improve robot manipulation policies across multiple agents. AI-Driven Collaborative Assembly Line Inspection System Integration and Deployment Challenges This work discusses the challenges of integrating and deploying AI systems for collaborative assembly line inspection. Human-Guided Planning for Complex Manipulation Tasks Using the Screw Geometry of Motion This paper shows how human guidance can be used to plan complex manipulation tasks by leveraging screw geometry. A Multi-modal Tactile Fingertip Design for Robotic Hands to Enhance Dexterous Manipulation This paper proposes a multi-modal tactile fingertip design aimed at improving dexterous manipulation skills. Unifying Deep Predicate Invention with Pre-trained Foundation Models This work connects deep predicate invention with pre-trained foundation models to create more capable systems. Helical Tendon-Driven Continuum Robot with Programmable Follow-the-Leader Operation This paper describes a continuum robot that uses helical tendons for programmable follow the leader operation. LogicEnvGen Task-Logic Driven Generation of Diverse Simulated Environments for Embodied AI This tool generates diverse simulated environments based on task logic to train embodied AI agents. Learning Geometrically-Grounded Amodal 3D Representations for View-Generalizable Robotic Manipulation This method learns 3D representations that are geometrically grounded and view-generalizable for robot manipulation. Self-Evolutionary Replanning for Failure-Aware Motion Planning This work focuses on self-replanning to handle failures in motion planning tasks. Dynamic Model Identification and Gravity Compensation for the dVRK-Si Patient Side Manipulator This paper deals with identifying dynamic models and compensating for gravity in a specific patient side manipulator. Task-Driven Co-Design of Heterogeneous Multi-Robot Systems This work addresses how to co-design heterogeneous multi-robot systems based on specific task requirements. Learning Control Policies to Provably Satisfy Hard Affine Constraints for Black-Box Hybrid Dynamical Systems This paper learns control policies that can provably satisfy hard affine constraints in complex hybrid dynamical systems. PrefMoE Robust Preference Modeling with Mixture-of-Experts Reward Learning This method uses a mixture of experts approach to create robust preference modeling for reward learning. FlyMirage A Fully Automated Generation Pipeline for Diverse and Scalable UAV Flight Data via Generative World Model This pipeline automatically generates diverse and scalable UAV flight data using a generative world model. IDOL Inverse-Dynamics-Guided Future Prediction for End-to-End Autonomous Driving This work uses inverse dynamics to guide future prediction for end-to-end autonomous driving systems. DriveAnchor Progressive Anchor-based Flow Learning for Autonomous Driving Planning This method uses progressive anchor learning to improve flow estimation and planning in autonomous driving. PACE Phase-Aware Chunk Execution for Robot Policies with Action Chunking This method improves robot policies by chunking actions based on phase awareness during execution. Efficient-WAM A 1B-Parameter World-Action Model with Low-Cost Future Imagination This paper introduces a small world action model that uses low cost imagination for efficient future planning. Assistron Bayesian Shared Autonomy with Off-the-shelf Vision Language Action Models This work proposes a shared autonomy framework using off-the-shelf vision language action models in a bayesian setting. The paper is not provided, so I cannot provide an output for it. [paper]
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
- The paper is not provided, so I cannot provide an output for it.
The papers
- Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment — Finetuning vision-language models (VLMs) for robot manipulation via behavior cloning often leads to catastrophic forgetting and language-action misalignment, which this work addresses by proposing Anchor-Align, a method that augments standard behavior cloning with two objectives [episode]
- InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation — InCoM is an intent-driven perception and structured coordination framework for mobile manipulation that jointly considers stage-adaptive perception and coordinated action generation, addressing two key challenges: strong coupling between base and arm actions complicating control [episode]
- Vault: One-Step Latent Generation with Positive-Anchored Rewards for Autonomous Driving — One-step latent generation with positive-anchored rewards for autonomous driving addresses the latency constraints of iterative planning models by replacing multi-step denoising chains with a single forward pass conditioned on learned semantic features. [episode]
- Guava: Distilling Frontier VLMs into a Compact Agent through a Robotic Manipulation Harness — Guava presents a harness framework for embodied manipulation that identifies three key ingredients for effective embodied agents: iterative reasoning, semantic action abstractions, and multimodal observations. [episode]
- Robust and Efficient MuJoCo-based Model Predictive Control via Web of Affine Spaces Derivatives — This paper introduces a method to accelerate model derivative computations within MuJoCo-based Model Predictive Control (MPC) by replacing finite differencing (FD) with Web of Affine Spaces (WASP) derivatives, aiming to address computational bottlenecks in high-DOF systems. [episode]
- WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos — WALA is a framework designed to jointly learn executable latent actions from both action-labeled demonstrations and action-free videos, addressing the limitations of current vision-language models by leveraging future scene evolution to ground latent actions in task-relevant sema [episode]
- Sound Compilation of Weighted Event Signal Temporal Logic to Timeless Geometric Control — A new control synthesis paradigm is introduced that overcomes vulnerabilities in traditional, time-indexed Cyber-Physical Systems (CPS) controllers by translating temporal logic specifications directly into timeless geometric constraints. [episode]
- SkillWrapper: Generative Predicate Invention for Task-level Robot Planning — This research introduces SkillWrapper, a novel method designed to learn provably sound and complete symbolic models for robot planning by leveraging large Foundation Models (FMs) to actively collect data and generate human-interpretable, plannable representations using only RGB i [episode]
- Interaction Dynamics MPC for Knee Rehabilitation Exoskeletons: A Closed-Loop SEA Outer-Loop Study — Safe rehabilitation is an interaction-dynamics problem where a controller must regulate prescribed motion while absorbing involuntary spasm, voluntary effort, actuator compliance, and model mismatch as disturbances. [episode]
- RedFlow: Redirect Failure into Action-level Corrections for Flow-matching VLA Policy — Flow-matching Vision-Language-Action (VLA) policies often suffer from compounding errors during deployment due to distribution shifts, and this paper introduces RedFlow, a fine-grained offline RL framework that redirects failure experiences into high-fidelity action-level correct [episode]
- Wire-Level Interrupt-to-Decision Latency of On-Sensor MLC versus Host Inference on the NVIDIA Jetson Orin Nano: A Pre-Registered Measurement Study — Wire-level interrupt-to-decision latency measurements reveal that host inference exhibits lower median wire-level latency than on-sensor Machine Learning Core (MLC) pipelines across all tested conditions. [episode]
- Toward Proactive RF Charging Scheduling: Generative AI for Decision Support — Radio frequency wireless power transfer (RFWPT) is an enabling technology for supporting uninterrupted communications in future Internet of Things systems by reducing the need for battery replacement and mitigating battery-waste-related issues. [episode]
- Source-Lifted Flow Matching for Intervenable Multimodal Imitation — Source-Lifted Flow Matching (SL-FM) proposes a novel flow-matching policy that transforms passive source randomness into an actionable intervention variable for multimodal imitation learning. [episode]
- Risk-Bounded Multi-Agent Visual Navigation via Iterative Risk Allocation — Safe navigation for autonomous systems operating in hazardous environments, especially when multiple agents must coordinate using only high-dimensional visual observations, is addressed by introducing a framework that dynamically distributes per-agent risk budgets during planning [episode]
- Temporal Forcing: 4D Representation Alignment for Vision-Language-Action Models — Temporal Forcing is a 4D representation alignment method for Vision-Language-Action (VLA) models designed to improve manipulation performance by aligning model representations with 3D scene geometry, specifically addressing limitations in long-horizon manipulation and observation [episode]
- Efficient streaming dynamic mode decomposition — Dynamic mode decomposition (DMD) is a widely used technique for revealing the discrete spectrum in complex dynamical systems, and this work proposes an efficient streaming variant that reduces computational redundancy by maintaining only a single orthonormal basis. [episode]
- Diffusion-Guided Multi-Arm Motion Planning — Multi-arm motion planning is fundamental for enabling arms to complete complex long-horizon tasks in shared spaces efficiently but current methods struggle with scalability due to exponential state-space growth and reliance on large training datasets for learned models. [episode]
- ViTacWorld: Scaling Visuo-Tactile World Models for Contact-Rich Robot Manipulation — Contact-rich robot manipulation requires physical interaction cues that are often invisible to cameras, making tactile sensing essential for robust control. [episode]
- UniBYD: A Unified Framework for Learning Robotic Manipulation Across Embodiments Beyond Imitation of Human Demonstrations — UniBYD proposes a unified reinforcement learning framework that learns manipulation policies across diverse robotic hand morphologies by transitioning from imitation-based learning to online-adaptive exploration, achieving significant improvements in success rates and task perfor [episode]
- SafeManip: A Property-Driven Benchmark for Temporal Safety Evaluation in Robotic Manipulation — Robotic manipulation safety evaluation often fails because task success does not guarantee safe execution, leading to temporal failures that are missed by traditional metrics. [episode]
- GIFT: Glove-Inferred Force Transfer: Force-Aware Human-to-Robot Skill Transfer from a Wearable Sensing Glove to a Robot Hand Without Tactile Sensors — GIFT (Glove-Inferred Force Transfer) is an end-to-end pipeline for human-to-robot skill transfer of force without tactile hardware on the robot, where fingertip force is measured in newtons on the human side and estimated in newtons on the robot side. [episode]
- Interaction Dynamics Modeling and Predictive Control for Safe Steerable Catheter--Tissue Interaction — Safe steerable catheter control is fundamentally a problem of interaction dynamics: the tip must follow a planned motion, remain compliant against moving tissue, reject friction and hysteresis, and respect a clinically meaningful never-exceed contact-force bound. [episode]
- BEV-ODOM2: Enhanced BEV-based Monocular Visual Odometry with PV-BEV Fusion and Dense Flow Supervision for Ground Robots — Scale-consistent ego-motion estimation is fundamental for autonomous ground robots, and Bird’s-Eye-View (BEV) representation naturally addresses scale drift by providing a metric-scaled planar workspace, simplifying 6-DoF ego-motion to a more robust 3-DoF model. [episode]
- Enhanced Sampled-Data Model Predictive Control via Nonlinear Lifting — This paper introduces a novel nonlinear model predictive control (NMPC) framework that incorporates a lifting technique to enhance control performance for nonlinear systems, addressing a gap where lifting has been widely employed in linear systems but its application to nonlinear [episode]
- HumanHalo: Safe and Efficient 3D Navigation Among Humans via Minimally Conservative MPC — HumanHalo is a Model Predictive Control (MPC) framework for 3D Micro Air Vehicle (MAV) navigation among humans that combines theoretical safety guarantees with data-driven models for realistic human motion forecasting. [episode]
- DiffCVaR: Reinforcement Learning for Risk Adaptation via Differentiable CVaR Barrier Functions — Planning through crowded environments under uncertain obstacle motions remains difficult, as stochastic interactions often induce overly conservative behavior or reduced efficiency. [episode]
- Model-Free Output Feedback Stabilization via Policy Gradient Methods — Stabilizing dynamical systems is a fundamental problem in control theory, and this work proposes an algorithmic framework for learning stabilizing static output feedback (SOF) controllers for open-loop unstable discrete-time linear systems without requiring a system model. [episode]
- Signal Temporal Logic Evaluation and Synthesis Using Deep Reachability Analysis and Layered Control Architecture — We propose a signal temporal logic (STL)-based framework that rigorously verifies the feasibility of a mission described in STL and synthesizes control to safely execute it. [episode]
- Spatial Load Correlation in AI Data-Center-Dominated Power Systems — The proliferation of large-scale data centers introduces spatially correlated demand profiles that challenge the long-standing assumption of statistical independence of loads in power system analysis. [episode]
- STeP: Signal Temporal Logic for Precise Specifications for Action Generation with Vision Language Models — Vision-language-action (VLA) models often lack interpretability and struggle to follow precise natural language instructions that encode spatial, temporal, and logical requirements. [episode]
- Adversarial Vulnerabilities of Learned Telesurgery Policies — This paper presents "the first study of adversarial threats to learning-based policies in surgical robotics." It investigates two threat modes: "(a) disruptive attacks, where imperceptible visual perturbations interrupt policy execution, and (b) steering attacks, where such pertu [episode]
- eVGGT: An Efficient Geometry-Aware Vision Encoder for Visuomotor Policies — Existing RGB-based imitation learning approaches typically employ traditional vision encoders such as ResNet or ViT, which lack explicit 3D reasoning capabilities. [episode]
- Two-Stage Optimization for Dynamic Line Rating and Energy Storage Deployment — The increasing penetration of distributed energy resources (DER) and weather-driven variability has intensified congestion and reliability stress in transmission networks, making strategies that enhance utilization of existing infrastructure, such as static line ratings (SLR) and [episode]
- Offset-free Data-Driven Predictive Control for Grid-Connected Power Converters in Weak Grid Faults — Grid-connected power converters encounter significant stability challenges during weak grid faults, when conventional PI-based controllers exhibit an oscillatory response and poor fault-ride-through performance. [episode]
- Meta-Optimization and Program Search using Language Models for Task and Motion Planning — Intelligent interaction with the real world requires robotic agents to jointly reason over high-level plans and low-level controls, and this paper addresses this challenge by introducing a novel meta-optimization framework that integrates foundation models with trajectory optimiz [episode]
- Learning New Tasks via Reusable Skills: Skill-Compositional Experts for Embodied Continual Learning — Embodied Continual Learning (ECL) aims to enable robots to continually acquire new manipulation tasks while retaining previously learned behaviors under closed-loop control, but existing methods suffer from severe catastrophic forgetting due to feature drift. [episode]
- Preview-Based Relative-Motion Control of an Insertion Tool for Neural-Thread Placement in Pulsating Tissue — Flexible neural electrode threads must be placed at a prescribed depth while the tissue they enter is not stationary. [episode]
- MILE: A Mechanically Isomorphic Hand Exoskeleton and Visuotactile Robotic Hand for Data Collection in Dexterous Manipulation — Dexterous robotic hands are expected to perform complex, contact-rich object manipulation, but learning such skills remains challenging because high-dimensional hands require high-fidelity demonstrations. [episode]
- TransforMARS: Fault-Tolerant Self-Reconfiguration for Arbitrarily Shaped Modular Aerial Robot Systems — Modular Aerial Robot Systems (MARS) are flexible, adaptive agents that can respond to environmental changes through disassembly and reassembly. [episode]
- Model Predictive Communication for Timely Status Updates in Low-Altitude Networks — Timely information delivery in low-altitude networks is critical for many time-sensitive applications, such as unmanned aerial vehicle (UAV) navigation, inspection, and surveillance. [episode]
- A Survey on Reinforcement Learning Applications in SLAM — A Survey on Reinforcement Learning Applications in SLAM explores how Reinforcement Learning (RL) methodologies are being integrated into Simultaneous Localization and Mapping (SLAM) to enhance robot decision-making and navigation skills in complex, dynamic environments. [episode]
- Fading Expert Guidance: Bridging Model-Based and Learning-Based Control for Abortable Autonomous Overtaking — Overtaking maneuvers on two-lane roads present a significant challenge for autonomous vehicles because oncoming traffic requires dynamic decision-making and potential aborts, making it crucial to balance traditional control engineering with learning systems. [episode]
- CoinFT: A Coin-Sized, Capacitive 6-Axis Force Torque Sensor for Robotic Applications — CoinFT introduces a compact, light, and low-cost capacitive 6-axis force/torque (F/T) sensor designed for various robotic applications. [episode]
- LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models — Vision-Language-Action (VLA) and World Action (WAM) models have achieved remarkable performance in robotic manipulation, but success under ideal conditions does not imply real-world robustness because robots must recognize and recover from failures to continue tasks. [episode]
- Traffic Characterization of Event-Triggered Control Systems: A Geometric-Algebraic Perspective — This paper characterizes triggering behaviors of event-triggered control systems from a geometric–algebraic perspective, providing necessary and sufficient conditions for transition relations to be feasible. [episode]
- Simulating Robotic Locomotion in Sand: Resistive Force Theory in an Open-Source Physics Engine — Recent advancements in Resistive Force Theory (RFT) enable approximation of ground reaction forces for locomotion in sand without the computational expense of modeling interactions with individual grains. [episode]
- Safety for Weakly-Hard Control Systems via Graph-Based Barrier Functions — This article addresses safety verification and controller synthesis for a class of control systems subject to weakly-hard constraints (WH constraints), which model phenomena such as packet dropouts in networked control systems or computational overruns in real-time applications. [episode]
- Reinforcement Learning to Initialize Newton-Raphson for AC Power Flow with Quantum Annealing-Based Environment Updates — The proposed work addresses limitations in traditional Newton-Raphson (NR) methods for power flow (PF) analysis by integrating Reinforcement Learning (RL) with quantum annealing to optimize initial conditions, thereby accelerating convergence and enhancing robustness in complex p [episode]
- SecuLEx: a Secure Limit Exchange Market for Dynamic Operating Envelopes — SecuLEx (Secure Limit Exchange) is a new market-based paradigm introduced to allocate power injection and withdrawal limits, called dynamic operating envelopes (DOEs), which guarantee network security during time periods. [episode]
- Constrained finite-time stabilization by model predictive control: an infinite control horizon framework — An infinite-horizon Model Predictive Control (MPC) framework is proposed to achieve constrained finite-time stabilization for discrete-time systems, overcoming limitations in existing methods by replacing short-horizon terminal costs with an infinite sum of stage costs. [episode]
- Safe Multi-Agent Navigation guided by Goal-Conditioned Safe Reinforcement Learning — Safe navigation is essential for autonomous systems operating in hazardous environments, and this work introduces a novel hierarchical framework that combines high-level multi-agent planning using MAPF approaches with low-level control through safe GCRL policies. [episode]
- Tunable Leg Stiffness in a Monopedal Hopper for Energy-Efficient Vertical Hopping Across Varying Ground Profiles — We present the design and implementation of HASTA (Hopper with Adjustable Stiffness for Terrain Adaption), a vertical hopping robot with real-time tunable leg stiffness, aimed at optimizing energy efficiency across various ground profiles (a pair of ground stiffness and damping c [episode]
- Graph approach for observability analysis in power system dynamic state estimation — The proposed approach yields a numerical method that provably executes in linear time with respect to the number of nodes and edges in a graph, offering a scalable solution for observability analysis in power system dynamic state estimation. [episode]
- BR-MPPI: Barrier-Rate Guided MPPI for Enforcing Multiple Inequality Constraints with Learned Signed Distance Fields — Model Predictive Path Integral (MPPI) control is integrated with Control Barrier Function (CBF) conditions to solve unconstrained optimal control problems while enforcing multiple inequality constraints. [episode]
- Towards Drone-based Mapping of Volcanic Gases using Gas Tomography — Volcanoes emit large amounts of CO2, directly influencing human lives, and mapping volcanic gas emissions helps to forecast eruptions and understand their impact on climate and the environment. [episode]
- Radar Sensing Based on 1-Bit Quantized Reconfigurable Intelligent Surfaces — We present a radar sensing framework based on a low-complexity, quantized reconfigurable intelligent surface (RIS) that enables programmable manipulation of electromagnetic wavefonts for enhanced detection in non-specular and shadowed regions. [episode]
- Inertia-Corrected Newton Method For Generalized Nash Equilibria in Dynamic Games with Optimality Verification —
- Low Cost Eye Tracking for Vision Screening —
- Reducing Combinatorial Redundancy in Mixed-Integer MPC via Ranking-Based Feasible-Set Restriction for Reconfigurable Batteries —
- RFSR-MI-MPC: Ranking-Based Feasible-Set Restriction for Real-Time Control of Reconfigurable Battery Packs —
- Grasp2Twist: Learning Bimanual Dexterous Jar Opening by Reinforcement Learning —
- Find Something You Can't Do: Agentic Real-World Reinforcement Learning for Self-Improving VLA Models —
- Fiber-Normalized Manipulability and Determinant Proxies: Intrinsic Redundancy Optimization Across and Within Task Fibers —
- Active 3D weaves for load-bearing and damage-resilient locomotion —
- Residual Denoising Enables Sample-Efficient Multi-Agent Coordination on Demand —
- GAUGE: Planner-Conditioned Active Calibration of Opaque Quadruped Velocity Interfaces —
- RecastVLA: From Past Interaction to Future Control with Adaptive Policy States —
- AquaBEV-Nav: Learned BEV Occupancy for Underwater Navigation and Exploration —
- FutureRay: Control-Aligned Future Range for Agile Quadruped Navigation —
- A Priority-Aware Dual-Channel Feature Fusion Method for Urban Rail Service Traffic Classification —
- Neural Decentralized Conditions for Power System Stability Analysis —
- Neural-Analytica Energy Functions for Power System Stability Analysis —
- Physics-Informed Learning of Feedback-Linearizing Representations —
- RoboFFT: Finetuning generative robot policy via online reinforcement learning with forward process —
- Federated Subspace Guided Vision-Language-Action Policy Distillation for Non-IID Multi-Robot Manipulation —
- DS-VLA: A Dendritic-inspired Vision-Language-Action Model for Robust Action Control —
- Neural ODEs Meet Concurrent Learning: Stable Online Learning with Lyapunov Guarantees —
- Affordance-Conditioned Decision Making: Bridging the Semantic-Spatial Gap in Zero-Shot Cross-Floor Vision-and-Language Navigation —
- MemTransfer: Benchmarking Memory Beyond Matched Experience in Embodied Decision-Making —
- WSM-Aware HRI: An IoT-Enhanced Framework for Early Detection and Norm-Guided Repair of Failures with LLM Guidance —
- Human Motion Prediction for Human-Robot Collaboration —
- Proactive Motion Planning for Human-Robot Cooperation —
- RE-0: Verified Recursive Improvement of Embodied Code-as-Policy Agents through Local On-Policy Distillation —
- DRAM: Delta-rule Recurrent Associative Memory for Robot Manipulation Policies —
- GlowTact: Simple and Compact Vision-Based Tactile Sensing with High Sensitivity and Spatial Resolution —
- Assisting for Open-Ended Tasks: Goal-Oriented Shared Autonomy as a Particle Filter —
- Think Fast, Plan Selectively: Adaptive Deliberation for Efficient Data-Driven MPC —
- RECAST: Recasting Vision-Language Semantics into an Actionable Cost Map for Robot Navigation —
- World SLAM Model: Joint World Modeling for SLAM and Navigation —
- PF-RL: Progress Field Reinforcement Learning via Goal-Conditioned Value Geometry for Vision-Language-Action Models —
- Schur-Neural KF: Learned Schur-Consistent Corrections to the Extended Kalman Filter —
- SEES: A Self-Evolving Embodied System via Failure-Guided VLA Policy Adaptation —
- Characterization and Monitoring of Nonlinear Dynamics and Chaos in Complex Physiological Systems —
- MORPH: Self-Organising Multi-Robot Task Allocation via Neuroplasticity-Inspired Adaptive Topology —
- Path Invariance of a Quadrotor System under Cyber Attacks with Theoretical Guarantees —
- An Empirical Study on What Matters for Viewpoint-Generalizable Policies in Visual Imitation Learning —
- CLAP: Closed-Loop Alignment with Pressure for Precise Suction Manipulation —
- Copper-Policy: Focus on the Representation for Robust Robot Manipulation —
- CollisionGAT: Controller-Agnostic One-Step Collision Screening for Multi-Agent Motion —
- Towards Kinematic Actionable Infeasibility Detection in Motion Planning —
- Scanning While Imagining: A Scene-Graph World Model for Robotic Ultrasound Navigation —
- FINE: Future-Informed Navigation Encoding for Data-Efficient Vision-Language Navigation —
- RoboFoundry: System-as-Policy Evolution for Self-Learning Embodied Agents —
- Communication-Aware Heterogeneous Graph Learning for Decentralized Multi-Human Multi-Robot Task Allocation —
- Learning Geometry-Aware Virtual Fixtures From Sparse Demonstrations —
- TriDrive: Joint Driver, Vehicle, and Road Modeling for Forecasting and Driver Monitoring —
- CAPEX: Efficiently Distilling Foundation Model Behavior into Deployable Robot Policies through Experience-Adaptive Reasoning —
- Is Online Interaction Necessary for Recovery? A Minimalist Approach to Robust Planning via Perturbation —
- SwingRL: Adaptive Observation Reinforcement Learning with World-Model Prediction for Cable-Suspended Hoisting Control —
- REALM: A Coarse-to-Fine Generative Framework for Embodied Reactive Listening —
- Humanoids for Robot-Assisted Surgery: Bimanual Base Placement and Tool-Mount Optimization via Capability Maps —
- Evolving Dexterous Robots from Scratch —
- VPTwin: Real-Sim-Real Video Prediction for Robotic Manipulation Planning —
- Multi-Modal Non-Prehensile Estimation of Physical Parameters via Press-and-Pull Tipping —
- Beyond State-as-Action: Exploiting Command-State Discrepancy for Robot Imitation Learning —
- TimelyDAgger: Timing-Aware Expert Querying for VLA Policy Improvement —
- Beyond Tasks: A Vision for Reproducing an Animal-like Behavioral Substrate Using Modern Robot Learning Techniques —
- Dynamic Manipulation with World-Action Models via Counterfactual Planning —
- Multi-Terrain Mastery: A Comprehensive Controller for Bipedal Locomotion —
- DeltaWAM: Change-Centric Visual Foresight via Delta Tokens for an Efficient World-Action Model —
- TAO-DA: Towards Autonomous Operation--A Dual-Arm Vision-Language-Action Model for Coordinated Manipulation —
- Learning with Object-centric Representations of Tactile Interactive Perception for Robot Manipulation —
- SurgFlow: 3D Object-Centric Contact Flow for Surgical Robot Manipulation —
- ActionGround: Training-Free Runtime Refinement of Frozen VLA Policies —
- PORTER: Edge-Cloud Residency for Persistent 3D Scene Graph Memory —
- Q-WAM: 4-Bit Quantization of World Action Models with Action-Subspace Protection —
- AquaWAM: A Dynamics-aware World Action Model for Underwater Embodied Agents —
- CompliantWBC: Whole-Body Compliance for Heavy Humanoids via Force Latent Estimation and Residual Impedance Targets —
- SocialHumanoid: Towards Expressive Humanoid Behavior via One-Step Co-Speech Motion Generation —
- Traceable Human-to-Humanoid Sign Language Benchmarking —
- Recursive Harness Distillation across Agents for Robot Manipulation —
- Large Language Models for Model-Based Robot Design —
- DAOCP: a dual active set solver for optimal control problems —
- VIDEAS: Distilling Explicit Action Semantics from Demonstration Videos for World Models via Prior-Guided Simulation —
- AMBIT: Anticipatory Multimodal Body Recruitment for Bimanual Tracking on a Humanoid —
- AI-Driven Collaborative Assembly Line Inspection: System Integration and Deployment Challenges —
- A Disk-Shaped Magnetoelastic Torque Sensor for Robotic Joints Using Permanent Magnetization —
- Steer2Grasp: Inference-Time Embodiment-Aware Steering for Diverse Physically Feasible Grasp Diffusion —
- FoLD: Force-Informed Learning for Dexterous Articulated Object Manipulation —
- Grid-Forming E-STATCOMs for Stable Integration of Large-Scale Data Centers: Modeling and Control —
- SLIP-VLA: Single-Step Latent Imagination for Policy Learning in Vision-Language-Action Models —
- Control Data Scheduling over Shared Communication Channels: A Sparse and Collision-Free Mechanism —
- Beyond One-Step Accuracy: State-Affine Latent Transition for Reliable Visual Planning —
- A Minimum-Order Functional Observer Beyond the Darouach and Luenberger Constructions —
- Hierarchical Multi-agent Reinforcement Learning for Warehouse Robot Coordination under Communication Loss —
- InfraVLA: Extending Vision-Language-Action Navigation with Infrastructure Cameras —
- Demonstration-Free Success-Probability Reward Learning for Generalist Robot Policies —
- Nonlinear Participation Factor-based Power System Model Reduction Addressing Near-Resonance Conditions —
- Observability-Informed Optimal Sensor Placement for Soft Robots —
- MomWorld: Momentum-Aware Latent World Model for Long-Horizon Autonomous Driving —
- AnyStep-WAM: Budget-Aligned Distillation and Adaptive Inference for World Action Models —
- Principal Steering Subspaces for Online Adaptation of Frozen Generative Robot Policies —
- Deep Behaviour Cloning of Model Predictive Control for Real-Time Operation of a Hydrogen-Diesel Dual-Fuel Engine —
- CodeActionBench: Evaluating Agentic Code-as-Policy for Embodied Manipulation —
- Achieve What You Imagined: Learning to Align Actions with Visual Plans —
- DeltaSeek: Toward Active Perception in Evolving Construction Environments —
- Robot-GST: geometry-aware spatial-temporal robot policy representation and evaluation —
- DexTaG: Tactile-as-Guidance in Reinforcement Learning for Dexterous Manipulation —
- Residual Learning-Based Control of Vehicle Platoons with 2 Stability Guarantees via Recurrent Equilibrium Networks —
- ArticulateArena: A Metric for Articulated Kinematics —
- EpiTransfer: Sparse, Training-Free Long-Range Depth Estimation from Temporal Monocular Aerial Frames —
- ReSync: Re-Aligning the Two Clocks of Asynchronous World-Action Models —
- Integrity Detection and Characterization of Malicious Injections in RAVEN II —
- FINGR: Learning Dexterous Hand Control for Real-World Rubik's Cube Solving —
- Test-Time Spatial Reasoning for Robot Manipulation Using Generative Real-to-Sim —
- Regime-Dependent Value of CVaR in Preventive Maintenance Scheduling under RUL Uncertainty —
- TacGooseBumps (TacGB): Retrofitting Normal-Only Tactile Sensors with Shear Encoding for Learning Contact-Rich Manipulation —
- ZeroBot: Learning from Scratch in Minutes with Generative Real2Sim —
- Estimate, Don't Imitate: Reusing Differentiable State-Based Policies for Visuomotor Control —
- Quantile Head for Vision-Language-Action Models —
- AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving —
- Integrated Thermal and Power Management for Wave-Powered Subsea Data Centers via Nonlinear Model Predictive Control —
- Beyond Retrieval Relevance: Scene-Grounded Risk Entailment for Vision-Language Driving —
- Reliability-Aware Sparse Route Memory for Round-Trip Vision-Language Navigation —
- RAVEL: Asynchronous Rolling Inference for Flow-Based Vision-Language-Action Models —
- A Weak Notion of Symmetry for Control Systems —
- FailPatch: Failure Residual Patching for Vision-Language-Action Models —
- Unified Visual-Tactile-Action Modeling from Human Demonstrations for Dexterous Manipulation —
- WB-WAM: Heterogeneous Body-Hand Pre-training for Humanoid Loco-Manipulation —
- RLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning Engineers —
- mmHRI: Towards Privacy-Preserving Human-Robot Interaction with Millimeter-Wave Radar —
- Proprioceptive Force Estimation for Quadruped Locomotion and Human-Robot Interaction —
- GAE: General Action Expert for Real-Time Humanoid Teleoperation —
- WAM-OPD: Sharpening World Action Models via On-Policy Distillation —
- UMR: Universal Manipulation Representation —
- RoboICL: Embodied In-Context Learning with GPT-6 Astra —
- SAGE: Symbolic Action-Gating and Editing for LLM Task Planners —
- Bayesian Active Learning for Intent Disambiguation in Interactive Robot Planning —
- NavHarness: Towards Lifelong Embodied Navigation —
- TLC-DiT: Task-Aligned Local Visual Conditioning for Robust Multitask Robot Manipulation —
- When World Models Lie: Adaptive Safety Analysis Under Wrong Imaginations —
- Predictive Semantic Safety: From Visual Physical Reasoning to Safety-Critical Control —
- FutureDuet: Decoupling Observation Access from Future Supervision in World Action Models —
- Derivative-Free Generalized Multivariable Super-Twisting Control for Constrained Euler-Lagrange Systems —
- RoboIRGBench: Benchmarking Implicit Referential Grounding in Vision-Language-Action Models —
- Contact-Aware Impedance Controller for Robot-Assisted Ultrasound Imaging —
- From Language to Task Maps: Compiling Semantic Relations While Preserving Task-Relevant Freedom —
- From World Models to World Action Models: Rethinking Next-State Prediction —
- ARS: Agentic Reward System for Robot Learning —
- Trajectory-Safe Orienteering for Human-Robot Shared Environments —
- Model-Informed Safe Reinforcement Learning for Bipedal Locomotion via Step-to-Step Prediction —
- Robot-Assisted Deployment and Maintenance of Inflatable Modules for Lunar Habitation: A Field Demonstration —
- MonoEgo: Monocular Metric Egocentric Demonstration Capture with Passive Wrist Constellations and Sparse Workstation Anchors —
- Frequency Measurement Practices for Inertia Assessment in Inverter Based Resources Dominated Power Systems —
- Gaze Prompts: Temporally Dense Human Attention for Vision-Language-Action Fine-Tuning —
- Where Memory Belongs: Ledger, an Object Ledger for Memory-Augmented VLAs —
- Robust Variable-Horizon MPC for Landing a Multirotor UAV on a Moving Platform —
- On the Achievable Inertia Constant of Inverter Based Resources —
- Aerial GRIPPER: A Gradient-based Real-time Inverse-game Predictor and Planner —
- Efficient World Action Model Inference with Adaptive Intermediate States —
- On the Numerical Reliability of Differentiable Physics-Based Optimization for Robotic Material Manipulation —
- HOI-Retarget: Contact-Centric Retargeting for Human-Object Interaction —
- Natural State-Prediction Accuracy can Hide Weak Controlled Responsiveness in VLA Readouts —
- Sufficiency of Zeroth-Order Reward Shaping for Policy Gradient in Stabilization Control —
- MarsLab: A Martian Rover Simulator for Planetary Rover Autonomous Navigation —
- SOR-Nav: Search or Relocate? Context-Gated Exploration and Cross-Region Relocation for Object Navigation —
- ExcavaTwin: Training-Free Geometry-Guided Semantic Elevation Mapping for Autonomous Excavation —
- DexWeave: Learning Dexterous Humanoid Loco-Manipulation from Human Demonstrations —
- Action Sequence Transfer via LLMs for Heterogeneous Environments —
- Simulation for Planetary Robotic Perception and Autonomy: A Concise Survey of Recent Capabilities and Gaps —
- CoHuB: A Simulation Benchmark for Multi-Humanoid Collaboration —
- AGRO-SUVIDE: Agentic Robotics for Surgical Viscoelastic Debridement —
- Don't Throw Away the Tail: Action Upcycling for Policy Acceleration —
- RoboFL: Federated Expert Assembly for World Action Models —
- NavJev: Efficient Vision-Language Navigation via Action-Centric Visual Compression and Discriminative Action-Semantic Memory —
- Graph-Based Simultaneous Path and Foothold Planning for Multi-Limbed Intra-Vehicular Robots in Space Stations —
- Learning to Act under Visual Interruptions with Vision-Language-Action Models —
- Do Not Cut When Uncertain: Rejectable and Calibrated Decision Heads for VLA Policies in Robotic Harvesting —
- EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning —
- QuadHand: A Compact Quadrotor Aerial Manipulator with MRC-SDF-Based Whole-Body Motion Planning —
- Two-Timescale Reinforcement Learning for Real-Time Optimization and Economic NMPC: Experimental Validation —
- Strategically Robust Game-Theoretic Multi-Agent Trajectory Optimization —
- ReCAT: Remember, Count, and Time: Structured Recurrent Memory for Robot Manipulation —
- Zero-Shot Reactive Obstacle Avoidance for Generative Robot Policies —
- Hard-Constrained Probabilistic Factor Graph Neural Network for Distribution System State Estimation under Non-Gaussian Uncertainty —
- Spatial Grafting: Grounding 3D Features for Flow-Matching Robot Policies —
- Repair Before You Fuse: Frozen-Host Adaptation for Corrupted-but-Present Sensors —
- GuardPIBT: Counterfactually Gated Neural Guidance for Ultra-Large-Scale 3D Multi-Agent Path Finding —
- DexAgent: An Agentic Human2Sim2Robot Framework for Dexterous Manipulation with Self-Evolving Tool Library —
- From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations —
- Adaptive Safety Filtering for Frozen ACC Policies via Conformal Residual Calibration —
- Scalable Small-Signal Stability Assessment of Power Systems Based on Frequency-Domain Quadratic Constraints —
- Memory in the Sky: Low-Altitude Question Answering with Multi-Agent Memory Aggregation —
- Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence —
- Revision, Not Restart: Revisable Visual Plans for Closed-Loop World-Action Models —
- Uni-VLaT: Whole-Body Tactile Adaptation of VLA Policies for Humanoid Loco-Manipulation —
- Rethinking Causal Action Tokenization with Conditional Annealing in Flow Matching —
- CoBrush: A Hierarchical Planning Framework for Human-Robot Co-Painting —
- Robot Tool Design from Scratch via Behavior-Aware Hierarchical Optimization —
- ForVis: An In-Field Dataset and Benchmark for VIO Using Under-Canopy UAV Flights in Forests —
- Terrain-Aware Autonomous Planetary Exploration for Exteroceptive-Proprioceptive Mapping with Quadruped Scouts —
- Inspection-SPARS: Task-Oriented Sparse Roadmaps for Inspection Planning —
- EdgeVLN: Runtime-Aware Deployment Ready Quantized Vision Language Navigation Model —
- F4R: Failure-Driven Recognition, Reconstruction, Refinement, and Redeployment for Continual Robot Self-Improvement —
- CollisionSplatting: Collision-Aware Motion Planning in 3DGS Scenes with Image-Conditioned Objectives and Adjustable Conservatism —
- Denoising Multi-Robot Trajectories —
- MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining —
- Agent Priors-guided Policy Learning —
- LQR-ArUco Fusion: Robust Hierarchical Control for Navigation and Asymmetric Manipulation in Two-Wheeled Robots —
- Humanoid Loco-Manipulation With Discrete VLA Model —
- RoboCompiler: Graph-Native Compilation of Closed-Chain Robots for Consistent Modeling, Control, and Simulation —
- Statistical Learning of Contractive Dynamical Representations for Composite Adaptive Control —
- Human-Guided Planning for Complex Manipulation Tasks Using the Screw Geometry of Motion —
- DexRoam: Learning Mobile Bimanual Dexterous Manipulation from Egocentric Whole-Body Human Demonstrations —
- Easier Said Than Done: Unpacking Intent-Behavior Gap in Jailbreaking LLM-Based Robots —
- On Uniformly Time-Varying Control Barrier Functions —
- Dynamic Buffers: Cost-Efficient Planning for Tabletop Rearrangement with Stacking —
- A multi-modal tactile fingertip design for robotic hands to enhance dexterous manipulation —
- SPIDER: Scalable Physics-Informed Dexterous Retargeting —
- Unifying Deep Predicate Invention with Pre-trained Foundation Models —
- Helical Tendon-Driven Continuum Robot with Programmable Follow-the-Leader Operation —
- LogicEnvGen: Task-Logic Driven Generation of Diverse Simulated Environments for Embodied AI —
- From Instruction to Event: Sound-Triggered Mobile Manipulation —
- Learning Geometrically-Grounded Amodal 3D Representations for View-Generalizable Robotic Manipulation —
- Self-Evolutionary Replanning for Failure-Aware Motion Planning —
- Input Dexterity and Output Negotiation in Feedback-Linearizable Nonlinear Systems —
- Dynamic Model Identification and Gravity Compensation for the dVRK-Si Patient Side Manipulator —
- A Tutorial on Learning-Based Radio Map Construction: Data, Paradigms, and Physics-Awareness —
- A Modular Platooning and Vehicle Coordination Simulator for Research and Education —
- Safety-Constrained Optimal Control and Hamiltonian-Gradient Learning for Systems with Unknown Dynamics —
- A Two-Stage Optimization Framework for Validating Electric Vehicle Charging Infrastructure under Grid Constraints —
- FastGrasp: Learning-based Whole-Body Control Method for Fast Dexterous Grasping with Mobile Manipulators —
- Energetic Resilience under Temporal Logic Specifications —
- Task-Driven Co-Design of Heterogeneous Multi-Robot Systems —
- Learning Control Policies to Provably Satisfy Hard Affine Constraints for Black-Box Hybrid Dynamical Systems —
- Region of Attraction of a Backstepping Observer for the Quasilinear Heat Equation —
- PrefMoE: Robust Preference Modeling with Mixture-of-Experts Reward Learning —
- The Potential Welfare Gains from Curtailment Trading Under Non-Firm Interconnection —
- RoboAlign-R1: Distilled Multimodal Reward Alignment for Robot Video World Models —
- When to Trust Imagination: Adaptive Action Execution for World Action Models —
- ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation —
- StereoPolicy: Improving Robotic Manipulation Policies via Stereo Perception —
- FlyMirage: A Fully Automated Generation Pipeline for Diverse and Scalable UAV Flight Data via Generative World Model —
- Scalable Distributed Learning-Based MPC for Heterogeneous Consumer-Prosumer Building Aggregations: MPC-Aware Feature Selection and Convex Constraint-Coupled Optimization —
- IDOL: Inverse-Dynamics-Guided Future Prediction for End-to-End Autonomous Driving —
- DriveAnchor: Progressive Anchor-based Flow Learning for Autonomous Driving Planning —
- PACE: Phase-Aware Chunk Execution for Robot Policies with Action Chunking —
- Enhancing Offshore Wind Simulations: Interpolated Switching via DLL Black-Boxes —
- Tac2Pix: Image-Space Visuo-Tactile Fusion for Dexterous Manipulation —
- Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination —
- Robust Zonotopic Control —
- Assistron: Bayesian Shared Autonomy with Off-the-shelf Vision-Language-Action Models —
- SurgVIL: Scaling Surgical Robot Imitation Learning with Open-source Surgical Videos —
- RoboEdit: Turning Human Manipulation Videos into Scalable Robot Experience —
- What Matters for Latent Actions in Robot Learning —
- What Stops Recursive Self-Improvement in Robotics? Lessons from 123 Rounds of Agentic Skill Discovery —
- Robot Manipulation with GPT-6-Astra: Body Knowledge, Experience Reuse, Emergent Skills, and Sim2Real Transfer —
- Reduced Cartesian Kinetostatics for Tendon-Driven Continuum Robots: Residual-Stabilized Full-Shape Propagation —
- Timed Rule-Based Supervision of an End-to-End Autonomous Parking Policy —
- Path Planning with Motion Primitives in Dynamic Environments: SIPP on Lattices —
- Humanoid Badminton: Learning Dynamic Racket Skills from Limited Human Motion Data —
- An Exact Lyapunov Characterization of Regional Rate Performance for Rational Stability —
- PHIRL: Aligning Learned Rewards with Task Progress for Inverse Reinforcement Learning —
- Efficient Bezier Velocity Optimization for Free-Floating Space Manipulators —
- Differentiable Dynamics for Autonomous Micro-Mobility Navigation —
- Game-Theoretic Control with Constrained Potential Surgery —
- GT-VLA: Target-Conditioned Trace Guidance for Generalizable Robotic Manipulation —
- End-to-end QP-based policies: A unified perspective on robust control and robot learning —
- HapticWorld: an Interactive World Simulator with Real-time Torque Feedback —
Important terms
- PHIRL
- This method aligns learned rewards with actual task progress in inverse reinforcement learning. It helps agents learn optimal behaviors directly from demonstrations instead of relying solely on trial and error.
- DS-VLA
- A dendritic-inspired model that combines vision, language, and action for robust control. It aims to make complex models more reliable when interacting with the physical world.
- VPTwin
- This focuses on real-sim-real video prediction specifically for robotic manipulation planning. It seeks to bridge the gap between what happens in simulation and what actually occurs physically.
- AquaBEV-Nav
- A system designed to learn occupancy in underwater environments for navigation. This is crucial for mapping unknown spaces and allowing autonomous vehicles to explore underwater conditions.