Robotics papers — 2026-10-02
Today's focus is on how we can make robotic foundation models generate actions more reliably, which is crucial for truly general intelligence in physical systems. The work that matters most involves Kinematic MeanFlow, which attempts to create a one-step action generation policy by focusing on kinematic mean flow. This approach simplifies the complex decision-making process for robots by looking at average motion patterns.
This offers a potentially more stable way to plan movement than purely reactive methods. Then there is FuncBridge, which tackles functional tool use generalization by using keypoint trajectory reasoning. This means it tries to understand the path a specific body part takes when interacting with a tool instead of just seeing what an object looks like.
This is important for making robots capable of using tools in novel ways. SlotVLA builds on this idea by focusing on modeling object-relation representations during robotic manipulation tasks. This helps the model grasp how different objects relate to each other in space before it can plan an action involving those objects.
DynamicVLA also contributes to this area by developing a vision-language-action model specifically for dynamic object manipulation. This work addresses the challenge of handling moving parts, which is a common hurdle in real-world robotics. It shows how to integrate perception and action planning for these scenarios.
Finally, Rewind-IL focuses on online failure detection and state respawning within imitation learning frameworks. This mechanism allows a system to recover gracefully when an initial plan fails during execution. This provides robustness to the learned policies developed by models like Guide, Think, Act which explore interactive embodied reasoning.
The most significant progress today involves the development of World Motion Models, which focuses on flexible sequence modeling of SE(3) trajectories. This work is crucial because it allows systems to understand and predict complex movements in three-dimensional space over time. A key component explored was the integration of these models with other planning frameworks to ensure smooth and coherent motion generation.
Another important piece of work is UniTrackPLA, which presents a unified panorama language action model designed for instruction-guided navigation and dynamic person tracking. This means it aims to create a single system that can follow complex verbal commands while simultaneously keeping track of moving individuals in an environment. This approach builds upon the foundational ideas seen in NarrativeFlow, which uses flow-based vision-language-action models employing robot velocity fields to achieve similar instruction following.
Then there is DexPolicy, which deals with scheduled exploration for trajectory-guided dexterous manipulation. This work investigates how to plan a sequence of actions for a robotic arm by scheduling when and where it should explore different states. This is vital for complex physical tasks, contrasting with the more general navigation focus of UniTrackPLA.
ALFRED addresses long-term plant monitoring through requirement-driven development of an open-source mobile manipulator. This system is designed to autonomously monitor plants over extended periods based on specific operational needs, moving beyond short-term reactive control.
Finally, we see work on deployment focused protocols for interpreting tracking evaluation in pedestrian-centric environments. This moves beyond simple leaderboard scores to assess how well a tracking system performs in real-world human interaction scenarios. This contrasts with the more theoretical trajectory modeling efforts.
The most significant development today centers on how we can make robots learn complex physical tasks directly from experience, which is crucial for autonomous systems. InterEvolve explored test-time evolution of reward programs for humanoid locomotion and manipulation. They were trying to dynamically change the goals a robot pursues while it is operating, aiming to give the robot flexibility in real-world scenarios where pre-programmed instructions might fail.
This builds upon work like AdaptManip, which focuses on learning adaptive whole-body object lifting and delivery using online recurrent state estimation. This essentially teaches the robot how to adjust its movements as it interacts with an object.
A related piece, FAME, introduced force-adaptive reinforcement learning for expanding the manipulation envelope of a full-scale humanoid by adjusting its behavior based on sensed forces. This suggests that robots can be more robust when handling physical interactions than if they rely solely on fixed control policies.
A key challenge addressed is bridging the sim-to-real gap, which is tackled by multipanda ros2, a real-time ROS2 framework designed for multimanual systems to handle the complexity of interacting with multiple tools simultaneously. This framework connects directly to constant-time planning for chaining collision-free motion to manipulation behaviors.
This deals with generating safe paths through complex obstacle spaces while executing intricate actions. Finally, the work on region based SLAM aware exploration presents a strategy for efficient and robust autonomous mapping that can scale. This provides the necessary spatial awareness for these complex physical tasks to operate in unknown environments.
The most significant development today involves PACE, which explores how to adapt robot personas through conversational elicitation. Understanding how robots interact socially is key to building useful human-robot systems. This work investigates the process of shaping a robot's personality by having humans guide that shaping through conversation.
A related piece of research focuses on G2-Nav, which creates grounded and guarded vision-language costmaps for robot social navigation. This aims to help robots move around people safely while maintaining a certain level of social awareness. Safe and socially aware navigation is crucial for any real-world deployment.
Then there is the work on Dex-X, which learns visual-tactile dexterous manipulation from human videos using simulated interaction. This suggests a path toward teaching robots complex physical skills by observing humans perform them in a virtual setting, building upon the idea of learning from demonstration but adding tactile feedback.
IndoorBEV presents a lightweight real-time LiDAR BEV perception system for indoor mobile robots. This is vital because it provides the necessary spatial awareness for robots operating in cluttered indoor environments without needing heavy computational resources. This perception system underpins many other navigation tasks.
The optimization work on sequential object placement using convex decomposition addresses how to efficiently plan the order in which a robot should place objects. It breaks down complex placement problems into simpler, manageable convex shapes, which is a core planning technique that helps robots execute multi-step tasks smoothly.
Finally, TOAST introduces stochastic robot action tokenization for autoregressive vision-language-action models. This tackles how to structure the sequence of actions a robot takes when using these advanced models. This addresses the practical challenge of translating high-level plans into executable robot commands.
The most significant development today involves eRLT, which tackles the challenge of making large vision-language models more efficient for robotic control by using action relevant token routing. This directly addresses the computational bottleneck in deploying complex reinforcement learning policies to real-world visual tasks.
This work builds upon ideas from Divide-and-Remember, which focuses on creating recursive action relevant memory to handle long horizon policies. That memory structure is crucial because it allows the system to retain and reuse past experiences effectively during extended interactions.
Another key piece of research explores external photoreflective tactile sensing based on surface deformation measurement. This provides a way for robots to feel their environment without relying solely on vision, giving the model richer physical feedback about contact.
We also saw work on frequency aware decomposition learning for sensorless wrench estimation in vibration rich robotic contact. This helps estimate forces and torques even when direct sensors are unavailable, which is vital because accurate force information is necessary for safe and dexterous manipulation tasks.
BORA addresses the gap between offline reinforcement learning and online residual adaptation by bridging them to create more robust real-world dexterous vision language agent models. This method attempts to leverage pre-collected data while allowing the model to adapt incrementally in live operation.
FlashNav introduces a method for training deployable robot navigation policies incredibly quickly. This suggests a path toward rapid deployment of complex navigational skills, which is important for practical applications where quick adaptation is needed on the fly.
Finally, direct action head injection of a grounded three dimensional point unlocks spatial and task generalization by directly feeding actions into the model's spatial understanding. This technique aims to improve how the model understands and executes actions in 3D space, which underpins many complex manipulation goals.
The most significant development today concerns ACE, which introduces agentic control for embodied manipulation through zero shot workflow reasoning. This suggests a new way to give complex AI agents the ability to plan and execute tasks without needing extensive retraining for every single new scenario.
This is built upon work that explores how to make simulation demonstrations useful for governance benchmarks, specifically Bounded-Fidelity Sim-as-Demo-Stage, which focuses on mocap handoff. This helps ground the agent's actions in real movement data rather than just abstract goals.
Another key piece involves multi reference path tracking control for an agricultural tractor using nonlinear model predictive control. This addresses how a machine can follow a desired path even when its physical model isn't perfectly linear, which is crucial for real-world farming applications.
We also saw research into probabilistic plan legibility when using off the shelf planners. This looks at making the AI's intended plan understandable to humans, connecting to humanoidttt, which focuses on test time capability reuse for efficient humanoid control. This shows how to make control systems faster during operation.
Finally, there is work on decentralized safe path following for multiple quadrotors navigating intersecting paths with theoretical guarantees. This speaks directly to the safety challenges in multi agent environments where collision avoidance needs mathematical proof.
The most significant development this morning concerns the work on token world modeling, which aims to build a representation of the physical world directly within the vision-language model's token space. This is crucial because it suggests a more integrated way for robots to understand and interact with their environment.
Researchers explored how to achieve this by training models using data that links visual observations directly to manipulation actions. This essentially teaches the model what objects look like in a way that is useful for planning physical tasks, contrasting with previous methods that might treat vision and language as separate inputs.
Another important piece of research focused on whole-body aerial grasping and lifting using only partial visual observations. This addresses the practical challenge of robotic dexterity in real-world settings by demonstrating a method where a robot can successfully grasp and lift objects by relying on incomplete visual data.
This contrasts with the work on humanoid locomotion models that pretrain using egocentric whole-body human data to achieve general manipulation capabilities. That approach focuses more on learning complex movement patterns from human demonstrations rather than direct vision-to-action mapping for grasping.
ScaffoldM3C presented a multimodal sequential Monte Carlo framework designed for generative stable construction planning. This is significant because it tackles the complex sequencing required for building things in a way that maintains stability throughout the process. It builds upon prior work by integrating visual and sequential planning into a single coherent structure.
The findings from same scene, different task show how skill alignment can improve compositional generalization in vision-language models. This means these models can transfer skills learned in one context to perform novel tasks in similar contexts, which is a step toward making AI systems more adaptable.
Finally, the investigation into admissibility-preserving control for multi-input systems with joint capacity constraints addresses the safety and reliability of complex robotic systems. This ensures that control actions remain within safe operational limits even when multiple inputs are involved, providing necessary guardrails for deploying these increasingly capable models.
Today's papers
- Kinematic MeanFlow: One-Step Action Generation Policy for Robotic Foundation Models. [paper]
- FuncBridge: Towards Functional Tool-Use Generalization via Keypoint Trajectory Reasoning. [paper] [episode]
- UrbanVLA: A Vision-Language-Action Model for Urban Micromobility. [paper] [episode]
- SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation. [paper] [episode]
- DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation. [paper] [episode]
- Rewind-IL: Online Failure Detection and State Respawning for Imitation Learning. [paper] [episode]
- Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models. [paper] [episode]
- DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies. [paper]
- Retrospective Open-Vocabulary Memory for Long-Term Object Search. [paper]
- DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation. [paper]
- UniTrackPLA: Unified Panorama-Language-Action Model for Instruction-Guided Navigation and Dynamic Person Tracking. [paper]
- A Survey on End-to-End Autonomous Driving Training from the Perspectives of Data, Strategy, and Platform. [paper]
- NarrativeFlow: Flow-Based Vision-Language-Action Model Using Robot Velocity Fields. [paper]
- ALFRED: Requirement-driven development of an open-source mobile manipulator for long-term plant monitoring. [paper]
- Beyond Leaderboard Scores: A Deployment-Focused Protocol for Interpretable Tracking Evaluation in Pedestrian-Centric Environments. [paper]
- World Motion Models: Flexible Sequence Modeling of SE(3) Trajectories. [paper]
- End-to-End Learning vs. Modular Architectures: Comparative Insights into Autonomous Driving Systems. [paper]
- InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation. [paper]
- Region Based SLAM-Aware Exploration: Efficient and Robust Autonomous Mapping Strategy That Can Scale. [paper] [episode]
- Constant-Time Planning for Chaining Collision-free Motion to Manipulation Behaviors. [paper] [episode]
- Bridging the Sim-to-Real Gap with multipanda ros2: A Real-Time ROS2 Framework for Multimanual Systems. [paper] [episode]
- AdaptManip: Learning Adaptive Whole-Body Object Lifting and Delivery with Online Recurrent State Estimation. [paper] [episode]
- FAME: Force-Adaptive RL for Expanding the Manipulation Envelope of a Full-Scale Humanoid. [paper] [episode]
- Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models. [paper] [episode]
- PACE: Persona Adaptation through Conversational Elicitation in Human-Robot Interaction. [paper] [episode]
- G2-Nav: Grounded and Guarded Vision-Language Costmaps for Robot Social Navigation. [paper] [episode]
- Sequential Object Placement Optimization with Convex Decomposition. [paper] [episode]
- Dex-X: Learning Visual-Tactile Dexterous Manipulation From Human Videos with Simulated Interaction. [paper] [episode]
- IndoorBEV: A Lightweight Real-Time LiDAR BEV Perception System for Indoor Mobile Robots. [paper] [episode]
- Real-Time Human-Adaptive Task Allocation for Multi-Human Multi-Robot Supervision. [paper] [episode]
- TOAST: Stochastic Robot Action Tokenization for Autoregressive Vision-Language-Action Models. [paper] [episode]
- Screw Attention: Rigid-Body Algebra Inside a Transformer. [paper] [episode]
- eRLT: Efficient VLA Reinforcement Learning via Action-Relevant Token Routing. [paper] [episode]
- Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies. [paper] [episode]
- External Photoreflective Tactile Sensing Based on Surface Deformation Measurement. [paper] [episode]
- Frequency-aware decomposition learning for sensorless wrench estimation in vibration-rich robotic contact. [paper] [episode]
- BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models. [paper] [episode]
- Deformable In-Hand Slip-Aware Tactile Sensor with Integrated Velocity Sensing, Force/Torque and Pressure Map Estimation. [paper] [episode]
- FlashNav: Training Deployable Robot Navigation Policies in Seconds. [paper] [episode]
- Direct Action-Head Injection of A Grounded 3D Point Unlocks Spatial and Task Generalization. [paper] [episode]
- ACE: Agentic Control for Embodied Manipulation via Zero-shot Workflow Reasoning. [paper] [episode]
- Bounded-Fidelity Sim-as-Demo-Stage: Mocap Handoff for Governance Benchmarks. [paper] [episode]
- Multi-Reference Path Tracking Control for an Agricultural Tractor with Nonlinear Model Predictive Control. [paper] [episode]
- Probabilistic Plan Legibility with Off-the-shelf Planners. [paper] [episode]
- HumanoidTTT: Test-Time Capability Reuse for Efficient Humanoid Control. [paper] [episode]
- The Geometry of Time: Horizon-Independent Feasibility and Repair for STL.
- Decentralized Safe Path Following for Multiple Quadrotors on Intersecting Paths with Theoretical Guarantees. [paper] [episode]
- Real-Time Whole-Body Safe Motion Generation for Multi-Segment Tendon-Driven Continuum Robots. [paper] [episode]
- DeepJEPA: Scaling World Models from Within. [paper] [episode]
- Whole-Body Aerial Grasping and Lifting via Partial Visual Observations. [paper] [episode]
- Towards a General Humanoid Loco-Manipulation Model via Egocentric Whole-Body Human Data Pretraining. [paper] [episode]
- ScaffoldM3C: A Multimodal Sequential Monte Carlo Framework for Generative Stable Construction Planning. [paper] [episode]
- Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs. [paper] [episode]
- Admissibility-Preserving Control for Multi-Input Systems with Joint Capacity Constraints. [paper] [episode]
- Does Continual Imitation Learning Remain Grounded? A Language-Perturbed Benchmark for Robotic Task Retention. [paper] [episode]
- Token-World: World Modeling in Vision-Language Model Token Space for Robot Manipulation. [paper] [episode]
- When Reasoning Helps Action: Monitoring and Steering Chain-of-Thought in Vision-Language-Action Policies. [paper] [episode]
- TacDyn-WAM: Learning Implicit Tactile Dynamics in a Heterogeneous Visuo-Tactile World Action Model. [paper] [episode]
- Toward Humanoid Robots in Construction: A Teleoperation Feasibility Study. [paper] [episode]
- CF-JEPA: Improving Robustness of JEPA World Models via Controllability Factorization. [paper] [episode]
The papers
- A Geometric Decision Procedure for STL Feasibility and Repair — Signal Temporal Logic (STL) control synthesis frequently encounters physical infeasibility due to actuator limits or flawed task deadlines, and this paper presents a geometric decision procedure that evaluates physical feasibility completely independently of the temporal horizon [episode]
- FlashDexRetarget: Accelerating Dexterous Manipulation Data Generation through Multi-Motion Retargeting — FlashDexRetarget introduces an RL-based framework for high-success, efficient dexterous motion retargeting by jointly training a single policy across multiple human hand-object demonstrations. [episode]
- FlashNav: Training Deployable Robot Navigation Policies in Seconds — FlashNav presents a GPU-first framework for ultra-fast range-based robot navigation training, achieving seconds-level policy training by aligning simulation with the navigation MDP. [episode]
- ActiveWAM: Evidence-Aware Active Vision for World-Action Models — ActiveWAM introduces a unified world-action model that learns observation and manipulation jointly by formulating active vision manipulation as an evidence-aware retain–acquire problem. [episode]
- Decentralized Power-Optimal Coordination for Spacecraft Swarms Using Time-Varying Magnetorquer Actuation — Decentralized power-optimal coordination for magnetically actuated spacecraft swarms using time-varying magnetorquer actuation addresses the challenge of forming large space structures from spacecraft by developing a framework that jointly derives interaction graphs, frequency gr [episode]
- Minimal Experiments for Robust Stabilization: Information, Spectral Geometry, and Duration — Short input sequences can provide robust data-driven stabilization for broad classes of linear systems, and this paper investigates how minimal experiments—defined by information content and duration—can achieve a fixed fraction of the robustness guaranteed by an optimal plan [episode]
- Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies — Vision–language–action (VLA) models struggle on history-dependent manipulation tasks where current observations alone do not determine action, necessitating a memory mechanism to retain relevant past information. [episode]
- Vulnerability-Weighted Routing of Timing-Critical Nets for Configuration-Upset-Resilient SRAM-Based FPGAs — Conventional FPGA routing optimizes timing, congestion, and routability but does not distinguish routes with similar nominal performance and substantially different susceptibility to configuration-induced delay degradation. [episode]
- Recova: Agent-Guided Failure Recovery for Autonomous Robotic Manipulation — Manipulation failures can leave scenes from which a task policy cannot recover, and Recova presents an agent-guided framework that jointly develops task execution and recovery in a reconstructed digital twin, then verifies and refines both through real-world experience. [episode]
- Continual Learning for 6-DoF Grasp Synthesis via Experience and Demonstrations — Continual learning for 6-DoF grasp synthesis addresses the limitation where fixed grasping models fail in novel deployment environments by introducing an adaptive framework that updates grasp scores and recalls user demonstrations without retraining network weights. [episode]
- UniWAM: Unified World-Action Model — UniWAM introduces a unified architecture that integrates a physical reasoner, a world generator, and an action predictor to jointly learn semantic understanding of the physical world, visual generation, and action prediction. [episode]
- Cross-entropy optimization with prioritized constraints — When constraints conflict, an optimizer must determine which requirements to preserve and which to relax. [episode]
- Dex-X: Learning Visual-Tactile Dexterous Manipulation From Human Videos with Simulated Interaction — Human videos are an abundant source of dexterous manipulation behaviors, but they lack tactile information that is crucial for contact-rich interaction. [episode]
- Mixed Bernstein-Fourier Approximants for Optimal Trajectory Generation with Periodic Behavior — Mixed Bernstein-Fourier approximation methodology provides a robust, theoretically grounded, and computationally efficient approach for advanced optimal trajectory planning in autonomous systems. [episode]
- Identifiability Limits of Forced Oscillation Sources in Power Systems — Whether a forced-oscillation source can be uniquely localized depends jointly on the available measurements and the candidate intervention dictionary. [episode]
- PACE: Persona Adaptation through Conversational Elicitation in Human-Robot Interaction — Equipping humanoid robots with coherent and adaptable personas is crucial for fostering natural, engaging, and trustworthy human-robot interaction (HRI), as existing static approaches lack flexibility in adapting to individual user contexts. [episode]
- SkeleWAM: Skeleton World-Action Modeling for Efficient Robotic Manipulation — World action models (WAMs) combine robot action generation with future state prediction, and SkeleWAM introduces a compact WAM that represents manipulation scenes as sparse 3D skeletons composed of robot joints, object centers, and interaction points. [episode]
- TacDyn-WAM: Learning Implicit Tactile Dynamics in a Heterogeneous Visuo-Tactile World Action Model — World action models improve robotic manipulation by conditioning actions on predicted futures, yet existing tactile variants largely inherit video-generation pipelines that reconstruct future tactile observations through iterative denoising. [episode]
- AdaptManip: Learning Adaptive Whole-Body Object Lifting and Delivery with Online Recurrent State Estimation — AdaptManip presents a fully autonomous framework for humanoid robots to perform integrated navigation, object lifting, and delivery by training a robust loco-manipulation policy via reinforcement learning without human demonstrations or teleoperation data. [episode]
- Interactive Power Flow in the Browser — This paper introduces tellegen, an open source framework for interactive power flow (PF) and optimal power flow (OPF) studies that run in the browser. This provides intuitive and democratized access to power system analysis tools compiled to WebAssembly. [episode]
- BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models — Vision-Language-Action (VLA) models face significant challenges in real-world dexterous manipulation due to high degrees of freedom and compounding execution errors, necessitating a framework that bridges offline learning with online adaptation for reliable deployment. [episode]
- Continue, Abort, or Fall: Viability-Aware Policy Selection (VAPS) for Safe Humanoid Acrobatics — Dynamic humanoid motions, such as flips, risk hardware damage due to suboptimal policies or disturbances, and this work presents Viability-Aware Policy Selection (VAPS), which treats safety as a policy-conditioned, receding-horizon decision framework to select between continuing [episode]
- Real-Time Whole-Body Safe Motion Generation for Multi-Segment Tendon-Driven Continuum Robots — Real-time motion control for multi-segment tendon-driven continuum robots remains challenging due to spatially nonuniform structural properties and distributed collision risks across the entire continuous body. [episode]
- OrbitTAMP: Grounding Language Models for Task and Motion Planning in Spacecraft Rendezvous — Spacecraft rendezvous and proximity operations (RPO) are currently planned through an expertise-intensive process that creates a bottleneck for scalable operations, and this paper introduces a hierarchical framework to ground Large Language Model (LLM) reasoning in orbital dynami [episode]
- Whole-Body Aerial Grasping and Lifting via Partial Visual Observations — Aerial grasp-and-lift tasks require whole-body coordination across approach, acquisition, and lifting under partial target observations. [episode]
- ScaffoldM3C: A Multimodal Sequential Monte Carlo Framework for Generative Stable Construction Planning — Autonomously constructing physically realizable 3D structures remains a significant challenge due to combinatorial action spaces, interchangeable components, equifinal assembly sequences, and strict stability requirements during construction. [episode]
- TRACE: Privacy-Preserving Next-Best-View Selection over Distributed 3D Gaussian-Splat Maps — Share the light, not the map. This work introduces TRACE, a distributed next-best-view selection protocol designed for teams of robots to select optimal viewpoints over private 3D Gaussian Splatting maps without sharing any raw map data. [episode]
- DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication — DuoMind introduces a distributed hierarchical framework for multi-robot coordination that leverages vision-language models and vision-language-action models to enable robots to perform long-horizon tasks through semantic communication. [episode]
- Towards a General Humanoid Loco-Manipulation Model via Egocentric Whole-Body Human Data Pretraining — Humanoid whole-body manipulation has rapidly advanced, but existing supervision methods often lack coverage for whole-body coordination and hand–object interaction. [episode]
- eRLT: Efficient VLA Reinforcement Learning via Action-Relevant Token Routing — Vision-Language-Action (VLA) models provide strong behavioral priors for robotic manipulation, yet efficiently adapting them to downstream tasks remains challenging. [episode]
- Evolutionary Analysis of Continuous-time Finite-state Mean Field Games with Discounted Payoffs — We consider a class of continuous-time dynamic games involving a large number of players where individual state evolution and population-wide effects are explicitly modeled, which is necessary for many real-world applications. [episode]
- WBAG: A Whole-Body and Attached-Geometry Safety Framework for Vision-Language-Action Manipulation — Vision-language-action (VLA) policies have demonstrated impressive capabilities in generalizable robotic manipulation, but their deployment in the real world remains challenging due to potential collisions involving different parts of the robot, manipulated objects, and the surro [episode]
- HumanoidTTT: Test-Time Capability Reuse for Efficient Humanoid Control — HumanoidTTT introduces a framework for test-time capability reuse in continual humanoid control, addressing the challenges of reliable motion reuse under changing robot states and managing validated capabilities within finite storage. [episode]
- Decentralized Safe Path Following for Multiple Quadrotors on Intersecting Paths with Theoretical Guarantees — This research presents a novel decentralized controller for multiple quadrotors operating on intersecting paths, providing theoretical guarantees for collision avoidance and strict adherence to pre-assigned routes. [episode]
- Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents — Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. [episode]
- FAME: Force-Adaptive RL for Expanding the Manipulation Envelope of a Full-Scale Humanoid — Maintaining balance under external hand forces is critical for humanoid bimanual manipulation, where interaction forces propagate through the kinematic chain and constrain the feasible manipulation envelope. [episode]
- A Theory-Guided Advanced Regulatory Control Synthesis for Cooling-Limited Exothermic Semi-Batch Reactors — This paper develops a theory-guided approach to synthesize Advanced Regulatory Control (ARC) architectures for cooling-limited exothermic semi-batch reactors, addressing the design gap where existing methods lack systematic guidance for changing active constraints. [episode]
- ReCo: Response-Consistent Locomotion with Policy-Aware MPC for Legged Manipulation — Continuous legged manipulation requires accurate end-effector tracking while the base keeps walking, and this paper presents ReCo, a framework that couples response-consistent locomotion with policy-aware MPC for legged manipulation. [episode]
- Screw Attention: Rigid-Body Algebra Inside a Transformer — Learned manipulation policies can rediscover spatial relations from data, but they often lack robustness to geometric changes in the scene. [episode]
- Rewind-IL: Online Failure Detection and State Respawning for Imitation Learning — Rewind-IL is a training-free online safeguard framework designed for generative action-chunked imitation learning policies that provides two capabilities: (1) zero-shot real-time failure detection based on the internal self-consistency of the policy’s action and (2) state respa [episode]
- ACE: Agentic Control for Embodied Manipulation via Zero-shot Workflow Reasoning — Open-ended tabletop manipulation requires agents to adapt to dynamic environments and execution failures, which ACE addresses by introducing an agentic workflow reasoning framework that decouples high-level semantic planning from low-level physical control. [episode]
- Training-Free Diffusion Planning with Analytical Local Scores — Motion planning requires trajectories that are smooth, goal-directed, and collision-free in complex environments, and existing diffusion planners are limited by their requirement for large collections of feasible trajectories for training. [episode]
- GlassGuard: Verified Glass Plane Mapping for Robot Navigation — Transparent and specular surfaces pose a serious challenge to LiDAR-based SLAM and navigation because laser returns may pass through glass, leaving collision boundaries absent from the map. [episode]
- OpenSpace Lab Solution to the IROS 2026 Indoor Exploration Competition — This report presents OpenSpace Lab’s solution to the Competition on Intelligent Information Gathering for Single and Multi-Robot Systems Workshops at IROS 2026, detailing strategies that successfully achieved top rankings in both single-robot and multi-robot exploration tracks. [episode]
- Experience-Based Feasibility-Aware Generative Adversarial Imitation from Observation under Embodiment Mismatch — Imitation from observation (IfO) learning robot behaviors from state-only demonstrations is enhanced by introducing Experience-Based Feasibility-Aware Generative Adversarial Imitation from Observation (EF-GAIfO), a framework that estimates whether demonstrated state transitions a [episode]
- MASkillBlender: Decentralized Whole-Body Coordination for Multi-Humanoid Loco-Manipulation via Skill Blending — MASkillBlender proposes a general multi-agent reinforcement learning framework that enables decentralized whole-body coordination for multiple humanoids by learning a shared high-level policy over reusable pre-trained single-humanoid skills, which is significant because it achiev [episode]
- Getting Out and Getting Back: World and Behavior Grounding in Real2Sim2Real Co-Training — Simulation can expand scarce real demonstrations for co-training, yet how world fidelity and similarity to human behavior affect policy performance remains unclear. The gist: Grounded simulation remains beneficial when co-training foundation models. [episode]
- Reactive Humanoid Multi-Contact Using Learned Stability Models — Reactive humanoids can be stabilized in low-stability scenarios by reactively using hand contacts, which is crucial for maximizing reliability in real-world applications. The gist: A reactive (≤10 ms) contact planner that models recovery for arbitrary (e.g. [episode]
- CF-JEPA: Improving Robustness of JEPA World Models via Controllability Factorization — Controlling an agent with vision requires separating useful task information from irrelevant background noise, and this work introduces Controllability Factorized JEPA (CF-JEPA), a world model that splits its latent space into controllable and uncontrollable subspaces to improve [episode]
- DeepJEPA: Scaling World Models from Within — World-model planners typically scale outward by rolling farther, sampling more trajectories, or optimizing longer, while assigning the same computation to every imagined transition. [episode]
- Query-Conditioned Articulation Estimation from a Single Image — QueryArt is a model designed to estimate kinematic parameters of articulated objects from only a single RGB image and a 2D query point, which enables robots to infer object structure before physical interaction. [episode]
- Cost-Informed Learning for Aggregating Building HVAC Flexibility — This research develops a cost-informed learning framework to aggregate building HVAC flexibility by jointly learning surrogate parameters and an inner-approximation objective from downstream utilization costs, which addresses the limitation that existing methods fail to preserve [episode]
- TOAST: Stochastic Robot Action Tokenization for Autoregressive Vision-Language-Action Models — Autoregressive Vision-Language-Action models often represent continuous robot actions as discrete token sequences, and this work introduces TOAST, a novel stochastic action tokenization method that samples alternative tokenizations of the same quantized action sequence during pol [episode]
- Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks — Vision-Language-Action (VLA) models have achieved high task success rates on robot manipulation benchmarks, but this work proposes a benchmark-agnostic evaluation framework to measure behavioral robustness by characterizing how successful trajectories are executed under input per [episode]
- PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots — PROMO introduces Preference-Conditioned Multi-Objective Reinforcement Learning (MORL) for quadrupedal robots, addressing the limitation of fixed scalar rewards by allowing operator intent to explicitly condition locomotion trade-offs at runtime. [episode]
- Direct Action-Head Injection of A Grounded 3D Point Unlocks Spatial and Task Generalization — Direct Action-Head Injection of A Grounded 3D Point Unlocks Spatial and Task Generalization. [episode]
- Bounded-Fidelity Sim-as-Demo-Stage: Mocap Handoff for Governance Benchmarks — Sim-to-real research prioritizes physics fidelity, but for governance benchmarking of LLM-driven robots, contact fidelity during object handoffs becomes a liability because contact-force integration noise injects audit-chain divergence unrelated to the governance property under t [episode]
- FuncBridge: Towards Functional Tool-Use Generalization via Keypoint Trajectory Reasoning — Functional generalization in robotic tool-use—the ability to repurpose tools for novel functions despite differing motor patterns—is a critical challenge that current policies fail to address, and this work proposes a method called FORGE to bridge the perception-to-action gap [episode]
- Bridging the Sim-to-Real Gap with multipanda ros2: A Real-Time ROS2 Framework for Multimanual Systems — Multipanda ros2 presents a novel, open-source ROS2 architecture designed for real-time multi-robot control of Franka Robotics robots, addressing critical challenges in torque control, interaction control, and robot-environment modeling. [episode]
- ThermE: Predictive Management of Shared Thermal Headroom for Sustained LLM Inference on Thermally Constrained Edge SoCs — Compact edge system-on-chip (SoC) platforms increasingly run sustained LLM inference under thermal constraints, while their CPU, GPU, and RAM share a cooling path. [episode]
- ECoMEM: Explicit Concept Memory for Memory-Dependent Robot Control — Explicit Concept Memory (ECoMEM) introduces an explicit memory channel for Vision-Language-Action (VLA) policies, separating the maintenance of an evidence-grounded account of past history from how that history is used to generate actions. [episode]
- SonarVoxNet: Diver Detection in 3D Bounding Box using 3D Sonar — Autonomous underwater vehicles require continuous tracking of a diver's 3D position and full-body orientation for safe human-robot interaction, but existing forward-looking sonar methods discard elevation information, creating a fundamental geometric bottleneck. [episode]
- Token-World: World Modeling in Vision-Language Model Token Space for Robot Manipulation — Token-World introduces an action-conditioned world model simulator that models dynamics directly in a compact, policy-aligned VLM visual-token space, avoiding intermediate RGB generation. [episode]
- A Unified Family-optimal Solution to Covariance Intersection Problems with Semidefinite Programming — Covariance intersection (CI) methods provide a principled approach to fusing estimates with unknown crosscorrelations by minimizing a worst-case measure of uncertainty that is consistent with the available information. [episode]
- Region Based SLAM-Aware Exploration: Efficient and Robust Autonomous Mapping Strategy That Can Scale — Autonomous exploration for mapping unknown large scale environments remains a fundamental challenge in robotics, requiring solutions that are efficient in time, robust against map corruption, and computationally feasible. [episode]
- SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation — SlotVLA introduces an object-relation-centric framework for robotic manipulation that addresses the limitations of dense visual embeddings by compressing input into compact, interpretable representations. [episode]
- Communication-Induced Bifurcation and Collective Dynamics in Power Packet Networks: A Thermodynamic Approach to Information-Constrained Energy Grids — This paper investigates nonlinear dynamics and phase transitions in power packet networks by conceptualizing routers as macroscopic information-ratchets, providing a thermodynamic framework for understanding operational limits in information-constrained energy grids. [episode]
- Admissibility-Preserving Control for Multi-Input Systems with Joint Capacity Constraints — This paper introduces an Anisotropic Joint-Admissibility-Preserving Input Realization (AJ-APIR) framework to control multi-input strict-feedback nonlinear systems subject to coupled joint capacity constraints. [episode]
- Does Continual Imitation Learning Remain Grounded? A Language-Perturbed Benchmark for Robotic Task Retention — Continual imitation learning evaluates whether a robot can learn new knowledge without forgetting previously learned skills, but this paper introduces a benchmark protocol to study how language-guided behavior changes as robotic policies learn successive tasks, revealing that str [episode]
- When Reasoning Helps Action: Monitoring and Steering Chain-of-Thought in Vision-Language-Action Policies — Reasoning-enabled Vision-Language-Action (VLA) policies expose chain-of-thought (CoT) traces that can be monitored and steered at runtime to potentially improve safety. [episode]
- Measuring Asset and Scene Reconstruction Effects in Real-to-Sim Robot Evaluation — Simulated evaluation is increasingly used alongside real-world evaluation of robot policies because it is cheaper and easier to repeat; however, its value depends on how closely its outcomes track the real robot’s. [episode]
- DITTO-X: Forward and Reverse Teleoperation for Dexterous Manipulation and Human Intervention — Teleoperated demonstrations and human interventions are crucial for robot manipulation, yet existing systems often fail in shared autonomy due to limitations in handling dexterous hands and closing the feedback loop through vision alone. [episode]
- HumanoidToolBench: Benchmarking Humanoid Tool Use from Selection to Mobile Execution — HumanoidToolBench introduces an 18-task benchmark and a corresponding dataset to evaluate how humanoid policies select and use tools for various tasks, spanning selection, stationary use, and mobile execution. [episode]
- Feasibility of Simultaneous Input-Output Constraints for Tracking in a Class of LTI Systems: Part I — Simultaneous satisfaction of input and output constraints for tracking in linear time-invariant (LTI) systems with multiple inputs and integral action is addressed by deriving necessary and sufficient conditions for a Control Barrier Function (CBF)-based governor to guarantee saf [episode]
- Watch, Infer, Coordinate: Inferring Robot Partner Constraints for Zero-Shot Coordination — Robots operating in physical environments increasingly require coordination, especially when tasks involve objects too large or heavy for a single robot to manage, and this paper introduces a framework to enable one robot (the helper) to infer its partner's unknown physical const [episode]
- Robot Learning on Discrete Surfaces: Theory and Applications — All objects are enclosed within surfaces, yet most robot learning and motion generation frameworks treat surfaces as constraints ignoring their intrinsic geometry, creating a gap for polyhedral meshes whose discrete geometric structure remains unexploited. [episode]
- AFD-CAMLs: Agile Force-Distribution-Aware Planning and Control for Cable-Suspended Aerial Multi-Lifting Systems — Multiple UAVs can cooperatively transport heavy payloads while controlling their position and orientation. [episode]
- Toward Self-Organizing Production Logistics: A Multi-Agent Approach — Production logistics faces significant challenges due to increasing variability, dynamic interdependencies, and operational disturbances, particularly within complex circular production systems. [episode]
- ChunkVLA-AM: Parallel Action Chunking for Vision-Language-Action Robot Control in Additive Manufacturing — Vision–language–action (VLA) models offer a promising route toward flexible robotic systems in additive manufacturing (AM), but deploying them in real AM workcells remains challenging due to issues related to adapting models to new robot embodiments and maintaining robustness [episode]
- IndoorBEV: A Lightweight Real-Time LiDAR BEV Perception System for Indoor Mobile Robots — Efficient indoor LiDAR perception for mobile robots requires balancing prediction accuracy, latency, and memory constraints in cluttered environments. [episode]
- Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs — Fine-tuned Vision-Language-Action (VLA) models often struggle to generalize to skill combinations absent from their fine-tuning demonstrations, even when every constituent skill has been demonstrated. [episode]
- Token Economy Design for Fair and Efficient Highway Congestion Management with Express Lanes — We study the design of a token economy for highway lane allocation that aims to improve fairness without sacrificing traffic efficiency, which matters because it provides a promising alternative to monetary congestion pricing for fairer management of scarce road capacity. [episode]
- Differentiable Hybrid-Action Neural Feedback Control for District Heating Networks — Many cyber-physical systems require control policies that combine continuous setpoints with discrete operational decisions, such as equipment switching, mode selection, or resource scheduling. [episode]
- TouchTherm: Building Multimodal Digital Twins of Objects for Tactile and Thermal Rendering — TouchTherm introduces a framework for constructing simulation-ready multimodal digital twins of real-world objects by integrating visual geometry, contact-aligned tactile microgeometry, and observation-driven dynamic thermal fields. [episode]
- Learning a Resolution-Consistent Jacobian Field for Bio-Inspired Rigid-Soft Finger — Bio-inspired tendon-driven rigid-soft coupled dexterous fingers exhibit strong nonlinearity and configuration-dependent sensitivity, making accurate modeling challenging. [episode]
- Building Seasonal Highways for Residential Energy Hubs: Sizing, planning and operating thermal energy storage — Building seasonal highways for residential energy hubs addresses the challenge of managing energy storage differences in time-constants and efficiencies across electricity, heat, and mobility carriers by proposing a data-driven framework to steer short-term daily control towards [episode]
- Online Planning for Sparse Ground Target Search from a High-Altitude UAV under Partial Observability — Unmanned aerial vehicles (UAVs) searching for sparse ground targets from high altitudes face a unique challenge when targets are small and unobservable, necessitating novel sensing strategies. [episode]
- Extreme Length Generalization in a Compact Recurrent Architecture for One-Shot Exploration — Autonomous robots on one-shot missions run over horizons far longer than training data, and this work introduces FRANK, a compact recurrent architecture that demonstrates extreme length generalization. [episode]
- UrbanVLA: A Vision-Language-Action Model for Urban Micromobility — UrbanVLA introduces a route-conditioned Vision-Language-Action (VLA) framework designed for scalable urban navigation by explicitly aligning noisy route waypoints with visual observations during execution and subsequently planning trajectories to drive the robot. [episode]
- Completion Aware Guidance for World Action Models — World Action Models (WAMs) predict visual futures and robot actions, yet they remain susceptible to task-incomplete imagination, where plausible, action-consistent predictions omit the transition needed for task completion. [episode]
- Finite-Data Safety Informativity Under Dynamic Asymmetric Actuation — When system models are unknown and measurements are finite, ensuring safety under dynamic asymmetric actuation requires developing a certificate that validates commands against all data-consistent models. [episode]
- ShowerFlex: Achieving Pseudo-Static Balancing in a Continuum Shower Hose — With the global population rapidly aging, maintaining independence in Activities of Daily Living (ADLs)—particularly bathing or showering—has become a critical challenge, and this research introduces ShowerFlex, a highly articulated continuum mechanism designed for accessible [episode]
- A Dynamic Generalized Kalman Consensus Filter for Switching Sensor Networks — Distributed state estimation is critical for applications such as surveillance, autonomous navigation, and wide-area monitoring, where sensor agents must cooperatively track targets using only local measurements and neighbor-to-neighbor communication. [episode]
- BLT*: Informed Belief Localization Trees for Uncertainty-Aware Planning on Digital Twins — Informed Belief Localization Trees (Informed BLT) are presented as a sampling-based belief space planning algorithm designed to scale to large outdoor digital twins by efficiently connecting sampled belief states while accounting for available information and probabilistic collis [episode]
- DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation — Manipulating dynamic objects remains an open challenge for Vision-Language-Action (VLA) models, which struggle in dynamic scenarios requiring rapid perception, temporal anticipation, and continuous control. [episode]
- Probabilistic Plan Legibility with Off-the-shelf Planners — Legible planning is addressed by proposing a method to generate plans that best disambiguate their goals from other candidates from an observer’s perspective, which matters because it provides a mechanism for implicit communication in human-robot teaming scenarios by allowing a [episode]
- EIDA: Execution-Interface Dynamics Adaptation for Real-to-Sim-to-Real Robot Navigation — Simulation-to-robot transfer can fail when velocity commands produce motion and feedback that differ from those modeled during policy training, leading to navigation failures in physical execution. [episode]
- Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models — Reinforcement Learning's efficacy in refining robotic manipulation policies is currently bottlenecked by the difficulty of designing generalizable, dense reward functions, which this paper addresses by proposing Large Reward Models (LRMs) that adapt foundation Vision-Language Mod [episode]
- Interaction-Stiffness-Guided Basis Allocation in Dynamic Movement Primitives for Efficient Skill Transfer — Dynamic Movement Primitives (DMPs) are a compact framework for trajectory representation in robot skill learning, but their fixed basis layout limits precision allocation according to stage-dependent requirements. [episode]
- Are Frontier VLM Agents Ready to Be Robot Generalists? An Empirical Study with the Embodied Agent Arena — Frontier vision-language models (VLMs) combine scene estimation, interaction grounding, and executable actions; understanding how these abilities support complete robotic tasks is central to evaluating their readiness as robot generalists. [episode]
- Ageing-aware Energy Management for Residential Multi-Carrier Energy Systems — Ageing-aware Energy Management for Residential Multi-Carrier Energy Systems presents an optimization-based nonlinear Model Predictive Control (MPC) framework that integrates physics-based battery ageing models into energy management systems for multi-carrier buildings. [episode]
- A two-stage approach to satellite constellation optimization: classical and QUBO formulations — The design of satellite constellations for Earth observation requires balancing spatial coverage, revisit time, cost, and operational complexity. [episode]
- Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models — GTA-VLA (Guide, Think, Act) is an interactive Vision-Language-Action (VLA) framework that enables spatially steerable embodied reasoning by allowing users to guide robot policies with explicit visual cues. [episode]
- quARtet Marker: A 3D-Printable Multi-Tag Fiducial for Robust Near-Frontal Pose Estimation — Coded planar fiducials are being enhanced by tilting multiple AprilTags within one compact footprint to improve pose estimation accuracy in near-frontal views while maintaining graspability for robotic manipulation. [episode]
- Constant-Time Planning for Chaining Collision-free Motion to Manipulation Behaviors — A family of algorithms called Constant-Time Motion Planning (CTMP) has been introduced, which leverages a preprocessing phase to enable collision-free motion queries in a fixed, user-specified time budget (e.g., 10 milliseconds). [episode]
- A Time-invariant Network Flow Model for Ride-pooling in Mobility-on-Demand Systems — A framework is presented to incorporate ride-pooling into time-invariant network flow models for Mobility-on-Demand systems, transforming a microscopic combinatorial phenomenon into a solvable linear problem. [episode]
- Aging-Aware Online Distributed Scheduling for Lifecycle Carbon Reduction in Geo-Distributed Data Centers — The rapid proliferation of data centers has led to massive energy demand and carbon emissions, posing significant sustainability challenges. [episode]
- Deformable In-Hand Slip-Aware Tactile Sensor with Integrated Velocity Sensing, Force/Torque and Pressure Map Estimation — This paper introduces a novel tactile sensor that integrates velocity, force/torque, and pressure map sensing into a single device with a deformable contact pad to enable slip-aware control for in-hand manipulation. [episode]
- Distributed Adaptive Neural Interval Observers for Unknown Nonlinear Systems — This paper develops a distributed adaptive neural interval observer for unknown nonlinear systems with locally incomplete measurements, addressing the challenge of preserving state enclosures across spatially distributed sensor nodes when nonlinear dynamics are unknown. [episode]
- ColoACT: Multi-Cue Action Chunking for Smooth Autonomous Colon Navigation on a Self-Propelled Endoscopic Robot — Autonomous colonoscopic navigation remains challenging due to deformable anatomy, weak-texture and specular endoscopic visuals, and contact-rich viscoelastic interactions. [episode]
- Sequential Object Placement Optimization with Convex Decomposition — Robotic object packing faces significant challenges due to combinatorial search complexity and difficulties in handling dynamic constraints for irregularly shaped objects. [episode]
- Development of an EMT model of the Balearic power system — Detailed ElectroMagnetic Transient (EMT) simulation studies are necessary to analyze the stability of power systems like the Balearic Islands due to their high integration of inverter-based resources and reduced synchronous generation. [episode]
- Real-Time Human-Adaptive Task Allocation for Multi-Human Multi-Robot Supervision — Real-Time Human-Adaptive Task Allocation for Multi-Human Multi-Robot Supervision proposes an adaptive method to dynamically regulate supervisory capacity and allocate robot supervision tasks to multiple human operators based on their real-time cognitive states. [episode]
- Communication-aware Synthesis of Safe Controllers for Discrete-Time Linear Multi-Agent Systems with Distributed k-Hop Observation — A distributed k-hop observer and LMI-based optimization framework are developed to jointly synthesize safe controllers, distributed observers, and local robust safe invariant sets for discrete-time linear multi-agent systems operating under limited communication. [episode]
- Identification of the Steering and Speed Systems of a Four-Wheel-Steering Tractor for Optimal Control — Model-based control design necessitates an accurate system model, and this paper addresses system identification for steering and speed control of a four-wheel-steering agricultural tractor to facilitate path tracking control. [episode]
- Multi-Reference Path Tracking Control for an Agricultural Tractor with Nonlinear Model Predictive Control — Guiding agricultural tractors along predefined paths is crucial for precision agriculture, and this study develops an advanced path tracking controller using Nonlinear Model Predictive Control (NMPC) that incorporates multiple segments of a piecewise-linear reference path directl [episode]
- Structural Sign Herdability in Temporal Networks: A Sufficient Condition via pi p-Graphs — A temporal network study investigates herdability, a relaxed form of controllability, in systems where the switching sequence is fixed over time. [episode]
- G2-Nav: Grounded and Guarded Vision-Language Costmaps for Robot Social Navigation — Social navigation requires robots to reason and respond in complex real-world environments, and G2-Nav addresses this by grounding abstract social reasoning into reliable, interpretable vision-language costmaps while incorporating safety checks for deployment. [episode]
- Fixed-Time Voltage Regulation in Distribution Networks with Impedance Awareness — This letter introduces an optimization-based fixed time control algorithm for solving the voltage regulation problem of a radial and balanced power distribution network, offering a method to guarantee voltage convergence within a predefined fixed time window even when exact netwo [episode]
- LiDARFlow: Real-Time Panel-Based MAV Guidance in Unknown Environments — This paper presents a guidance algorithm for micro aerial vehicles operating in unknown, cluttered environments using only onboard sensing. [episode]
- A local recursive least squares approach for discrete-time adaptive fuzzy control — A local recursive least squares approach for discrete-time adaptive fuzzy control proposes a membership-weighted RLS law with a forgetting factor to approximate unknown nonlinearities in quasi-Linear Parameter Varying/Takagi–Sugeno (qLPV/TS) systems, providing LMI synthesis con [episode]
- H-SPAR: Hydrodynamic-aware Simulation for Particle Transport and Autonomous Robots — H-SPAR is an open-source, hydrodynamic-aware simulation framework designed to evaluate autonomous marine sampling missions by jointly modeling spatio-temporally varying flow fields, Lagrangian particle transport, and closed-loop Uncrewed Surface Vehicle (USV) autonomy. [episode]
- External Photoreflective Tactile Sensing Based on Surface Deformation Measurement — An externally attachable photoreflective module reads surface deformation of silicone skin to estimate contact force without embedding tactile transducers, offering a practical and robust route to equip soft robots with force perception while preserving structural flexibility and [episode]
- New VSC-HVDC interconnection between the Iberian Peninsula and Balearic Archipelago to enable energy transition — A new High Voltage Direct Current (HVDC) interconnection between the Iberian Peninsula power system and the Balearic Islands power system is planned to facilitate the decarbonisation of the Balearic Archipelago. [episode]
- Toward Humanoid Robots in Construction: A Teleoperation Feasibility Study — We present a teleoperation system that enables a single operator to perform construction tasks on a Unitree G1 humanoid, combining extended reality (XR) based upper body control with pedal-based locomotion to enable simultaneous manipulation and locomotion. [episode]
- Overlapping Covariance Intersection: Fusion with Partial Structural Knowledge of Correlation from Multiple Sources — Emerging large-scale engineering systems require distributed fusion to achieve situational awareness, but tracking crosscorrelations becomes infeasible at scale, necessitating methods that incorporate partial structural knowledge of correlation from multiple sources. [episode]
- Learning-Based Predictive Control Method for Vehicle Lateral Control with a Multi-Step Gaussian Process Regression Prediction — A novel approach to model predictive control that incorporates multi-step uncertainty prediction for safely controlling systems characterized by uncertainties dependent on both state and control variables addresses the challenge of accumulating and propagating modeling errors ove [episode]
- Emulation-based Neuromorphic Control for the Stabilization of LTI Systems — Neuromorphic control for Linear Time-Invariant (LTI) systems is addressed by presenting a systematic, two-step emulation-based design procedure that ensures practical closed-loop stability. [episode]
- Composite learning control with modular backstepping and high-order tuners — A composite learning backstepping control (CLBC) strategy, utilizing modular backstepping and high-order tuners, is proposed to achieve closed-loop exponential stability for strict-feedback uncertain nonlinear systems under relaxed excitation conditions. [episode]
- Event-Triggered Practical Fixed-Time Integral Reinforcement Learning for Unknown Nonlinear Systems — Event-triggered fixed-time integral reinforcement learning for unknown nonlinear systems addresses the challenge of designing optimal controllers for complex, unknown nonlinear systems by combining system identification with an inverse-optimal formulation and an event-triggered m [episode]
- Congestion-aware Ride-pooling in Mixed Traffic for Autonomous Mobility-on-Demand Systems — This paper presents a modeling and optimization framework to study congestion-aware ride-pooling Autonomous Mobility-on-Demand (AMoD) systems, where self-driving robotaxis share vehicles for part of their journey. [episode]
- Local Consistency Does Not Guarantee Global Conservation: Auditing Zero-Shot Composition of Airway Flow Operators — Neural operators approximate Partial Differential Equation (PDE) solutions, but independently learned local operators need not form a consistent global simulator. [episode]
- Closed-Loop Refinement and Execution for Learned Driving Planners — Learning-based driving planners are typically trained and evaluated in open loop against logged trajectories, but this approach fails to guarantee reliable closed-loop execution because planning errors accumulate when actions change subsequent observations. [episode]
- Towards Physical Underwater Robotic Assistance for Scuba Diver Movement in Confined Spaces — A novel wearable robotic system, RADMCS, is introduced to assist scuba divers in maintaining safe standoff distances from subsea structures in confined or hazardous underwater environments. [episode]
- Ex vivo breach detection using electrical conductivity during robotic pedicle drilling in the spine — Pedicle screw placement (PSP) is a technically demanding spinal procedure where high precision is crucial due to limited visibility and anatomical variability, and this study proposes using robotic pedicle drilling combined with real-time preventive breach detection via electrica [episode]
- Frequency-aware decomposition learning for sensorless wrench estimation in vibration-rich robotic contact — Force and torque (F/T) sensing is critical for robot-environment interaction, but physical F/T sensors impose constraints in size, cost, and fragility. [episode]
- NarrativeFlow: Flow-Based Vision-Language-Action Model Using Robot Velocity Fields —
- ALFRED: Requirement-driven development of an open-source mobile manipulator for long-term plant monitoring —
- Beyond Leaderboard Scores: A Deployment-Focused Protocol for Interpretable Tracking Evaluation in Pedestrian-Centric Environments —
- A Survey on End-to-End Autonomous Driving Training from the Perspectives of Data, Strategy, and Platform —
- InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation —
- DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation —
- UniTrackPLA: Unified Panorama-Language-Action Model for Instruction-Guided Navigation and Dynamic Person Tracking —
- Kinematic MeanFlow: One-Step Action Generation Policy for Robotic Foundation Models —
- World Motion Models: Flexible Sequence Modeling of SE(3) Trajectories —
- Retrospective Open-Vocabulary Memory for Long-Term Object Search —
- End-to-End Learning vs. Modular Architectures: Comparative Insights into Autonomous Driving Systems —
- DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies —
Important terms
- Kinematic MeanFlow
- This approach simplifies robot action generation by focusing on average motion patterns, offering a more stable way to plan movement compared to purely reactive methods.
- FuncBridge
- This technique improves functional tool use generalization by reasoning about the trajectory of specific body parts when interacting with a tool, rather than just object appearance.
- World Motion Models
- These flexible sequence models focus on predicting complex movements in 3D space over time, which is essential for understanding and planning intricate physical motions.
- eRLT
- This method makes large vision-language models more efficient for control by using action relevant token routing, tackling the computational bottleneck in deploying complex policies.