Robotics papers — 2026-10-10
Today's focus is on building reliable systems where the underlying rules are not perfectly known, which is crucial for things like robotic planning. We explored how to make inference safer during complex robotic task planning using SafeInferCom, which involves inserting a verifier-guided intervention mid-generation to ensure the plan remains safe. This builds upon work in action consistency through inverse dynamics for planning with world models called ACID, which uses those world models to check the actions taken.
Another big piece of work involved scaling dexterity for robotic hands in cluttered environments with OmniDex, aiming to make hand grasping more robust across diverse scenes. We also looked at how 3D point world models improve dynamics learning by enabling better point completion, which feeds into the action consistency checks we discussed earlier. Furthermore, we touched upon navigation runtime improvements with NavGPT-3, which focuses on harnessing context within a hierarchical navigation structure to make decision-making more informed.
Finally, we looked at visual scene understanding challenges by investigating whether dynamic point filtering helps when texture is scarce in synthetic indoor scenes using ORB-SLAM2 front-ends. This study connects the need for robust perception with the overall goal of creating scalable and safe enclosures for uncertain systems.
The most significant development today concerns the Latent World-Action Model Being M0.7, which attempts to create a model that can understand and act within a world context for humanoid robots. This is important because it moves beyond simple reactive control toward genuine world understanding for embodied AI.
Being M0.7 was developed by integrating latent space representations with action policies to allow the robot to reason about its environment before executing movements, which is a big step forward from previous purely visual approaches. A related effort involved WARP-VLA, which focused on improving policy execution in vision-language-action models specifically for wrist cameras, aiming to make the robot's actions more robust when adapting to different viewing angles.
Another piece of work addressed perception in autonomous driving with CoCam4D, which is designed for camera-only systems and focuses on geometry awareness within cooperative 4D perception. This work builds upon the need for accurate environmental understanding by providing a method that considers spatial relationships directly from visual data.
On the manipulation side, there was research into neural networks for temporal pattern recognition and dynamic arm gesture speed estimation, which is crucial for robot control tasks requiring precise timing of movements. This contrasts with another study that explored a minimal optical-flow representation to classify tactile rotation in robotic manipulation across various gravity domains.
The most critical development is the work on ExecVLA because it directly addresses the challenge of ensuring robots follow fine-grained execution constraints in vision-language-action models. This means making sure the robot's actions are precise enough to meet complex instructions, which is vital for real-world deployment. The team explored using a bi-level action representation within ExecVLA to achieve this level of constraint adherence.
This work builds upon the foundation laid by LeWAM, which introduced a JEPA World Action Model using diffusion steering-based model predictive control. That model helps the robot understand and predict world actions in a way that is grounded in experience. Furthermore, DreamTrue offers an action-faithful robot world model enhanced with counterfactual post-training to improve how these models interact with the environment.
Dex-One2Many tackles the problem of learning dexterous manipulation from just a single human demonstration, which is a significant step toward general robotic skill acquisition. This contrasts with ImagiNav, which focuses on scalable embodied navigation by using generative visual prediction and inverse dynamics for surveying surfaces.
The Kinetics Observer provides a tightly coupled estimator specifically for legged robots, offering improved state estimation crucial for dynamic locomotion.
The most important development today concerns the sliding-window filter approach to online continuous-time continuum robot state estimation, which matters because it offers a robust way to track robot positions in real time without needing perfect prior knowledge. Researchers tried implementing this filter on a continuum robot system using sensor data, and the results showed that it maintained better tracking accuracy than traditional methods when dealing with noisy measurements. This improved estimation capability is foundational for any reliable autonomous operation of such complex machinery.
Another significant piece of work addresses census-based population autonomy for distributed robotic teaming, which is crucial because it tackles how groups of robots can manage tasks without a central controller. The study explored using census data to assign roles within a team, and the findings suggest that this method helps in achieving greater operational independence among the robots. This idea connects to how other systems might organize themselves, such as those dealing with spatial awareness in navigation.
The work on RoboAug is important because it provides a way to rapidly annotate hundreds of scenes using just one annotation through region-contrastive data augmentation. This speeds up the training of robotic manipulation models, and it relates to how action encoding might be improved in other systems.
ActionCodec investigated what makes a good action tokenizer, focusing on creating better representations for robot actions by analyzing existing datasets. They found that certain tokenization strategies significantly improve the model's ability to understand and execute complex sequences of movements. This refinement in understanding actions is vital when planning complex behaviors, which ties into how progressive action plan refinement works in latent space.
Seed2Scale introduced a self-evolving data engine with parallel worlds expansion to allow for scalable robot learning. This means it can adapt its knowledge base as it encounters new situations. This engine's ability to scale and evolve is interesting when considering the temporal ensemble advantage modeling explored in STEAM, which looks at how different temporal models combine their strengths for real-world learning.
The most significant progress on the day involves GeniWorld, which attempts to create a generalizable interactive world model for robotic manipulation. This suggests a leap toward more adaptable robot intelligence. This model is built by integrating visual actions with attention mechanisms from action, aiming to help policies learn better by identifying emergent visual bottlenecks.
This work is important because it tackles the core challenge of making robots understand how to interact with the physical world in a flexible way. A related effort explored how attention from action can reveal these visual bottlenecks, which is crucial for policy learning. Furthermore, CoToGrasp synthesized dexterous grasps by conditioning them on contact topology within a canonical workspace learning framework.
Another piece of research focused on real-time estimation of actuator control and robot health using REACH on an eel-inspired soft robot. This demonstrated the ability to monitor the physical state of a soft manipulator in real time, which is vital for safe operation. This contrasts with work like RoboRacer Arena, which focused on specification-driven track construction for autonomous racing, showing a different kind of system design challenge.
Then there was RA-VLA, which introduced retrieval-augmented vision language agents designed for test-time adaptation. This means the agent can adapt its behavior during testing by retrieving relevant knowledge when needed. Finally, TacHair addressed tactile contact distribution guided online correction for robotic hair stroking and perception. This showed how fine tactile feedback can guide real-time adjustments to a delicate task.
The most significant piece of work today involves Skill-SLM, which attempts to make small language models more reliable for robot operation by grounding their skills in actual agent experience. This matters because it moves beyond just generating plausible actions toward creating skills that are robust enough for real-world deployment.
A related effort explored the variability in dynamic cloth manipulation, showing that the same action can lead to different outcomes depending on the environment. This suggests a need for more adaptable planning methods. This connects to how Skill-SLM tries to handle skill variation by using language models trained on diverse experiences.
TAPNAV focused on humanoid navigation through tactile active perception, aiming for better movement in complex spaces by incorporating touch into the perception loop. This is important because it addresses the challenge of navigating environments where visual data alone is insufficient for safe locomotion.
Another line of research looked at diagnosing and recovering from observation-space shift when dealing with long-horizon skill seams. This means figuring out why a robot's learned skills break down over time or in new situations. This diagnostic work informs how Skill-SLM might need to adapt its learned policies.
Cross-Embodiment Robot Foundation World Models with Latent Actions is also crucial because it tries to build world models that work across different physical bodies. This is a big hurdle for general robot intelligence. This contrasts with the more immediate tactile focus of TAPNAV and Skill-SLM.
The work on reconfigurable fabric based pneumatic actuators with button fastened constraint modules is significant because it offers a flexible way to create diverse actuation modes. This research explored how these actuators can achieve multi mode actuation by incorporating these constraint modules, aiming for adaptability in physical tasks.
A key piece of this effort involved developing the USDCraft system, which provides geometrically grounded programmatic modeling of articulated three dimensional assets specifically for simulation purposes. This modeling approach helps define the physical structure before deployment, and it feeds into how other systems might plan movement.
Moving down the list, there is work on PlanWAM which focuses on planning shaped future representations for end-to-end autonomous driving tasks. This means creating a way for a vehicle to plan its entire journey based on predicted future states.
Another area of investigation is RAGNAROK, which deals with radar aided gravity normalized alignment for robust open keyframe based radar visual kinematic inertial simultaneous localization and mapping. This technique aims to make the robot's positioning more reliable by fusing data from radar, vision, kinematics, and inertia.
This relates to distributed relative localization for homogeneous multi robot systems through ultra wide band ranging and limited communications. This work addresses how multiple robots can figure out their relative positions even when communication is scarce by using UWB ranging measurements.
Furthermore, there is research into distributed relative localization based on ultra wide band and LiDAR for multi robot navigation with limited communication. This combines the benefits of both UWB and LiDAR to achieve better spatial awareness across a group of robots under challenging communication constraints.
Towards path creative navigation through embodied interaction addresses how robots can navigate complex environments by interacting with them directly. This suggests a focus on learning movement through physical engagement rather than just pre-programmed paths.
Finally, there is FOCUS which deals with moving from privileged states to RGB D with controlled modality switching and representation alignment. This suggests a method for intelligently switching between different types of sensor data while keeping the resulting information consistent across those different data types.
The most significant development today involves the work on WAND, which tackles learning robust navigation under complex wind disturbances and dense obstacles for quadrotors. This is crucial because reliable flight in unpredictable environments is a major hurdle for autonomous aerial systems. The researchers explored how to train these systems without needing extensive real-world testing by focusing on learning from simulated data, which helps bridge the gap between lab work and actual deployment.
A related piece of work addresses path planning within 2D Gaussian Splatting maps through 2DGS-Planner, which uses rasterization to figure out paths in these complex visual representations. This is important because accurate pathfinding is fundamental for any robot navigating a mapped space. This planning method builds upon the concept of using spatial representations to guide movement, similar in spirit to how other works approach environmental understanding.
We also saw progress on hierarchical frameworks for composable multi-agent human-object interaction, which moves beyond single agents by creating systems where different agents can cooperate. This framework suggests a way to build more sophisticated interactions by organizing simpler agent behaviors into a larger structure. This idea connects to the need for robust navigation, as coordinating multiple agents requires careful management of their individual capabilities and interactions.
Another piece of research focuses on ultra-light luma, which introduces an edge-deployable perception network specifically for segmenting crop rows in agricultural robotics. This work is important because it aims to make high-level vision tasks practical for deployment on resource-constrained devices in the field. It suggests a pathway toward deploying sophisticated perception directly onto the machinery performing the task.
Finally, there was exploration into path-time decoupling within PathTime-VLA, which focuses on factorizing post-training for vision language action policies. This technique is significant because it seeks to separate the temporal aspects of decision-making from the visual and linguistic understanding components. This factorization approach offers a different angle on how agents can learn complex behaviors compared to purely end-to-end methods.
Today's papers
- Certified Scalable Enclosures for Uncertain Underdetermined Systems. [paper]
- NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime. [paper]
- SafeInferCom: Safe Inference-Time Compute via Verifier-Guided Mid-Generation Intervention for Robotic Task Planning. [paper]
- 3D Point World Models: Point Completion Enables More Accurate Dynamics Learning. [paper] [episode]
- ACID: Action Consistency via Inverse Dynamics for Planning with World Models. [paper] [episode]
- Does Dynamic-Point Filtering Help When Texture Is Scarce? A Controlled Study of ORB-SLAM2 Front-Ends in Synthetic Indoor Scenes. [paper]
- When Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAs. [paper]
- OmniDex: Scaling Dexterous Hand Grasping to Diverse Cluttered Scenes. [paper]
- How Firm Should a Grasp Be?. [paper]
- WAM-Cache: Staleness-Bounded KV Reuse for Efficient World Action Models. [paper]
- Being-M0.7: A Latent World-Action Model for Humanoid Robots. [paper]
- WARP-VLA: Wrist-Camera Adaptation for View-Robust Policy Execution in Vision-Language-Action Models. [paper]
- CoCam4D: Geometry-Aware Cooperative 4D Perception for Camera-Only Autonomous Driving. [paper]
- Neural Networks for Temporal Pattern Recognition and Dynamic Arm Gesture Speed Estimation for Robot Control. [paper]
- A Minimal Optical-Flow Representation for Vision-Based Tactile Rotation Classification in Robotic Manipulation Across Gravity Domains. [paper]
- SuperNav: An Agentic Navigation System for Any Task in Any Scene. [paper]
- Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement. [paper]
- LeWAM: A JEPA World Action Model with Diffusion-Steering-Based MPC. [paper]
- DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training. [paper]
- Dex-One2Many: Learning Dexterous Manipulation from a Single Human Demonstration. [paper]
- ImagiNav: Scalable Embodied Navigation via Generative Visual Prediction and Inverse Dynamics. [paper] [episode]
- ExecVLA: Following Fine-Grained Execution Constraints in Vision-Language-Action Models with Bi-Level Action Representation. [paper] [episode]
- Multi-Depth Uniform Coverage Path Planning for Unmanned Surface Vehicle Surveying. [paper] [episode]
- The Kinetics Observer: A Tightly Coupled Estimator for Legged Robots. [paper] [episode]
- A Sliding-Window Filter for Online Continuous-Time Continuum Robot State Estimation. [paper] [episode]
- Census-Based Population Autonomy For Distributed Robotic Teaming. [paper] [episode]
- RoboAug: One Annotation to Hundreds of Scenes via Region-Contrastive Data Augmentation for Robotic Manipulation. [paper] [episode]
- ActionCodec: What Makes for Good Action Tokenizers. [paper] [episode]
- Seed2Scale: A Self-Evolving Data Engine with Parallel Worlds Expansion for Scalable Robot Learning. [paper] [episode]
- SPAN-Nav: Generalized Spatial Awareness for Versatile Embodied Navigation. [paper] [episode]
- PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space. [paper] [episode]
- STEAM: Self-Supervised Temporal Ensemble Advantage Modeling for Real-World Robot Learning. [paper] [episode]
- GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions. [paper] [episode]
- Attention from Action, for Action: Emergent Visual Bottlenecks for Policy Learning. [paper] [episode]
- Real-time Estimator of Actuator Control and Health (REACH) on an Eel-Inspired Soft Robot. [paper] [episode]
- CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning. [paper] [episode]
- RoboRacer Arena: Specification-Driven Track Construction for Autonomous Racing. [paper] [episode]
- RA-VLA: Retrieval-Augmented VLA for Test-Time Adaptation. [paper] [episode]
- Teaching a Robot Dog New Tricks: Diverse Quadruped Skills via Combined Reinforcement and Imitation Learning with Adversarial Task Selection. [paper] [episode]
- TacHair: Tactile Contact-Distribution Guided Online Correction for Robotic Hair Stroking and Perception. [paper] [episode]
- Masked Generative Motion Planning with Geometry-Guided Token Search. [paper] [episode]
- TAPNAV: Humanoid Navigation through Tactile Active Perception. [paper] [episode]
- Same Action, Different Outcome: Variability in Dynamic Cloth Manipulation. [paper] [episode]
- Diagnosing and Recovering from Observation-Space Shift at Long-Horizon Skill Seams. [paper] [episode]
- Skill-SLM: Agent Skill-driven Small Language Models for Reliable Robot Operation. [paper] [episode]
- Cross-Embodiment Robot Foundation World Models with Latent Actions. [paper] [episode]
- OmniHOI: Dexterous Hand-Object Interaction from Monocular Human Video. [paper] [episode]
- Informationally Decoupled Trajectory Design for Sim-to-Real System Identification. [paper] [episode]
- A Reconfigurable Fabric Based Pneumatic Actuator with Button Fastened Constraint Modules for Multi Mode Actuation. [paper] [episode]
- Towards Path-Creative Navigation: Robot Navigation through Embodied Interaction. [paper] [episode]
- FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment. [paper] [episode]
- Distributed Relative Localization Based on Ultra-WideBand and LiDAR for Multi-robot with Limited Communication. [paper] [episode]
- Distributed Relative Localization for Homogeneous Multi-Robot Systems through UWB Ranging and Limited Communications. [paper] [episode]
- USDCraft: Geometrically Grounded Programmatic Modeling of Articulated 3D Assets for Simulation. [paper] [episode]
- PlanWAM: Planning-Shaped Future Representations for End-to-End Autonomous Driving. [paper] [episode]
- RAGNAROK: Radar-Aided Gravity-Normalized Alignment for Robust Open Keyframe-based Radar-Visual-Kinematic-Inertial SLAM. [paper] [episode]
- Scalable LEO Conjunction Screening using Adaptive Synthetic-Covariance Thresholds. [paper]
- YOCO: You Only Calibrate Once! Fast Mocap Calibration for Dexterous Teleoperation. [paper] [episode]
- From Solo to Ensemble: A Hierarchical Framework for Composable Multi-Agent Human-Object Interaction. [paper] [episode]
- 2DGS-Planner: Rasterization-based Path Planning in 2D Gaussian Splatting Map. [paper] [episode]
The papers
- RA-VLA: Retrieval-Augmented VLA for Test-Time Adaptation — The gist The RA-VLA framework is a retrieval-augmented VLA system that integrates behavior-aligned context retrieval with a grounded execution pipeline to facilitate seamless task adaptation while preserving inference efficiency, establishing a robust framework for training-free [episode]
- STEAM: Self-Supervised Temporal Ensemble Advantage Modeling for Real-World Robot Learning — The gist The STEAM framework is a self-supervised advantage modeling framework for real-world robot learning that learns advantage prediction offline from expert demonstrations without manual annotations or hand-crafted rewards. [episode]
- Integrating Active Damping with Shaping-Filtered Reset Tracking Control for Piezo-Actuated Nanopositioning — The gist Piezoelectric nanopositioning systems are often limited by lightly damped structural resonances and the gain–phase constraints of linear feedback, which restrict achievable bandwidth and tracking performance This paper presents a dualloop architecture that combines an [episode]
- SPAN-Nav: Generalized Spatial Awareness for Versatile Embodied Navigation — The gist The SPAN-Nav end-to-end foundation model infuses embodied navigation with universal 3D spatial awareness using RGB video streams to achieve robust generalization across complex environments. How it works 1. [episode]
- RoboAug: One Annotation to Hundreds of Scenes via Region-Contrastive Data Augmentation for Robotic Manipulation — The gist The proposed RoboAug framework, a region-contrastive data augmentation framework, significantly minimizes reliance on large-scale pretraining and perfect visual recognition by requiring only bounding box annotation of a single image during training <ref:2602.14032#pg8>. [episode]
- Compositional and Equilibrium-Free Stability Certification for Power Systems--Part II: Algorithms and Applications — The gist This two-part paper proposes a compositional and equilibrium-free approach to analyzing power system stability. [episode]
- ExecVLA: Following Fine-Grained Execution Constraints in Vision-Language-Action Models with Bi-Level Action Representation — The gist The KineVLA framework introduces a kinematics-rich vision-language-action (VLA) task that explicitly decouples goal-level invariance from kinematics-level variability through a bi-level action representation and bi-level reasoning tokens to serve as explicit, supervised [episode]
- CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning — The gist The proposed framework synthesizes diverse, stable grasps strictly conditioned on specific contact topologies by projecting local object features into a feature-based canonical workspace, effectively decoupling semantic functional intent from arbitrary object geometry Co [episode]
- Can Julia land on the Moon? On the development of a GNC simulation framework for the Argonaut lunar lander — The gist: No, the Julia programming language cannot land on the Moon — but it can play a crucial role in designing and analysing the Guidance, Navigation, and Control (GNC) algorithms required for doing so. [episode]
- PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space — The gist PearlVLA proposes a VLA framework that moves deliberation into the latent space of a vision-language model to improve action planning while maintaining low-latency execution. [episode]
- GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions — The gist The paper introduces GeniWorld, an interactive world model that generalizes robustly across unseen scenarios by conditioning on visual actions to enable closed-loop interaction with policies and human operators. How it works 1. [episode]
- Seed2Scale: A Self-Evolving Data Engine with Parallel Worlds Expansion for Scalable Robot Learning — The gist: Seed2Scale is a self-evolving data engine that breaks the data bottleneck in Embodied AI through a heterogeneous synergy of “small-model collection, large-model evaluation, and target-model learning”. [episode]
- Census-Based Population Autonomy For Distributed Robotic Teaming — The gist The census-based population autonomy model introduces a layered framework combining nonlinear opinion dynamics for collective decision-making and multi-objective behavior optimization for individual actions to enhance distributed robotic teaming. How it works 1. [episode]
- Modal Analysis of Spatial Load Correlation in AI Data Center-Dominated Power Systems — The gist: This paper proposes a data-driven modal framework based on Dynamic Mode Decomposition (DMD) to analyze nonstationary spatial load correlation in AI data center-dominated power systems, which addresses the limitations of classical methods that assume stationarity. [episode]
- Koopman operator theory: fundamentals, control, and applications — Detailed Research Summary: Koopman Operator Theory (Fundamentals, Control, and Applications) This research paper provides a comprehensive overview of the Koopman operator framework, detailing its theoretical foundations in dynamical systems theory, its application in control syst [episode]
- ImagiNav: Scalable Embodied Navigation via Generative Visual Prediction and Inverse Dynamics — The gist The ImagiNav framework introduces a novel modular paradigm that decouples visual planning from robot actuation, enabling robots to navigate open-world environments via natural language by synthesizing future egocentric videos conditioned on instructions and interpreting [episode]
- 3D Point World Models: Point Completion Enables More Accurate Dynamics Learning — The gist: 3D Point World Models (3DPWM) are a task-agnostic world model that operates entirely in 3D space by first completing partial point clouds and then learning action-conditioned dynamics in this completed 3D scene, enabling reliable long-horizon rollouts and more accurate [episode]
- Attention from Action, for Action: Emergent Visual Bottlenecks for Policy Learning — The gist: Seeker, an action-supervised module that learns where visual evidence is needed for visuomotor control, turns observation–action data into a progression-aware ROI, exposing an explicit spatial bottleneck without relying on semantic boxes, gaze, language grounding, or [episode]
- ActionCodec: What Makes for Good Action Tokenizers — The gist The introduction establishes that action tokenization design remains unanswered, and ActionCodec introduces a high-performance action tokenizer guided by information-theoretic insights to enhance VLA optimization. [episode]
- RoboRacer Arena: Specification-Driven Track Construction for Autonomous Racing — The gist The RoboRacer Arena system creates 3D racing environments directly from occupancy maps by using an automated map-to-environment builder that converts ROS occupancy grids into collision-ready, textured USD environments for Isaac Sim in 1.18 s to 2.48 s System Overview Rob [episode]
- ACID: Action Consistency via Inverse Dynamics for Planning with World Models — The gist The proposed ACID framework introduces cycle action consistency into decision-time planning for action-conditioned world models to ensure predicted trajectories are realizable, which consistently improves planning performance across diverse world models and tasks. [episode]
- Transfer Learning of Multiobjective Indirect Low-Thrust Trajectories Using Diffusion Models and Markov Chain Monte Carlo — The gist: This work introduces a transfer-learning framework that combines homotopy in a mission parameter with Markov chain Monte Carlo (MCMC) to generate training data more efficiently for diffusion models, enabling them to learn a global representation of the underlying soluti [episode]
- Multi-Depth Uniform Coverage Path Planning for Unmanned Surface Vehicle Surveying — The gist: The proposed Multi-Depth Non-revisiting Uniform Coverage (MDNUC) algorithm introduces a novel, template-free coverage path planning method that adapts to varying seafloor depths by dynamically adjusting the sensing beam aperture to achieve optimized, uniform seafloor co [episode]
- Real-time Estimator of Actuator Control and Health (REACH) on an Eel-Inspired Soft Robot — The gist The architecture employs a soft robot model, sigma point filter, and a formal statistical hypothesis test to adequately capture the nonlinearities and changes over time; REACH Algorithm The REACH algorithm uses a Sigma Point Kalman Filter (SPKF) to estimate actuator heal [episode]
- A Sliding-Window Filter for Online Continuous-Time Continuum Robot State Estimation — The gist: This work presents the first stochastic sliding-window filter specifically designed for continuum robots, which improves accuracy over filtering methods while enabling continuous-time methods to operate online and at faster-than-realtime speeds. [episode]
- An Information Theory of Finite Abstractions and their Fundamental Scalability Limits — The gist The work derives a statistical, quantitative theory of abstractions’ size-accuracy tradeoff and uncovers fundamental limits on their scalability through rate-distortion theory. [episode]
- Towards Ultra-Reliable 6G in-X Subnetworks: Dynamic Link Adaptation by Deep Reinforcement Learning — The gist: The proposed method dynamically adjusts transmission power and blocklength based on SINR using SAC to optimize both consecutive outages and energy efficiency in 6G in-X subnetworks. [episode]
- Optimal Battery Bidding under Decision-Dependent State-of-Charge Uncertainties — The gist: The uncertainty-aware formulation outperforms other constraint-tightening approaches in maximizing revenue while ensuring reliable frequency reserve provision by treating SOC uncertainty as an endogenous process within the operational strategy. [episode]
- The Kinetics Observer: A Tightly Coupled Estimator for Legged Robots — The Kinetics Observer proposes a novel, tightly coupled estimator that simultaneously estimates contact and perturbation forces along with robot kinematics for real-time proprioceptive odometry. [episode]
- A User-Triggered UAV Dispatching System for Precise and Timely Mountain Search Missions — The gist: This paper presents a user-triggered UAV dispatching system that integrates mobile information collection, ground station mission management, and UAV search into a single workflow. [episode]
- Distributed Relative Localization Based on Ultra-WideBand and LiDAR for Multi-robot with Limited Communication — The gist Distributed relative localization based on Ultra-WideBand and LiDAR for multi-robot with limited communication proposes a fully distributed relative position estimation approach for homogeneous robots using onboard UWB and LiDAR sensors to overcome challenges in GPS-deni [episode]
- PMTRM: Pseudo-Memory Temporal Re-encoding Module for Embodied Policy Learning — The gist: PMTRM, a lightweight bounded-history temporal representation module, improves phase disambiguation in robotic manipulation without retrieval or extra policy tokens. [episode]
- Higher-Order Action Supervision Makes A Strong Policy Class — The gist Higher-Order Action Supervision Makes A Strong Policy Class shows that simultaneously supervising both zeroth- and first-order actions can dramatically enhance policies’ performance and control robustness Problem Statement Modern data-driven decision-making methods oft [episode]
- OmniDex: Scaling Dexterous Hand Grasping to Diverse Cluttered Scenes —
- Demonstrating Arena 5.0: A Photorealistic ROS2 Simulation Framework for Developing and Benchmarking Social Navigation — The gist Arena 5.0 provides three main contributions: 1) The complete integration of NVIDIA Isaac Gym, enabling photorealistic simulations and more efficient training, 2) A comprehensive benchmark of state-of-the-art social navigation strategies, evaluated on a diverse set of gen [episode]
- How Firm Should a Grasp Be? —
- SafeInferCom: Safe Inference-Time Compute via Verifier-Guided Mid-Generation Intervention for Robotic Task Planning —
- Reference-Filter-Driven Transition Probabilities for IMM-Based Satellite Maneuver Detection — The gist: The proposed Reference-Driven IMM (RD-IMM) filter eliminates feedback loops in adaptive-TPM IMM filters by driving the transition probability matrix from a statistic computed outside the IMM, leading to better maneuver detection and reduced position error compared to ex [episode]
- SimVLA: Zero-Shot Sim-to-Real VLA Learning for Mobile Manipulation — The gist SimVLA introduces an end-to-end framework that trains Vision-Language Models for mobile manipulation entirely on synthetic simulation data without requiring teleoperation, demonstrating zero-shot transfer to real home environments. [episode]
- Being-M0.7: A Latent World-Action Model for Humanoid Robots —
- Distributed Relative Localization for Homogeneous Multi-Robot Systems through UWB Ranging and Limited Communications — The gist Accurate and reliable relative localization is crucial for multi-robot applications like exploration, search, and rescue missions <ref:2610.11308#pg6>. [episode]
- USDCraft: Geometrically Grounded Programmatic Modeling of Articulated 3D Assets for Simulation — The gist USDCraft enables generation, reconstruction, and real-to-sim-to-real manipulation because it formulates articulated asset reconstruction as programmatic modeling grounded in partial geometric evidence and introduces a framework where a pretrained LLM writes and revises e [episode]
- PlanWAM: Planning-Shaped Future Representations for End-to-End Autonomous Driving — The gist The authors propose PlanWAM, a Planning-Shaped World Action Model, which reformulates future modeling in end-to-end autonomous driving as a planning-oriented predictive representation learning problem to enable foresighted planning. How it works 1. [episode]
- WAM-Cache: Staleness-Bounded KV Reuse for Efficient World Action Models —
- Rewiring Semantics, Dynamics, and Control: A Simple yet Effective Action-Centric Tri-Stream Transformer — The gist: ACT3 introduces a simple yet effective Action-Centric Tri-Stream Transformer that fuses semantic and dynamics information into control actions while preserving the distinct roles of context streams. [episode]
- Experience-Guided Initiation Search for Learned Skills in Skill Composition — The gist: EVIS, an Experience-Guided and BehaviorValidated Initiation Search framework, discovers reliable initiation configurations for frozen learned skills in new environments by using historical execution experience to prioritize candidates and target-environment rollouts to [episode]
- Causal-fate dynamics of unrealized influence — The gist: Causal-fate dynamics provide a minimal language for a form of history that is absent from the present trajectory yet remains consequential for future evolution. [episode]
- RoboAware: Learning to Coordinate Embodied Skills from Counterfactual Outcomes — The gist The RoboAware framework learns state-dependent responsibility between modular composition and end-to-end control while keeping the coding agent and API library fixed. [episode]
- DAMP: Humanoid Locomotion via Denoised Belief Learning and Adversarial Motion Priors — The gist: DAMP introduces a reinforcement learning framework for robust and naturalistic humanoid locomotion over challenging terrains by implicitly inferring privileged and other task-relevant latent information using recurrent neural networks. [episode]
- WARP-VLA: Wrist-Camera Adaptation for View-Robust Policy Execution in Vision-Language-Action Models —
- RAGNAROK: Radar-Aided Gravity-Normalized Alignment for Robust Open Keyframe-based Radar-Visual-Kinematic-Inertial SLAM — The gist The RAGNAROK framework presents a novel radar-visual-kinematic-inertial SLAM system designed for robust operation in challenging environments by integrating slip- and rolling-contact-aware leg velocity estimation, a kinematics-aware radar factor, and degradation-aware im [episode]
- CoCam4D: Geometry-Aware Cooperative 4D Perception for Camera-Only Autonomous Driving —
- SDPAD: A Fully Spike-Driven Pipeline for End-to-End Autonomous Driving — The gist: SDPAD, a fully spike-driven end-to-end planning pipeline, achieves planning accuracy comparable to mainstream ANN planners while consuming less than 2% of their energy by converting perception and planning into integer spike representations in a single feed-forward pass [episode]
- Acting from Belief, Looking When Needed: A Bayesian Spatial World Model for Navigation under Intermittent Perception — The gist The robot navigation system ALONE acts from an internal spatial belief and looks again only when execution needs a new observation, potentially freeing the shared sensor for other tasks between navigation observations. [episode]
- Learning Language-Conditioned Traversability Representations for Adaptive Visual Navigation — The gist The framework presents LaTraNav, a dual-system architecture that couples a slow VLM with a fast flow-matching planner through agent- and language-conditioned traversability representations. [episode]
- Neural Networks for Temporal Pattern Recognition and Dynamic Arm Gesture Speed Estimation for Robot Control —
- Battery Second Life: A Review of Experimental Studies — As a meticulous AI researcher, I have thoroughly analyzed both provided texts (A and B) concerning the paper "Battery Second Life: A Review of Experimental Studies." My objective is to synthesize these summaries into a single, comprehensive, and highly detailed description of the [episode]
- Scalable LEO Conjunction Screening using Adaptive Synthetic-Covariance Thresholds —
- YOCO: You Only Calibrate Once! Fast Mocap Calibration for Dexterous Teleoperation — The gist: YOCO presents a fast few-shot, fine-tuning-free calibration framework that corrects biased hand-pose streams from a small set of paired raw and target poses to improve downstream dexterous teleoperation. [episode]
- Certified Scalable Enclosures for Uncertain Underdetermined Systems —
- From Solo to Ensemble: A Hierarchical Framework for Composable Multi-Agent Human-Object Interaction — The gist: This paper proposes a hierarchical framework that converts single-agent human-object interaction policies into reusable Object-oriented Motion Skills, enabling high-level policies to coordinate multiple agents through compact object-level proxy motions. How it works 1. [episode]
- 2DGS-Planner: Rasterization-based Path Planning in 2D Gaussian Splatting Map — The gist—2DGS-Planner proposes a path planner for ground robots that reads planning-relevant geometry from a 2D Gaussian splatting (2DGS) map through rasterization, rather than treating individual Gaussian primitives as obstacles. How it works 1. [episode]
- UltraLight Luma: A Novel Edge-Deployable Perception Network for Crop-Row Segmentation in Agricultural Robotics — The gist The proposed approach uses an ultra-lightweight segmentation model to predict crop-row regions from RGB images in real time on constrained hardware, improving parameter efficiency over U-Net, YOLOv8, and YOLOv26 while maintaining reliable crop-row detection performance C [episode]
- Narrow and Deep: An Ontology Tower as the Knowledge of an LLM Agent for an Industrial Equipment System — The gist The ontology tower, a narrow-and-deep ontology of a single equipment system whose knowledge deepens through physical derivations and through lessons incorporated from the operating journal, was proposed as the knowledge that an LLM agent receives about a plant and it was [episode]
- PathTime-VLA: Path-Time Decoupling for Factorized Post-Training of Vision-Language-Action Policies — The gist The PathTime-VLA framework introduces a path-time decoupled action representation for Vision-Language-Action policies, separating geometric guidance from execution timing to enable distinct post-training interfaces for geometry and timing. How it works 1. [episode]
- WAND: Learning Robust Navigation under Complex Wind Disturbances and Dense Obstacles for Quadrotors — The gist Robust navigation in cluttered environments remains a fundamental challenge for quadrotors, particularly when strong wind disturbances arise, which perturb vehicle dynamics, limit control authority, and substantially increase collision risk. [episode]
- Instrumentation and Stabilization of Electric Arcs for Plasma Smelting Reduction —
- Redefining fuel poverty: Introducing the temporal equity framework (TEF) — The gist: The Temporal Equity Framework (TEF) is proposed as a budget standard-based approach paired with a proportional energy expenditure indicator to address deficiencies in current fuel poverty definitions, arguing that responsive definitions are paramount to ensuring the equ [episode]
- STAG: A Sparse Traversability-Aware Graph Representation from Grid-Based Costmaps for Robotic Navigation — The gist The STAG Sparse Traversability-Aware Graph representation converts grid-based traversability costmaps into compact graphs for efficient global path planning, reducing planning time and memory while preserving essential terrain information> (Page 1). [episode]
- TACROSS: An Efficient and Low-Cost Scalable Human Touch System Across Heterogeneous Tactile Sensors for Dexterous Robot Learning — The gist: TACROSS presents an efficient and low-cost system for learning from human touch and transferring it to robots by aligning tactile streams at the level of contact events rather than raw sensor values. [episode]
- Tell Robot What Not to Do: A Negation Understanding Perspective — The gist The proposed NegaAlign framework enables vision-language-action models (VLAs) to follow negated instructions by extending pretrained VLAs through image–language supervision alone, achieving significant improvements in negated instruction following success rates while p [episode]
- Reliability-Aware Future Conditioning for Temporally Robust Robot Manipulation — The gist: Reliability-Aware Future Conditioning (RAFC) treats temporal misalignment as a control problem by estimating how far to trust received clips and which nearby temporal hypothesis to prefer, recovering most of the loss incurred by generated futures under off-grid phase sh [episode]
- From Asymptotic to Designer-Assigned-Time Control: A Review of Stability Notions, Design Mechanisms, and Controller Architectures — It is not a continuous narrative summary but rather an analytical framework presented through structured definitions, layered analysis, comparative tables, and concluding theses. [episode]
- CAPABLE: Capability-Aware Policy Adaptation via Behavioral Latent Encoding — The gist The CAPABLE framework introduces a unified capability-aware adaptation framework for frozen Vision-Language-Action (VLA) policies that integrates self-supervised capability inference with residual reinforcement learning to enable fault recovery without requiring explicit [episode]
- Predicting Cable Dynamics with Physical Attention Bias — The gist A physical attention bias improves prediction on unseen cables by allowing learned simulators to choose between arc-length and Euclidean distances for attention, which is most effective when attention is the only mechanism connecting distant segments How it works The pap [episode]
- REACT: Rolling Denoising and Dual Decoupling for Reactive Robot Control with VLA Models — The gist REACT reformulates flow-based action generation into a receding-horizon rolling denoising process to make flow-based VLAs more reactive while preserving long-horizon context How it works REACT introduces a rolling-denoising framework that maintains a persistent action bu [episode]
- Humanoid World Action Model With Joint State--Action Generation — The gist The HWAM introduces a Humanoid World Action Model with joint state–action generation, which makes the robot’s post-execution proprioceptive state an explicit prediction target to address the action–execution gap in hierarchical humanoid systems. [episode]
- Policy Synthesis for Finite Populations of MDP Agents under Aggregate Reach-Avoid Chance Constraints — The gist: This research presents a policy synthesis framework for chance-constrained reach-avoid control of finite-N swarms on discrete MDPs using Sequential Convex Approximation to optimize density-feedback policies under aggregate reach–avoid chance constraints. [episode]
- Traceable World State: A Provenance-Aware State Representation and Deterministic Replay Framework for Robotic Systems — The gist The Traceable World State (TWS) presents a middleware-neutral semantic representation and reference runtime for provenance-aware robot world state, which is crucial for determining where facts came from, reproducing earlier decision contexts, or detecting corruption in r [episode]
- SkillWeave: Weaving Heterogeneous Demonstrations into Long-Horizon Manipulation Skills — The gist: SkillWeave introduces a framework that learns from multi-modal demonstration by combining teleoperation for long-horizon coverage and kinesthetic teaching for precise, contact-rich interaction to achieve 27% average end-to-end success across three real-world long-horizo [episode]
- A Minimal Optical-Flow Representation for Vision-Based Tactile Rotation Classification in Robotic Manipulation Across Gravity Domains —
- ManiUnit: A Manipulation Skill Dataset and Benchmark for Long-Horizon Tasks — The gist The work introduces ManiUnit, a manipulation skill dataset and benchmark built from 50 BEHAVIOR-1K activities, to address challenges in learning and evaluating long-horizon mobile manipulation tasks by treating individual operations as units of training and evaluation Sk [episode]
- UNITAS: A 3D-Native World Action Model for Embodied Manipulation — The gist The first 3D-native world action model that unifies observations, actions, and scene dynamics in a shared metric 3D frame within each interaction, using a common representation across robot embodiments and human hands. [episode]
- Predefined-Time Integral Reinforcement Learning for Unknown Nonlinear Systems via Inverse-Optimal Design — The gist This paper develops a new predefined-time integral reinforcement learning framework for optimal control of unknown nonlinear systems via inverse-optimal design. [episode]
- PIER: An Evidence-Gated Execution Interface for Robotic Manipulation — The gist The authors present PIER, an execution-authorization interface that separates evidence checks, decision provenance, and stage-scoped re-observation from hardware control. [episode]
- SuperNav: An Agentic Navigation System for Any Task in Any Scene —
- Leveraging Human-In-The-Loop Demonstrations in Reinforcement Learning for Digital Twin-Driven Robot Flexibility — The gist The proposed framework combines a digital twin, reinforcement learning, and human demonstrations to create an online training system for collaborative robots that can adapt to changes in their physical workspace. [episode]
- Stochastic Distribution Network Reconfiguration under Load Uncertainty — The gist This paper investigates distribution network reconfiguration under demand uncertainty using a twostage stochastic formulation, showing that while reconfiguration reduces expected losses, the incremental benefit from explicitly modeling demand uncertainty is small when sc [episode]
- Instance-anchored interaction evidence: Grounding robot plans in human pointing and handling — The gist The authors propose instance-anchored interaction evidence (IAE), which registers every object of the final scene to its public identifier, keeps each identity through the video by backward mask propagation, and describes every frame by the geometry between hands, forear [episode]
- Unifying Policy Learning and State Prediction through Spatial Language Modeling — The gist Learning how actions change scene geometry can provide complementary supervision for goal-directed manipulation through Spatial Language Modeling. [episode]
- RESETTLE: Robotic Recovery through Disagreement-Triggered Retrieval and Efficient Corrective Control — The gist: RESETTLE introduces a model-agnostic framework that provides computationally efficient recovery at the action-execution interface of frozen robot policies by triggering recovery when two action proposals independently sampled under identical conditioning persistently di [episode]
- MiniWAM: Learning Compact Future Targets for Efficient World-Action Modeling — The gist The introduction introduces MiniWAM, which instead predicts compact future representations learned from privileged current–future transitions, demonstrating that effective world–action modeling does not require predicting native visual futures and that compact predic [episode]
- MAP2: Model- and Acceleration-Based Pursuit with MPC and Gaussian Process Residual Learning for Autonomous Racing — The gist Autonomous racing requires accurate trajectory tracking near handling limits while maintaining low computational latency, and MAP2 addresses this by combining curvature-based kinematic MPC with Sparse Gaussian Process residual learning to enhance performance compared to [episode]
- Sim-to-Real RL for ASVs using SysID — The gist: This work presents an ASV simulator and accompanying pipeline that enables training policies starting from unknown vehicle dynamics. [episode]
- Stabilization of Unidirectional First-Order PDE-ODE Coupled Systems with Boundary and Distributed Input Delays — The gist The proposed method can simultaneously address distinct delays of boundary control and distributed control without requiring ordering of the delays <ref:2610.12226#pg6>. [episode]
- Residual Modeling Closes the Regression and Generative Policy Gap in Robot Learning — The gist Learning from demonstration has enabled impressive robot behaviors, but a gap exists between generative policies and direct action regression due to residual modeling mismatches, which this paper addresses by introducing heteroscedastic Student-t action regression (HT-Po [episode]
- From Language to Motion: Task-Conditioned Focal-Stack Trajectory Integration for Microscopic Robots —
- Fixed-Reference Pose Residuals for Measuring Cross-Dataset Cue Transfer in Human-Robot Interaction Anticipation — The gist The fixed-reference pose residual (FRPR) model provides an exact split between geometry and pose terms, which serves as a measurement instrument to quantify how different cues transfer between cross-dataset environments. [episode]
- Real-Time Motion Planning with Dynamic Hazards: Classical vs. Learning-Based Methods — The gist: Uncertainty in obstacle evolution, rather than partial observability, is the key factor that changes planning difficulty and determines which paradigm is practically effective in real time. [episode]
- Walking on Roofs: Exploring the Potential of Walking Robots for Construction Work on Roofs — The gist The first systematic experimental evaluation of quadruped locomotion in tiled roof environments establishes a baseline for future research. [episode]
- Toward Lunar Legged Robots: Field Deployment Lessons at LUNA — The gist Legged robots are promising candidates for future lunar surface missions because they can traverse steep, loose, and obstacle-rich terrain that challenges conventional wheeled rovers> Campaign Overview and Setup The LUNA testing campaign involved operating two quadrupeda [episode]
- PLaW-VLA: Predictive Latent World Modeling for Vision-Language-Action Policies — The gist: PLaW-VLA introduces a framework that models task-relevant future states in a prediction-oriented representation space and conditions action generation on these predicted futures, enabling improved long-horizon control and generalization. [episode]
- Convex Safety Filtering via Spectral Selection for Nonconvex Safe Sets — The gist The discrete-time control barrier function condition for a safe set that is a union of convex sets is nonconvex in the control input, and this paper shows that selecting eigenvectors of the matrix function at the current state yields a convex input constraint that acts o [episode]
- LiteNWM: Efficient Latent World Models for Onboard Visual Navigation in the Wild — The gist The LiteNWM model selects trajectories from a frozen policy without RGB reconstruction by combining action-conditioned multi-horizon latent prediction with incumbent-relative scoring How it works LiteNWM is designed to address the bottleneck of candidate-wise RGB rollout [episode]
- Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement —
- ARC: A Reasoning Recipe for Robot Foundation Models — Text A is a detailed extraction of key findings, contributions, and summaries from the paper "ARC: A Reasoning Recipe for Robot Foundation Models." Text B correctly identifies that it is an inference-time prompt rather than the paper itself. [episode]
- A Physics-Informed Collision Learning Framework for Collaborative Robot Motion Generation — The gist The Physics-Informed Unified Differentiable Framework (PI-UDF) is a compact framework for body-to-body collision distance learning between articulated robots that provides a differentiable collisiondistance representation suitable for closed-loop collision-aware collabor [episode]
- LeWAM: A JEPA World Action Model with Diffusion-Steering-Based MPC —
- GLIO2: A GPU-Parallelized Tightly-Coupled LiDAR-Inertial-GNSS System for Robust and Real-Time Global Localization and Mapping — The gist Globally consistent, real-time state estimation in large-scale, perceptually degraded environments is essential for autonomous vehicles and aerial robots, and requires fusing LiDAR, inertial, and GNSS measurements. [episode]
- RoboRSI: Stable, efficient, and reusable robot self-evolution in complex real-world environments — The gist The introduction states that execution experience becomes reusable only when it is organized around the task structure that gives it meaning, and this structure must serve both human-friendly steering and agent-friendly structure. [episode]
- Control-Ready Uncertainty for Trajectory Diffusion —
- FAITH: Feasibility-Aware Safety-Filtered RL for High-Dimensional Systems — The gist The FAITH framework introduces a feasibility-aware, model-free reinforcement learning approach that separates safety enforcement from task policy optimization by approximating an optimal state-action safety value and training the task policy through a learned feedforward [episode]
- VioLA: Learning Generalist Humanoid Control Policies from Human Data — The gist The VioLA generalist humanoid policy learns to predict body and hand motion latents instead of joint commands, enabling zero-shot locomotion and manipulation on a real robot without task-specific fine-tuning How it works VioLA introduces a generalist humanoid policy that [episode]
- Generative Neural Retargeting for Human-to-Robot Dexterous Manipulation — The gist The proposed Generative Neural Retargeting (GNR) is a scalable and generalizable retargeting method that retargets large-scale, long-horizon and high-precision human demonstrations for five-fingered dexterous manipulation. [episode]
- SpatialHarness: Test-Time Spatial Scaffolding for Fine Robotic Manipulation — The gist The introduction argues that an important source of failure in fine robotic manipulation is insufficient spatial observability, which can be addressed by introducing SpatialHarness, a test-time embodied harness that provides test-time spatial scaffolding for fine robotic [episode]
- A Balanced Data Diet: Addressing the Exploration Bottleneck in Mega-Scale RL for Robot Control — The gist: Success Guided Sampling (SGS) is introduced as a simple adaptive sampler that concentrates Reinforcement Learning training on task configurations around the frontier of the policy’s capabilities, enabling large-scale simulated RL to make the most out of experience in [episode]
- CSF: Contextual Safety Filtering for Motion Generators — The gist The Contextual Safety Filtering (CSF) introduces a training-free filter that grounds natural-language safety rules in safe and unsafe reference trajectories produced by motion generators, reducing danger-event rates by up to 90% while preserving 88–100% of benign motio [episode]
- DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training —
- Dex-One2Many: Learning Dexterous Manipulation from a Single Human Demonstration —
- Does Dynamic-Point Filtering Help When Texture Is Scarce? A Controlled Study of ORB-SLAM2 Front-Ends in Synthetic Indoor Scenes —
- Teaching a Robot Dog New Tricks: Diverse Quadruped Skills via Combined Reinforcement and Imitation Learning with Adversarial Task Selection — The gist This work proposes a three-stage method that trains a single policy to perform distinct tasks such as walking, digging, and hopping, and compose them into novel behaviors such as crawling. [episode]
- TacHair: Tactile Contact-Distribution Guided Online Correction for Robotic Hair Stroking and Perception — The gist: TacHair proposes a tactile contact-distribution guided online correction framework for robotic hair stroking that improves task success and contact maintenance by explicitly representing local hair–sensor interaction as a spatial tactile contact distribution. [episode]
- Masked Generative Motion Planning with Geometry-Guided Token Search — The gist The Masked Generative Motion Planning (MGMP) introduces a method that extends learned trajectory priors from efficient parallel generation to structural repair, enabling route-level restructuring beyond local trajectory deformation How it works MGMP reframes motion plann [episode]
- Safe Learning of Adaptive Control Policies for Remote Patient Monitoring — The gist Remote Patient Monitoring (RPM) enables continuous observation of patients in their daily environments, improving both health outcomes and quality of life, while this paper develops a learning-based control framework that estimates system parameters and adapts monitoring [episode]
- TAPNAV: Humanoid Navigation through Tactile Active Perception — The gist Navigation in vision-denied environments is challenging for humanoid robots because proprioceptive odometry drifts and localization uncertainty accumulates rapidly, and TAPNAV presents a tactile active-perception framework that enables humanoid navigation toward a goal b [episode]
- NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime —
- Compute-Constrained Safety Filters with Neuromorphic Event Triggering — The gist The proposed dual-LIF controller uses separate safety and performance states to schedule CBF-filter computations under delayed input application, guaranteeing robust forward invariance, completeness, and non-Zeno execution How it works The core mechanism involves a dual [episode]
- Same Action, Different Outcome: Variability in Dynamic Cloth Manipulation — The gist The authors quantify and characterize variability in dynamic cloth manipulation in 269 collected executions across four dynamic tasks with a total of 27 different trial configurations > Interaction Variability Interaction variability is defined as "the non-repeatability [episode]
- High-Fidelity Baseline Design and Station Keeping Analyses for Earth-Moon Vertical Orbits — The gist: This work develops an end-toend framework for high-fidelity reference generation under rotating-frame geometry constraints and receding-horizon model predictive control for Earth–Moon vertical orbits. [episode]
- Diagnosing and Recovering from Observation-Space Shift at Long-Horizon Skill Seams — The gist: The authors diagnose long-horizon skill seam failures as Observation-Space Shift (OSS), finding that they are primarily driven by displaced scene state, and propose a learned detect–restore–resume framework to recover from these failures. [episode]
- Skill-SLM: Agent Skill-driven Small Language Models for Reliable Robot Operation — The gist Skill-SLM proposes a framework that reformulates SLM-driven robot operation as a task-decomposition and skill-composition problem, which substantially outperforms distillation-oriented baselines, especially on unseen tasks that require generalization of capabilities. [episode]
- Cross-Embodiment Robot Foundation World Models with Latent Actions — The gist: The Latent Action Conditioned Robot World Model (LAC-WM) improves downstream performance over an Explicit Action-Conditioned World Model (EAC-WM) by up to 46.7% on dexterous manipulation and 11.7% on LIBERO, due to its unified latent action space which allows its perfor [episode]
- OmniHOI: Dexterous Hand-Object Interaction from Monocular Human Video — The gist: OmniHOI presents a pipeline that turns an RGB video of hand-object interaction into an interaction-faithful trajectory on dexterous hands by enforcing physical consistency across reconstruction, retargeting, and physics-in-the-loop refinement. [episode]
- Moving Horizon Estimation of Hybrid Systems Using Hybrid Zonotopes and Mixed-Integer Quadratic Programs —
- Deceptive Stochastic Patrolling via Markov Chain Lifting — The gist In this paper we propose lifted Markov Chains as a new paradigm for deriving patrol strategies for mobile agents on an environment represented as a graph > How it works 1. [episode]
- Informationally Decoupled Trajectory Design for Sim-to-Real System Identification — The gist The Informationally Decoupled Trajectory Design framework (IDTD) formulates an objective for exploration policy built on the Schur complement score derived from the Fisher information matrix to ensure every parameter leaves a unique effect on the trajectory, thereby impr [episode]
- When Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAs —
- Higher-Order Morphology Priors for Quadruped Reinforcement Learning Under Actuator Degradation — The gist The node-edge-face Hodge actor achieves the highest return on unseen actuator degradations, with higher survival and lower velocity-tracking error. [episode]
- EnergyNet in Practice: Long Grid, Short Grids and Mobility-Based Energy Peering — The gist: The Short Grid proposes that local electrification needs not automatically require a second investment in proportionate upstream capacity by allowing locally controlled buildings, blocks, or campuses to coordinate their demand, generation, storage, and local market behi [episode]
- iAm.md: Robot Skill Self-Assessment through Agentic Introspection for Unknown Open-Vocabulary Domains — The gist The paper presents iAm.md, a Markdown standard and generation framework, that allows anchoring agentic introspection in robot behavior generation through open-vocabulary semantic mapping and persistent object records. [episode]
- Constructive Safety-Critical Control for a Class of Underactuated Systems: A Hierarchical Approach — The gist Underactuated systems with nontrivial geometry are abundant in practical robotic problems, and this work provides a constructive framework for synthesizing safe control architectures and control barrier functions for quadrotor-like underactuated systems by studying their [episode]
- Centrality-based Structural-Electrical Contingency Screening and Transmission Reinforcement Prioritization —
- ActiveReg: Information-Driven Active Regional Probing for Partial-to-Full Bone Registration — The gist: ActiveReg presents a closed-loop framework that recommends probing regions rather than individual points, allowing flexibility in contact location and achieving accurate bone registration with substantially fewer acquired points. [episode]
- A Reconfigurable Fabric Based Pneumatic Actuator with Button Fastened Constraint Modules for Multi Mode Actuation — The gist: This study presents a reconfigurable fabric-based pneumatic actuator that enables rapid, toolfree, and reusable switching among multiple actuation modes by attaching constraint fabric modules to a buttoned textile sleeve. [episode]
- A Latent Space Optimization Approach for Symbolic Discovery of Dynamical Models — The gist The proposed framework uses a hybrid latent-space Bayesian Optimization (BO) and global optimization framework for solving symbolic regression tasks to discover dynamical models from data. [episode]
- Towards Path-Creative Navigation: Robot Navigation through Embodied Interaction — The gist Autonomous navigation in cluttered and constrained environments typically assumes a fixed environment and searches only for paths within existing free space, however, reaching the goal may therefore require appropriate embodied interaction with the environment. [episode]
- FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment — The gist FOCUS, a single-stage PPO framework, trains an actor to use privileged state information during training while automatically regulating whether it collects rollouts from RGB-D or privileged state latents, thereby improving sample efficiency and achieving superior test pe [episode]
Important terms
- SafeInferCom
- This system inserts a verifier-guided intervention during robot planning to ensure that even when rules aren't perfectly known, the generated plan remains safe for complex tasks.
- Latent World-Action Model Being M0.7
- This is a major development creating a model that understands and acts within a world context for humanoid robots by integrating latent space with action policies, moving beyond simple reactive control.
- ExecVLA
- This work directly tackles ensuring robots follow very precise execution constraints in vision-language-action models by using bi-level action representations for better adherence to instructions.
- Sliding-window filter approach
- A robust method for tracking a robot's state in real time, especially with noisy sensor data, which is foundational for reliably operating complex continuum robots.