Robotics papers — 2026-10-05
Today’s focus is squarely on improving how vision language action models interact with the real world, which is crucial because we need these systems to move beyond simulation and actually perform fine-grained robot manipulation. We explored World-to-Wrist, which tackles task-conditioned future wrist modeling for robot manipulation, aiming to give the model a better sense of what the end effector needs to do next based on the current situation. This builds upon GeoScaffold, where we learned how to get compact geometric latents through reconstruction for more efficient vision language navigation in complex environments.
A key piece of work involves World-Calibrated Proposal-to-Action Flow, which focuses on calibrating the flow between proposal and action for vision language models. This is important because it helps ground the abstract planning in concrete actions that are actually executable by a robot. We also looked at FastOPD, which uses on-policy distillation to create lightweight versions of vision language action models, making them faster for deployment.
Finally, we touched upon PointWAM, which deals with 3D world action modeling specifically for dexterous robotic manipulation. This work connects the high-level planning ideas from World-Calibrated Proposal-to-Action Flow with the low-level physical requirements addressed by PointWAM.
The most important development is the work on learning low-frequency motion control for robust and dynamic robot locomotion because it directly tackles making robots move better in real, unpredictable environments. This involved developing methods to learn how to control these robots without needing extensive pre-programming. Specifically, one line of research focused on learning this low-frequency motion control, which suggests a way for robots to handle the subtle, slow movements needed for stable walking or crawling.
Another key piece is the work on accurate open-loop control of a soft continuum robot using visually learned latent dynamics. This means they figured out how to precisely command these flexible robots to move without constantly needing real-time feedback loops, by learning from visual data about their internal states. This contrasts with the previous work because it moves away from purely reactive control toward predictive movement based on what the robot sees.
Then there is the INSIGHT project, which focuses on inference-time sequence introspection for generating help triggers in vision-language-action models. This work is significant because it aims to make these complex AI systems more helpful by figuring out when and how to prompt them for assistance during operation. This builds upon the idea of using visual information to guide action, similar to how the soft robot control uses visual data.
Finally, there is the ROS Help Desk framework, which provides a GenAI powered, user-centric system for diagnosing and debugging ROS errors. While less focused on physical robotics than some of these papers, it matters because it improves the usability of the entire ecosystem by making troubleshooting much easier for developers. This tool complements the research by providing a practical way to debug the complex systems being developed in areas like motion control or vision-language models.
The most significant work today involved developing a method for dynamic robotic cloth folding using an efficient Koopman operator based model predictive control. This approach allows the robot to handle the complex, non-linear dynamics of fabric manipulation in real time by predicting future states.
This is important because it moves beyond pre-programmed motions toward genuine physical interaction with deformable objects, which is a key hurdle for practical manipulation in unstructured settings. The method leverages a Koopman operator model to predict how the cloth will behave under control inputs, enabling precise folding actions.
Another area of progress focused on long-term navigation through change robust online topological memory. This system aims to keep track of the environment's layout even when it undergoes significant changes over time, which is crucial for persistent robotic agents. It achieves this by maintaining a map that can be updated incrementally as new information is gathered.
We also saw work on agentic navigation where a zero-shot vision-and-language navigation system was framed as a tool-calling harness. This means the robot learns to use existing tools, like language models, to figure out how to navigate novel areas without explicit prior training for every possible scenario.
Finally, there is research into action expert pretraining which improves instruction generalization for vision-language-action policies. This work suggests that by pretraining experts on specific actions, the resulting policies become much better at following complex instructions in new situations.
The most significant development today centers on the work that addresses uncertainty quantification for flow-based generalist robot policies, which is crucial because it allows these robots to make safer decisions when they encounter situations outside their training data. This approach involves developing methods to measure how much the robot's predictions might be wrong, which helps in planning actions under novel conditions.
Building on this foundational work, there was progress on communication-aware robot execution for cloud inference under spatially heterogeneous connectivity; this tackles the real-world problem of robots needing reliable data transfer when their network connection is patchy and uneven across different areas. This is important because it moves AI from controlled lab settings to unpredictable environments where data transmission is a major hurdle.
Another area of focus was the development of a biomimetic myoelectric tentacle prosthesis that incorporates sensorless object detection and vibrotactile feedback, which aims to give users more intuitive control over their prosthetic limbs. This work connects directly to the need for better interaction, as it focuses on how the physical interface between human and machine can be made more natural.
Simultaneously, research into real-time sEMG-based telecontrol of an assistive robotic arm using a one-dimensional convolutional neural network showed promising results in controlling robotic arms with muscle signals. This demonstrates a practical application of deep learning for direct human control over physical machinery.
Furthermore, the concept of making a change of frame affect the capture point proprioception in humanoid single-leg balance suggests that manipulating how we perceive space can improve complex locomotion tasks. This is an interesting way to enhance the robot's internal sense of self and its interaction with its environment during movement.
Finally, there is ongoing work on awomo-simdataengine, which creates agentic simulation-ready worlds, providing a robust platform for training and testing these complex robotic systems before they ever touch the real world. This engine supports the broader goal of creating more capable agents through sophisticated simulation environments.
The most critical piece of work today involves developing a social perception gateway for human reaction based failure detection and recovery in visual language agent manipulation, which matters because it addresses the safety concerns when robots interact with people. This research explored SocialVLA, which aims to detect when a humanoid robot's actions are failing by observing how humans react to those actions.
This is supported by work on filter-aware fine-tuning for safe whole-body tracking, where researchers adjusted models based on specific filters to ensure the robot maintains stable tracking during movement. This relates to the broader effort in rethinking world-action models for compositional and in-context robotic manipulation, which seeks a more flexible way for robots to understand and execute complex tasks.
Another important direction is the development of programmable effect-to-execution world-action models, which allows the system to focus on achieving a desired outcome rather than rigidly following a pre-set actor path. This concept connects directly to OpenRUA, which investigates how robot use agents can achieve zero-shot visuomotor policies without prior training data.
Finally, there is the work on degradation-balanced motion planning for robotic manipulators, which focuses on motion planning that accounts for the expected degradation of the system over time to ensure reliable movement.
The most significant development centers on CriticHack, which attempts to evaluate visual rewards under robot policy optimization. This matters because understanding how robots learn from visual feedback is crucial for building more robust autonomous systems. The work involved setting up a framework where a robot's actions are assessed based on the resulting visual reward, and they found that this method provides a structured way to judge policy performance.
This is supported by DeltaWorld, which creates physically consistent interactive world simulators using action-conditioned latent increment learning. This simulation technique is important because it allows researchers to train agents in a realistic virtual environment before deploying them in the real world. Furthermore, AdaTempo focuses on learning shared relative tempo from demonstrations to speed up robot manipulation tasks.
A passive AI system for verifying physical state on automated liquid handlers was also explored, which is significant for ensuring safety in complex industrial settings. This system works by observing the environment to confirm if a physical state is correct, and it showed promise in this verification task. Finally, RoboBridge presents a self-evolving embodied agent framework designed specifically for sim-to-real transfer.
The most significant development today concerns the work on Skill2Real, which addresses the challenge of agentic skill learning for zero-shot sim to real robot manipulation. This research is crucial because it aims to bridge the gap between simulated training and real-world deployment for robotic skills. They explored an agentic approach where skills are learned directly from simulation without requiring extensive prior pretraining, suggesting a more efficient path to physical tasks.
Another important piece of progress involves Proprioceptive Sketches as Long-Horizon Intent for Generative Action Policies, which tackles how robots can maintain long-term goals by using internal sensory data. This method suggests that providing proprioceptive sketches allows generative action policies to anticipate future needs rather than just reacting to immediate stimuli. This builds upon the idea of learning complex behaviors, similar to the agentic skill learning discussed earlier.
The work on SimpleTouch investigates whether vision-language-action models can master contact-rich manipulation without needing tactile policy pretraining. This is significant because it tests if purely visual and language inputs are sufficient for intricate physical interactions, which is a major hurdle in dexterity tasks. This contrasts with the more foundational physics-based assessments like ManiPhysicsBench, which focuses on assessing object preservation during manipulation.
LOCUS provides a method for landmark-oriented container discrimination using spatial graphs to help robots identify objects based on their structural features. This contributes to perception capabilities, complementing the skill acquisition work by giving the robot better ways to understand its environment before attempting manipulation.
The work on SceneFactory-3D is particularly important because it tackles the challenge of making safety evaluations scalable by lifting two dimensional traffic scenes into three dimensional physical counterfactuals. This approach allows researchers to test how systems behave in real-world scenarios that are physically grounded, which is a significant step toward reliable autonomous system validation.
We also saw some progress on MixVLA, which focuses on the adaptive mixing of non invariant information for generalizable vision language action models. This method aims to make these models more robust by intelligently combining different types of data during training. This builds upon the earlier work concerning learning reflexive behavior for contact rich manipulation, which explored how agents can learn to interact physically with objects.
A related piece looked at register routed delayed fusion, which rewires shortcut prone observation fusion in visuomotor imitation tasks. This suggests a way to better process sensory input when an agent is trying to mimic physical actions through vision and motor commands. This contrasts with the work on permutation robustness in multiagent transformer policies, which found that simple permutation invariance is insufficient for preventing action collapse in those agents.
Finally, there was research into subject specific predictive musculoskeletal simulations of lower limb exoskeleton assistance, examining the metabolic and biomechanical effects of different joint assistance strategies. This provides a detailed look at how physical support systems impact human movement and energy expenditure.
Today's papers
- DriftWorld: Fast World Modeling through Drifting. [paper] [episode]
- World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation. [paper] [episode]
- World-Calibrated Proposal-to-Action Flow for Vision-Language-Action Models. [paper]
- GeoScaffold: Learning Compact Geometric Latents via Reconstruction for Efficient Vision-Language Navigation. [paper]
- FastOPD: On-Policy Distillation for Lightweight VLA Deployment. [paper]
- PointWAM: 3D World Action Modeling for Dexterous Robotic Manipulation. [paper]
- From Language Priors to Field Adaptation: Preference Learning for Traversability Estimation. [paper]
- HexVIO: Towards All-Day Stereo-Inertial Tracking Through Commodity DSPs. [paper]
- I2CD: Direct Image-to-Convex Decomposition for Simulation-Ready Collision Geometry. [paper]
- XGenAct: Geometry-Enhanced World Action Models through Cross-Task Generation. [paper]
- Learning Low-Frequency Motion Control for Robust and Dynamic Robot Locomotion. [paper] [episode]
- ROS Help Desk: GenAI Powered, User-Centric Framework for ROS Error Diagnosis and Debugging. [paper] [episode]
- INSIGHT: INference-time Sequence Introspection for Generating Help Triggers in Vision-Language-Action Models. [paper] [episode]
- A Learning-Free Characterization Framework for the Resilience and Sensitivity of Polyurethane Vision-Based Tactile Sensors. [paper] [episode]
- Robot Crash Course: Learning Soft and Stylized Falling. [paper] [episode]
- Accurate Open-Loop Control of a Soft Continuum Robot Using Visually Learned Latent Dynamics. [paper] [episode]
- Move-Then-Operate: Behavioral Phasing for Human-Like Robotic Manipulation. [paper] [episode]
- Predictive Spatio-Temporal Scene Graphs for Semi-Static Scenes. [paper] [episode]
- Change-Robust Online Topological Memory for Long-Term Relocalization and Semantic Navigation. [paper] [episode]
- AGT-CV: An Aerial-Ground Team Cross-View Dataset for Heterogeneous Robot Teams in Unstructured Environments. [paper] [episode]
- Dynamic robotic cloth folding with efficient Koopman operator-based model predictive control. [paper] [episode]
- S2M-Trek: From Single to Multi-Sphere Transport via Per-Frame Deep Sets on a Wheel-Legged Robot. [paper] [episode]
- AgenticNav: Zero-Shot Vision-and-Language Navigation as a Tool-Calling Harness. [paper] [episode]
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies. [paper] [episode]
- ROVE: Unlocking Human Interventions for Humanoid Manipulation via Reinforcement Learning. [paper] [episode]
- Uncertainty Quantification for Flow-Based Generalist Robot Policies. [paper] [episode]
- Communication-Aware Robot Execution for Cloud Inference under Spatially Heterogeneous Connectivity. [paper] [episode]
- A Biomimetic Myoelectric Tentacle Prosthesis with Sensorless Object Detection and Vibrotactile Feedback. [paper] [episode]
- Real-Time sEMG-Based Telecontrol of an Assistive Robotic Arm Using a 1D Convolutional Neural Network. [paper] [episode]
- A Change of Frame Makes the Capture Point Proprioceptive: Distillation-Free Humanoid Single-Leg Balance. [paper] [episode]
- Awomo-SimDataEngine: Agentic Simulation-ReadyWorld Generation. [paper]
- SoTa: Soft Tactile Skins for Dexterous Manipulation. [paper]
- NEEDLEWORK: Offline Rewriting of Robot Data with Verified Local Stitches. [paper]
- Filter-Aware Fine-Tuning for Safe Humanoid Whole-Body Tracking. [paper]
- SocialVLA: A Social Perception Gateway for Human-Reaction-Based Failure Detection and Recovery in VLA Manipulation. [paper]
- Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation. [paper]
- Keep the Effect, Drop the Actor: Programmable Effect-to-Execution World-Action Models. [paper]
- Autonomous mobile robot operations logistics: a dataset of jobs, dispatch events and robot states. [paper]
- OpenRUA: Robot-Use Agents Are Zero-Shot Visuomotor Policies. [paper]
- RUL-Aware RRT*: Degradation-Balanced Motion Planning for Robotic Manipulators. [paper]
- CriticHack: Evaluating Visual Rewards Under Robot Policy Optimization. [paper]
- Real-time Event-camera Stereo Visual Odometry via Keytime Gaussian Process Regression. [paper]
- A Passive AI System for Verifying Physical State on Automated Liquid Handlers. [paper]
- DeltaWorld: Physically Consistent Interactive World Simulators via Action-Conditioned Latent Increment Learning. [paper]
- Who Went Where When on the Lunar Surface: Forensic Trajectory Analysis to Identify Byzantine Rovers. [paper]
- AdaTempo: Learning Shared Relative Tempo from Demonstrations for Faster Robot Manipulation. [paper]
- RoboChemGym: A Protocol-Driven Generative Simulation Framework for Long-Horizon Chemical Manipulation. [paper]
- RoboBridge: A Self-Evolving Embodied Agent Framework for Sim-to-Real Transfer. [paper]
- Around the World: Unified Learned Locomotion on a 270 g Continuous-Rotation Quadruped. [paper]
- CSIR: Contextually and Socially Informed Robots for Efficient Person Goal Navigation. [paper]
- Proprioceptive Sketches as Long-Horizon Intent for Generative Action Policies. [paper]
- Localized Conformal Safety Monitoring with Vision-Language Models for Autonomous Driving. [paper]
- SimpleTouch: Can Vision-Language-Action Models Master Contact-Rich Manipulation Without Tactile Policy Pretraining?. [paper]
- Skill2Real: Agentic Skill Learning for Zero-Shot Sim-to-Real Robot Manipulation. [paper]
- ManiPhysicsBench: Physics-Based Assessment of Object Preservation in VLA Manipulation. [paper]
- LOCUS: Landmark-Oriented Container Discrimination Using Spatial Graphs. [paper]
- SARI: Phase-Split Sim-Real Co-Training for Contact-Rich Manipulation. [paper]
- Learning Reflexive Behavior for Contact-Rich Manipulation. [paper]
- Register-Routed Delayed Fusion: Rewiring Shortcut-Prone Observation Fusion in Visuomotor Imitation. [paper]
- Permutation Robustness Is Not Enough: Action Collapse in Multi-Agent Transformer Policies. [paper]
The papers
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies — Vision-Language-Action (VLA) models often struggle to generalize to out-of-distribution (OOD) language instructions because continuous action experts, when trained from random initialization on imbalanced data, develop visual shortcuts that corrupt the Vision Language Model's lan [episode]
- Move-Then-Operate: Behavioral Phasing for Human-Like Robotic Manipulation — Move-Then-Operate presents a Vision language action framework that explicitly decouples robotic manipulation into two distinct behavioral phases: coarse relocation (move) and contact-critical interaction (operate). [episode]
- Predictive Spatio-Temporal Scene Graphs for Semi-Static Scenes — Predictive Spatio-Temporal Scene Graphs for Semi-Static Scenes addresses the challenge of enabling robots to perform complex reasoning across geometry and semantics in environments where objects exhibit semi-static changes over time. [episode]
- Learning Low-Frequency Motion Control for Robust and Dynamic Robot Locomotion — Robotic locomotion can be achieved robustly and dynamically even when using motion controllers operating at very low frequencies. [episode]
- ROVE: Unlocking Human Interventions for Humanoid Manipulation via Reinforcement Learning — ROVE presents a reinforcement learning framework designed to improve Vision-Language-Action (VLA) policies for humanoid manipulation by learning from imperfect human interventions. [episode]
- A Neuromodulable Current-Mode Silicon Neuron for Robust and Adaptive Neuromorphic Systems — Neuromorphic engineering makes use of mixed-signal analog and digital circuits to directly emulate the computational principles of biological brains, and this work presents a novel current-mode neuron design that supports robust neuromodulation with minimal model complexity, comp [episode]
- INSIGHT: INference-time Sequence Introspection for Generating Help Triggers in Vision-Language-Action Models — Recent Vision-Language-Action (VLA) models lack introspective mechanisms for anticipating failures and requesting help from human supervisors, which limits their safety and reliability in unstructured settings. [episode]
- NN-ETM: Enabling safe neural network-based event-triggering mechanisms for consensus problems — Event-triggering mechanisms (ETM) have been developed for consensus problems to reduce communication while ensuring performance guarantees, but their design has grown increasingly complex by incorporating the agent’s local and neighbor information. [episode]
- Uncertainty Quantification for Flow-Based Generalist Robot Policies — Vision-language-action models (VLAs) lack mechanisms to quantify confidence in their predictions and to detect when their actions may be unreliable, which presents a critical limitation for real-world deployment in non-stationary environments. [episode]
- Change-Robust Online Topological Memory for Long-Term Relocalization and Semantic Navigation — A new representation for spatial-semantic reasoning enables autonomous robots to maintain localization and navigate effectively in dynamic, real-world environments despite significant changes in appearance and scene content. [episode]
- Fairness-Guaranteed Online Power Allocation Policies for EV Fast Charging Stations — The rapid expansion of electric vehicle (EV) fast charging station (FCS) infrastructure necessitates scalable and efficient power allocation policies to prevent user bias and secure equitable access to limited resources while maximizing infrastructure utilization. [episode]
- Dynamic robotic cloth folding with efficient Koopman operator-based model predictive control — Robotic cloth folding is addressed by integrating physics-based simulation with efficient, kernel-based Koopman operator regression within a model predictive control framework to generate fast, accurate trajectories for real robotic execution. [episode]
- DriftWorld: Fast World Modeling through Drifting — Predictive world models enable robots to plan by imagining the outcomes of their actions, but their value for control hinges on generating many rollouts quickly. [episode]
- Scalar Federated Learning for Linear Quadratic Regulator — SCALARFEDLQR proposes a communication-efficient federated algorithm for model-free learning of a common policy in linear quadratic regulator (LQR) control of heterogeneous agents, significantly reducing per-agent uplink communication from O(d) to O(1) while maintaining fast linea [episode]
- Robot Crash Course: Learning Soft and Stylized Falling — A reinforcement learning technique is proposed that balances user-guided stylized pose objectives and damage-minimizing soft falling objectives for bipedal and other legged robots, addressing the risk of uncontrolled falls in real-world operation. [episode]
- AgenticNav: Zero-Shot Vision-and-Language Navigation as a Tool-Calling Harness — Zero-shot vision-and-language navigation in continuous environments (VLN-CE) has recently become feasible with large vision-language models (VLMs), but existing methods suffer from limitations in action space, depth utilization, and memory management. [episode]
- World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation — Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. [episode]
- A Change of Frame Makes the Capture Point Proprioceptive: Distillation-Free Humanoid Single-Leg Balance — Unified humanoid policies struggle to maintain clean single-leg balance, often resorting to recovery actions like stepping or hopping rather than prevention. [episode]
- AGT-CV: An Aerial-Ground Team Cross-View Dataset for Heterogeneous Robot Teams in Unstructured Environments — Heterogeneous air-ground robot teams combine complementary sensing modalities, mobility characteristics, and spatial viewpoints that can significantly enhance perception in complex outdoor environments. [episode]
- Deception Against Data-Driven Linear-Quadratic Control — Deception is a common defense mechanism against adversaries with an information disadvantage, forcing them to select suboptimal policies for a defender’s benefit. [episode]
- S2M-Trek: From Single to Multi-Sphere Transport via Per-Frame Deep Sets on a Wheel-Legged Robot — Multiple identical free-rolling spheres form an unordered set whose slot assignments may change independently at each history frame, creating a per-frame permutation symmetry that standard history-concatenation set encoders do not explicitly enforce—these encoders impose only a [episode]
- Accurate Open-Loop Control of a Soft Continuum Robot Using Visually Learned Latent Dynamics — Accurate open-loop control of a soft continuum robot (SCR) from video-learned latent dynamics addresses the challenge of controlling complex, continuous systems without real-time camera feedback by leveraging interpretable latent representations. [episode]
- Real-Time sEMG-Based Telecontrol of an Assistive Robotic Arm Using a 1D Convolutional Neural Network — Real-time sEMG-based telecontrol of an assistive robotic arm using a 1D Convolutional Neural Network addresses the challenge of providing intuitive, reliable, and responsive control for individuals with upper limb motor impairments by developing a complete real-time pipeline that [episode]
- End-to-End Abstraction-Based Control with LLM-Enhanced NL-to-LTL Translation — Abstraction-Based Controller Design (ABCD) offers a principled framework for the safe control of complex CyberPhysical Systems (CPSs), but interfacing real-world requirements with its formal synthesis machinery remains a major bottleneck, which this paper addresses by leveraging [episode]
- ROS Help Desk: GenAI Powered, User-Centric Framework for ROS Error Diagnosis and Debugging — ROS Help Desk provides an accessible interface enabling operators of all expertise levels to proactively detect errors and participate in debugging processes within robotic environments. [episode]
- Subspace Consensus — This paper investigates subspace consensus for matrix-weighted multi-agent networks, which addresses a gap in traditional consensus theory by allowing agents to agree only on specific dimensions of their state vectors while maintaining desired relative configurations in the remai [episode]
- Input-to-state stabilization of linear systems under data-rate constraints — A communication and control strategy is proposed for feedback stabilization of linear systems under data-rate constraints in the presence of completely unknown disturbances, establishing input-to-state stability (ISS) with respect to the disturbance. [episode]
- Limited Preemption of the 3-Phase Task Model using Preemption Thresholds — Phased execution models are employed to manage complexity in modern multi-core platforms, and this research introduces preemption thresholds as a method to limit preemptions in 3-phase task models to minimize local memory usage while maintaining schedulability. [episode]
- Communication-Aware Robot Execution for Cloud Inference under Spatially Heterogeneous Connectivity — Cloud-hosted foundation models enable robots to use semantic reasoning beyond onboard computational limits, but this execution becomes fragile under spatially heterogeneous connectivity because the current primitive determines when the next result is needed, while wireless enviro [episode]
- From Inference to Control: Structure-Guided Control of Hypergraph Dynamics — Controllability determines whether a system’s state can be guided toward any desired configuration, making it a fundamental prerequisite for designing effective control strategies. [episode]
- Communication-Aware Synthesis of Safety Controller for Networked Control Systems — Networked control systems (NCS) are widely used in safety-critical applications, but they are often analyzed under the assumption of ideal communication channels. [episode]
- A Biomimetic Myoelectric Tentacle Prosthesis with Sensorless Object Detection and Vibrotactile Feedback — This research presents the design and evaluation of a myoelectric tentacle-shaped prosthesis integrating electromyographic (EMG) control, sensorless object detection, and vibrotactile feedback. [episode]
- A Learning-Free Characterization Framework for the Resilience and Sensitivity of Polyurethane Vision-Based Tactile Sensors — Vision-based tactile sensors (VBTSs) are promising for robots but existing silicone gels suffer from durability issues, prompting this study to characterize polyurethane rubber as a more resilient alternative, revealing a critical tradeoff between sensor resilience and sensitivit [episode]
- ManiPhysicsBench: Physics-Based Assessment of Object Preservation in VLA Manipulation —
- LOCUS: Landmark-Oriented Container Discrimination Using Spatial Graphs —
- SARI: Phase-Split Sim-Real Co-Training for Contact-Rich Manipulation —
- Learning Reflexive Behavior for Contact-Rich Manipulation —
- Register-Routed Delayed Fusion: Rewiring Shortcut-Prone Observation Fusion in Visuomotor Imitation —
- On BESS-Backed Trading on the Continuous Intraday Electricity Market —
- FastOPD: On-Policy Distillation for Lightweight VLA Deployment —
- PointWAM: 3D World Action Modeling for Dexterous Robotic Manipulation —
- Permutation Robustness Is Not Enough: Action Collapse in Multi-Agent Transformer Policies —
- SceneFactory-3D: Lifting 2D Traffic Scenes into 3D Physical Counterfactuals for Scalable Physically Grounded Safety Evaluation —
- Model-Based Disturbance Rejection via Sliding-Mode Control: Continuous-Time Formulation and Proper Implicit Discretization —
- MixVLA: Adaptive Mixing of Non-Invariant Information for Generalizable Vision-Language-Action Models —
- Systematic SDP Search for Decentralized Stability Certificates —
- TwinJEPA: Action-Preferred Predictive Representations for Goal-Conditioned Control —
- FARM: Fundamental Agentic Reward Model For Multi-task Wireless Network Optimization —
- Memory-Dependent Interval Markov Chain Abstractions of Stochastic Dynamics —
- Profile-Aware Trustworthy Recipe Generation with Planner-Critic Agentic Remediation —
- From Language Priors to Field Adaptation: Preference Learning for Traversability Estimation —
- On Representational Alignment among Embodied Agents —
- Subject-Specific Predictive Musculoskeletal Simulations of Lower-Limb Exoskeleton Assistance: Metabolic and Biomechanical Effects of Joint Assistance Strategies —
- An Analytical Review of Model Order Reduction Methodologies of Converter-Dominated Power Systems for Stability Studies —
- Safe Streaming Flow Planning by Aligning Sampling Dynamics with Execution Dynamics —
- A New Index for Quantifying Stability Adaptation Effort in Power Electronics-Dominated Power Systems Based on Frequency-Domain Impedance Identification —
- SIS Epidemic Containment under Game-theoretic Rational Learning —
- Beyond Reward Hacking: Proxy Divergence Across Four Layers of a Staged Humanoid Learning Pipeline —
- How suboptimal is my stochastic network controller allowed to be? Completion certificates with application to power grids hosting AI data centers —
- DexJoCo-X: Benchmarking Action Representations for Multi-Hand Dexterous Manipulation —
- HexVIO: Towards All-Day Stereo-Inertial Tracking Through Commodity DSPs —
- Equivariant Visual-Tactile Diffusion Policy for Contact-Rich Manipulation —
- KungfuAthleteBot: learning high-dynamic humanoid motion from video with unified robust recovery —
- I2CD: Direct Image-to-Convex Decomposition for Simulation-Ready Collision Geometry —
- MobiAgent: Dual-Loop Recursive Policy Self-Improvement for Long-Horizon Mobile Manipulation —
- Learning in Inverse Games: Tractable Training with Probabilistic Guarantees —
- Detect and Suppress: A Mechanistic Defense against Adversarial Patches in VLA Models —
- MagLearn 2: High-Fidelity, Saturation-Aware, and History-Efficient Sequence-to-sequence Modeling of Transient B-H Behaviour Under PWM Excitation —
- Hierarchical Control via MPC-RL for Multi-Timescale Battery Systems —
- XGenAct: Geometry-Enhanced World Action Models through Cross-Task Generation —
- DR-IPC: Disturbance-Resilient Integrated Planning and Control for LiDAR-Based Quadrotor Navigation —
- Autonomous Robotic Navigation for Endovascular Brain-Computer Interface Access —
- RATE: Risk-Aware Tactile Encoding for Contact-rich Robotic Manipulation —
- A Unified Framework for Empowerment and Predictive Control —
- AVL-JEPA: Preventing Causal Dynamics Information Collapse In Joint Embedding Predictive Architecture World Models —
- World Action Learning via Interaction-Centric Spectral Latent Guidance —
- Bridging Frontier Reasoning and Robot Execution: From Autonomous Demonstration Generation to Dense Language Supervision —
- LLA-MPC on Embedded Hardware: Rapid Adaptive Control with Thousands of Parallel Models —
- CORNAV: Construction-Aware Reasoning for Robot Navigation on Active Worksites —
- Risk-Aware Input-Constrained Safe Intercept Guidance Against Multiple Moving Defenders —
- EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras —
- Awomo-SimDataEngine: Agentic Simulation-ReadyWorld Generation —
- World-Calibrated Proposal-to-Action Flow for Vision-Language-Action Models —
- SoTa: Soft Tactile Skins for Dexterous Manipulation —
- NEEDLEWORK: Offline Rewriting of Robot Data with Verified Local Stitches —
- Filter-Aware Fine-Tuning for Safe Humanoid Whole-Body Tracking —
- SocialVLA: A Social Perception Gateway for Human-Reaction-Based Failure Detection and Recovery in VLA Manipulation —
- Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation —
- Keep the Effect, Drop the Actor: Programmable Effect-to-Execution World-Action Models —
- Towards a Theory of Quantitative Vulnerability Analysis for Engineering Systems —
- Autonomous mobile robot operations logistics: a dataset of jobs, dispatch events and robot states —
- OpenRUA: Robot-Use Agents Are Zero-Shot Visuomotor Policies —
- SD-DPC: Sparse Dictionary Differentiable Predictive Control —
- RUL-Aware RRT*: Degradation-Balanced Motion Planning for Robotic Manipulators —
- Modular Model Order Reduction for Inverter-Based Power Systems with Stability and Accuracy Guarantees —
- CriticHack: Evaluating Visual Rewards Under Robot Policy Optimization —
- Real-time Event-camera Stereo Visual Odometry via Keytime Gaussian Process Regression —
- A Passive AI System for Verifying Physical State on Automated Liquid Handlers —
- Generation and Transmission Expansion Planning with BESS-Based Virtual Transmission Lines —
- DeltaWorld: Physically Consistent Interactive World Simulators via Action-Conditioned Latent Increment Learning —
- Who Went Where When on the Lunar Surface: Forensic Trajectory Analysis to Identify Byzantine Rovers —
- GeoScaffold: Learning Compact Geometric Latents via Reconstruction for Efficient Vision-Language Navigation —
- AdaTempo: Learning Shared Relative Tempo from Demonstrations for Faster Robot Manipulation —
- RoboChemGym: A Protocol-Driven Generative Simulation Framework for Long-Horizon Chemical Manipulation —
- RoboBridge: A Self-Evolving Embodied Agent Framework for Sim-to-Real Transfer —
- Performance Limits and Tradeoffs of Power Systems Synchronization Recovery in Complex-Frequency Representation —
- Around the World: Unified Learned Locomotion on a 270 g Continuous-Rotation Quadruped —
- CSIR: Contextually and Socially Informed Robots for Efficient Person Goal Navigation —
- Proprioceptive Sketches as Long-Horizon Intent for Generative Action Policies —
- Localized Conformal Safety Monitoring with Vision-Language Models for Autonomous Driving —
- Mission-Centric Requirements Analysis of Model Predictive Control in Spacecraft Rendezvous and Proximity Operations —
- SimpleTouch: Can Vision-Language-Action Models Master Contact-Rich Manipulation Without Tactile Policy Pretraining? —
- Skill2Real: Agentic Skill Learning for Zero-Shot Sim-to-Real Robot Manipulation —
Important terms
- World-to-Wrist
- This technique focuses on predicting what the robot's end effector needs to do next based on its current situation, helping it plan tasks better in real-world manipulation.
- World-Calibrated Proposal-to-Action Flow
- This work calibrates the path between abstract planning and concrete robot actions, making sure the high-level plans can actually be executed by a physical robot.
- Koopman operator based model predictive control
- This method uses a mathematical model to predict how deformable objects, like cloth, will behave under control inputs in real time for dynamic folding.
- Skill2Real
- This research aims to learn complex robotic skills directly from simulation without needing extensive prior training data, bridging the gap to real-world deployment.