Robotics papers — 2026-09-17
Today we are diving into how we can make these vision language action models run much faster in real-world scenarios. This is crucial because inference speed directly impacts how smoothly a robot moves and responds to things. The main focus is on rMuscle, a new framework inspired by human muscle memory that uses a dual-phase cache to reuse visual tokens and neuron activation patterns. This has shown significant speedups across different tasks, meaning we are trying to make the complex reasoning of these models more efficient for practical robotic deployment.
Beyond this optimization, we are also looking at how to improve the underlying structure of generalized morphology control using RecMorph. This architecture uses recurrent sequences to handle cross-limb communication and transformation. It has shown strong performance on various tasks even when generalizing to larger bodies, suggesting a path toward more flexible and efficient control policies for robots with varying physical structures.
Finally, we are exploring how to make these models better at perceiving the world dynamically through ActiveScale. This framework augments VLA models with historical video observations and explicit camera pose supervision to allow them to reason across changing viewpoints. Fixed viewpoints often hide necessary information during manipulation tasks, making this dynamic perception important for real-world use.
The most critical work right now is building smart and adaptive agents for active sensing because current deep learning methods struggle with environmental dynamics. This means they can't just adapt to data drift; they need to understand the environment itself. This paper introduces an agentic system that uses active inference to allow on-device perception and planning, enabling real-time action in environments with a very small memory footprint of about three hundred megabytes.
This system is demonstrated by a saccade agent controlling an IoT camera on an NVIDIA Jetson device, simulating human eye movements for surveillance or robotics. This concept builds upon the idea of using large foundation models for embodied AI, specifically looking at how different approaches like Robot Foundation Models and Vision-Language Action models can work together.
Another significant piece explores how to diagnose and direct adaptation in vision-language-action models by figuring out which parts of the model need fine-tuning based on what kind of shift is happening. This suggests a way to make adaptation much more efficient than uniform fine-tuning. A diagnostic pipeline ranks regions within the model based on cost, showing that this structured approach can match full fine-tuning with very few trainable parameters.
Furthermore, there is work focused on improving long-horizon planning in vision-language-action models by adding an explicit language memory module to maintain temporal consistency during complex tasks. This architecture decouples high-level semantic reasoning from low-level control. The high-level model can recursively update its instructions using past memory as context, which showed that this explicit memory significantly boosts the success rate and robustness of VLA models on difficult, long-term tasks while also offering a clear explanation of how decisions are made.
The work that matters most is the CALOS safety layer because it directly addresses the fundamental problem of guaranteeing safety in deep reinforcement learning policies for quadrotors during training and deployment. This layer works by formulating attitude constraints as a single quadratic program whose solution provides the minimum-norm correction to the policy's nominal torque output. It allows it to enforce four tilt-angle inequalities while maintaining a computational cost low enough for real-time use across thousands of parallel simulations.
This approach significantly reduces lateral tracking error, showing a reduction of fifty-five to sixty percent compared to an unconstrained proximal policy optimization baseline. Importantly, it achieves zero attitude constraint violations on the training trajectory, meaning by restricting exploration to safe state regions, the safety layer simultaneously accelerates training convergence and improves data efficiency without sacrificing policy quality.
Another important area is learning contact dynamics through touching using action-conditional graph neural networks to predict end effector motion and reaction forces in contact-rich manipulation scenarios. This model represents the robot and environment as interacting meshes in a graph structure, predicting object-level pose updates directly while deriving reaction torque from a per-vertex force field. In simulation, this model successfully transfers to peg insertion with unseen concave geometry, reaching up to a ninety-eight percent success rate when used by an MPC agent.
This learning of contact dynamics is further validated because the model outperforms the system-identified muJoCo model in real-world tests by forty-five percent in position and seventy-four percent in force and torque error. This suggests that this physics-based model provides a much more accurate representation of physical interaction than traditional simulators.
Moving toward perception, there is work on task-aware evaluation of gan-based synthetic sonar data for robotic perception. This addresses the gap between pixel fidelity and actual downstream performance. This research found that while conventional image-fidelity metrics like SSIM and PSNR can be misleading, PatchGAN configurations often yield stronger object detection results even when their pixel scores are not the highest.
This points toward needing task-oriented evaluation rather than just looking at image similarity scores for synthetic sensor data. Finally, there is a focus on making vision language action models reliable in real-time through a framework called Real-Time EXPO-FT. This decouples slow action generation from fast reactive edits. This method allows a large pretrained vision language model to propose action chunks while a lightweight edit policy performs rapid decision-making based on the latest state changes.
This technique has been shown to improve average policy performance from forty-two percent to ninety-seven percent across several dynamic real-world tasks when using online robot data. The Visual Perception Engine work matters because it directly tackles the computational bottleneck when running multiple vision models on limited robotic hardware, which is crucial for real-time operation.
The VPEngine framework introduces a shared foundation model backbone that extracts image representations once. This allows several task-specific heads to run in parallel without redundant GPU memory transfers. This design cuts down on the inefficiency seen when deploying traditional sequential models and enables dynamic task prioritization based on what the application needs most at any given moment.
Building upon this efficiency gain, Mem2Ego aims to improve how vision-language models navigate complex spaces by bridging global context with local perception. This method involves adaptively retrieving relevant cues from a global memory module and merging them with the agent's immediate visual inputs. This is designed to boost spatial reasoning in long-horizon tasks, though it still faces the challenge of making optimal decisions when relying only on a first-person perspective.
Finally, the Mixed-Integer Nonlinear Differentiable Predictive Control work shows how advanced control methods can handle complex physical systems with high precision. This method extends MI-DPC to manage multi-modal discrete decisions and non-convex dynamics found in underground pumped hydro energy storage systems. The framework achieves only a one point six percent suboptimality compared to a standard baseline, while simultaneously providing five orders of magnitude speedup in the time required for online scheduling.
Today's papers
- rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference This framework uses a dual-phase cache inspired by muscle memory to speed up vision-language-action model inference. [paper]
- RecMorph: Topology-Guided Spatial Recurrence for Generalized Morphology Control This architecture uses recurrent computation to jointly handle cross-limb communication and representation transformation in generalized morphology control. [paper]
- Learning Multi-Humanoid Pickup and Transport via Decentralized Object-Centric Control This approach allows different sized humanoid robots to cooperatively pick up and transport objects using a shared object attachment interface. [paper]
- ActiveScale: Scaling Active Perception for Robots across Model, Data, and Hardware This framework enhances active perception by augmenting VLA models with historical video observations and explicit camera-pose supervision. [paper]
- Characterizing Replay Retention Under Dynamics Shift in Model-Based Reinforcement Learning This study characterizes how replay retention depends on the magnitude of dynamic changes when adapting model-based reinforcement learning policies. [paper]
- ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models This method introduces physical rank consistency to action tokenization to preserve local physical distance rankings during reconstruction. [paper]
- Solving Conic Programs over Sparse Graphs using a Variational Quantum Approach: The Case of the AC Optimal Power Flow This paper proposes a variational quantum approach to solve conic programs arising in physics and engineering problems like optimal power flow. [paper]
- Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation This pipeline uses generated video and audio to derive force profiles, enabling zero-shot manipulation tasks requiring contact forces. [paper]
- Towards smart and adaptive agents for active sensing on edge devices This agentic system incorporates active inference to enable real-time perception and planning on resource-constrained edge devices. [paper]
- Language-Guided Grasping under Partial Observation for Mobile Manipulation in Field Inspection and Maintenance This pipeline uses language guidance to ground object detection into a grasp selection process for legged mobile manipulators under partial observation. [paper]
- PACT-WAM: Predicting Actions and Visual Foresight with Compact Temporal Encoding for Robot Manipulation This model jointly generates an action trajectory and a visual forecast using hierarchical history encoding to improve robot manipulation success rates. [paper]
- Not All Layers Need Tuning: Diagnosing and Directing Adaptation in Vision-Language-Action Models This pipeline diagnoses adaptation costs per network region to direct parameter-efficient fine-tuning of VLA models. [paper]
- HINT-Plan: Human Intention-Aware Robot Task Planning in Context-Rich Environments using Vision Language Models This approach uses VLMs to predict human intentions and integrate them into formal task planning for proactive human-aware robot decision-making. [paper]
- A Comprehensive Review of Generative Physical Artificial Intelligence This survey provides a taxonomy of five approaches for generative physical artificial intelligence systems, covering foundation models and diffusion policies. [paper]
- Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents This framework converts real-world object interactions into simulatable episodic twins using vision-language agents for downstream policy fine-tuning. [paper]
- Explicit Language Memory for Long-Horizon Planning in Vision-Language-Action Models This architecture uses an explicit language memory module to maintain temporal consistency and improve success rates in long horizon VLA tasks. [paper]
- CALOS: Control-Affine Lyapunov On-manifold Safety Layer for Safe Deep Reinforcement Learning for Quadrotors This runtime safety layer enforces attitude constraints on quadrotors using a quadratic program to ensure safety during training and deployment. [paper]
- Learning Contact Dynamics through Touching: Action-conditional Graph Neural Networks for Robotic Peg Insertion This model predicts end effector motion and reaction forces by representing the robot and environment as interacting meshes in a graph structure. [paper]
- Synthetic Electric Vehicle Charging Session Generation Using a Conditional Variational Autoencoder This CVAE generates synthetic EV charging sessions that preserve key statistical properties of real transaction data. [paper]
- Acting in Meters: Learning Metric Interactions for Precise Robotic Manipulation This framework models object and scene interactions at a shared metric scale using Interaction-Centric Tokens and Metric Action Interaction Fields. [paper]
- Reinforcement Learning for Real-Time Vision-Language-Action Policies This framework enables sample-efficient fine-tuning of VLA policies by decoupling slow action generation from fast, reactive action edits in real time. [paper]
- Beyond Pixel Similarity: Task-Aware Evaluation of GAN-Based Synthetic Sonar Data for Robotic Perception This work suggests that task-oriented evaluation is necessary because pixel fidelity metrics alone do not capture the realism relevant to downstream perception tasks. [paper]
- M2Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models This tokenizer uses multiple heads and codebooks to minimize reconstruction error and boost the success rate of VLA models. [paper]
- VLA-ULAP: Interleaving Cloud VLA Calls with Ultra-Lightweight Local Action Prediction at the Edge This framework interleaves remote VLA calls with a lightweight local predictor to reduce inference time and energy consumption on edge devices. [paper]
- Visual Perception Engine: Fast and Flexible Multi-Head Inference for Robotic Vision Tasks This modular framework enables efficient GPU usage by sharing foundation models across multiple task-specific heads for robotic vision tasks. [paper]
- Mem2Ego: Empowering Vision-Language Models with Global-to-Ego Memory for Long-Horizon Embodied Navigation This VLM framework adaptively retrieves global context to enhance spatial reasoning and decision-making in long horizon embodied navigation. [paper]
- Mixed-Integer Nonlinear Differentiable Predictive Control for Underground Pumped Hydro Energy Storage Systems This extension applies differentiable control to model discrete decisions and nonconvex dynamics for optimal scheduling of energy storage systems. [paper]
The papers
- Towards smart and adaptive agents for active sensing on edge devices —
- Mem2Ego: Empowering Vision-Language Models with Global-to-Ego Memory for Long-Horizon Embodied Navigation —
- Visual Perception Engine: Fast and Flexible Multi-Head Inference for Robotic Vision Tasks —
- Solving Conic Programs over Sparse Graphs using a Variational Quantum Approach: The Case of the AC Optimal Power Flow —
- Learning Contact Dynamics through Touching: Action-conditional Graph Neural Networks for Robotic Peg Insertion —
- PACT-WAM: Predicting Actions and Visual Foresight with Compact Temporal Encoding for Robot Manipulation —
- Language-Guided Grasping under Partial Observation for Mobile Manipulation in Field Inspection and Maintenance —
- Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents —
- Explicit Language Memory for Long-Horizon Planning in Vision-Language-Action Models —
- CALOS: Control-Affine Lyapunov On-manifold Safety Layer for Safe Deep Reinforcement Learning for Quadrotors —
- HINT-Plan: Human Intention-Aware Robot Task Planning in Context-Rich Environments using Vision Language Models —
- Synthetic Electric Vehicle Charging Session Generation Using a Conditional Variational Autoencoder —
- Learning Multi-Humanoid Pickup and Transport via Decentralized Object-Centric Control —
- Mixed-Integer Nonlinear Differentiable Predictive Control for Underground Pumped Hydro Energy Storage Systems —
- Not All Layers Need Tuning: Diagnosing and Directing Adaptation in Vision-Language-Action Models —
- Beyond Pixel Similarity: Task-Aware Evaluation of GAN-Based Synthetic Sonar Data for Robotic Perception —
- A Comprehensive Review of Generative Physical Artificial Intelligence —
- Characterizing Replay Retention Under Dynamics Shift in Model-Based Reinforcement Learning —
- Reinforcement Learning for Real-Time Vision-Language-Action Policies —
- Acting in Meters: Learning Metric Interactions for Precise Robotic Manipulation —
- M2Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models —
- RecMorph: Topology-Guided Spatial Recurrence for Generalized Morphology Control —
- ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models —
- ActiveScale: Scaling Active Perception for Robots across Model, Data, and Hardware —
- VLA-ULAP: Interleaving Cloud VLA Calls with Ultra-Lightweight Local Action Prediction at the Edge —
- rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference —
- Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation —
Important terms
- rMuscle
- A new framework inspired by human muscle memory that uses a dual-phase cache to reuse visual tokens and neuron activation patterns, significantly speeding up vision language action models for real-world robot deployment.
- RecMorph
- An architecture for generalized morphology control that uses recurrent sequences to handle communication between limbs and transformations, allowing robots with varying physical structures to have more flexible control policies.
- ActiveScale
- A framework that augments VLA models with historical video and camera pose supervision to enable dynamic perception, helping models reason across changing viewpoints for real-world manipulation.
- CALOS safety layer
- A safety layer that guarantees quadrotor safety by formulating attitude constraints as a quadratic program, enforcing tilt angle inequalities while maintaining low computational cost for real-time use.