Robotics papers — 2026-09-17

Today we are diving into how we can make these vision language action models run much faster in real-world scenarios. This is crucial because inference speed directly impacts how smoothly a robot moves and responds to things. The main focus is on rMuscle, a new framework inspired by human muscle memory that uses a dual-phase cache to reuse visual tokens and neuron activation patterns. This has shown significant speedups across different tasks, meaning we are trying to make the complex reasoning of these models more efficient for practical robotic deployment.

Beyond this optimization, we are also looking at how to improve the underlying structure of generalized morphology control using RecMorph. This architecture uses recurrent sequences to handle cross-limb communication and transformation. It has shown strong performance on various tasks even when generalizing to larger bodies, suggesting a path toward more flexible and efficient control policies for robots with varying physical structures.

Finally, we are exploring how to make these models better at perceiving the world dynamically through ActiveScale. This framework augments VLA models with historical video observations and explicit camera pose supervision to allow them to reason across changing viewpoints. Fixed viewpoints often hide necessary information during manipulation tasks, making this dynamic perception important for real-world use.

The most critical work right now is building smart and adaptive agents for active sensing because current deep learning methods struggle with environmental dynamics. This means they can't just adapt to data drift; they need to understand the environment itself. This paper introduces an agentic system that uses active inference to allow on-device perception and planning, enabling real-time action in environments with a very small memory footprint of about three hundred megabytes.

This system is demonstrated by a saccade agent controlling an IoT camera on an NVIDIA Jetson device, simulating human eye movements for surveillance or robotics. This concept builds upon the idea of using large foundation models for embodied AI, specifically looking at how different approaches like Robot Foundation Models and Vision-Language Action models can work together.

Another significant piece explores how to diagnose and direct adaptation in vision-language-action models by figuring out which parts of the model need fine-tuning based on what kind of shift is happening. This suggests a way to make adaptation much more efficient than uniform fine-tuning. A diagnostic pipeline ranks regions within the model based on cost, showing that this structured approach can match full fine-tuning with very few trainable parameters.

Furthermore, there is work focused on improving long-horizon planning in vision-language-action models by adding an explicit language memory module to maintain temporal consistency during complex tasks. This architecture decouples high-level semantic reasoning from low-level control. The high-level model can recursively update its instructions using past memory as context, which showed that this explicit memory significantly boosts the success rate and robustness of VLA models on difficult, long-term tasks while also offering a clear explanation of how decisions are made.

The work that matters most is the CALOS safety layer because it directly addresses the fundamental problem of guaranteeing safety in deep reinforcement learning policies for quadrotors during training and deployment. This layer works by formulating attitude constraints as a single quadratic program whose solution provides the minimum-norm correction to the policy's nominal torque output. It allows it to enforce four tilt-angle inequalities while maintaining a computational cost low enough for real-time use across thousands of parallel simulations.

This approach significantly reduces lateral tracking error, showing a reduction of fifty-five to sixty percent compared to an unconstrained proximal policy optimization baseline. Importantly, it achieves zero attitude constraint violations on the training trajectory, meaning by restricting exploration to safe state regions, the safety layer simultaneously accelerates training convergence and improves data efficiency without sacrificing policy quality.

Another important area is learning contact dynamics through touching using action-conditional graph neural networks to predict end effector motion and reaction forces in contact-rich manipulation scenarios. This model represents the robot and environment as interacting meshes in a graph structure, predicting object-level pose updates directly while deriving reaction torque from a per-vertex force field. In simulation, this model successfully transfers to peg insertion with unseen concave geometry, reaching up to a ninety-eight percent success rate when used by an MPC agent.

This learning of contact dynamics is further validated because the model outperforms the system-identified muJoCo model in real-world tests by forty-five percent in position and seventy-four percent in force and torque error. This suggests that this physics-based model provides a much more accurate representation of physical interaction than traditional simulators.

Moving toward perception, there is work on task-aware evaluation of gan-based synthetic sonar data for robotic perception. This addresses the gap between pixel fidelity and actual downstream performance. This research found that while conventional image-fidelity metrics like SSIM and PSNR can be misleading, PatchGAN configurations often yield stronger object detection results even when their pixel scores are not the highest.

This points toward needing task-oriented evaluation rather than just looking at image similarity scores for synthetic sensor data. Finally, there is a focus on making vision language action models reliable in real-time through a framework called Real-Time EXPO-FT. This decouples slow action generation from fast reactive edits. This method allows a large pretrained vision language model to propose action chunks while a lightweight edit policy performs rapid decision-making based on the latest state changes.

This technique has been shown to improve average policy performance from forty-two percent to ninety-seven percent across several dynamic real-world tasks when using online robot data. The Visual Perception Engine work matters because it directly tackles the computational bottleneck when running multiple vision models on limited robotic hardware, which is crucial for real-time operation.

The VPEngine framework introduces a shared foundation model backbone that extracts image representations once. This allows several task-specific heads to run in parallel without redundant GPU memory transfers. This design cuts down on the inefficiency seen when deploying traditional sequential models and enables dynamic task prioritization based on what the application needs most at any given moment.

Building upon this efficiency gain, Mem2Ego aims to improve how vision-language models navigate complex spaces by bridging global context with local perception. This method involves adaptively retrieving relevant cues from a global memory module and merging them with the agent's immediate visual inputs. This is designed to boost spatial reasoning in long-horizon tasks, though it still faces the challenge of making optimal decisions when relying only on a first-person perspective.

Finally, the Mixed-Integer Nonlinear Differentiable Predictive Control work shows how advanced control methods can handle complex physical systems with high precision. This method extends MI-DPC to manage multi-modal discrete decisions and non-convex dynamics found in underground pumped hydro energy storage systems. The framework achieves only a one point six percent suboptimality compared to a standard baseline, while simultaneously providing five orders of magnitude speedup in the time required for online scheduling.

Today's papers

The papers

Important terms

rMuscle
A new framework inspired by human muscle memory that uses a dual-phase cache to reuse visual tokens and neuron activation patterns, significantly speeding up vision language action models for real-world robot deployment.
RecMorph
An architecture for generalized morphology control that uses recurrent sequences to handle communication between limbs and transformations, allowing robots with varying physical structures to have more flexible control policies.
ActiveScale
A framework that augments VLA models with historical video and camera pose supervision to enable dynamic perception, helping models reason across changing viewpoints for real-world manipulation.
CALOS safety layer
A safety layer that guarantees quadrotor safety by formulating attitude constraints as a quadratic program, enforcing tilt angle inequalities while maintaining low computational cost for real-time use.