DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation".
Dev: Manipulating dynamic objects remains an open challenge for Vision-Language-Action (VLA) models, which struggle in dynamic scenarios requiring rapid perception, temporal anticipation, and continuous control.
Rosa: First, who's behind it and why it matters.
Paper summary: Rosa: Looking at the conclusion of DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation, it seems the authors are really focused on how their specific combination of components solves the core issue of dynamic object manipulation. They emphasize that this framework offers better performance in speed and accuracy compared to prior methods when dealing with moving objects.
Dev: I agree, and the authors are quite specific about how their design choices—the compact VLA, continuous inference, and latent-aware action streaming—work together to mitigate the latency problems inherent in dynamic environments.
Taro: The implication here is that we can move closer to systems that can handle complex physical interactions in real-time, even when those objects are changing their motion during the task.
Rosa: So, in simpler terms for our listeners, this paper is about creating an AI system that can manipulate physical objects dynamically without getting caught between what it sees and what it does.
Dev: That makes sense; they are focusing on making the loop rate fast enough to keep up with fast-moving things, and their methods aim to maintain that alignment even when the model has some processing lag.
Taro: The future work mentioned points toward extending this capability to multi-stage tasks involving persistent object motion and integrating planning and memory while keeping things real-time.
Rosa: It's exciting because it shows how we can build models that are efficient enough for practical use, moving beyond just static manipulation scenarios.
Conclusion: Rosa: So, we've been talking about this DynamicVLA framework that handles dynamic object manipulation, and now we need to look at what this whole thing is really trying to achieve with its title and authors.
Dev: Yeah, Rosa, it's crucial to understand that the authors are trying to solve the problem of making AI systems capable of interacting with physical objects in a changing environment. That’s the core focus here.
Taro: I think what they’re aiming for is a system that doesn't just follow pre-planned paths but can actually adapt its actions on the fly when things get messy, which is pretty important for real autonomy.
Rosa: Exactly, and their name tells us immediately that this AI model isn't designed for static objects sitting on a table; it’s built to handle movement and change in real-time.
Dev: From an engineering standpoint, the implication is that we might see robotic systems operating outdoors or in complex indoor settings where things are constantly shifting, rather than just controlled lab environments.
Taro: That would be huge for real-world deployment because it means the autonomy isn't crippled by simple movements; it can react to unexpected changes.
Rosa: So, when you look at the authors and their approach, it seems they’re focusing heavily on closing that gap between what the AI perceives and what it actually executes in motion.
Dev: That execution gap is where we see the real technical challenge; if the loop rate isn't fast enough or the perception lags, even a good model fails in a dynamic scenario.
Taro: And that’s why their design choices about continuous inference and action streaming are so compelling; they’re specifically targeting those timing issues you mentioned, Dev.
Rosa: It really paints a picture of an AI that has learned not just *what* to do, but *how* to do it fluidly when the world keeps moving around it.
Dev: Precisely, and that fluidity is what makes the system potentially useful beyond simple demonstrations; we're looking at systems that can manage continuous tasks.
Taro: It suggests a future where autonomous agents can handle more complex, unpredictable physical interactions without constant human supervision during the execution phase.
Rosa: That’s a big picture shift, moving from pre-programmed actions to truly adaptive manipulation in dynamic settings.
Nanyang Technological University
cs.RO, cs.CV
Submitted: 2026-01-29
Updated: 2026-10-01
Comments: NeurIPS 2026. Project Page: https://www.infinitescript.com/project/dynamic-vla/
Code: https://github.com/kakaobrain/coyo-dataset
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 80/100
The gist: Manipulating dynamic objects remains an open challenge for Vision-Language-Action (VLA) models, which struggle in dynamic scenarios requiring rapid perception, temporal anticipation, and continuous
Key concepts
- Continuous Inference
- This design allows the model to reason about and execute actions simultaneously rather than waiting for one sequence to finish. By overlapping prediction and execution, it reduces latency significantly, allowing the model to adapt quickly when objects are moving, avoiding delays between steps in a task.
- Latent-aware Action Streaming
- This mechanism fixes the gap between seeing an object and acting on it by enforcing temporal consistency. It discards old actions that are outdated relative to the current observation and prioritizes newer actions when sequences overlap, ensuring the model always reacts to the most recent environment state.
- FastViT Vision Encoder
- This is a convolutional vision encoder used in DynamicVLA for efficient visual processing. It compresses visual information spatially while preserving structural details, which helps the model process complex, multi-frame visual inputs quickly without increasing computational load quadratically.
- DOM Benchmark
- The Dynamic Object Manipulation benchmark provides a large dataset of synthetic and real-world dynamic manipulation episodes. It tests models on interaction (reactivity), perception (motion understanding), and generalization to ensure the model can handle various complex, moving object scenarios.
Terminology
Summary
Manipulating dynamic objects remains an open challenge for Vision-Language-Action (VLA) models, which struggle in dynamic scenarios requiring rapid perception, temporal anticipation, and continuous control. DynamicVLA is presented as a framework that integrates temporal reasoning and closed-loop adaptation through three key designs to address the perception–execution gap inherent in existing VLA models.
The gist: DynamicVLA is a framework for dynamic object manipulation that integrates temporal reasoning and closed-loop adaptation through three key designs: 1) a compact 0.4B VLA using a convolutional vision encoder for spatially efficient, structurally faithful encoding, enabling fast multimodal inference; 2) Continuous Inference, enabling overlapping reasoning and execution for lower latency and timely adaptation to object motion; and 3) Latent-aware Action Streaming, which bridges the perception–execution gap by enforcing temporally aligned action execution.
The DynamicVLA Architecture
The framework is built around a compact 0.4B VLA model designed for fast and spatially efficient multimodal reasoning. It utilizes a convolutional vision encoder, FastViT, for efficient spatial compression and stronger structural preservation,
which avoids quadratic token growth when processing multiframe visual inputs. The language backbone is truncated to its first 16 layers of SmolLM2-360M to significantly reduce inference latency with minimal impact on multimodal reasoning. Action generation is handled by a Diffusion-Based Action Expert,
instantiated as a conditional Flow Matching Transformer, trained using the objective where it learns to match the denoising vector field.
Continuous Inference
This design addresses inter-chunk waiting by enabling overlapping reasoning and execution for lower latency and timely adaptation to object motion. Unlike prior models where inference is triggered only after a sequence is fully executed, Continuous Inference overlaps prediction and action execution,
allowing actions from the current sequence to be executed while the next one is being inferred. This eliminates inter-chunk waiting
under dynamic object motion.
Latent-aware Action Streaming
This mechanism restores temporal alignment by enforcing temporally consistent control despite inference delay, which manifests as a Perception–Execute Gap.
It resolves this through an explicit execution strategy: actions in the current sequence corresponding to timesteps earlier than the observation time are discarded as outdated,
and execution proceeds with the subsequence of newer actions. Furthermore, where action chunks overlap, actions from the newer sequence are prioritized, overwriting those from At,
ensuring adaptation to the most recent environment state.
The Dynamic Object Manipulation (DOM) Benchmark
To fill the missing foundation of dynamic manipulation data, a large-scale benchmark was introduced. DOM features 200K synthetic episodes across 2.8K scenes and 206 objects, and enables fast collection of 2K real-world episodes without teleoperation. The benchmark organizes evaluation along three principal dimensions: Interaction (Closed-loop reactivity, Dynamic adaptation, Long-horizon sequencing), Perception (Visual understanding, Spatial reasoning, Motion perception), and Generalization (Visual generalization to unseen objects/scenes and Motion generalization).
Experimental Evaluation
DynamicVLA was evaluated across dynamic manipulation tasks using the DOM benchmark. Extensive evaluations demonstrated remarkable improvements in response speed, perception, and generalization,
outperforming baselines such as SmolVLA by significant margins across all interaction settings. Ablation studies confirmed that the 360M model strikes the optimal balance between inference efficiency and model capacity, yielding the highest overall performance in dynamic object manipulation. The dominant failure mode identified is temporal misalignment between observation and action execution.
Ablation Findings
Ablation studies revealed that increasing LLM depth leads to a modest increase in latency, which can be amortized by Continuous Inference and Latent-aware Action Streaming. Aggressively truncating the backbone improves inference speed but degrades success rate due to reduced reasoning capacity. The integration of CI and LAAS into existing models, such as SmolVLA, shows consistent performance improvements under moderate inference latency, whereas larger backbones in models like π0.5 limit the effectiveness of these mechanisms due to high underlying latency.
Future Work Directions
Future work should focus on extending dynamic manipulation to multi-stage tasks with persistent object motion,
integrating planning and memory while maintaining real-time constraints. Additionally, extending the model to handle non-rigid or fluid dynamics
remains an open challenge for both the VLA models and data pipelines.
Acknowledgments
The work was supported by the Singapore Ministry of Education under its Academic Research Fund Tier 2 (MOE-T2EP20221-0012, MOE-T2EP2023-0001), and by cash and in-kind contributions from NTU S-Lab and industry partner(s).
References
[1] DeepSeek AI. DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning. arXiv, 2501.12948, 2025.
Improvements for AI systems
Based on the DynamicVLA paper, here are specific improvements that can be made to AI systems, along with what those improved systems will be able to do:
-
Improve Real-Time Responsiveness and Stability in Dynamic Environments by Integrating Latency-Aware Action Streaming (LAAS).
-
Improve Perception-Execution Alignment by Implementing Continuous Inference (CI).
-
Enhance Generalization to Unseen Objects, Scenes, and Motion Regimes by Utilizing the Dynamic Object Manipulation (DOM) Benchmark for Training.
-
Increase Model Efficiency and Throughput while Maintaining High Performance by Employing a Compact 0.4B VLA Architecture with FastViT Vision Encoder.
-
Enable Robust Closed-Loop Reactivity to Fast-Moving Objects and Abrupt Changes in Motion through Temporal Reasoning Integration (e.g., using the specific temporal context window of two frames: 3 and 2).
-
Achieve Superior Spatial Reasoning and Motion Perception by Training Models on Multi-faceted DOM Sub-tasks (Interaction, Perception, Generalization).
-
Develop Autonomous Data Collection Pipelines for Dynamic Manipulation by Utilizing the Automated Simulation and Real-world Data Collection Framework (using Isaac Sim and a
real-world simulator
pipeline) to generate 200K synthetic and 2K real-world episodes without reliance on slow teleoperation. -
Create More Efficient, Adaptable VLA Models by Performing Ablation Studies that identify the optimal balance between LLM Backbone Capacity (e.g., 16 layers vs. 32 layers) and Inference Latency, leading to smaller, faster models with high success rates.
This improved AI system can perform the following specific actions:
-
Navigate and interact with objects in real-time environments where objects are moving at varying speeds (e.g., grasping a rolling coffee can or placing an object onto a frisbee) while maintaining stable, closed-loop control that adjusts immediately to sudden disturbances.
-
Execute complex, long-horizon manipulation sequences (e.g., gathering multiple items into a container) without stalling due to inference delays between action steps.
-
Identify and grasp objects in cluttered or visually similar scenes (e.g., picking up a specific tennis ball among several others) by accurately perceiving and grounding visual cues in dynamic, changing environments.
-
Generalize its manipulation skills to novel objects with unseen appearances, novel 3D scenes, and unpredictable motion patterns (e.g., grasping an apple with an irregular trajectory), demonstrating robustness beyond the training distribution.
-
Operate reliably in complex physical settings (like real-world Franka or PiPER arms) where human reaction times are too slow to track fast-moving objects, effectively acting as a high-speed, closed-loop control system.
-
Perform precise spatial reasoning tasks, such as placing an object relative to a specific location or within a dynamically defined boundary (e.g., placing a bottle onto a blue tape) despite continuous object motion and visual noise.
-
Self-sustain learning by autonomously generating large, diverse datasets for dynamic manipulation through automated simulation and real-world sensing, reducing the dependency on costly human teleoperation for data labeling.
Sources
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- Achieving Human Level Competitive Robot Table Tennis
- VIMA: General Robot Manipulation with Multimodal Prompts
- OpenVLA: An Open-Source Vision-Language-Action Model
- BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation
- Sim-to-Real Dynamic Object Manipulation on Conveyor Systems via Optimization Path Shaping
- Running VLAs at Real-time Speed
- SmolVLM: Redefining small and efficient multimodal models
- Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning
- GPT-4 Technical Report
- SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
- VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference
- The Llama 3 Herd of Models
- Qwen2.5-1M Technical Report
- Qwen2.5-VL Technical Report
- VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
- 3D Scene Generation: A Survey
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving