AnyStep-WAM: Budget-Aligned Distillation and Adaptive Inference for World Action Models
cs.RO
Submitted: 2026-09-27
Updated: 2026-09-30
Terminology
Sources
- Flash-WAM: Modality-Aware Distillation for World Action Models
- Adaptive-WAM: Quality-Guided Early-Exit Planning from Intermediate Video-Diffusion Features
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- Let It Be Simple: One-Step Action Generation for Vision-Language-Action Models
- Improved Mean Flows: On the Challenges of Fastforward Generative Models
- AdaVLA: Adaptive Step Flow Matching for Training-free Acceleration of Vision-Language-Action Models
- LoRA: Low-Rank Adaptation of Large Language Models
- DSA: Dynamic Step Allocation for Fast Autoregressive Video Generation
- ELASTIC: Efficiently Learning to Adaptively Scale Test-Time Compute for Generative Control Policies
- Causal World Modeling for Robot Control
- Flow Matching for Generative Modeling
- Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference
- The Geometry of Flow-Matching Uncertainty: A Cost-free Uncertainty Proxy and Its Application in Flow-based VLA Failure Detection
- Elastic Queries Reinforcement Learning: Self-Aware Policy Execution for VLA Models
- ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training
- Stable Velocity: A Variance Perspective on Flow Matching
- D3P: Dynamic Denoising Diffusion Policy via Reinforcement Learning
- Fast-WAM: Do World Action Models Need Test-time Future Imagination?
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving