ATI-VLA: Action-Centric Predictive Vision-Language-Action Models via Actionable Alignment Then Adaptive Injection
cs.CV
Submitted: 2026-10-01
Updated: 2026-10-01
Project page: https://jiutian-vl.github.io/ATI-VLA-page
Terminology
Sources
- Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
- WorldVLA: Towards Autoregressive Action World Model
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- ActionCodec: What Makes for Good Action Tokenizers
- The Llama 3 Herd of Models
- Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- OpenVLA: An Open-Source Vision-Language-Action Model
- Behavior Generation with Latent Actions
- STAR: Learning Diverse Robot Skill Abstractions through Rotation-Augmented Vector Quantization
- Unified Video Action Model
- ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation
- CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification
- Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation
- OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models