Quantile Head for Vision-Language-Action Models
cs.RO, cs.CV
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/xwangrs/Quantile-Head-for-VLA
Terminology
Sources
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- ReconVLA: An Uncertainty-Guided and Failure-Aware Vision-Language-Action Framework for Robotic Control
- Let It Be Simple: One-Step Action Generation for Vision-Language-Action Models
- Geometric Action Model for Robot Policy Learning
- SUREFlow: State-space Uncertainty-aware REsidual Flow Matching for Robust Robot Manipulation
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- MolmoAct: Action Reasoning Models that can Reason in Space
- SnapFlow: One-Step Action Generation for Flow-Matching VLAs via Progressive Self-Distillation
- SimVLA: A Simple VLA Baseline for Robotic Manipulation
- InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization
- FAST: Efficient Action Tokenization for Vision-Language-Action Models
- Learning Policies through Quantile Regression
- MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models
- VLANeXt Family: A Systematic Study of VLA Models from Core Recipes to Emerging Paradigms
- Continuous Reasoning for Vision-Language-Action
- World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis
- CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention
- LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving