Demonstration-Free Success-Probability Reward Learning for Generalist Robot Policies
cs.RO, cs.AI, cs.LG
Submitted: 2026-09-27
Updated: 2026-09-27
Project page: https://duowuyms.github.io/evta0
Terminology
Sources
- Qwen3-VL Technical Report
- EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- WorldVLA: Towards Autoregressive Action World Model
- LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies
- $\pi_\texttt{RL}$: Online RL Fine-tuning for Flow-based Vision-Language-Action Models
- SARM: Stage-Aware Reward Modeling for Long Horizon Robot Manipulation
- TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics
- ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy
- Vision Language Action Models in Robotic Manipulation: A Systematic Review
- Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model Learning
- LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- LoRA: Low-Rank Adaptation of Large Language Models
- RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance
- NORA-1.5: A Vision-Language-Action Model Trained using World Model- and Action-based Preference Rewards
- $\pi^{*}_{0.6}$: a VLA That Learns From Experience
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- OpenVLA: An Open-Source Vision-Language-Action Model
- Offline Reinforcement Learning with Implicit Q-Learning
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving