Learning Process Rewards via Success Visitation Matching for Efficient RL
Raymond Tsao, Andrew Wagenmaker, Sergey Levine
cs.LG, cs.AI, cs.RO, stat.ML
Submitted: 2026-06-22
Project page: https://success-visitation-matching.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Video-Language Critic: Transferable Reward Functions for Language-Conditioned Robotics
- $\pi^{*}_{0.6}$: a VLA That Learns From Experience
- Update-Free On-Policy Steering via Verifiers
- Efficient Online Reinforcement Learning with Offline Data
- Vision-Language Models as a Source of Rewards
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- Exploration by Random Network Distillation
- Learning Generalizable Robotic Reward Functions from "In-The-Wild" Human Videos
- SARM: Stage-Aware Reward Modeling for Long Horizon Robot Manipulation
- ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy
- Process Reward Models for LLM Agents: Practical Framework and Directions
- Process Reinforcement through Implicit Rewards
- EXPO: Stable Reinforcement Learning with Expressive Policies
- What Matters for Batch Online Reinforcement Learning in Robotics?
- Vision-Language Models as Success Detectors
- A Connection between Generative Adversarial Networks, Inverse Reinforcement Learning, and Energy-Based Models
- Learning Robust Rewards with Adversarial Inverse Reinforcement Learning
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Improving Vision-Language-Action Model with Online Reinforcement Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks