AGM: Achievement-Grounded Memory for Closed-Loop Agents with Frozen VLA Policies
cs.RO, cs.AI
Submitted: 2026-08-30
Updated: 2026-09-28
Comments: 19 pages, 9 figures
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- RT-1: Robotics Transformer for Real-World Control at Scale
- RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies
- PaLM-E: An Embodied Multimodal Language Model
- Vision-Language Models as Success Detectors
- AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
- SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation
- DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment
- Inner Monologue: Embodied Reasoning through Planning with Language Models
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos
- OpenVLA: An Open-Source Vision-Language-Action Model
- REFLECT: Summarizing Robot Experiences for Failure Explanation and Correction
- MemER: Scaling Up Memory for Robot Control via Experience Retrieval
- Octo: An Open-Source Generalist Robot Policy
- RoboTAP: Tracking Arbitrary Points for Few-Shot Visual Imitation
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Any-point Trajectory Modeling for Policy Learning
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving