Simple Agentic Memory for Generalist Robot Policies
cs.RO, cs.AI
Submitted: 2026-09-29
Updated: 2026-09-29
Project page: https://simplearm.github.io/ABSTRACT
Terminology
Sources
- Qwen3-VL Technical Report
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- RMBench: Memory-Dependent Robotic Manipulation Benchmark with Insights into Policy Design
- RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies
- ReWind: Understanding Long Videos with Instructed Learnable Memory
- VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
- Chameleon: Control-Indexed Prospective Memory for Visuomotor Manipulation
- RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark
- Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
- A Survey on Self-Improving Test-Time Intelligence: Feedback-Driven Adapting, Learning, and Scaling at Inference
- DINOv2: Learning Robust Visual Features without Supervision
- MemGPT: Towards LLMs as Operating Systems
- SAM 2: Segment Anything in Images and Videos
- A Simple Baseline for Streaming Video Understanding
- MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
- MemER: Scaling Up Memory for Robot Control via Experience Retrieval
- MEM: Multi-Scale Embodied Memory for Vision Language Action Models
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving