MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning
cs.RO, cs.AI, cs.CL
Submitted: 2026-07-15
Updated: 2026-08-31
Project page: https://yuzihaowashu.github.io/MEMORA
Terminology
Sources
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- EgoAVFlow: Robot Policy Learning with Active Vision from Human Egocentric Videos via 3D Flow
- Selectively Answering Ambiguous Questions
- Memory for Autonomous LLM Agents:Mechanisms, Evaluation, and Emerging Frontiers
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- Ego4D: Around the World in 3,000 Hours of Egocentric Video
- EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
- Inner Monologue: Embodied Reasoning through Planning with Language Models
- Selective Question Answering under Domain Shift
- RoboMemory: A Brain-inspired Multi-memory Agentic Framework for Interactive Environmental Learning in Physical Embodied Systems
- Masquerade: Learning from In-the-wild Human Videos using Data-Editing
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
- Code as Policies: Language Model Programs for Embodied Control
- ActiveMimic: Egocentric Video Pretraining with Active Perception
- EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations
- BrainMem: Brain-Inspired Evolving Memory for Embodied Agent Task Planning
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving