PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control
cs.RO, cs.AI
Submitted: 2026-08-25
Updated: 2026-09-17
Comments: Project page: https://worv-ai.github.io/ponderpounce/
Project page: https://worv-ai.github.io/ponderpounce
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- RT-H: Action Hierarchies Using Language
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation
- Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation
- See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
- Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
- StreamVLA: Breaking the Reason-Act Cycle via Completion-State Gating
- Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control
- OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation
- RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies
- SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation
- ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation
- In-Context Imitation Learning via Next-Token Prediction
- ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
- A Dual Process VLA: Efficient Robotic Manipulation Leveraging VLM
- Training Large Language Models to Reason in a Continuous Latent Space
- What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents
- VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving