ENPIRE: Agentic Robot Policy Self-Improvement in the Real World
cs.AI
Submitted: 2026-06-18
Updated: 2026-09-20
Comments: 2026 Conference on Robot Learning
Code: https://github.com/huggingface/gym-pusht
Project page: https://trinkle23897.github.io/learning-beyond-gradients
License: http://creativecommons.org/licenses/by/4.0/
The gist: Achieving dexterous robotic manipulation in the real world heavily relies on human supervision and algorithm engineering, which becomes a central bottleneck in the pursuit of general physical
Terminology
Abstract
Achieving dexterous robotic manipulation in the real world heavily relies on human supervision and algorithm engineering, which becomes a central bottleneck in the pursuit of general physical intelligence. Although emerging coding agents can generate code to automate algorithm search, their successes remain largely confined in digital environments. We conjecture that the missing abstraction to automate robotics research is a repeatable feedback loop for real-world policy improvement: reset the scene, execute a policy, verify the outcome, and refine the next iteration. To bridge this gap, we introduce ENPIRE, a harness framework for coding agents that instantiates this physical feedback routine with four core modules: an Environment module (EN) for automatic reset and verification, a Policy Improvement module (PI) that launches policy refinement, a Rollout module (R) to evaluate policies with one or multiple physical robots operating in parallel, and an Evolution module (E) in which coding agents analyze logs, consult literature, improve training infrastructure and algorithm code to address failure modes. This closed-loop system transforms real-world manipulation learning into a controllable optimization procedure, minimizing human effort while allowing fair ablations across training recipe and agent variants. Powered by ENPIRE, frontier coding agents can autonomously train a policy to achieve a 99% success rate on challenging, dexterous manipulation tasks, such as organizing a pin box, fastening a zip tie, and tool use, a process that further accelerates when we dispatch an agent team on a robot fleet. Our results suggest a practical and scalable path toward deploying coding agents to autonomously advancing robotics in the physical world.
Sources
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- OpenAI Gym
- RT-1: Robotics Transformer for Real-World Control at Scale
- CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
- Learning to Walk via Deep Reinforcement Learning
- $\pi^{*}_{0.6}$: a VLA That Learns From Experience
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- RL-100: Performant Robotic Manipulation with Real-World Reinforcement Learning
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist Robots
- AgentRxiv: Towards Collaborative Autonomous Research
- ProgPrompt: Generating Situated Robot Task Plans using Large Language Models
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation
- Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
- ReAct: Synergizing Reasoning and Acting in Language Models
- Language to Rewards for Robotic Skill Synthesis
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection