ENPIRE: Agentic Robot Policy Self-Improvement in the Real World
- Published
- Source
- arXiv
- Paper number
- 466
- Field
- AI / General
- arXiv ID
- 2606.19980
Key points
- It is the first integrated framework that lets coding agents such as Codex, GPT-5.5, Claude Code, and Kimi K2.6 autonomously improve policies in real robot environments.
- It has a four-stage structure: Environment with automatic reset and verification, Policy Improvement with heuristics, BC, and RL, Rollout on the real robot, and Evolution where the agent analyzes logs and improves code.
- In Pin Insertion, it reaches 100% success faster than human-in-the-loop state-of-the-art methods, and 99% on tasks such as PushT and zip tie cutting.
- When scaled to a multi-robot fleet from 1 to 8 agents, it shows shorter wall-clock time, but token cost grows superlinearly and spikes at MTU 8 agents.
- It introduces Mean Robot Utilization (MRU) and Mean Token Utilization (MTU) as two physical resource efficiency metrics.
- It automatically discovers the synergy between VLA, GR00T, and procedural tool calls, improving VLA accuracy in RoboCasa and successfully transferring to the real world.
Paper links
External research summaries. These are not HDATF publications or measured product results.