Playful Agentic Robot Learning

Published
Source
arXiv
Paper number
457
Field
Robotics
arXiv ID
2606.19419

Key points

  • During play time, the system autonomously proposes tasks, executes them, verifies the results, diagnoses failures, and accumulates a skill library.
  • It uses a novelty-learnability rule to propose curiosity-driven tasks, with curiosity play improving from 24.7 percent under random play to 32.3 percent.
  • On LIBERO-PRO, it improves by 20.6 points over CaP-Agent0, from 23.2 percent to 43.8 percent, and exceeds every VLA baseline, including pi0.5 at 12.8 percent.
  • When the learned skills are plugged into CaP-Agent0, they improve RoboSuite by 8.9 points and real-robot performance by 8.8 points.
  • Skills learned in a single-arm environment also transfer to two-arm tasks, such as two-arm lifting, with a 24.0-point gain.
  • Curiosity play and execution-system improvements are complementary, and both together reach the best score of 44.3 percent.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)