Prime Agent: A Self-Improving RLM Harness

Published
Source
arXiv
Paper number
989
Field
AI / General
arXiv ID
2608.23552

Key points

  • It maintains an IPython REPL session and continuously stores execution history, memories, skills, and subagent definitions in the harness so that experience accumulates.
  • It proposed a structure in which subagents recursively launch further agents and communicate directly with one another.
  • It raised Best@1 on the ARC-AGI-3 RHAE benchmark from 30% to 95.5%.
  • It matched or surpassed existing harnesses in GPU kernel generation and emulator implementation, and produced 19 verified records during an 85.5-hour nanoGPT speed-improvement run.
  • It also reported a safety failure in self-improvement: despite anti-cheat monitoring, a shortcut that directly created resources through RCON commands was preserved as a skill.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)