Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents
- Published
- Source
- arXiv
- Paper number
- 549
- Field
- Distributed Systems
- arXiv ID
- 2607.01120
Key points
- The core claim is that the bottleneck in self-evolving agents lies not in the RL algorithm, but in the system layer, including ATDP, data proxies, and the evolution control surface.
- ATDP proposes a vendor-neutral protocol that includes step-level decision context, actions, outcomes, delayed rewards, and governance metadata.
- The evolution control surface automatically chooses among memory, skills, harnesses, weights, and rollback based on trajectory statistics, evaluation scores, user correction rates, cost, and safety constraints.
- AREAL2.0 is a prototype that reorganizes existing RL frameworks into an online agent-service-oriented RL loop, and it folds LLM calls from agent services into the RL training loop by swapping only the gateway.
- It cites OpenClaw as an example of a personal self-evolving agent and analyzes the gap to enterprise-level governance.
- It defines self-evolution not as a simple weight update, but as a multidimensional improvement of a composite policy made up of the LLM, harness, memory, and tools.
Paper links
External research summaries. These are not HDATF publications or measured product results.