Agent Lightning v1.0: Towards Harnessed Agentic RL
- Published
- Source
- arXiv
- Paper number
- 948
- Field
- AI / General
- arXiv ID
- 2608.17528
Key points
- In harnessed agentic reinforcement learning, the deployment harness, rather than the training engine, owns the environment-interaction loop, while the trainer observes only LLM request-response pairs. OpenClaw, Claude Code, and OpenHands are cited as examples of such harnesses.
- The paper identifies four main challenges: retokenization, where identical text cannot be merged when token boundaries differ; a variable number of samples per rollout; loss normalization; and backend scheduling.
- Computing both advantages and normalization at the rollout level rather than the sample level performed best, reaching a verification reward of 38.2%.
- After the SWE-smith curation pipeline retained just 6,000 examples, reinforcement learning alone improved SWE-bench Verified performance from 41.8% to 56.4%.
- Safeguards against reward hacking, including blocking Git commands and allowlisting network access, were necessary to stabilize training.
Paper links
External research summaries. These are not HDATF publications or measured product results.