OpenForgeRL: Train Harness-native Agents in Any Environment

Published
Source
arXiv
Paper number
707
Field
AI / General
arXiv ID
2607.21557

Key points

  • It records harness model calls as proxies so standard RL code, such as veRL, can train on them.
  • Each rollout is isolated in a Kubernetes container, allowing complex environments to run safely.
  • Training and deployment share the same harness, eliminating train-deploy mismatch.
  • OpenForge-Claw reaches 55.9 pass@3 on ClawEval, and OpenForge-GUI reaches 37.7 on OSWorld.
  • RL improves self-verification and tool coverage, but error recovery remains weak.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)