ClawGym II: Exploring Black-Box RL on Agent Harness
- Published
- Source
- arXiv
- Paper number
- 926
- Field
- LLMs / NLP
- arXiv ID
- 2608.16798
Key points
- It isolates the task environment and execution tools in separate temporary environments. Model requests are handled by a middle server that records input tokens, output tokens, and generation probabilities.
- ClawGym-Bench uses either code inspection alone or a weighted sum of code at 0.7 and evaluation criteria at 0.3. GPT-5.4 is used for evaluation-criterion judgments, and PinchBench evaluates 30 tasks after removing multimodal ones.
- PinchBench scores for the Qwen3-30A3B family improve from 75.61 to 87.32 in OpenClaw cold-start initial models. In Claude Code, they rise from 54.14 to 71.42 in the original model.
- Under the same execution budget, PPO runs 256 tasks once each, while GRPO runs 32 tasks eight times each. Both methods show stable gains over about 200 to 400 training steps.
- JobBench-Easy improves from 20.46 to 27.20, and OfficeQA-Full from 8.53 to 21.54. The score of a model trained in WhiteBox AgentLoop and transferred to OpenClaw is 50.33, lower than 62.62 for a model trained directly in OpenClaw.
Paper links
External research summaries. These are not HDATF publications or measured product results.