EnvHarness: Awakening Static Worlds for Agent Learning
- Published
- Source
- arXiv
- Paper number
- 967
- Field
- AI / General
- arXiv ID
- 2608.19880
Key points
- It proposed an EnvHarness layer that changes only behavior by wrapping an existing environment at the reset/step interface with Stage, Contract, and Chain plugins, rather than constructing a new environment.
- By inheriting the original verifier, the ground-truth scorer, unchanged, it inherently prevents the ground-truth reliability problem of LLM-generated environments.
- It created an automated loop in which EnvRigger diagnoses policy-failure trajectories, automatically synthesizes environments targeting those weaknesses, and validates them with new rollouts.
- Across 5 benchmarks in 4 domains, including SWE-bench Verified, it achieved gains of up to +9.0 points and a 9.8% reduction in execution steps over learning in the original environments.
- Through iterative policy-environment co-evolution, performance continued to rise from 47.67 to 54.79 even with 300 environments, going beyond the point where conventional environment scaling falters.
Paper links
External research summaries. These are not HDATF publications or measured product results.