Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement
- Published
- Source
- arXiv
- Paper number
- 1062
- Field
- AI / General
- arXiv ID
- 2609.01481
Key points
- It secured both agent autonomy and verifiability by leaving the harness implementation untouched and enforcing only the formats of artifacts (plans and test reports), triggering retries on violations.
- Across GameCraft-Bench, FrontierSWE, and ProgramBench, the Codex+GPT-5.5, OpenCode+DeepSeek-V4-Pro, and Pi+MiniMax-M3 combinations all improved after just 3 loops, with an average relative gain of 52.25% and a maximum of 82.86%.
- On FrontierSWE, Codex+GPT-5.5 continued improving from 22% to 72.67% over 10 loops, showing that gains accumulated with more iterations.
- Given only a requirements document in an empty working directory, it autonomously completed 'Fusepoint,' a first-person shooter with a story, combat AI, presentation, and audio, over more than 70 loops. Human intervention was limited to restoring network connectivity.
- It closed 65 of 81 recorded issues on its own and used version history to recover from 17 regression issues in which previously verified features had broken, showing that regression management is central to long-term development.
- It used a progressive-disclosure structure that assigned skills for Godot engine integration, asset generation, UI/UX, and testing by role to encourage reuse.
Paper links
External research summaries. These are not HDATF publications or measured product results.