Living-Harness Is an Interactive-Agent Evolver
- Published
- Source
- arXiv
- Paper number
- 763
- Field
- Agents
- arXiv ID
- 2607.26598
Key points
- It turns agent failures into persistent procedural repair rather than one-off lessons and stores them permanently.
- It accumulates knowledge along two axes: episode memory, including trigger conditions, failure patterns, and recovery actions, and a state graph, including state nodes, repair edges, and transition rules.
- It structures what to change and how to change it through a domain-level guide called Evolution-SOP.
- It improves by an average of 10.07 points on tau^2-Bench and 9.91 points on MultiWOZ-2.4.
- Even when the harness state evolved by GPT-5.2 is transferred to other models such as Gemini 3 Pro, GLM-5, and Qwen3-max through retrieval alone, performance still improves.
Paper links
External research summaries. These are not HDATF publications or measured product results.