Self-Harness: Harnesses That Improve Themselves

Published
Source
arXiv
Paper number
380
Field
LLMs / NLP
arXiv ID
2606.09498

Key points

  • It proposes Self-Harness, a paradigm in which the agent improves its own harness through weakness mining, harness proposal, and proposal validation.
  • Across three models, MiniMax M2.5, Qwen3.5-35B-A3B, and GLM-5, the held-out pass rates improve from 40.5% to 61.9%, from 23.8% to 38.1%, and from 42.9% to 57.1%, respectively.
  • Different harness changes are produced for different models, so the updates are model-specific rather than generic instruction additions.
  • A regression-test-based acceptance rule prevents overfitting, and the edits remain small and auditable.
  • The work is directly related to existing harness systems such as SemaClaw and OpenClaw.
  • It is a meaningful step forward in the self-improvement research line for agents.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)