HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry

Published
Source
arXiv
Paper number
419
Field
AI / General
arXiv ID
2606.14249

Key points

  • We compose harnesses in a type-safe way using a 9-dimensional processor taxonomy, covering context, tools, skills, control, memory, and more, plus a substitution algebra.
  • AEGIS is a four-stage evolution pipeline, Digester, Planner, Evolver, Critic, that defends by mapping RL pathologies into symbolic space.
  • Variant isolation prevents interference through independent evolution per task cluster on heterogeneous benchmarks, yielding +13.6 on GAIA.
  • Cross-harness GRPO jointly loops harness evolution and model learning, adding another +4.7 over harness-only training.
  • On ALFWorld, Qwen3.5-9B improves by +44.0%, showing that harness improvements matter more when the model is weaker.
  • The average gain is +14.5% across 15 settings, with positive improvements in 14 of them, and the lower the baseline, the larger the gain.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)