Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning

Published
Source
arXiv
Paper number
826
Field
AI / General
arXiv ID
2608.05144

Key points

  • It proposes a verification-gated runtime self-evolution structure that approves goal and method changes only through a verification gate.
  • The Manager, Planner, Engineer, and Reviewer roles each carry out bounded missions on durable project state.
  • On SWE-Bench Pro, it reaches 78% compared with 59% for Direct Copilot, while using 1.41x as many tokens.
  • At the mature stage, it uses 21% fewer solve tokens and shortens workflow time by 15% compared with the new Wave stage.
  • It is validated on seven GPT-5.5 benchmarks and produces practical results, including an upstream merge of the RWKV6 kernel.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)