Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning
- Published
- Source
- arXiv
- Paper number
- 826
- Field
- AI / General
- arXiv ID
- 2608.05144
Key points
- It proposes a verification-gated runtime self-evolution structure that approves goal and method changes only through a verification gate.
- The Manager, Planner, Engineer, and Reviewer roles each carry out bounded missions on durable project state.
- On SWE-Bench Pro, it reaches 78% compared with 59% for Direct Copilot, while using 1.41x as many tokens.
- At the mature stage, it uses 21% fewer solve tokens and shortens workflow time by 15% compared with the new Wave stage.
- It is validated on seven GPT-5.5 benchmarks and produces practical results, including an upstream merge of the RWKV6 kernel.
Paper links
External research summaries. These are not HDATF publications or measured product results.