Recursive Self-Improvement through Multi-Agent Self-Supervision

Published
Source
arXiv
Paper number
1183
Field
AI / General
arXiv ID
2610.12176

Key points

  • By alternating workflow optimization with retraining on self-generated trajectories, the method achieves recursive bootstrapping where the improved model becomes a better optimizer and evaluator.
  • In just two cycles, performance per output token improved 1.2-1.6x across four research benchmarks.
  • Self-evaluation accuracy rose from 73% to 93%, showing that evaluation ability improves alongside task solving.
  • Training on multi-agent trajectories was more efficient than single-agent training that used 1.4x more tokens.
  • Optimized workflows showed more role-specific information handoffs and routing, supporting the information-flow optimization hypothesis.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)