Recursive Self-Improvement through Multi-Agent Self-Supervision
- Published
- Source
- arXiv
- Paper number
- 1183
- Field
- AI / General
- arXiv ID
- 2610.12176
Key points
- By alternating workflow optimization with retraining on self-generated trajectories, the method achieves recursive bootstrapping where the improved model becomes a better optimizer and evaluator.
- In just two cycles, performance per output token improved 1.2-1.6x across four research benchmarks.
- Self-evaluation accuracy rose from 73% to 93%, showing that evaluation ability improves alongside task solving.
- Training on multi-agent trajectories was more efficient than single-agent training that used 1.4x more tokens.
- Optimized workflows showed more role-specific information handoffs and routing, supporting the information-flow optimization hypothesis.
Paper links
External research summaries. These are not HDATF publications or measured product results.