SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness

Published
Source
arXiv
Paper number
1093
Field
AI Agents
arXiv ID
2609.20519

Key points

  • Applies RSI (recursive self-improvement) at the harness layer — the AI observes its execution traces and automatically improves the harness code.
  • Stage-gated exploration from broad to deep: 152 proposed directions → 500 execution environments → 3,000+ runs → 4 surviving mechanisms.
  • The 4 survivors: Action Fusion (command merging), Online Context Compaction, ObservationPack (observation bundling), and Evidence-Preserving Delegated Reading.
  • On EdgeBench's 51 tasks, it matched Pi's performance (42.0 vs 44.8) while cutting tokens by 44.7-49.0% and reducing API costs to about one-third.
  • Consistent improvements across both GPT-5.6 Sol and Opus 5 — confirming generalization across models.
  • Proposes the concept of 'harness pretraining' — a recursive efficiency vision where a more efficient harness lowers the cost of the next research cycle.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)