SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness
- Published
- Source
- arXiv
- Paper number
- 1093
- Field
- AI Agents
- arXiv ID
- 2609.20519
Key points
- Applies RSI (recursive self-improvement) at the harness layer — the AI observes its execution traces and automatically improves the harness code.
- Stage-gated exploration from broad to deep: 152 proposed directions → 500 execution environments → 3,000+ runs → 4 surviving mechanisms.
- The 4 survivors: Action Fusion (command merging), Online Context Compaction, ObservationPack (observation bundling), and Evidence-Preserving Delegated Reading.
- On EdgeBench's 51 tasks, it matched Pi's performance (42.0 vs 44.8) while cutting tokens by 44.7-49.0% and reducing API costs to about one-third.
- Consistent improvements across both GPT-5.6 Sol and Opus 5 — confirming generalization across models.
- Proposes the concept of 'harness pretraining' — a recursive efficiency vision where a more efficient harness lowers the cost of the next research cycle.
Paper links
External research summaries. These are not HDATF publications or measured product results.