RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement
- Published
- Source
- arXiv
- Paper number
- 746
- Field
- Software Engineering
- arXiv ID
- 2607.25886
Key points
- It is the first benchmark that fixes the training, evaluation, and serving infrastructure so that only pure data research ability is measured.
- Among four frontier agents, including Codex and Claude Code, none is consistently dominant. Codex wins on non-SWE tasks, while three SWE tasks are won by different agents.
- After feedback, 78% of runs still drop in score on the final attempt even after reaching a peak, which reveals a discovery-reliability gap.
- Increasing max effort helps find stronger candidates earlier, but it reduces the number of exploration attempts, which reveals a depth-width trade-off.
- A self-improvement experiment on the same model family, kimi-k2.6, improves from 8% to 21% but still falls short of the original 33%.
Paper links
External research summaries. These are not HDATF publications or measured product results.