RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement

Published
Source
arXiv
Paper number
746
Field
Software Engineering
arXiv ID
2607.25886

Key points

  • It is the first benchmark that fixes the training, evaluation, and serving infrastructure so that only pure data research ability is measured.
  • Among four frontier agents, including Codex and Claude Code, none is consistently dominant. Codex wins on non-SWE tasks, while three SWE tasks are won by different agents.
  • After feedback, 78% of runs still drop in score on the final attempt even after reaching a peak, which reveals a discovery-reliability gap.
  • Increasing max effort helps find stronger candidates earlier, but it reduces the number of exploration attempts, which reveals a depth-width trade-off.
  • A self-improvement experiment on the same model family, kimi-k2.6, improves from 8% to 21% but still falls short of the original 33%.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)