SWE-Explore: Benchmarking How Coding Agents Explore Repositories

Published
Source
arXiv
Paper number
377
Field
Software Engineering
arXiv ID
2606.07297

Key points

  • A new benchmark that evaluates repository exploration independently from patch generation.
  • It derives trajectory-grounded, line-level ground truth from successful agent trajectories.
  • It covers diverse coding settings with 848 issues, 10 languages, and 203 repositories.
  • Agent exploration forms a clear top tier compared with classical retrieval.
  • File-level hit rates are strong, but core line recall remains only 0.14 to 0.19, making it the bottleneck.
  • Missing key context is more damaging to patch success than noisy context.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)