SWE-Explore: Benchmarking How Coding Agents Explore Repositories
- Published
- Source
- arXiv
- Paper number
- 377
- Field
- Software Engineering
- arXiv ID
- 2606.07297
Key points
- A new benchmark that evaluates repository exploration independently from patch generation.
- It derives trajectory-grounded, line-level ground truth from successful agent trajectories.
- It covers diverse coding settings with 848 issues, 10 languages, and 203 repositories.
- Agent exploration forms a clear top tier compared with classical retrieval.
- File-level hit rates are strong, but core line recall remains only 0.14 to 0.19, making it the bottleneck.
- Missing key context is more damaging to patch success than noisy context.
Paper links
External research summaries. These are not HDATF publications or measured product results.