SWE-Together: Evaluating Coding Agents in Interactive User Sessions
- Published
- Source
- arXiv
- Paper number
- 535
- Field
- Software Engineering
- arXiv ID
- 2606.29957
Key points
- It extracts 109 verifiable tasks from 11,260 real sessions, for a conversion rate of 0.97 percent.
- The state-conditioned LLM user simulator preserves the original intent while adapting feedback to the trajectory.
- It evaluates along multiple dimensions using final accuracy, user correction, meaning the number of correction turns, and intent coverage.
- Stronger agents achieve higher success rates with fewer interventions, suggesting an improved user experience.
- It highlights the limits of static SWE-Bench style evaluation, which misses real interaction by scoring only single-turn commands and final code.
- It uses four upstream data sources: DataClaw, Pi-staging, Hyperswitch, and SWE-chat.
Paper links
External research summaries. These are not HDATF publications or measured product results.