SWE-Together: Evaluating Coding Agents in Interactive User Sessions

Published
Source
arXiv
Paper number
535
Field
Software Engineering
arXiv ID
2606.29957

Key points

  • It extracts 109 verifiable tasks from 11,260 real sessions, for a conversion rate of 0.97 percent.
  • The state-conditioned LLM user simulator preserves the original intent while adapting feedback to the trajectory.
  • It evaluates along multiple dimensions using final accuracy, user correction, meaning the number of correction turns, and intent coverage.
  • Stronger agents achieve higher success rates with fewer interventions, suggesting an improved user experience.
  • It highlights the limits of static SWE-Bench style evaluation, which misses real interaction by scoring only single-turn commands and final code.
  • It uses four upstream data sources: DataClaw, Pi-staging, Hyperswitch, and SWE-chat.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)