TRIM: Reducing AI-Generated CodeSlop via Agent Trajectory Minimization

Published
Source
arXiv
Paper number
680
Field
Software Engineering
arXiv ID
2607.18161

Key points

  • It defines CodeSlop by functional behavior rather than surface style, such as verbosity or duplication. Any edit that can be removed while tests still pass counts as CodeSlop.
  • The cause is the agent's search process. Once it finds a correct solution, there is no reason to roll back the attempted edits, so they remain in the final patch.
  • Telling the agent to reduce its own patch was unstable. In 3.8% to 44.9% of cases it broke behavior or made the patch larger.
  • TRIM uses the task history, not the final patch, as the search space because the temporal order of edits approximates dependency structure.
  • It uses a hierarchical approach that tries to remove large chunks first, confirms success, and then splits more finely if needed, so the verification cost is about half that of Delta Debugging.
  • On SWE-Bench-Verified, it preserved correctness on 327 of 330 successful patches, or 99.1%, and 18 patches became exactly identical to the human-written reference patch.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)