TRIM: Reducing AI-Generated CodeSlop via Agent Trajectory Minimization
- Published
- Source
- arXiv
- Paper number
- 680
- Field
- Software Engineering
- arXiv ID
- 2607.18161
Key points
- It defines CodeSlop by functional behavior rather than surface style, such as verbosity or duplication. Any edit that can be removed while tests still pass counts as CodeSlop.
- The cause is the agent's search process. Once it finds a correct solution, there is no reason to roll back the attempted edits, so they remain in the final patch.
- Telling the agent to reduce its own patch was unstable. In 3.8% to 44.9% of cases it broke behavior or made the patch larger.
- TRIM uses the task history, not the final patch, as the search space because the temporal order of edits approximates dependency structure.
- It uses a hierarchical approach that tries to remove large chunks first, confirms success, and then splits more finely if needed, so the verification cost is about half that of Delta Debugging.
- On SWE-Bench-Verified, it preserved correctness on 327 of 330 successful patches, or 99.1%, and 18 patches became exactly identical to the human-written reference patch.
Paper links
External research summaries. These are not HDATF publications or measured product results.