Do AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and Execution
- Published
- Source
- arXiv
- Paper number
- 636
- Field
- AI / General
- arXiv ID
- 2607.13034
Key points
- It defined ACRR to measure unnecessary agent work and proposed E3, which estimates scope, executes the minimal path, and expands only when verification fails.
- Across 121 edits in MSE-Bench's controlled simulator, it maintained the same 100% success rate as the strongest baseline.
- E3 reduced cost by 85%, tokens by 91%, and files inspected by 92%, and was also 16% better than the adaptive-search baseline.
- It can be applied as an execution policy that reduces the waste of reading an entire repository for a small change while retaining verification as a safeguard.
- The main figures come from a controlled simulator rather than a deployed agent, and real-model validation was limited to a gpt-4o case study, so broader validation is needed.
Paper links
External research summaries. These are not HDATF publications or measured product results.