An Empirical Study of Harness Design for Coding Agents
- Published
- Source
- arXiv
- Paper number
- 1099
- Field
- AI Agents
- arXiv ID
- 2609.20804
Key points
- 176 matched settings across 4 models (Nemotron-3 30B/120B/550B, Mistral-Medium-3.5) on SWE-Bench Verified and Terminal-Bench 2.1
- Most context management benefit comes from preventing context overflow failures
- Staged rule-based elision before LLM summarization is the most efficient strategy
- Planning shifts from accuracy scaffold for weaker models to cost saver for stronger ones
- Bash-capable models achieve equal performance with lower cost using bash-only interface
Paper links
External research summaries. These are not HDATF publications or measured product results.