An Empirical Study of Harness Design for Coding Agents

Published
Source
arXiv
Paper number
1099
Field
AI Agents
arXiv ID
2609.20804

Key points

  • 176 matched settings across 4 models (Nemotron-3 30B/120B/550B, Mistral-Medium-3.5) on SWE-Bench Verified and Terminal-Bench 2.1
  • Most context management benefit comes from preventing context overflow failures
  • Staged rule-based elision before LLM summarization is the most efficient strategy
  • Planning shifts from accuracy scaffold for weaker models to cost saver for stronger ones
  • Bash-capable models achieve equal performance with lower cost using bash-only interface

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)