Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

Published
Source
arXiv
Paper number
600
Field
AI / General
arXiv ID
2607.08964

Key points

  • It covers 46 long-horizon terminal tasks across 9 categories, including experimental reproduction, software engineering, multimodal analysis, games, and scientific computation.
  • It uses subtask-based dense rewards that provide partial scores instead of a binary pass or fail signal, making intermediate progress visible.
  • The benchmark averages 9.9 million tokens, 231 episodes, and 85.3 minutes per task, making it far more demanding than Terminal-Bench or SWE-Bench.
  • Even the strongest model, GPT-5.5, reaches only 15.2 percent pass@1 at R ≥ 0.95, while the average pass@1 is just 4.3 percent, leaving substantial room for improvement.
  • The main failure mode is partial progress followed by timeouts and early false finishes, with weak self-verification emerging as the key bottleneck.
  • Dense rewards are essential for exposing false-finish patterns, whereas binary evaluation misses them.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)