Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
- Published
- Source
- arXiv
- Paper number
- 600
- Field
- AI / General
- arXiv ID
- 2607.08964
Key points
- It covers 46 long-horizon terminal tasks across 9 categories, including experimental reproduction, software engineering, multimodal analysis, games, and scientific computation.
- It uses subtask-based dense rewards that provide partial scores instead of a binary pass or fail signal, making intermediate progress visible.
- The benchmark averages 9.9 million tokens, 231 episodes, and 85.3 minutes per task, making it far more demanding than Terminal-Bench or SWE-Bench.
- Even the strongest model, GPT-5.5, reaches only 15.2 percent pass@1 at R ≥ 0.95, while the average pass@1 is just 4.3 percent, leaving substantial room for improvement.
- The main failure mode is partial progress followed by timeouts and early false finishes, with weak self-verification emerging as the key bottleneck.
- Dense rewards are essential for exposing false-finish patterns, whereas binary evaluation misses them.
Paper links
External research summaries. These are not HDATF publications or measured product results.