On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length

Published
Source
arXiv
Paper number
175
Field
Agents / Reinforcement Learning
arXiv ID
2605.02572

Key points

  • Large language models deployed as interactive agents struggle on multi-step, long-horizon real-world tasks.
  • Prior work has largely overlooked intrinsic task horizon length as a fundamental factor shaping LLM training dynamics, often conflating it with other complexity factors.
  • Longer horizons worsen challenges such as error accumulation, exponential growth in state-action mapping complexity, and ambiguous credit assignment caused by sparse and delayed feedback.
  • The authors perform a systematic empirical study using controlled task environments, Sudoku and Rush Hour, designed to isolate and vary intrinsic task horizon length while keeping reasoning complexity constant.
  • The study evaluates the horizon-shortening principle, specifically macro actions that let the agent produce multiple atomic actions in one step, and subgoal decomposition that splits the task into verifiable subgoals with dense intermediate rewards.
  • After supervised fine-tuning, the LLM agent is trained with reinforcement learning using a re-evaluated REINFORCE algorithm combined with off-policy stabilization techniques for robust training.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)