On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length
- Published
- Source
- arXiv
- Paper number
- 175
- Field
- Agents / Reinforcement Learning
- arXiv ID
- 2605.02572
Key points
- Large language models deployed as interactive agents struggle on multi-step, long-horizon real-world tasks.
- Prior work has largely overlooked intrinsic task horizon length as a fundamental factor shaping LLM training dynamics, often conflating it with other complexity factors.
- Longer horizons worsen challenges such as error accumulation, exponential growth in state-action mapping complexity, and ambiguous credit assignment caused by sparse and delayed feedback.
- The authors perform a systematic empirical study using controlled task environments, Sudoku and Rush Hour, designed to isolate and vary intrinsic task horizon length while keeping reasoning complexity constant.
- The study evaluates the horizon-shortening principle, specifically macro actions that let the agent produce multiple atomic actions in one step, and subgoal decomposition that splits the task into verifiable subgoals with dense intermediate rewards.
- After supervised fine-tuning, the LLM agent is trained with reinforcement learning using a re-evaluated REINFORCE algorithm combined with off-policy stabilization techniques for robust training.
Paper links
External research summaries. These are not HDATF publications or measured product results.