LADDER: Self-Improving LLMs Through Recursive Problem Decomposition

Published
Source
arXiv
Paper number
043
Field
Math Reasoning
arXiv ID
2503.00735

Key points

  • Reinforcement learning for large language models (LLMs) requires a progressive curriculum of difficulty in the training tasks, but that is often unavailable and leads to training stagnation or catastrophic performance collapse.
  • Complex reasoning domains such as mathematics pose a substantial challenge because there is a large gap between problems the model can solve and those it cannot solve yet.
  • Existing LLM training and improvement methods often rely heavily on large human-curated datasets or continuous human supervision and feedback.
  • The LADDER framework enables recursive problem decomposition, allowing the LLM to autonomously generate a tree of progressively simpler variants of a complex problem it cannot currently solve, using transformation libraries and induced prompting.
  • A robust numerical verification framework evaluates answer correctness and provides a clear, objective reward signal for reinforcement learning.
  • Test-Time Reinforcement Learning (TTRL) extends LADDER by generating problem-specific focused variants at inference time for previously unsolved test problems, enabling targeted and dynamic learning on individual hard questions.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)