LADDER: Self-Improving LLMs Through Recursive Problem Decomposition
- Published
- Source
- arXiv
- Paper number
- 043
- Field
- Math Reasoning
- arXiv ID
- 2503.00735
Key points
- Reinforcement learning for large language models (LLMs) requires a progressive curriculum of difficulty in the training tasks, but that is often unavailable and leads to training stagnation or catastrophic performance collapse.
- Complex reasoning domains such as mathematics pose a substantial challenge because there is a large gap between problems the model can solve and those it cannot solve yet.
- Existing LLM training and improvement methods often rely heavily on large human-curated datasets or continuous human supervision and feedback.
- The LADDER framework enables recursive problem decomposition, allowing the LLM to autonomously generate a tree of progressively simpler variants of a complex problem it cannot currently solve, using transformation libraries and induced prompting.
- A robust numerical verification framework evaluates answer correctness and provides a clear, objective reward signal for reinforcement learning.
- Test-Time Reinforcement Learning (TTRL) extends LADDER by generating problem-specific focused variants at inference time for previously unsolved test problems, enabling targeted and dynamic learning on individual hard questions.
Paper links
External research summaries. These are not HDATF publications or measured product results.