LIMO: Less is More for Reasoning
- Published
- Source
- arXiv
- Paper number
- 027
- Field
- Reasoning / Data Efficiency
- arXiv ID
- 2502.03387
Key points
- Conventional wisdom says that complex reasoning in LLMs requires vast amounts of training data.
- Collecting and training on large datasets is expensive in compute and resources.
- Models trained on large, uncurated reasoning datasets have limited generalization and may memorize.
- LIMO is a data-efficient supervised fine-tuning (SFT) approach that uses a minimal dataset.
- It builds a highly curated dataset of 800 math reasoning problems through multi-stage filtering for problem difficulty and detailed rule-based scoring of reasoning-chain quality.
- The training process allows the base model to use pretrained knowledge and integrate meta-reasoning tasks into coherent chains.
Paper links
External research summaries. These are not HDATF publications or measured product results.