The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models
- Published
- Source
- arXiv
- Paper number
- 039
- Field
- Training Efficiency
- arXiv ID
- 2503.02875
Key points
- Existing methods for improving LLM reasoning require either massive labeled data or computationally expensive sampling.
- Traditional fine-tuning approaches struggle with efficiency and scalability on complex reasoning tasks.
- The method exploits the Prefix Self-Consistency phenomenon, where different solution paths share a common early reasoning prefix.
- We develop UPFT, a method that fine-tunes the model on extracted prefix substrings while preserving the reasoning structure.
- It integrates structure tuning to prevent catastrophic forgetting of reasoning ability.
- The initial tokens in a reasoning process carry key information and remain highly consistent across different solutions.
Paper links
External research summaries. These are not HDATF publications or measured product results.