The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models

Published
Source
arXiv
Paper number
039
Field
Training Efficiency
arXiv ID
2503.02875

Key points

  • Existing methods for improving LLM reasoning require either massive labeled data or computationally expensive sampling.
  • Traditional fine-tuning approaches struggle with efficiency and scalability on complex reasoning tasks.
  • The method exploits the Prefix Self-Consistency phenomenon, where different solution paths share a common early reasoning prefix.
  • We develop UPFT, a method that fine-tunes the model on extracted prefix substrings while preserving the reasoning structure.
  • It integrates structure tuning to prevent catastrophic forgetting of reasoning ability.
  • The initial tokens in a reasoning process carry key information and remain highly consistent across different solutions.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)