Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090

Published
Source
arXiv
Paper number
1040
Field
LLMs / NLP
arXiv ID
2608.27370

Key points

  • FP8 low-precision training on a consumer RTX 5090 cluster reduced the pretraining cost of a 2B model to approximately $6,900.
  • Training on 1.4 trillion tokens over 22,514 GPU hours and 17.6 days surpassed Qwen2-1.5B and approached Qwen2.5-1.5B performance.
  • It proposed the 'Puro cost scaling law,' fitted to the relationship between training cost and performance, and estimated that approximately $4,400 would suffice to reach Qwen2-1.5B-level performance.
  • Five factors produced the cost efficiency: hardware selection, low-precision training, Hyperball optimization, curriculum model averaging, and the data recipe.
  • It released all 10 sets of weights, data manifests, and training and preprocessing code under Apache 2.0.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)