Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090
- Published
- Source
- arXiv
- Paper number
- 1040
- Field
- LLMs / NLP
- arXiv ID
- 2608.27370
Key points
- FP8 low-precision training on a consumer RTX 5090 cluster reduced the pretraining cost of a 2B model to approximately $6,900.
- Training on 1.4 trillion tokens over 22,514 GPU hours and 17.6 days surpassed Qwen2-1.5B and approached Qwen2.5-1.5B performance.
- It proposed the 'Puro cost scaling law,' fitted to the relationship between training cost and performance, and estimated that approximately $4,400 would suffice to reach Qwen2-1.5B-level performance.
- Five factors produced the cost efficiency: hardware selection, low-precision training, Hyperball optimization, curriculum model averaging, and the data recipe.
- It released all 10 sets of weights, data manifests, and training and preprocessing code under Apache 2.0.
Paper links
External research summaries. These are not HDATF publications or measured product results.