Efficient Pre-Training with Token Superposition
- Published
- Source
- arXiv
- Paper number
- 188
- Field
- LLMs / Pretraining / Efficiency
- arXiv ID
- 2605.06546
Key points
- Pretraining large language models is often extremely expensive and inefficient at scale, and it usually requires complex and invasive modifications to achieve high data throughput.
- The paper presents Token Superposition Training, or TST, a simple drop-in method that greatly improves data throughput per FLOP during pretraining without changing parallelism, optimizers, tokenizers, data, or model architecture.
- TST has two stages: a highly efficient superposition stage that combines many consecutive tokens into one bundle and trains with a multi-hot cross-entropy, or MCE, objective, followed by a recovery stage that returns to standard training.
- The authors evaluate TST extensively at the 270M and 600M parameter scales and validate it on 3B and 10B A1B mixture-of-experts models, which demonstrates strong robustness across settings.
- Ultimately, TST consistently outperforms baseline loss and downstream evaluation and reduces the total pretraining time of the 10B A1B scale by up to 2.5 times at matched loss.
Paper links
External research summaries. These are not HDATF publications or measured product results.