Efficient Pre-Training with Token Superposition

Published
Source
arXiv
Paper number
188
Field
LLMs / Pretraining / Efficiency
arXiv ID
2605.06546

Key points

  • Pretraining large language models is often extremely expensive and inefficient at scale, and it usually requires complex and invasive modifications to achieve high data throughput.
  • The paper presents Token Superposition Training, or TST, a simple drop-in method that greatly improves data throughput per FLOP during pretraining without changing parallelism, optimizers, tokenizers, data, or model architecture.
  • TST has two stages: a highly efficient superposition stage that combines many consecutive tokens into one bundle and trains with a multi-hot cross-entropy, or MCE, objective, followed by a recovery stage that returns to standard training.
  • The authors evaluate TST extensively at the 270M and 600M parameter scales and validate it on 3B and 10B A1B mixture-of-experts models, which demonstrates strong robustness across settings.
  • Ultimately, TST consistently outperforms baseline loss and downstream evaluation and reduces the total pretraining time of the 10B A1B scale by up to 2.5 times at matched loss.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)