Improved Large Language Diffusion Models

Published
Source
arXiv
Paper number
501
Field
LLMs / NLP
arXiv ID
2606.25331

Key points

  • It is the first case of scaling a masked diffusion language model to 12T tokens, trained from scratch with 8B parameters.
  • Grouped-query attention reduces KV-cache memory and tied embeddings slim the parameters down to 7.62B.
  • Confidence-based scoring improves PIQA by +1.3, ARC-C by +0.6, and HellaSwag by +2.3 over likelihood-based scoring on multiple-choice evaluation.
  • Extending SFT to 12 epochs yields continuous gains on GSM8K, MATH, and MMLU-Pro, confirming data reuse effects in diffusion LLMs.
  • The base model reaches MMLU 74.8, BBH 71.3, and GSM8K 81.9, surpassing Qwen2.5 7B across multiple benchmarks.
  • The instruct model is competitive without RL alignment, with MMLU-Redux 76.4 and GSM8K 89.0, but it still trails Qwen2.5 Instruct on code and math.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)