Improved Large Language Diffusion Models
- Published
- Source
- arXiv
- Paper number
- 501
- Field
- LLMs / NLP
- arXiv ID
- 2606.25331
Key points
- It is the first case of scaling a masked diffusion language model to 12T tokens, trained from scratch with 8B parameters.
- Grouped-query attention reduces KV-cache memory and tied embeddings slim the parameters down to 7.62B.
- Confidence-based scoring improves PIQA by +1.3, ARC-C by +0.6, and HellaSwag by +2.3 over likelihood-based scoring on multiple-choice evaluation.
- Extending SFT to 12 epochs yields continuous gains on GSM8K, MATH, and MMLU-Pro, confirming data reuse effects in diffusion LLMs.
- The base model reaches MMLU 74.8, BBH 71.3, and GSM8K 81.9, surpassing Qwen2.5 7B across multiple benchmarks.
- The instruct model is competitive without RL alignment, with MMLU-Redux 76.4 and GSM8K 89.0, but it still trails Qwen2.5 Instruct on code and math.
Paper links
External research summaries. These are not HDATF publications or measured product results.