Large Language Diffusion Models

Published
Source
arXiv
Paper number
034
Field
LLMs / Diffusion
arXiv ID
2502.09992

Key points

  • The common belief is that the capabilities of large language models, such as scalability and in-context learning, come only from autoregressive architectures.
  • Autoregressive models have limitations, including the curse of reversal, where they struggle with tasks that require bidirectional or reverse reasoning.
  • There has been no large language diffusion model that can match the performance of billion-parameter autoregressive models.
  • We developed LLaDA, a masked diffusion model for discrete token sequences, using a Transformer backbone without causal masking.
  • We optimized a principled cross-entropy loss that upper-bounds the negative log-likelihood and trained an 8B-parameter LLaDA model from scratch on a 2.3-trillion-token dataset.
  • We implemented a flexible iterative sampling strategy at inference time, including low-confidence remasking to balance generation quality and speed.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)