Large Language Diffusion Models
- Published
- Source
- arXiv
- Paper number
- 034
- Field
- LLMs / Diffusion
- arXiv ID
- 2502.09992
Key points
- The common belief is that the capabilities of large language models, such as scalability and in-context learning, come only from autoregressive architectures.
- Autoregressive models have limitations, including the curse of reversal, where they struggle with tasks that require bidirectional or reverse reasoning.
- There has been no large language diffusion model that can match the performance of billion-parameter autoregressive models.
- We developed LLaDA, a masked diffusion model for discrete token sequences, using a Transformer backbone without causal masking.
- We optimized a principled cross-entropy loss that upper-bounds the negative log-likelihood and trained an 8B-parameter LLaDA model from scratch on a 2.3-trillion-token dataset.
- We implemented a flexible iterative sampling strategy at inference time, including low-confidence remasking to balance generation quality and speed.
Paper links
External research summaries. These are not HDATF publications or measured product results.