Unlocking Lossless Speedups in LLMs via Discrete Diffusion

Published
Source
arXiv
Paper number
1067
Field
Machine Learning
arXiv ID
2609.04010

Key points

  • Adding only a lightweight training stage that equips an existing LLM with diffusion weights, allowing multiple tokens to be filled in simultaneously, delivered up to a 3× speedup without major architectural changes.
  • Unlike speculative decoding methods such as EAGLE-3 and DFlash, which require a small draft-only model, it handles drafting and verification in a single model and provides lossless acceleration with an answer distribution exactly matching the original.
  • Whereas earlier diffusion LLMs lost their gains as batch size increased, Uno achieved higher throughput than speculative decoding at every evaluated batch size and maintained a 2× speedup even at the largest batch.
  • The 8B Uno outperformed the 26B DiffusionGemma and commercial Mercury 2 across all benchmarks for agentic tool use, coding, and long-document reasoning.
  • It also accelerated rollout generation, the slowest stage of RL post-training, showing that it can speed up reinforcement-learning training itself as well as inference.
  • It can be built on already released open-weight models, so there is no need for a new large-scale training run.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)