Unlocking Lossless Speedups in LLMs via Discrete Diffusion
- Published
- Source
- arXiv
- Paper number
- 1067
- Field
- Machine Learning
- arXiv ID
- 2609.04010
Key points
- Adding only a lightweight training stage that equips an existing LLM with diffusion weights, allowing multiple tokens to be filled in simultaneously, delivered up to a 3× speedup without major architectural changes.
- Unlike speculative decoding methods such as EAGLE-3 and DFlash, which require a small draft-only model, it handles drafting and verification in a single model and provides lossless acceleration with an answer distribution exactly matching the original.
- Whereas earlier diffusion LLMs lost their gains as batch size increased, Uno achieved higher throughput than speculative decoding at every evaluated batch size and maintained a 2× speedup even at the largest batch.
- The 8B Uno outperformed the 26B DiffusionGemma and commercial Mercury 2 across all benchmarks for agentic tool use, coding, and long-document reasoning.
- It also accelerated rollout generation, the slowest stage of RL post-training, showing that it can speed up reinforcement-learning training itself as well as inference.
- It can be built on already released open-weight models, so there is no need for a new large-scale training run.
Paper links
External research summaries. These are not HDATF publications or measured product results.