DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
- Published
- Source
- arXiv
- Paper number
- 575
- Field
- AI / General
- arXiv ID
- 2607.05147
Key points
- A semi-autoregressive architecture injects token-to-token dependence with a parallel backbone and a lightweight RNN head, while keeping per-token probabilities as exact softmax values.
- Confidence-scheduled verification uses position-wise prefix survival probability estimates and a hardware-aware scheduler to adjust verification length dynamically.
- On Qwen3-4B, 8B, and 14B, it improves accepted length by 30.9%, 26.7%, and 30.0% over Eagle3, and by 16% to 18% over DFlash.
- On real traffic for DeepSeek-V4-Flash, it delivers 60% to 85% speedups per user over MTP-1, and 57% to 78% in V4-Pro.
- It automatically shortens verification length in highly concurrent environments, preventing wasted batch capacity and expanding the Pareto frontier.
- It releases the DSpark checkpoint and the DeepSpec training repository as open source.
Paper links
External research summaries. These are not HDATF publications or measured product results.