DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation

Published
Source
arXiv
Paper number
575
Field
AI / General
arXiv ID
2607.05147

Key points

  • A semi-autoregressive architecture injects token-to-token dependence with a parallel backbone and a lightweight RNN head, while keeping per-token probabilities as exact softmax values.
  • Confidence-scheduled verification uses position-wise prefix survival probability estimates and a hardware-aware scheduler to adjust verification length dynamically.
  • On Qwen3-4B, 8B, and 14B, it improves accepted length by 30.9%, 26.7%, and 30.0% over Eagle3, and by 16% to 18% over DFlash.
  • On real traffic for DeepSeek-V4-Flash, it delivers 60% to 85% speedups per user over MTP-1, and 57% to 78% in V4-Pro.
  • It automatically shortens verification length in highly concurrent environments, preventing wasted batch capacity and expanding the Pareto frontier.
  • It releases the DSpark checkpoint and the DeepSpec training repository as open source.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)