AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding

Published
Source
arXiv
Paper number
767
Field
LLMs / NLP
arXiv ID
2607.25852

Key points

  • The authors empirically show that the best drafter depends on the task type in real service settings, such as chat versus code or math.
  • DFly improves block diffusion quality with a hybrid target-conditioned backbone and an autoregressive head conditioned on predecessors.
  • D-cut treats verification as a shared batch-level resource and adjusts verification depth adaptively based on load.
  • On Hy3-A21B, it increases average accepted length by about 30 percent and achieves the highest throughput across concurrency levels from 4 to 64.
  • It records 1.98 to 2.40 times the throughput of autoregressive decoding and 10.5 to 11.8 percent higher throughput than DFlash.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)