AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding
- Published
- Source
- arXiv
- Paper number
- 767
- Field
- LLMs / NLP
- arXiv ID
- 2607.25852
Key points
- The authors empirically show that the best drafter depends on the task type in real service settings, such as chat versus code or math.
- DFly improves block diffusion quality with a hybrid target-conditioned backbone and an autoregressive head conditioned on predecessors.
- D-cut treats verification as a shared batch-level resource and adjusts verification depth adaptively based on load.
- On Hy3-A21B, it increases average accepted length by about 30 percent and achieves the highest throughput across concurrency levels from 4 to 64.
- It records 1.98 to 2.40 times the throughput of autoregressive decoding and 10.5 to 11.8 percent higher throughput than DFlash.
Paper links
External research summaries. These are not HDATF publications or measured product results.