ExpertPlex: A High-Goodput Disaggregated Serving System for MoE LLMs with Adaptive Persistent Kernels
- Published
- Source
- arXiv
- Paper number
- 679
- Field
- Distributed Systems
- arXiv ID
- 2607.18002
Key points
- More than 95 percent of the weights in an MoE language model belong to the expert part, so duplicating the full model for both prefill and decode wastes a lot of memory.
- The key idea is a hybrid structure in which the heavy experts are shared across the two stages, while less than 5 percent attention is assigned separate GPUs by stage.
- An adaptive persistent kernel, or APK, schedules at matrix-multiplication tile boundaries so that urgent decode work can take over and free capacity can return to prefill without relaunching kernels or involving the CPU.
- A one-way MoE communication path initiated from the attention side removes MoE polling, avoids deadlock and cross-stage network interference, and overlaps communication in one stage with computation in the other.
- In serving experiments with MiniMax-M2.7 and GLM-5.1-FP8, effective throughput rises by 5.65 times over chunked prefill, by 2.01 times over instance-level separation, and by 1.66 times over Green Context colocation.
- Because the preemption interval depends on tile execution time rather than token count, even a prefill with thousands of input tokens can be interrupted by short decode jobs.
Paper links
External research summaries. These are not HDATF publications or measured product results.