ExpertPlex: A High-Goodput Disaggregated Serving System for MoE LLMs with Adaptive Persistent Kernels

Published
Source
arXiv
Paper number
679
Field
Distributed Systems
arXiv ID
2607.18002

Key points

  • More than 95 percent of the weights in an MoE language model belong to the expert part, so duplicating the full model for both prefill and decode wastes a lot of memory.
  • The key idea is a hybrid structure in which the heavy experts are shared across the two stages, while less than 5 percent attention is assigned separate GPUs by stage.
  • An adaptive persistent kernel, or APK, schedules at matrix-multiplication tile boundaries so that urgent decode work can take over and free capacity can return to prefill without relaunching kernels or involving the CPU.
  • A one-way MoE communication path initiated from the attention side removes MoE polling, avoids deadlock and cross-stage network interference, and overlaps communication in one stage with computation in the other.
  • In serving experiments with MiniMax-M2.7 and GLM-5.1-FP8, effective throughput rises by 5.65 times over chunked prefill, by 2.01 times over instance-level separation, and by 1.66 times over Green Context colocation.
  • Because the preemption interval depends on tile execution time rather than token count, even a prefill with thousands of input tokens can be interrupted by short decode jobs.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)