Full-Pipeline Inference Optimization for MiMo-V2.5 Series: Pushing Hybrid SWA Efficiency to the Limit

Published
Source
arXiv
Paper number
656
Field
AI / General
arXiv ID
2607.13095

Key points

  • It separates the KV cache for full attention and sliding window attention, compressing SWA storage to the theoretical limit of O(W), which is about 7x more efficient than before.
  • It builds GCache, an RDMA-based distributed cache infrastructure that shares KV caches efficiently across multiple nodes.
  • It redesigns the prefix cache tree to be SWA-aware and fixes accuracy bugs that the old prefix matching approach caused under SWA.
  • It also optimizes the multimodal pipeline end to end, including GPU image preprocessing, parallel video decoding from 156 seconds to 23 seconds for a 1-hour video, and cross-request batching for the encoder.
  • It improves MTP, or multi-token prediction, so that it also works in prefill and speeds up the first 128 decoded tokens by 2.3x.
  • It shows that resolving only the NUMA contention improves end-to-end performance by about 10%, which means system-level details matter a lot.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)