Full-Pipeline Inference Optimization for MiMo-V2.5 Series: Pushing Hybrid SWA Efficiency to the Limit
- Published
- Source
- arXiv
- Paper number
- 656
- Field
- AI / General
- arXiv ID
- 2607.13095
Key points
- It separates the KV cache for full attention and sliding window attention, compressing SWA storage to the theoretical limit of O(W), which is about 7x more efficient than before.
- It builds GCache, an RDMA-based distributed cache infrastructure that shares KV caches efficiently across multiple nodes.
- It redesigns the prefix cache tree to be SWA-aware and fixes accuracy bugs that the old prefix matching approach caused under SWA.
- It also optimizes the multimodal pipeline end to end, including GPU image preprocessing, parallel video decoding from 156 seconds to 23 seconds for a 1-hour video, and cross-request batching for the encoder.
- It improves MTP, or multi-token prediction, so that it also works in prefill and speeds up the first 128 decoded tokens by 2.3x.
- It shows that resolving only the NUMA contention improves end-to-end performance by about 10%, which means system-level details matter a lot.
Paper links
External research summaries. These are not HDATF publications or measured product results.