Kwai Keye-VL-2.0 Technical Report
- Published
- Source
- arXiv
- Paper number
- 382
- Field
- Computer Vision
- arXiv ID
- 2606.10651
Key points
- It is a 30B MoE model with 3B active parameters, designed as an efficiently deployable multimodal foundation model.
- It applies DeepSeek Sparse Attention to a GQA-based architecture for the first time, enabling 256K lossless context.
- Cross-Modal MOPD integrates video, code, tools, and search into a single MoE backbone.
- Context-RL and Video-RL reduce visual hallucination and stabilize long-horizon decision making.
- It achieves state-of-the-art results on TimeLens, with 58.5 mIoU on ActivityNet, and on LongVideoBench, with 74.1.
- It is also strong on agent benchmarks such as LiveCodeBench v6 at 64.2 and OJBench at 71.5.
Paper links
External research summaries. These are not HDATF publications or measured product results.