Kwai Keye-VL-2.0 Technical Report

Published
Source
arXiv
Paper number
382
Field
Computer Vision
arXiv ID
2606.10651

Key points

  • It is a 30B MoE model with 3B active parameters, designed as an efficiently deployable multimodal foundation model.
  • It applies DeepSeek Sparse Attention to a GQA-based architecture for the first time, enabling 256K lossless context.
  • Cross-Modal MOPD integrates video, code, tools, and search into a single MoE backbone.
  • Context-RL and Video-RL reduce visual hallucination and stabilize long-horizon decision making.
  • It achieves state-of-the-art results on TimeLens, with 58.5 mIoU on ActivityNet, and on LongVideoBench, with 74.1.
  • It is also strong on agent benchmarks such as LiveCodeBench v6 at 64.2 and OJBench at 71.5.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)