Kimi K3: Open Frontier Intelligence

Published
Source
arXiv
Paper number
727
Field
LLMs / NLP
arXiv ID
2607.24653

Key points

  • Kimi K3 is a mixture-of-experts model that activates 104 billion of its 2.8 trillion total parameters and supports visual input and a 1-million-token context.
  • It combines Kimi Delta Attention, Attention Residuals, and Stable LatentMoE, using 16 of 896 experts for each token.
  • Reinforcement learning across general reasoning, agentic work, coding, and multiple reasoning-effort levels strengthened long-horizon work and compositional generalization.
  • Its overall scaling efficiency is approximately 2.5 times higher than Kimi K2's, and all weights are released for use in frontier-level model research and deployment.
  • Although it was stronger than other open and commercial models in the evaluation suite, it still lagged behind top closed models such as Claude Fable 5 and GPT-5.6 Sol.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)