The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

Published
Source
arXiv
Paper number
241
Field
AI / General
arXiv ID
2605.26494

Key points

  • A MoE architecture with about 10 billion active parameters achieves both cost efficiency and competitive performance.
  • It builds a large synthetic-data pipeline for a wide range of agent tool environments, including code, search, spreadsheet, and slide generation.
  • Interleaved-thinking SFT trains trajectories where reasoning and action alternate, enabling the model to handle complex agent loops.
  • CISPO policy optimization and agent RL strengthen agent capability by modeling interaction with the environment as an MDP.
  • It demonstrates strong performance across practical domains such as financial analysis, document management, and role playing.
  • A three-axis scaling strategy for reasoning data, across query side, response side, and training side, improves out-of-distribution generalization.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)