MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement

Published
Source
arXiv
Paper number
1181
Field
LLMs / NLP
arXiv ID
2610.11959

Key points

  • Asynchronous training that consumes 1,568 samples and 2.7-3.7B tokens per step handled long tasks with context lengths up to 1M tokens.
  • Environments across coding, general workflows, visual tasks, and cybersecurity were mixed with diverse agent harnesses to improve generalization.
  • Instead of binary pass/fail grading, groupwise grading that compares answers within a group produced more precise reward signals.
  • Freezing the MoE router and building multi-layer defenses against reward hacking kept large-scale RL training stable.
  • Training dynamics, RL environments, and the RL framework were fully open-sourced, creating a reproducible baseline for large-scale agentic RL research.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)