MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
- Published
- Source
- arXiv
- Paper number
- 1181
- Field
- LLMs / NLP
- arXiv ID
- 2610.11959
Key points
- Asynchronous training that consumes 1,568 samples and 2.7-3.7B tokens per step handled long tasks with context lengths up to 1M tokens.
- Environments across coding, general workflows, visual tasks, and cybersecurity were mixed with diverse agent harnesses to improve generalization.
- Instead of binary pass/fail grading, groupwise grading that compares answers within a group produced more precise reward signals.
- Freezing the MoE router and building multi-layer defenses against reward hacking kept large-scale RL training stable.
- Training dynamics, RL environments, and the RL framework were fully open-sourced, creating a reproducible baseline for large-scale agentic RL research.
Paper links
External research summaries. These are not HDATF publications or measured product results.