Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale Thinking Model
- Published
- Source
- arXiv
- Paper number
- 090
- Field
- LLMs / Reasoning
- arXiv ID
- 2510.18855
Key points
- The lack of an open-source one-trillion-parameter thinking model has made advanced reasoning capabilities less accessible to the broader AI research community.
- Scaling reinforcement learning (RL) to a one-trillion-parameter Mixture-of-Experts (MoE) model led to substantial instability caused by training-inference mismatch.
- Inefficiencies in handling long reasoning trajectories and system bottlenecks in RL infrastructure created major challenges at an unprecedented scale.
- They developed Ring-1T, a one-trillion-parameter MoE model trained through a multi-stage pipeline that includes Long Chain-of-Thought supervised fine-tuning and large-scale reinforcement learning.
- They introduced IcePop, an RL algorithm variant that stabilizes training by suppressing noisy gradient updates through bilateral masking correction for probability drift.
- They designed C3PO++ for efficient trajectory processing with dynamic, budget-controlled partitioning, and they built ASystem, a high-performance distributed RL infrastructure tuned for the trillion-parameter scale.
Paper links
External research summaries. These are not HDATF publications or measured product results.