Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale Thinking Model

Published
Source
arXiv
Paper number
090
Field
LLMs / Reasoning
arXiv ID
2510.18855

Key points

  • The lack of an open-source one-trillion-parameter thinking model has made advanced reasoning capabilities less accessible to the broader AI research community.
  • Scaling reinforcement learning (RL) to a one-trillion-parameter Mixture-of-Experts (MoE) model led to substantial instability caused by training-inference mismatch.
  • Inefficiencies in handling long reasoning trajectories and system bottlenecks in RL infrastructure created major challenges at an unprecedented scale.
  • They developed Ring-1T, a one-trillion-parameter MoE model trained through a multi-stage pipeline that includes Long Chain-of-Thought supervised fine-tuning and large-scale reinforcement learning.
  • They introduced IcePop, an RL algorithm variant that stabilizes training by suppressing noisy gradient updates through bilateral masking correction for probability drift.
  • They designed C3PO++ for efficient trajectory processing with dynamic, budget-controlled partitioning, and they built ASystem, a high-performance distributed RL infrastructure tuned for the trillion-parameter scale.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)