DeepLoop: Depth Scaling for Looped Transformers

Published
Source
arXiv
Paper number
653
Field
Machine Learning
arXiv ID
2607.13491

Key points

  • The authors mathematically prove that the gradient accumulation in looped transformers with shared layers makes the usual DeepNorm scaling insufficient.
  • Simply increasing the exponent from one quarter to one half, with alpha equal to (2N) to the one-half power and beta equal to (8N) to the negative one-half power, stabilizes training as the number of loops grows.
  • On GPT-2 small and medium scales, increasing the number of loops consistently improves performance over prior methods.
  • When applied to the hierarchical reasoning model, or HRM, on the ARC-AGI reasoning benchmark, accuracy rises from 36.5 percent to 39.75 percent, which is a gain of 3.25 percentage points.
  • Because it changes only the scaling formula and adds no extra parameters, the method is essentially free to apply.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)