xHC: Expanded Hyper-Connections

Published
Source
arXiv
Paper number
648
Field
Machine Learning
arXiv ID
2607.14530

Key points

  • It is the first method to extend Hyper-Connection beyond N=4 to N=16, because prior mHC methods suffered strong diminishing returns above N=4.
  • Temporal feature augmentation enriches write-back with nearby token information, and sparse residual updates at k=4 or N=16 balance cost and efficiency.
  • On an 18B MoE model, it delivers a +4.0 point gain over mHC while adding only 4.1% more FLOPs than vanilla.
  • The scaling law shows that vanilla requires 1.50x compute and mHC requires 1.19x compute to reach the same loss, using xHC as the reference.
  • xHC-Flash keeps sublayer memory traffic at 40C even with N=16, which is similar to mHC at N=4 with 34C.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)