Modular TTT: Rethinking Test-Time Training as Composable Modules

Published
Source
arXiv
Paper number
854
Field
Machine Learning
arXiv ID
2608.07110

Key points

  • It expresses the internal learner of TTT as a DAG and builds a modular framework where loss functions, learning rates, weight decay, and normalization can be adjusted independently.
  • Small learning-rate initialization, weight decay, and single-layer nonlinearity, specifically SiLU, consistently improve performance.
  • In contrast, deep fast-weight networks and normalization hurt performance because activation values become too large. Residual connections and gating have little effect.
  • The best variant trains 410M and 1.45B models on 100B tokens and reaches performance on par with Gated DeltaNet.
  • Developed at ByteDance, it releases an optimized implementation that improves training throughput several times over prior TTT systems.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)