Rethinking Expressivity and Efficiency in Test-Time Training

Published
Source
arXiv
Paper number
987
Field
Machine Learning
arXiv ID
2608.21308

Key points

  • It derived a closed-form state-transition equation that enables chunk-parallel training while retaining per-token learning rates, momentum, and decay.
  • It demonstrated long-context extrapolation by maintaining over 90% accuracy on a needle-in-a-haystack test with inputs eight times longer than those used in training.
  • Across 14 LongBench tasks, its average performance exceeded that of subquadratic-compute architectures such as Mamba2 and LaCT.
  • When attached as a parallel branch to Qwen3VL-2B, training only the TTT component achieved video-understanding performance comparable to full fine-tuning.
  • Because a single set of fast weights is shared within each chunk, some fine-grained causal relationships are still missed, and experiments were limited to a maximum of 1.3 billion parameters, requiring validation on larger models.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)