Rethinking Expressivity and Efficiency in Test-Time Training
- Published
- Source
- arXiv
- Paper number
- 987
- Field
- Machine Learning
- arXiv ID
- 2608.21308
Key points
- It derived a closed-form state-transition equation that enables chunk-parallel training while retaining per-token learning rates, momentum, and decay.
- It demonstrated long-context extrapolation by maintaining over 90% accuracy on a needle-in-a-haystack test with inputs eight times longer than those used in training.
- Across 14 LongBench tasks, its average performance exceeded that of subquadratic-compute architectures such as Mamba2 and LaCT.
- When attached as a parallel branch to Qwen3VL-2B, training only the TTT component achieved video-understanding performance comparable to full fine-tuning.
- Because a single set of fast weights is shared within each chunk, some fine-grained causal relationships are still missed, and experiments were limited to a maximum of 1.3 billion parameters, requiring validation on larger models.
Paper links
External research summaries. These are not HDATF publications or measured product results.