Fast Weight Attention for Continual Learning

Published
Source
arXiv
Paper number
1050
Field
Machine Learning
arXiv ID
2608.27763

Key points

  • It reinterpreted state-update rules as online learning and developed a framework that separately analyzes temporal alignment, plasticity, forgetting, and rehearsal.
  • It showed that, in next-token prediction, the pair the state should learn is offset by one step: (previous key, current value).
  • It derived normalized first-order update equations for regression objectives (Falcon-1/2/3) and inner-product objectives (Falcon-1A/2A/3A), and provided recurrent, masked-parallel, and chunk-parallel forms.
  • It remained competitive with existing recurrent models in language-modeling evaluations.
  • In extrapolation to 33–48-digit addition, Falcon-3A.3 achieved 87.2%, outperforming both RetNet and Transformer baselines.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)