Fast Weight Attention for Continual Learning
- Published
- Source
- arXiv
- Paper number
- 1050
- Field
- Machine Learning
- arXiv ID
- 2608.27763
Key points
- It reinterpreted state-update rules as online learning and developed a framework that separately analyzes temporal alignment, plasticity, forgetting, and rehearsal.
- It showed that, in next-token prediction, the pair the state should learn is offset by one step: (previous key, current value).
- It derived normalized first-order update equations for regression objectives (Falcon-1/2/3) and inner-product objectives (Falcon-1A/2A/3A), and provided recurrent, masked-parallel, and chunk-parallel forms.
- It remained competitive with existing recurrent models in language-modeling evaluations.
- In extrapolation to 33–48-digit addition, Falcon-3A.3 achieved 87.2%, outperforming both RetNet and Transformer baselines.
Paper links
External research summaries. These are not HDATF publications or measured product results.