Parallax: Parameterized Local Linear Attention for Language Modeling

Published
Source
arXiv
Paper number
267
Field
Machine Learning
arXiv ID
2605.29157

Key points

  • Perplexity: On datasets such as LAMBADA and WikiText, Parallax achieved consistently lower perplexity than Transformer baselines, a measure of how well the model predicts samples.
  • Downstream tasks: On zero-shot benchmarks such as HellaSwag and ARC, which cover question answering and commonsense reasoning, Parallax showed clear performance gains.
  • Control experiments: To confirm that the gains were not simply due to more parameters or more compute, the authors built parameter-matched and compute-matched versions of the standard Transformer. Parallax outperformed both, suggesting that the improvement comes from the parameterized local linear mechanism itself.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)