ICA Lens: Interpreting Language Models Without Training Another Dictionary

Published
Source
arXiv
Paper number
390
Field
Machine Learning
arXiv ID
2606.11722

Key points

  • With GPU-parallel FastICA, row normalization, and a p95-LIM convergence criterion, it increases accepted layers in GPT-2 Small by 400 percent and reduces iterations by 21.5 percent.
  • ICA directions have significantly higher non-Gaussianity than random projections, and deeper layers show increasing context dependence, confirming the expected pattern of effective receptive field growth.
  • It is competitive with open SAE models on SAEBench Sparse Probing, and it outperforms SAE under small to medium budgets on Targeted Probe Perturbation.
  • ICA components capture a wide range of linguistic structures, including lexical form, local syntax, discourse context, semantic ambiguity, and long-range repetition.
  • It proposes ICA as a complementary first lens rather than a replacement for SAE and positions it as a cost-effective exploratory tool.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)