ICA Lens: Interpreting Language Models Without Training Another Dictionary
- Published
- Source
- arXiv
- Paper number
- 390
- Field
- Machine Learning
- arXiv ID
- 2606.11722
Key points
- With GPU-parallel FastICA, row normalization, and a p95-LIM convergence criterion, it increases accepted layers in GPT-2 Small by 400 percent and reduces iterations by 21.5 percent.
- ICA directions have significantly higher non-Gaussianity than random projections, and deeper layers show increasing context dependence, confirming the expected pattern of effective receptive field growth.
- It is competitive with open SAE models on SAEBench Sparse Probing, and it outperforms SAE under small to medium budgets on Targeted Probe Perturbation.
- ICA components capture a wide range of linguistic structures, including lexical form, local syntax, discourse context, semantic ambiguity, and long-range repetition.
- It proposes ICA as a complementary first lens rather than a replacement for SAE and positions it as a cost-effective exploratory tool.
Paper links
External research summaries. These are not HDATF publications or measured product results.