Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects

Published
Source
arXiv
Paper number
744
Field
Machine Learning
arXiv ID
2607.24645

Key points

  • It measures the geometry of how output logits change when a feature is removed.
  • Most features do not provide one stable control direction.
  • Value-like features more often show low-dimensional structure, but their effects are spread across multiple directions.
  • Pointer-like features produce mostly scattered effects.
  • The paper compares ReLU, TopK, and Matryoshka Batch TopK on Gemma-2-2B.
  • Even when a feature is interpretable and causal, that does not guarantee stable control.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)