Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects
- Published
- Source
- arXiv
- Paper number
- 744
- Field
- Machine Learning
- arXiv ID
- 2607.24645
Key points
- It measures the geometry of how output logits change when a feature is removed.
- Most features do not provide one stable control direction.
- Value-like features more often show low-dimensional structure, but their effects are spread across multiple directions.
- Pointer-like features produce mostly scattered effects.
- The paper compares ReLU, TopK, and Matryoshka Batch TopK on Gemma-2-2B.
- Even when a feature is interpretable and causal, that does not guarantee stable control.
Paper links
External research summaries. These are not HDATF publications or measured product results.