Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings
- Published
- Source
- arXiv
- Paper number
- 375
- Field
- LLMs / NLP
- arXiv ID
- 2606.07502
Key points
- It discovers representation collapse in LLM embeddings, where they align with high-frequency meaningless tokens.
- It identifies the edge-spectrum subspace of the unembedding matrix as the cause.
- EmbedFilter removes that subspace with a simple linear transform and does not require additional training.
- It improves MTEB performance by up to 14.1% and works across different LLM backbones.
- It also saves storage and speeds up retrieval through distance-preserving dimensionality reduction.
- It outperforms existing correction methods such as BERT-whitening in unsupervised settings.
Paper links
External research summaries. These are not HDATF publications or measured product results.