Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings

Published
Source
arXiv
Paper number
375
Field
LLMs / NLP
arXiv ID
2606.07502

Key points

  • It discovers representation collapse in LLM embeddings, where they align with high-frequency meaningless tokens.
  • It identifies the edge-spectrum subspace of the unembedding matrix as the cause.
  • EmbedFilter removes that subspace with a simple linear transform and does not require additional training.
  • It improves MTEB performance by up to 14.1% and works across different LLM backbones.
  • It also saves storage and speeds up retrieval through distance-preserving dimensionality reduction.
  • It outperforms existing correction methods such as BERT-whitening in unsupervised settings.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)