Transformer²: Self-adaptive LLMs

Published
Source
arXiv
Paper number
016
Field
Architecture / Adaptation
arXiv ID
2501.06252

Key points

  • SVF replaces additive adapters with learned scaling on the singular values of each weight matrix, changing task behavior while preserving the base singular vectors.
  • Transformer² uses dispatch-and-apply inference, and experts can be selected directly, by a learned classifier expert, or mixed through few-shot CEM optimization.
  • The paper reports that SVF and Transformer² outperform LoRA with less than 10% of LoRA's parameter count, including gains on math, code, reasoning, and vision-language benchmarks.
  • The strongest evidence is not just higher scores, but expert composability and transfer across models, which points to reusable skill vectors rather than one-off fine-tuning.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)