Transformer²: Self-adaptive LLMs
- Published
- Source
- arXiv
- Paper number
- 016
- Field
- Architecture / Adaptation
- arXiv ID
- 2501.06252
Key points
- SVF replaces additive adapters with learned scaling on the singular values of each weight matrix, changing task behavior while preserving the base singular vectors.
- Transformer² uses dispatch-and-apply inference, and experts can be selected directly, by a learned classifier expert, or mixed through few-shot CEM optimization.
- The paper reports that SVF and Transformer² outperform LoRA with less than 10% of LoRA's parameter count, including gains on math, code, reasoning, and vision-language benchmarks.
- The strongest evidence is not just higher scores, but expert composability and transfer across models, which points to reusable skill vectors rather than one-off fine-tuning.
Paper links
External research summaries. These are not HDATF publications or measured product results.