Normalized Low-Rank Adaptation
- Published
- Source
- arXiv
- Paper number
- 1052
- Field
- Machine Learning
- arXiv ID
- 2608.31036
Key points
- It identified that LoRA's early learning is governed by the down-projection A and showed that simply normalizing A along the rank dimension improves optimization.
- It raised the supervised fine-tuning average from 37.93 (LoRA) to 43.37, and even NoRA-init, which applies normalization only at initialization, reached 42.38.
- In reinforcement learning for mathematical reasoning (RLVR), it achieved an overall average of 44.4%, exceeding LoRA (42.8%), and remained stable in contrast to SVD-based PiSSA, which collapsed to 0.2% under the same settings.
- Post-adaptation retention of existing knowledge (average change across 3 evaluations including MMLU) was +0.02, showing much less forgetting than LoRA (-0.56) or MiSS (-0.70).
- Applying the same normalization principle to DoRA also raised its score from 38.30 to 41.40, confirming that the idea is not limited to the LoRA family.
Paper links
External research summaries. These are not HDATF publications or measured product results.