Normalized Low-Rank Adaptation

Published
Source
arXiv
Paper number
1052
Field
Machine Learning
arXiv ID
2608.31036

Key points

  • It identified that LoRA's early learning is governed by the down-projection A and showed that simply normalizing A along the rank dimension improves optimization.
  • It raised the supervised fine-tuning average from 37.93 (LoRA) to 43.37, and even NoRA-init, which applies normalization only at initialization, reached 42.38.
  • In reinforcement learning for mathematical reasoning (RLVR), it achieved an overall average of 44.4%, exceeding LoRA (42.8%), and remained stable in contrast to SVD-based PiSSA, which collapsed to 0.2% under the same settings.
  • Post-adaptation retention of existing knowledge (average change across 3 evaluations including MMLU) was +0.02, showing much less forgetting than LoRA (-0.56) or MiSS (-0.70).
  • Applying the same normalization principle to DoRA also raised its score from 38.30 to 41.40, confirming that the idea is not limited to the LoRA family.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)