arXiv cs.LG
7/30/2026

Between Gradient and Natural Gradient: A Continuum of LoRA Initializations
Short summary
This paper unifies existing LoRA initialization schemes into a two-parameter family (ULoRA) governed by spectral whitening and diagonal exponents, showing they are points on a continuum rather than distinct methods. The best operating point is task-dependent and often lies strictly inside the family. A search-free variant, ULoRA-Auto, selects per-layer exponents from spectral statistics and matches or exceeds full fine-tuning on all five GLUE tasks with RoBERTa-base while remaining competitive on GSM8K with LLaMA-2-7B.
- •Gradient-based and curvature-whitened LoRA initializations are points on a single two-parameter continuum
- •Optimal preconditioning strength is task-dependent, not a fixed design choice
- •ULoRA-Auto selects per-layer exponents from spectral statistics with no search cost, matching full fine-tuning on GLUE
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
