arXiv cs.LG
7/29/2026

Mechanisms of Width Scaling in Normalized Residual Networks: The Effective Alignment Dimension
Short summary
This paper introduces the effective alignment dimension, a measurable quantity describing signal-noise geometry of activation gradients in residual network width scaling. The authors derive a finite-sample upper bound on misalignment probability between training and test gradients, requiring only finite second moments and nonzero population gradient. Experiments across LLaMA-style Transformers, Pythia, and ResNet-20 confirm that wider models exhibit larger alignment dimensions and lower misalignment, with the statistic predicting held-out loss changes.
- •Introduces effective alignment dimension for guiding residual network width expansion
- •Derives finite-sample misalignment probability bound without spectral assumptions
- •Experiments on LLaMA, Pythia, and ResNet-20 confirm alignment statistic predicts test-loss changes
Generated with AI, which can make mistakes.
Is this a good recommendation for you?