Back to feed
arXiv cs.LG
arXiv cs.LG
7/29/2026
Mechanisms of Width Scaling in Normalized Residual Networks: The Effective Alignment Dimension

Mechanisms of Width Scaling in Normalized Residual Networks: The Effective Alignment Dimension

Short summary

This paper introduces the effective alignment dimension, a measurable quantity describing signal-noise geometry of activation gradients in residual network width scaling. The authors derive a finite-sample upper bound on misalignment probability between training and test gradients, requiring only finite second moments and nonzero population gradient. Experiments across LLaMA-style Transformers, Pythia, and ResNet-20 confirm that wider models exhibit larger alignment dimensions and lower misalignment, with the statistic predicting held-out loss changes.

  • Introduces effective alignment dimension for guiding residual network width expansion
  • Derives finite-sample misalignment probability bound without spectral assumptions
  • Experiments on LLaMA, Pythia, and ResNet-20 confirm alignment statistic predicts test-loss changes

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more