
Fingerprint, Not Blueprint: How Positional Schemes Set the Default Spectral Algebra of Attention
Short summary
This paper analyzes the eigenspectrum of attention head operators (M = W_q^T W_k) across seven pretrained models using three positional schemes (RoPE, learned-absolute, ALiBi). It finds that positional encoding determines the default spectral algebra: RoPE produces rotational signatures in previous-token heads while learned-absolute and ALiBi produce non-rotational ones, with perfect model-level separation. Dynamically, these spectral signatures emerge after circuit formation during training, acting as a fingerprint of function rather than a hard constraint, though models can reroute around any single spectral ban at the cost of slower formation.
- •Positional scheme (RoPE vs learned-absolute vs ALiBi) determines the default spectral algebra of attention heads, with perfect separation across seven models
- •Spectral rotational signatures in previous-token heads emerge after circuit formation during training, confirming they are fingerprints of function not precursors
- •No single spectral channel is necessary—constrained training reroutes around bans with capability intact but significant formation delay
Generated with AI, which can make mistakes.
Is this a good recommendation for you?

