
Hierarchical Grading in Large Language Models
Short summary
This paper introduces Graded Large Language Models (GLLMs), an algebraic framework that adds a grading structure to transformer representations, propagating weighted scalar actions through embeddings, attention, and training objectives. The framework uses geometric invariant theory to identify optimal grades via a Kempf-Ness functional, with the standard transformer appearing as a boundary point of a larger graded family. The authors prove a minimax separation showing exponential advantage for graded priors under level-stratified targets, and the grading compiles to a standard transformer post-training with no inference overhead.
- •Introduces GLLMs: algebraic grading framework for transformer representations using geometric invariant theory
- •Proves minimax separation showing exponential advantage over uniform architecture for stratified targets
- •Grading absorbed into learned parameters post-training, yielding identical inference cost to standard transformers
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
