arXiv cs.LG
7/1/2026

Hierarchical Global Attention (HGA)
Short summary
HGA introduces hierarchical two-level routing for transformer attention, reducing memory requirements while maintaining 97% of dense attention quality with only 3% sparsity. A 30B model runs at 64K token context on a single 32GB GPU with no retraining—it's a drop-in replacement for existing checkpoints.
- •Novel hierarchical routing system reduces GPU memory without retraining
- •Preserves pretrained parameters; works with existing model checkpoints
- •Tested on 64K context with 3% sparsity overhead and minimal quality loss
Generated with AI, which can make mistakes.
Is this a good recommendation for you?

