arXiv cs.CL
7/21/2026

SpecLA: Efficient Speculative Decoding for Linear-Attention Models
Short summary
SpecLA is a speculative decoding runtime designed for stateful linear-attention models, addressing challenges that existing Transformer-focused speculative systems don't handle. It uses topology-aware kernels for chain and tree verification, compact state recovery factors, and confidence-pruned EAGLE-style drafters. On an NVIDIA H100 with a GDN-1.3B target, SpecLA achieves up to 1.70x end-to-end speedup over autoregressive decoding.
- •Speculative decoding runtime built specifically for stateful linear-attention models
- •Topology-aware verification kernels plus EAGLE-style confidence-pruned drafting
- •1.70x end-to-end speedup over autoregressive decoding on NVIDIA H100 with GDN-1.3B
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



