Back to feed
arXiv cs.CL
arXiv cs.CL
7/21/2026
SpecLA: Efficient Speculative Decoding for Linear-Attention Models

SpecLA: Efficient Speculative Decoding for Linear-Attention Models

Short summary

SpecLA is a speculative decoding runtime designed for stateful linear-attention models, addressing challenges that existing Transformer-focused speculative systems don't handle. It uses topology-aware kernels for chain and tree verification, compact state recovery factors, and confidence-pruned EAGLE-style drafters. On an NVIDIA H100 with a GDN-1.3B target, SpecLA achieves up to 1.70x end-to-end speedup over autoregressive decoding.

  • Speculative decoding runtime built specifically for stateful linear-attention models
  • Topology-aware verification kernels plus EAGLE-style confidence-pruned drafting
  • 1.70x end-to-end speedup over autoregressive decoding on NVIDIA H100 with GDN-1.3B

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more