arXiv cs.LG
7/13/2026

Sticky Routing: Training MoE Models for Memory-Efficient Inference
Short summary
Researchers propose StickyMoE, a differentiable routing consistency loss that penalizes abrupt expert switches between adjacent tokens in Mixture-of-Experts models. This encourages routers to maintain expert assignments across semantically coherent spans, reducing weight swapping between slow storage and fast memory on edge devices. Experiments show StickyMoE reduces expert switch rate by up to 60% with less than 4% perplexity degradation, Pareto-dominating post-hoc fine-tuning approaches.
- •StickyMoE adds a routing consistency loss to reduce expert switching between adjacent tokens
- •Reduces expert switch rate by up to 60% with under 4% perplexity degradation
- •Pareto-dominates post-hoc fine-tuning on quality-locality frontier with no architectural changes
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
