Back to feed
arXiv cs.LG
arXiv cs.LG
7/13/2026
Sticky Routing: Training MoE Models for Memory-Efficient Inference

Sticky Routing: Training MoE Models for Memory-Efficient Inference

Short summary

Researchers propose StickyMoE, a differentiable routing consistency loss that penalizes abrupt expert switches between adjacent tokens in Mixture-of-Experts models. This encourages routers to maintain expert assignments across semantically coherent spans, reducing weight swapping between slow storage and fast memory on edge devices. Experiments show StickyMoE reduces expert switch rate by up to 60% with less than 4% perplexity degradation, Pareto-dominating post-hoc fine-tuning approaches.

  • StickyMoE adds a routing consistency loss to reduce expert switching between adjacent tokens
  • Reduces expert switch rate by up to 60% with under 4% perplexity degradation
  • Pareto-dominates post-hoc fine-tuning on quality-locality frontier with no architectural changes

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more