Back to feed
arXiv cs.LG
arXiv cs.LG
7/13/2026
Director: Accelerating Distributed MoE Serving via Online Proactive Expert Placement

Director: Accelerating Distributed MoE Serving via Online Proactive Expert Placement

Short summary

Director is a distributed MoE serving system that minimizes end-to-end latency through prediction-driven, online expert placement. It uses lightweight predictors or quantized replicas to forecast expert activation patterns for incoming requests, then migrates experts during compute-bound phases with near-zero downtime. Experiments show 11-55% latency reduction for popular MoE models like Mistral, DeepSeek, and Qwen compared to existing approaches.

  • Prediction-driven online expert placement reduces MoE serving latency 11-55%
  • Near-zero downtime migration executed during compute-bound phases
  • Polynomial-time optimizer achieves (1+ε) approximation under capacity constraints

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more