arXiv cs.LG
7/13/2026

Director: Accelerating Distributed MoE Serving via Online Proactive Expert Placement
Short summary
Director is a distributed MoE serving system that minimizes end-to-end latency through prediction-driven, online expert placement. It uses lightweight predictors or quantized replicas to forecast expert activation patterns for incoming requests, then migrates experts during compute-bound phases with near-zero downtime. Experiments show 11-55% latency reduction for popular MoE models like Mistral, DeepSeek, and Qwen compared to existing approaches.
- •Prediction-driven online expert placement reduces MoE serving latency 11-55%
- •Near-zero downtime migration executed during compute-bound phases
- •Polynomial-time optimizer achieves (1+ε) approximation under capacity constraints
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
