arXiv cs.LG
7/9/2026

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation
Short summary
TriRoute introduces a single lightweight controller that jointly manages three inference-cost axes—attention resolution, expert selection, and KV-cache bit-width—for every token at every layer. The controller trains end-to-end under a Lagrangian budget constraint and addresses cross-axis routing-collapse cascades via per-axis normalization and coupling-aware balancing loss. On 160M–1.3B parameter models, TriRoute Pareto-dominates independent MoD+MoE+KV-quantization combinations at matched FLOPs and memory while preserving robustness on rare entities, code, and arithmetic.
- •Single controller jointly routes attention mode, FFN experts, and KV-cache precision per token per layer
- •Identifies and fixes cross-axis routing-collapse cascade in naive joint training
- •Pareto-dominates independent axis-optimization at matched compute and memory on 160M–1.3B models
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
