Back to feed
arXiv cs.LG
arXiv cs.LG
7/9/2026
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation

Short summary

TriRoute introduces a single lightweight controller that jointly manages three inference-cost axes—attention resolution, expert selection, and KV-cache bit-width—for every token at every layer. The controller trains end-to-end under a Lagrangian budget constraint and addresses cross-axis routing-collapse cascades via per-axis normalization and coupling-aware balancing loss. On 160M–1.3B parameter models, TriRoute Pareto-dominates independent MoD+MoE+KV-quantization combinations at matched FLOPs and memory while preserving robustness on rare entities, code, and arithmetic.

  • Single controller jointly routes attention mode, FFN experts, and KV-cache precision per token per layer
  • Identifies and fixes cross-axis routing-collapse cascade in naive joint training
  • Pareto-dominates independent axis-optimization at matched compute and memory on 160M–1.3B models

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more