arXiv cs.CL
7/17/2026

Polestar: Drift-Aware Cache Calibration and Token Commitment for Efficient Inference of Diffusion LLMs
Short summary
Polestar is a training-free inference framework for diffusion LLMs that uses token representation drift as a unified signal to address KV-cache reuse and decoding parallelism challenges. Polestar-Cache identifies stale cache positions for sparse refreshes, while Polestar-Commit detects sharp drift events to identify commit-ready tokens. Across math and coding benchmarks, it achieves up to 10.73% accuracy improvement, 3.7x higher throughput, and 3.67 tokens per forward pass, setting a new state-of-the-art on the accuracy-throughput Pareto frontier.
- •Polestar uses token representation drift to jointly solve KV-cache reuse and decoding parallelism for diffusion LLMs
- •Polestar-Cache refreshes stale KV-cache positions; Polestar-Commit identifies commit-ready tokens via drift events
- •Achieves up to 10.73% accuracy improvement and 3.7x throughput gain over existing baselines
Generated with AI, which can make mistakes.
Is this a good recommendation for you?

