Back to feed
arXiv cs.CL
arXiv cs.CL
7/17/2026
Polestar: Drift-Aware Cache Calibration and Token Commitment for Efficient Inference of Diffusion LLMs

Polestar: Drift-Aware Cache Calibration and Token Commitment for Efficient Inference of Diffusion LLMs

Short summary

Polestar is a training-free inference framework for diffusion LLMs that uses token representation drift as a unified signal to address KV-cache reuse and decoding parallelism challenges. Polestar-Cache identifies stale cache positions for sparse refreshes, while Polestar-Commit detects sharp drift events to identify commit-ready tokens. Across math and coding benchmarks, it achieves up to 10.73% accuracy improvement, 3.7x higher throughput, and 3.67 tokens per forward pass, setting a new state-of-the-art on the accuracy-throughput Pareto frontier.

  • Polestar uses token representation drift to jointly solve KV-cache reuse and decoding parallelism for diffusion LLMs
  • Polestar-Cache refreshes stale KV-cache positions; Polestar-Commit identifies commit-ready tokens via drift events
  • Achieves up to 10.73% accuracy improvement and 3.7x throughput gain over existing baselines

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more