arXiv cs.CL
7/17/2026

Token Time Continuous Diffusion for Language Modeling
Short summary
TTCD is a diffusion language model that maps Gaussian noise to a token canvas in continuous space, avoiding the inaccuracy of discrete-space parallel sampling at high speedups. It introduces per-token times, allowing some tokens to denoise faster than others, which improves conditional generation modeling. A 160M parameter model trained on OpenWebText and self-distilled matches or exceeds comparable discrete models in unconditional generation and outperforms them in conditional generation and Sudoku solving.
- •TTCD operates in continuous space, avoiding discrete parallel sampling inaccuracy at high speedups
- •Per-token times allow differentiated denoising rates, improving conditional generation
- •160M parameter model outperforms comparable discrete models in conditional generation and Sudoku solving
Generated with AI, which can make mistakes.
Is this a good recommendation for you?


