arXiv cs.CL
8/4/2026

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis
Short summary
DLLM-TTS formulates text-to-speech as conditional block discrete diffusion over neural audio codec tokens, decomposing sequences into blocks with masked diffusion within each block. A 0.6B-parameter model trained on 20K hours achieves competitive performance on Seed-TTS-eval with a real-time factor of 0.15. The approach balances the intelligibility of autoregressive models with the speed of non-autoregressive methods through parallel token prediction within blocks.
- •Block discrete diffusion enables parallel TTS generation with 0.15 real-time factor
- •0.6B model trained on only 20K hours matches larger autoregressive systems
- •Combines local acoustic coherence with global text-speech alignment
Generated with AI, which can make mistakes.
Is this a good recommendation for you?