Back to feed
arXiv cs.CL
arXiv cs.CL
8/4/2026
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis

Short summary

DLLM-TTS formulates text-to-speech as conditional block discrete diffusion over neural audio codec tokens, decomposing sequences into blocks with masked diffusion within each block. A 0.6B-parameter model trained on 20K hours achieves competitive performance on Seed-TTS-eval with a real-time factor of 0.15. The approach balances the intelligibility of autoregressive models with the speed of non-autoregressive methods through parallel token prediction within blocks.

  • Block discrete diffusion enables parallel TTS generation with 0.15 real-time factor
  • 0.6B model trained on only 20K hours matches larger autoregressive systems
  • Combines local acoustic coherence with global text-speech alignment

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more