r/MachineLearning
6/29/2026
![I'm trying to implement CALM paper, and I have some questions. [P]](https://preview.redd.it/kr4u22yfx8ah1.png?width=140&height=83&auto=webp&s=784c46c82400e669571b4d8a7dcdc997ad0fba57)
I'm trying to implement CALM paper, and I have some questions. [P]
Short summary
An ML engineer details their struggle implementing Kyutai's Pocket TTS from paper, facing poor inference quality despite low training losses on smaller datasets. Key challenges include gradient explosion, exposure bias, and a fundamental trade-off: placing voice tokens near target latents produces better voice cloning but worse text generation. Seeking guidance on data scaling, model architecture, and resource commitment before larger training runs.
- •Implementing Pocket TTS on LJSpeech/LibriSpeech yields poor inference quality despite low training losses
- •Gradient explosion and conflicting trade-offs between text generation and voice cloning fidelity are primary blockers
- •Seeking recommendations on data scaling strategy, architecture fixes, and resource commitment before scaling to larger clusters
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



