arXiv cs.LG
7/7/2026

GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech
Short summary
GRAFT enables zero-shot text-to-speech systems to control pronunciation on a per-word basis using reference audio samples. Testing across five languages shows 22–39% reductions in target-word pronunciation errors while preserving speaker identity and naturalness.
- •Per-word pronunciation control from reference audio without explicit phoneme conditioning
- •22–39% reduction in target-word pronunciation errors across 5 languages
- •Maintains speaker similarity and naturalness while fixing rare words and technical terms
Generated with AI, which can make mistakes.
Is this a good recommendation for you?