Back to feed
arXiv cs.LG
arXiv cs.LG
7/7/2026
GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech

GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech

Short summary

GRAFT enables zero-shot text-to-speech systems to control pronunciation on a per-word basis using reference audio samples. Testing across five languages shows 22–39% reductions in target-word pronunciation errors while preserving speaker identity and naturalness.

  • Per-word pronunciation control from reference audio without explicit phoneme conditioning
  • 22–39% reduction in target-word pronunciation errors across 5 languages
  • Maintains speaker similarity and naturalness while fixing rare words and technical terms

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more