Dev.to
7/2/2026

Voice cloning models, measured across five languages
Short summary
Comparative benchmark of four open-source voice-cloning models (OmniVoice, Chatterbox, VoxCPM2, Fish Audio S2 Pro) across five languages using speaker similarity, error rates, and inference speed. OmniVoice delivered strongest overall performance; VoxCPM2 excelled on Arabic; Fish Audio S2 showed high similarity but slower inference. Published as engineering benchmark using Google FLEURS, not human evaluation.
- •OmniVoice int8 was strongest all-around model
- •VoxCPM2 bf16 excelled on Arabic speaker matching
- •Fish Audio S2 Pro fastest on German/Arabic but slowest inference speed
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



