RT by @_akhaliq: Our Confucius4-TTS paper is now on arXiv — and the open-source model has just…
TL;DR - Confucius4-TTS is an arXiv paper and upgraded open-source model for transcript-free, cross-lingual zero-shot text-to-speech. It targets high-quality voice generation and practical multilingual use.
- Supports multilingual and cross-lingual voice generation.
- Uses zero-shot TTS for voice cloning from an audio prompt.
- Removes the need for transcripts of prompt audio.
- The accompanying open-source model received a major upgrade.