AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis
Merged summary
TL;DR - AuEmoChat is a conversational speech synthesis framework that learns discrete emotion tokens from real emotional speech rather than relying on a few predefined labels. It aims to generate more contextually consistent and authentic emotional speech.
- AuEmoCodec learns a discrete emotion space using finite scalar quantization.
- AuEmoToMe merges redundant multimodal dialogue-history tokens while preserving emotion-relevant context.
- An autoregressive text-speech model predicts both target emotion and speech tokens.
- Experiments on NCSSD-EmCap show improvements over state-of-the-art CSS baselines in expressiveness and emotional authenticity.
Sources (1)
AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis
TL;DR - AuEmoChat is a conversational speech synthesis framework that learns discrete emotion tokens from real emotional speech rather than relying on a few predefined labels. It aims to generate more contextually consistent and authentic emotional speech.
- AuEmoCodec learns a discrete emotion space using finite scalar quantization.
- AuEmoToMe merges redundant multimodal dialogue-history tokens while preserving emotion-relevant context.
- An autoregressive text-speech model predicts both target emotion and speech tokens.
- Experiments on NCSSD-EmCap show improvements over state-of-the-art CSS baselines in expressiveness and emotional authenticity.