🛰️ Daily AI Frontier
‹ back to 2026-07-20

AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis

Research Multimodal & Generative

Merged summary

TL;DR - AuEmoChat is a conversational speech synthesis framework that learns discrete emotion tokens from real emotional speech rather than relying on a few predefined labels. It aims to generate more contextually consistent and authentic emotional speech.

  • AuEmoCodec learns a discrete emotion space using finite scalar quantization.
  • AuEmoToMe merges redundant multimodal dialogue-history tokens while preserving emotion-relevant context.
  • An autoregressive text-speech model predicts both target emotion and speech tokens.
  • Experiments on NCSSD-EmCap show improvements over state-of-the-art CSS baselines in expressiveness and emotional authenticity.

Sources (1)

AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis

arXiv cs.SD Zhenqi Jia, Yuan Zhao, Aruukhan, Rui Liu, Haizhou Li 2026-07-17 arXiv:2607.15755

TL;DR - AuEmoChat is a conversational speech synthesis framework that learns discrete emotion tokens from real emotional speech rather than relying on a few predefined labels. It aims to generate more contextually consistent and authentic emotional speech.

  • AuEmoCodec learns a discrete emotion space using finite scalar quantization.
  • AuEmoToMe merges redundant multimodal dialogue-history tokens while preserving emotion-relevant context.
  • An autoregressive text-speech model predicts both target emotion and speech tokens.
  • Experiments on NCSSD-EmCap show improvements over state-of-the-art CSS baselines in expressiveness and emotional authenticity.
item →