Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis
Merged summary
TL;DR - FacialTalker uses facial-expression tokens to generate conversational speech that better reflects visual affect and context. It matters because facial cues are often neglected in expressive, empathetic speech synthesis.
- AUTokenizer compresses frame-level expressions into discrete tokens supervised by facial Action Unit combinations.
- DualDPO jointly applies preference constraints to visual and speech token sequences.
- VSDD-1K provides 1,033+ hours of synchronized real-world conversational video and speech.
- Objective and subjective experiments report improvements over strong baselines in expression perception, naturalness, expressiveness, and contextual alignment.
Sources (1)
Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis
TL;DR - FacialTalker uses facial-expression tokens to generate conversational speech that better reflects visual affect and context. It matters because facial cues are often neglected in expressive, empathetic speech synthesis.
- AUTokenizer compresses frame-level expressions into discrete tokens supervised by facial Action Unit combinations.
- DualDPO jointly applies preference constraints to visual and speech token sequences.
- VSDD-1K provides 1,033+ hours of synchronized real-world conversational video and speech.
- Objective and subjective experiments report improvements over strong baselines in expression perception, naturalness, expressiveness, and contextual alignment.