🛰️ Daily AI Frontier
‹ back to 2026-07-28

Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis

arXiv cs.HC Multimodal & Generative Yifan Hu, Shuwei He, Rui Liu, Haizhou Li 2026-07-27

TL;DR - FacialTalker uses facial-expression tokens to generate conversational speech that better reflects visual affect and context. It matters because facial cues are often neglected in expressive, empathetic speech synthesis.

  • AUTokenizer compresses frame-level expressions into discrete tokens supervised by facial Action Unit combinations.
  • DualDPO jointly applies preference constraints to visual and speech token sequences.
  • VSDD-1K provides 1,033+ hours of synchronized real-world conversational video and speech.
  • Objective and subjective experiments report improvements over strong baselines in expression perception, naturalness, expressiveness, and contextual alignment.

view merged work →