🛰️ Daily AI Frontier
‹ back to 2026-08-09

EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation

arXiv cs.CV Multimodal & Generative Bingyuan Wang, Baistan Zhyldyzbekov, Kunyu Feng, Zeyu Wang 2026-08-06
Representative image for EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation

TL;DR - EmoWorld is a training-free framework that decouples emotional atmosphere, semantic cues, and temporal progression in a frozen flow-matching Video DiT, enabling controllable emotional video generation. It matters because current generators collapse all affective factors into one text condition, leaving emotion unsteerable at inference.

  • Uses a one-time preparation stage that extracts layer-specific affect directions and a reusable cue library from geometry-preserving neutral vs. emotion-edited panoramas.
  • Three inference-time steering mechanisms: VAS (injects atmosphere directions into hidden states), SAS (separately scalable prompt residual for semantic cues), and TAS (interpolates endpoint residual fields across denoising and video time).
  • On Wan2.2: VAS gives +19% target-emotion alignment and −48% on a temporal-fluctuation proxy; SAS gives +37% alignment and +36% detected affect-bearing cues; TAS improves transition monotonicity by 15% over the strongest baseline.
  • Evaluated over 27 emotion categories in T2V and I2V, portable across multiple Video-DiT backbones, and supports camera-conditioned composition without updating generator parameters.

view merged work →