EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation
Ranking
Observed public metrics from 1 member.
Merged summary
TL;DR - EmoWorld is a training-free framework that decouples emotional atmosphere, semantic cues, and temporal progression in a frozen flow-matching Video DiT, enabling controllable emotional video generation. It matters because current generators collapse all affective factors into one text condition, leaving emotion unsteerable at inference.
- Uses a one-time preparation stage that extracts layer-specific affect directions and a reusable cue library from geometry-preserving neutral vs. emotion-edited panoramas.
- Three inference-time steering mechanisms: VAS (injects atmosphere directions into hidden states), SAS (separately scalable prompt residual for semantic cues), and TAS (interpolates endpoint residual fields across denoising and video time).
- On Wan2.2: VAS gives +19% target-emotion alignment and −48% on a temporal-fluctuation proxy; SAS gives +37% alignment and +36% detected affect-bearing cues; TAS improves transition monotonicity by 15% over the strongest baseline.
- Evaluated over 27 emotion categories in T2V and I2V, portable across multiple Video-DiT backbones, and supports camera-conditioned composition without updating generator parameters.
Sources (1)
EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation
TL;DR - EmoWorld is a training-free framework that decouples emotional atmosphere, semantic cues, and temporal progression in a frozen flow-matching Video DiT, enabling controllable emotional video generation. It matters because current generators collapse all affective factors into one text condition, leaving emotion unsteerable at inference.
- Uses a one-time preparation stage that extracts layer-specific affect directions and a reusable cue library from geometry-preserving neutral vs. emotion-edited panoramas.
- Three inference-time steering mechanisms: VAS (injects atmosphere directions into hidden states), SAS (separately scalable prompt residual for semantic cues), and TAS (interpolates endpoint residual fields across denoising and video time).
- On Wan2.2: VAS gives +19% target-emotion alignment and −48% on a temporal-fluctuation proxy; SAS gives +37% alignment and +36% detected affect-bearing cues; TAS improves transition monotonicity by 15% over the strongest baseline.
- Evaluated over 27 emotion categories in T2V and I2V, portable across multiple Video-DiT backbones, and supports camera-conditioned composition without updating generator parameters.