🛰️ Daily AI Frontier
‹ back to 2026-08-09

EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation

Research Multimodal & Generative

Ranking

Overall 58
Content 65
Popularity 42

Observed public metrics from 1 member.

Representative image for EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation

Merged summary

TL;DR - EmoWorld is a training-free framework that decouples emotional atmosphere, semantic cues, and temporal progression in a frozen flow-matching Video DiT, enabling controllable emotional video generation. It matters because current generators collapse all affective factors into one text condition, leaving emotion unsteerable at inference.

  • Uses a one-time preparation stage that extracts layer-specific affect directions and a reusable cue library from geometry-preserving neutral vs. emotion-edited panoramas.
  • Three inference-time steering mechanisms: VAS (injects atmosphere directions into hidden states), SAS (separately scalable prompt residual for semantic cues), and TAS (interpolates endpoint residual fields across denoising and video time).
  • On Wan2.2: VAS gives +19% target-emotion alignment and −48% on a temporal-fluctuation proxy; SAS gives +37% alignment and +36% detected affect-bearing cues; TAS improves transition monotonicity by 15% over the strongest baseline.
  • Evaluated over 27 emotion categories in T2V and I2V, portable across multiple Video-DiT backbones, and supports camera-conditioned composition without updating generator parameters.

Sources (1)

EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation

arXiv cs.CV Bingyuan Wang, Baistan Zhyldyzbekov, Kunyu Feng, Zeyu Wang 2026-08-06 arXiv:2608.06231
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-02 14:24:56.204013 UTC

TL;DR - EmoWorld is a training-free framework that decouples emotional atmosphere, semantic cues, and temporal progression in a frozen flow-matching Video DiT, enabling controllable emotional video generation. It matters because current generators collapse all affective factors into one text condition, leaving emotion unsteerable at inference.

  • Uses a one-time preparation stage that extracts layer-specific affect directions and a reusable cue library from geometry-preserving neutral vs. emotion-edited panoramas.
  • Three inference-time steering mechanisms: VAS (injects atmosphere directions into hidden states), SAS (separately scalable prompt residual for semantic cues), and TAS (interpolates endpoint residual fields across denoising and video time).
  • On Wan2.2: VAS gives +19% target-emotion alignment and −48% on a temporal-fluctuation proxy; SAS gives +37% alignment and +36% detected affect-bearing cues; TAS improves transition monotonicity by 15% over the strongest baseline.
  • Evaluated over 27 emotion categories in T2V and I2V, portable across multiple Video-DiT backbones, and supports camera-conditioned composition without updating generator parameters.
item →