🛰️ Daily AI Frontier
‹ back to 2026-09-11

World in World: Explore the World with World Models

arXiv cs.CV Multimodal & Generative Chenxi Song, Yanming Yang, Chi Zhang 2026-09-10
Representative image for World in World: Explore the World with World Models

TL;DR - World in World is a training-free control interface for frozen autoregressive video world models that enables viewpoint-controlled rerendering, long-horizon revisiting, and human-motion transfer. It matters because it unifies several control requirements without task-specific modules or additional model training.

  • Converts video observations, scene projections, geometry renderings, and retrieved generated states into camera- and time-labelled visual evidence.
  • Uses correspondence routing based on persistent point identities and geometry to align target-view queries with source-video tokens.
  • Introduces evidence-wise attention classifier-free guidance to regulate each auxiliary evidence channel independently within one denoising pass.
  • Evaluates perceptual quality, temporal consistency, and camera-following accuracy across diverse viewpoint changes.

view merged work →