World in World: Explore the World with World Models
TL;DR - World in World is a training-free control interface for frozen autoregressive video world models that enables viewpoint-controlled rerendering, long-horizon revisiting, and human-motion transfer. It matters because it unifies several control requirements without task-specific modules or additional model training.
- Converts video observations, scene projections, geometry renderings, and retrieved generated states into camera- and time-labelled visual evidence.
- Uses correspondence routing based on persistent point identities and geometry to align target-view queries with source-video tokens.
- Introduces evidence-wise attention classifier-free guidance to regulate each auxiliary evidence channel independently within one denoising pass.
- Evaluates perceptual quality, temporal consistency, and camera-following accuracy across diverse viewpoint changes.