🛰️ Daily AI Frontier
‹ back to 2026-08-17

Marionette: Predicting World States, Rendering Geometry, Painting Appearance

Research Multimodal & Generative

Ranking

Overall 83
Content 90
Popularity 68

Observed public metrics from 1 member.

Merged summary

TL;DR - Marionette is an interactive game world model that predicts explicit 3D character states, renders geometric controls, and uses video diffusion for photorealistic appearance. This separation enables direct control and rule-based correction of long-horizon errors without retraining the observation model.

  • Predicts a 276-dimensional state containing articulated skeletons, metric root trajectories, and rotations.
  • A zero-parameter renderer computes geometry and occlusion before control-conditioned video synthesis.
  • State-level terrain and separation rules reduced ground penetration by 66% and prevented character drift.
  • Structured state routing showed no detected fidelity loss, scoring 831 FVD versus 799 with recorded poses.

Sources (1)

Marionette: Predicting World States, Rendering Geometry, Painting Appearance

arXiv cs.CV Zian Meng, Zhen Li, Chuanhao Li, Qiang Li, Kaipeng Zhang 2026-08-14 arXiv:2608.14530
Public signals Hugging Face upvotes 34
Providers: Hugging Face · Upvotes 34 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-15 14:32:42.038256 UTC

TL;DR - Marionette is an interactive game world model that predicts explicit 3D character states, renders geometric controls, and uses video diffusion for photorealistic appearance. This separation enables direct control and rule-based correction of long-horizon errors without retraining the observation model.

  • Predicts a 276-dimensional state containing articulated skeletons, metric root trajectories, and rotations.
  • A zero-parameter renderer computes geometry and occlusion before control-conditioned video synthesis.
  • State-level terrain and separation rules reduced ground penetration by 66% and prevented character drift.
  • Structured state routing showed no detected fidelity loss, scoring 831 FVD versus 799 with recorded poses.
item →